VLDB 2026 Research / reviewers in the wild / expert
Wei Zhou 0011
dblp:69/5011-11
· DBLP profile ↗
98ranked-venue papers
4as first author
87since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 38 · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 1 first-author · 18 since 2021Security and privacy · 14 · 14 since 2021Systems, architecture and hardware · 8 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 7 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 7 · 7 since 2021Computer networks · 5 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MFmamba: A Multi-function Network for Panchromatic Image Resolution Restoration Based on State-Space ModelabstractRemote sensing images are becoming increasingly widespread in military, earth resource exploration. Because of the limitation of a single sensor, we can obtain high spatial resolution grayscale panchromatic (PAN) images and low spatial resolution color multispectral (MS) images. Therefore, an important issue is to obtain a color image with high spatial resolution when there is only a PAN image at the input. The existing methods improve spatial resolution using super-resolution (SR) technology and spectral recovery using colorization technology. However, the SR technique cannot improve the spectral resolution, and the colorization technique cannot improve the spatial resolution. Moreover, the pansharpening method needs two registered inputs and can not achieve SR. As a result, an integrated approach is expected. We designed a novel multi-function model (MFmamba) to realize the tasks of SR, spectral recovery, joint SR and spectral recovery through three different inputs. Firstly, MFmamba utilizes UNet++ as the backbone, and a Mamba Upsample Block (MUB) is combined with UNet++. Secondly, a Dual Pool Attention (DPA) is designed to replace the skip connection in UNet++. Finally, a Multi-scale Hybrid Cross Block (MHCB) is proposed for initial feature extraction. Many experiments show that MFmamba is competitive in evaluation metrics and visual results and performs well in the three tasks when only the input PAN image is used. Qianqian Wang 0013, Xin Jin 0005, Michal Wozniak 0001, Shaowen Yao 0001, Wei Zhou 0011 |
AAAI | 6 |
| 2026 | Drifting Away from Truth: GenAI-Driven News Diversity Challenges LVLM-Based Misinformation DetectionabstractThe proliferation of multimodal misinformation poses growing threats to public discourse and societal trust. While Large Vision-Language Models (LVLMs) have enabled recent progress in multimodal misinformation detection (MMD), the rise of generative AI (GenAI) tools introduces a new challenge: GenAI-driven news diversity, characterized by highly varied and complex content. We show that this diversity induces multi-level drift, comprising (1) model-level misperception drift, where stylistic variations disrupt a model’s internal reasoning, and (2) evidence-level drift, where expression diversity degrades the quality or relevance of retrieved external evidence. These drifts significantly degrade the robustness of current LVLM-based MMD systems. To systematically study this problem, we introduce DriftBench, a large-scale benchmark comprising 16,000 news instances across six categories of diversification. We design three evaluation tasks: (1) robustness of truth verification under multi-level drift; (2) susceptibility to adversarial evidence contamination generated by GenAI; and (3) analysis of reasoning consistency across diverse inputs. Experiments with six state-of-the-art LVLM-based detectors show substantial performance drops (average F1 ↓ 14.8%) and increasingly unstable reasoning traces, with even more severe failures under adversarial evidence injection. Our findings uncover fundamental vulnerabilities in existing MMD systems and suggest an urgent need for more resilient approaches in the GenAI era. Fanxiao Li, Tingchao Fu, Yunyun Dong, Bingbing Song, Wei Zhou 0011 |
AAAI | 6 |
| 2026 | Hierarchical Dual-Domain Fusion with Frequency-Guided Spatial Modeling for Pan-SharpeningabstractPan-sharpening aims to generate high-resolution multispectral images by integrating the spectral richness of low-resolution multispectral images with the spatial details of high-resolution panchromatic images. Although frequency-domain modeling shows great potential in this field, most existing methods are still limited to spatial-domain processing or fail to effectively capture the contextual interactions between frequency and spatial features. To address these issues, we propose a novel multi-scale frequency-spatial collaborative fusion approach. A Frequency-Spatial U-Net is introduced as the backbone network, in which frequency-spatial modeling blocks are embedded to progressively enhance the frequency-guided spatial contextual modeling capability across layers. To this end, we design a Dual Branch Frequency Attention module that adaptively enhances high- and low-frequency information. In addition, we introduce fine-mid-coarse resolution branches and devise a main-auxiliary multi-scale reconstruction loss to facilitate collaborative optimization. The effectiveness of the proposed model is validated through extensive experiments, demonstrating superior performance in both qualitative and quantitative evaluations. Moreover, our model achieves the fastest inference time among all compared methods, striking an excellent balance between accuracy and efficiency. Huangqimei Zheng, Chengyi Pan, Wei Zhou 0011, Xin Jin 0005 |
AAAI | 4 |
| 2026 | Correct When Paired, Wrong When Split: Decoupling and Editing Modality-Specific Neurons in MLLMsabstractTingchao Fu, Wenkai Wang, Fanxiao Li, Huadong Zhang, Jinhong Zhang, Dayang Li, Yunyun Dong, Renyang Liu, Wei Zhou. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Tingchao Fu, Fanxiao Li, Dayang Li, Yunyun Dong, Renyang Liu 0001, Wei Zhou 0011 |
ACL (1) | 9 |
| 2026 | What's Left Unsaid? Detecting and Correcting Misleading Omissions in Multimodal News PreviewsabstractEven when factually correct, social-media news previews (image-headline pairs) can induce interpretation drift: by selectively omitting crucial context, they lead readers to form judgments that diverge from what the full article supports.This covert harm is subtler than explicit misinformation, yet remains underexplored.To address this gap, we develop a multistage pipeline that simulates preview-based and context-based understanding, enabling construction of the MM-MISLEADING benchmark.Using MM-MISLEADING, we systematically evaluate open-source LVLMs and uncover pronounced blind spots in omission-based misleadingness detection.We further propose OM-GUARD, which combines (1) Interpretation-Aware Fine-Tuning for misleadingness detection and (2) Rationale-Guided Misleading Content Correction, where explicit rationales guide headline rewriting to reduce misleading impressions.Experiments show that OM-GUARD lifts an 8B model's detection accuracy to the level of a 235B LVLM while delivering markedly stronger end-to-end correction.Further analysis shows that misleadingness usually arises from local narrative shifts, such as missing background, instead of global frame changes, and identifies image-driven cases where text-only correction fails, underscoring the need for visual interventions. Fanxiao Li, Tingchao Fu, Dayang Li, Herun Wan, Wei Zhou 0011, Min-Yen Kan |
ACL (1) | 6 |
| 2026 | DPSA: Deception Pattern Learning and Sentiment-Aware Enhancement for Unseen Misinformation Detection
Yunyun Dong, Jinfeng Luo, Tingchao Fu, Fanxiao Li, Dayang Li, Viradeth Sixanonh, Wei Zhou 0011 |
DASFAA (5) | 7 |
| 2026 | Let It Try First: Uncertainty-Guided Retrieval Switching for Retrieval-Augmented Question Answering
Xiaohai Wang, Wei Zhou 0011, Sixing Wu |
WWW | 3 |
| 2026 | Estimating difficulty for the open-ended reading comprehension questions via Large Language Models
Sixing Wu, Weiqun He, Wei Zhou 0011 |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | Attention-guided network for infrared unmanned aerial vehicle target detection
Xin Jin 0005, Puming Wang, Shin-Jye Lee, Shaowen Yao 0001, Wangming Lan, Wei Zhou 0011 |
Eng. Appl. Artif. Intell. | 9 |
| 2026 | Memory poisoning attacks on retrieval-augmented Large Language Model agents via deceptive semantic reasoning
Fanxiao Li, Yunyun Dong, Wei Zhou 0011, Renyang Liu 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | Chat with one voice: mitigating semantic disparity across languages for multilingual open-domain dialogue response generation systems
Sixing Wu, Wei Zhou 0011 |
Expert Syst. Appl. | 4 |
| 2026 | CRFD: A novel face privacy preservation via fine-grained controllable and reversible de-identification
Jiabao Zhou, Wei Zhou 0011, Chao Yi, Bingbing Song |
Expert Syst. Appl. | 3 |
| 2026 | Dual-label guided unrestricted target attack with diffusion model
Jinming Cui, Jun Ji, Shaowen Yao 0001, Wei Zhou 0011 |
Neurocomputing | 6 |
| 2026 | RWP: a robust watermarking plugin for attribution and protection in stable diffusion models
Yunyun Dong, Bingbing Song, Wei Zhou 0011 |
Neural Networks | 5 |
| 2026 | TDFG-GAN: Top-down-feature guided GAN for thermal infrared image colorization
Hongyue Huang, Wei Zhou 0011, Xin Jin 0005 |
Pattern Anal. Appl. | 6 |
| 2026 | HyDK: A Hybrid DRL-KKT Framework for Latency-Critical Service Placement With Multi-Source SynchronizationabstractIn mission-critical IoT-MEC environments, jointly optimizing service placement and resource allocation is intractable due to the high-dimensional coupling of discrete topological decisions with continuous resource dimensioning. Furthermore, traditional methods oversimplify dependencies, overlooking multi-source “Wait-for-All” synchronization and the stochastic variance of bursty workloads. To bridge these gaps, we propose HyDK, a variance-aware framework synergizing Deep Reinforcement Learning (DRL) with convex optimization. The core innovation is our Action Space Pruning mechanism. We theoretically decompose the hybrid decision space by solving the continuous sub-problem to optimality via a KKT-based convex optimization routine. This acts as a deterministic optimality backstop, effectively pruning continuous dimensions and allowing the agent to focus exclusively on the complex discrete topological search. To address physical realities, we construct a finegrained Directed Acyclic Graph (DAG) model to capture data aggregation bottlenecks and integrate an M/G/1 queuing model incorporating the second moment of service time to mitigate longtail latency risks. Trace-driven simulations using the Edge-IIoTset demonstrate that HyDK improves system responsiveness by up to 25.6% with significantly tighter confidence intervals compared to existing baselines. Zhenli He, Xiaolong Zhai, Mingxiong Zhao 0001, Wei Zhou 0011, Keqin Li 0001 |
IEEE Trans. Serv. Comput. | 4 |
| 2025 | Rethinking Machine Unlearning in Image Generation ModelsabstractWith the surge and widespread application of image generation models, data privacy and content safety have become major concerns and attracted great attention from users, service providers, and policymakers. Machine unlearning (MU) is recognized as a cost effective and promising means to address these challenges. Despite some advancements, image generation model unlearning (IGMU) still faces remarkable gaps in practice, e.g., unclear task discrimination and unlearning guidelines, lack of an effective evaluation framework, and unreliable evaluation metrics. These can hinder the understanding of unlearning mechanisms and the design of practical unlearning algorithms. We perform exhaustive assessments over existing state-of-the-art unlearning algorithms and evaluation standards, and discover several critical flaws and challenges in IGMU tasks. Driven by these limitations, we make several core contributions, to facilitate the comprehensive understanding, standardized categorization, and reliable evaluation of IGMU. Specifically, (1) We design CatIGMU, a novel hierarchical task categorization framework. It provides detailed implementation guidance for IGMU, assisting in the design of unlearning algorithms and the construction of testbeds. (2) We introduce EvalIGMU, a comprehensive evaluation framework. It includes reliable quantitative metrics across five critical aspects. (3) We construct DataIGM, a high-quality unlearning dataset, which can be used for extensive evaluations of IGMU, training content detectors for judgment, and benchmarking the state-of-the-art unlearning algorithms. With EvalIGMU and DataIGM, we discover that most existing IGMU algorithms cannot handle the unlearning well across different evaluation dimensions, especially for preservation and robustness. Data, source code, and models are available at https://github.com/ryliu68/IGMU. Renyang Liu 0001, Wenjie Feng 0001, Tianwei Zhang 0004, Wei Zhou 0011, Xueqi Cheng 0001, See-Kiong Ng |
CCS | 4 |
| 2025 | Asking Questions with Thoughts: An Efficient Difficulty-Controllable Question Generation Method with Posterior Knowledge DistillationabstractDifficulty Controllable Question Generation (DCQG) for reading comprehension learns to generate questions for measuring the reading abilities of examinees, playing a crucial role in educational scenarios. This work studies answer-aware DCQG, a challenging task that requires the generated questions to remain faithful to the assigned answer and match the desired difficulty level at the same time. To this end, we first propose an effective two-stage framework, Asking Questions with Thoughts (AQT), to guide a backbone large language model (LLM) to generate questions that are both faithful and difficulty-aware through conducting in-depth self-thinking. Then, we introduce a novel Posterior Knowledge Distillation (PKD) to efficiently fine-tune AQT by distilling knowledge from posterior inference. Finally, to address the scarcity of DCQG datasets, we use an efficient LLM Pretest-based Difficulty Estimation (LP-DE) to automatically construct DCQG datasets from common QG/QA datasets. Extensive experiments prove that our methods have promising results in terms of both faithfulness and difficulty awareness. Sixing Wu, Yujue Zhou, Wei Zhou 0011 |
CIKM | 5 |
| 2025 | GCQ-ViT: Group-Aware Collaborative Post-Training Quantization for Vision TransformersabstractPost-training quantization (PTQ) is widely utilized in Vision Transformers (ViTs) for its computational efficiency and retraining elimination. However, the unique architecture of ViTs introduces significant quantization challenges. Dynamic fluctuations in channel activations, particularly post-LayerNorm, result in distributional mismatches. Additionally, the heavy-tailed nature of post-Softmax activations compromises the accurate representation of critical attention regions, vital for ViT performance. Moreover, weight quantization at low bit-widths leads to a loss of structural information, degrading global feature representation. To address these challenges, we introduce the Group-aware Collaborative Quantization framework (GCQ-ViT), which significantly improves both the accuracy and efficiency of ViT quantization. The GCQ-ViT framework integrates a novel dynamic perception grouping quantization mechanism to ensure distributional consistency within groups, thus reducing hardware expense. It also utilizes a self-adaptive displaced uniform log2 quantizer, optimizing shift factors and nonlinear intervals to enhance representation in high-density regions of post-Softmax activations. Additionally, we propose a dynamic dimension-aware error compensation method to correct quantization errors across channel dimensions using a residual mean compensation skill, ensuring robust feature preservation. Extensive experiments on image classification, object detection, and instance segmentation tasks demonstrate that GCQ-ViT outperforms the current leading PTQ methods, setting a new benchmark for ViT quantization. Pan Peng 0006, Wei Zhou 0011 |
ECAI | 4 |
| 2025 | D-Judge: How Far Are We? Assessing the Discrepancies Between AI-synthesized and Natural Images through Multimodal GuidanceabstractIn the rapidly evolving field of Artificial Intelligence Generated Content (AIGC), a central challenge is distinguishing AI-synthesized images from natural images. Despite the impressive capabilities of advanced AI generative models in producing visually compelling content, significant discrepancies remain when compared to natural images. To systematically investigate and quantify these differences, we construct a large-scale multimodal dataset named DANI, comprising 5,000 natural images and over 440,000 AI-generated image (AIGI) samples produced by nine representative models using both unimodal and multimodal prompts, including Text-to-Image (T2I), Text-and-Image-to-Image (I2I), and Text and Image-to-Image (TI2I). We then introduce D-Judge, a benchmark designed to answer the critical question: how far are AI-generated images from truly realistic images? Our fine-grained evaluation framework assesses DANI across five key dimensions: naive visual quality, semantic alignment, aesthetic appeal, downstream task applicability, and coordinated human validation. Extensive experiments reveal substantial discrepancies across these dimensions, highlighting the importance of aligning quantitative metrics with human judgment to achieve a comprehensive understanding of AI-generated image quality. The code and dataset are publicly available at: https://github.com/ryliu68/DJudge, and https://huggingface.co/datasets/Renyang/DANI. Renyang Liu 0001, Ziyu Lyu, Wei Zhou 0011, See-Kiong Ng |
ACM Multimedia | 3 |
| 2025 | IMRRF: Integrating Multi-Source Retrieval and Redundancy Filtering for LLM-based Fake News DetectionabstractDayang Li, Fanxiao Li, Bingbing Song, Li Tang, Wei Zhou. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Dayang Li, Fanxiao Li, Bingbing Song, Wei Zhou 0011 |
NAACL (Long Papers) | 5 |
| 2025 | IPAttack: imperceptible adversarial patch to attack object detectors
Yongming Wen, Peiyuan Si, Wei Zhou 0011, Zongheng Zhao, Chao Yi, Renyang Liu 0001 |
Appl. Intell. | 3 |
| 2025 | Transferable adversarial attacks for multi-model systems coupling image fusion with classification modelsabstractAbstract Image preprocessing models typically serve as the initial step in advanced visual tasks, aiming to enhance the performance of subsequent tasks. For example, multi-focus image fusion technology significantly improves the performance of downstream semantic classification tasks. However, with the advancement of adversarial attack techniques, these models are facing significant challenges. Previous research has only explored the impact of adversarial attacks on the performance of individual models, lacking an in-depth investigation into the robustness of tasks involving the combination of multiple models. This study aims to delve into the robustness issues of tasks that combine multi-focus image fusion and image classification. To address this challenge, we have designed a new adversarial attack generator specifically for scenarios that combine multi-focus image fusion with image classification. This attack method uses a decision map surrogate model and a binary weight map to precisely add adversarial perturbations to the effective information parts of multi-focus images. It also incorporates attention mechanisms and Grad-CAM technology to optimize the perturbation areas, aiming to disrupt the key features of the fused image to improve the transferability of the attack. Comprehensive experimental results show that this method significantly improves the efficiency of attacks on downstream classification tasks while maintaining the effectiveness of the fusion model. Xin Jin 0005, Xueshuai Gao, Puming Wang, Shaowen Yao 0001, Wei Zhou 0011 |
Cybersecur. | 7 |
| 2025 | MixEI: Mixing explicit and implicit commonsense knowledge in open-domain dialogue response generation
Sixing Wu, Wei Zhou 0011 |
Neurocomputing | 3 |
| 2025 | A comprehensive review of network pruning based on pruning granularity and pruning time perspectives
Kehan Zhu, Fuyi Hu, Yuanbing Ding, Wei Zhou 0011, Ruxin Wang 0002 |
Neurocomputing | 4 |
| 2025 | DMNet: A Dense Multiscale Feature Extraction Network With Two-Stage Training for Infrared-Visible Image FusionabstractWith the increasing need for intelligent and secure multimedia systems, infrared and visible image fusion (IVIF) has garnered a lot of attention due to its ability to overcome the limitations of a single sensor and integrate unique information from different modalities. However, it is common to overlook how the spatial frequency information of visible and infrared images differs. A less thorough feature extraction may result from many approaches’ inability to reconcile the extraction of both global and local information. To solve the aforementioned difficulties, we propose a dense multiscale fusion network DMNet. Through a dual-stream collaborative feature decoupling, the proposed network optimizes both the encoder–decoder network and the diffusion model to extract multimodal information more comprehensively. Specifically, the three-stage progressive encoder sequentially integrates dense transformer block (DTB) and dense invertible neural network block (DIB) to achieve global feature extraction and multimodal feature decoupling. Our proposed channel and spatial attention block (CSAB) selectively focuses on the important feature maps to better capture the critical information. Additionally, multiscale latent features are extracted by the diffusion module (DM) to enhance the representation of cross-modal latent features. As demonstrated by extensive experiments, DMNet outperforms representative state-of-the-art methods. Furthermore, we conduct sufficient ablation experiments to validate each module’s effectiveness, and we demonstrate that DMNet can enhance downstream infrared-visible object detection performance. Our fused results and code will be accessible athttps://github.com/Pancy9476/DMNet. Chengyi Pan, Huangqimei Zheng, Hongyue Huang, Xin Jin 0005, Keqin Li 0001, Wei Zhou 0011 |
IEEE Internet Things J. | 7 |
| 2025 | LODAP: On-device incremental learning via lightweight operations and data pruning
Biqing Duan, Di Liu 0002, Wei Zhou 0011, Zhenli He, Shengfa Miao |
J. Syst. Archit. | 4 |
| 2025 | Multimodal alignment augmentation transferable attack on vision-language pre-training models
Tingchao Fu, Fanxiao Li, Wei Zhou 0011 |
Pattern Recognit. Lett. | 6 |
| 2025 | Joint Computation Offloading and Resource Allocation in Mobile-Edge Cloud Computing: A Two-Layer Game ApproachabstractMobile-Edge Cloud Computing (MECC) plays a crucial role in balancing low-latency services at the edge with the computational capabilities of cloud data centers (DCs). However, many existing studies focus on single-provider settings or limit their analysis to interactions between mobile devices (MDs) and edge servers (ESs), often overlooking the competition that occurs among ESs from different providers. This article introduces an innovative two-layer game framework that captures independent self-interested competition among MDs and ESs, providing a more accurate reflection of multi-vendor environments. Additionally, the framework explores the influence of cloud-edge collaboration on ES competition, offering new insights into these dynamics. The proposed model extends previous research by developing algorithms that optimize task offloading and resource allocation strategies for both MDs and ESs, ensuring the convergence to Nash equilibrium in both layers. Simulation results demonstrate the potential of the framework to improve resource efficiency and system responsiveness in multi-provider MECC environments. Zhenli He, Ying Guo 0018, Xiaolong Zhai, Mingxiong Zhao 0001, Wei Zhou 0011, Keqin Li 0001 |
IEEE Trans. Cloud Comput. | 5 |
| 2025 | STBA: Towards Evaluating the Robustness of DNNs for Query-Limited Black-Box ScenarioabstractExtensive studies have revealed that deep neural networks (DNNs) are vulnerable to adversarial attacks, especially black-box ones, which can heavily threaten the DNNs deployed in the real world. Many attack techniques have been proposed to explore the vulnerability of DNNs and further help to improve their robustness. Despite the significant progress made recently, existing black-box attack methods still suffer from unsatisfactory performance due to the vast number of queries needed to optimize desired perturbations. Besides, the other critical challenge is that adversarial examples built in a noise-adding manner are abnormal and struggle to successfully attack robust models, whose robustness is enhanced by adversarial training against small perturbations. There is no doubt that these two issues mentioned above will significantly increase the risk of exposure and result in a failure to dig deeply into the vulnerability of DNNs. Hence, it is necessary to evaluate DNNs' fragility sufficiently under query-limited settings in a non-additional way. In this paper, we propose the Spatial Transform Black-box Attack (STBA), a novel framework to craft formidable adversarial examples in the query-limited scenario. Specifically, STBA introduces a flow field to the high-frequency part of clean images to generate adversarial examples and adopts the following two processes to enhance their naturalness and significantly improve the query efficiency: a) we apply an estimated flow field to the high-frequency part of clean images to generate adversarial examples instead of introducing external noise to the benign image, and b) we leverage an efficient gradient estimation method based on a batch of samples to optimize such an ideal flow field under query-limited settings. Compared to existing score-based black-box baselines, extensive experiments indicated that STBA could effectively improve the imperceptibility of the adversarial examples and remarkably boost the attack success rate under query-limited settings. Renyang Liu 0001, Kwok-Yan Lam, Wei Zhou 0011, Sixing Wu, Jun Zhao 0007, Dongting Hu, Mingming Gong |
IEEE Trans. Multim. | 3 |
| 2025 | Crafting imperceptible and transferable adversarial examples: leveraging conditional residual generator and wavelet transforms to deceive deepfake detection
Xin Jin 0005, Puming Wang, Shin-Jye Lee, Shaowen Yao 0001, Wei Zhou 0011 |
Vis. Comput. | 7 |
| 2025 | DS-GAN: a dual sub-structure GAN for thermal infrared image colorization using U-Net with ConvNeXt and multi-scale large kernel attention
Guoliang Yao, Xin Jin 0005, Michal Wozniak 0001, Shengfa Miao, Shaowen Yao 0001, Wei Zhou 0011 |
Vis. Comput. | 7 |
| 2024 | Deep Learning Service for Efficient Data Distribution Aware SortingabstractIn this paper, we present a neural network-enabled data distribution aware sorting method, coined as NN-sort. Our approach explores the potential of developing deep learning techniques to speed up large-scale sort operations, enabling data distribution aware sorting as a deep learning service. Compared to traditional pairwise comparison-based sorting algorithms, which sort data elements by performing pairwise operations, NN-sort leverages the neural network model to learn the data distribution and uses it to map large-scale data elements into ordered ones. Our experiments demonstrate the significant advantage of using NN-sort. Measurements on both synthetic and real-world datasets show that NN-sort yields 2.18× to 10× performance improvement over traditional sorting algorithms. Xiaoke Zhu, Qi Zhang 0009, Wei Zhou 0011, Ling Liu 0001 |
IEEE Big Data | 3 |
| 2024 | SSTA: Salient Spatially Transformed AttackabstractExtensive studies have demonstrated that deep neural networks (DNNs) are vulnerable to adversarial examples (AEs), which brings a huge security risk to the application of DNNs, especially for the AI models developed in the real world. To impede the process of fully exploiting the vulnerabilities of existing DNNs and further improving their robustness in the face of such malicious inputs, many attack methods have been proposed to build AEs. Despite the significant progress that has been made recently, existing attack methods still suffer from the unsatisfactory performance of escaping from being detected by naked human eyes due to the formulation of AE heavily relying on a noise-adding manner. Such mentioned challenges will significantly increase the risk of exposure and result in an attack to be failed. Therefore, in this paper, we propose the Salient Spatially Transformed Attack (SSTA), a novel framework to craft imperceptible AEs, which enhance the stealthiness of AEs by estimating a smooth spatial transform metric on a most critical area to generate AEs instead of adding external noise to the whole image. Compared to SOTA baselines, extensive experiments indicated that SSTA could effectively improve the imperceptibility of the AEs while maintaining a 100% attack success rate. Renyang Liu 0001, Wei Zhou 0011, Sixing Wu, Jun Zhao 0007, Kwok-Yan Lam |
ICASSP | 2 |
| 2024 | CNFA: Conditional Normalizing Flow for Query-Limited AttackabstractTraditional black-box attack methods rely on sufficient feedback from the victim model through a large number of queries until the attack is successful. This may not be acceptable in real applications, since the deployed system may be equipped with certain defense mechanisms and only return the final result (i.e., hard label) to the client. In contrast, one possible approach is formulating a hard label attack, which can be successfully executed within limited queries. To implement this idea, in this paper, we bypass the reliance on victim models and benefit from the intrinsic characteristics of adversarial examples (AEs) and the transferability of examples across different data-driven models. This motivates us to generatively reformulate the attack problem and propose a conditional normalized flow-based attack (CNFA), which builds up a statistical mapping from the benign example to its adversarial counterpart by tackling the conditional likelihood under the hard-label black-box setting. A well-trained CNFA model can directly and efficiently generate a batch of AEs for specific condition inputs. Extensive experiments validate the effectiveness of the proposed idea in a hard-label black-box setting and the superiority of CNFA over SOTA techniques. Renyang Liu 0001, Wei Zhou 0011, Haoran Li 0023, Ruxin Wang 0002 |
ICASSP | 2 |
| 2024 | Imperceptible Text Steganography based on Group ChatabstractText steganography is a technique for hiding secret messages within texts. Previous approaches neglect the contextual relevance of generated stego texts (texts containing secrets) and consistently transmitted secret messages unidirectionally. This behavior is considered anomalous and thus arouses the suspicion of potential attackers. In this paper, we first propose a text steganography framework grounded in the group chat scenario named GCStego, aimed at enhancing the behavior imperceptibility. Additionally, we employ a large language model (LLM) to generate stego texts according to the chatting history and thus boosts the contextual relevance. The proposed scheme is well-suited for secret transmission in group chatting, where multiple agents can pass secret messages through stego texts like regular conversations. Furthermore, we propose token index-based encoding, position filtering and sentence split strategies to deliver the performance. Experimental results demonstrate the superiority of our proposed framework in terms of text semantic controllability, behavioral imperceptibility, and anti-steganalysis ability. Fanxiao Li, Tingchao Fu, Wei Zhou 0011 |
ICME | 5 |
| 2024 | TA-ASF: Attention-Sensitive Token Sampling and Fusing for Visual Transformer Models on the EdgeabstractVision Transformers ($V$iTs) have made significant progress in achieving performance comparable to traditional convolutional neural networks in computer vision tasks. However, high computational complexity restricts their application to resource-constrained edge devices. Previous methods for pruning redundant tokens have shown that it is possible to balance performance and computational cost by reducing the number of tokens. Unfortunately, simply removing redundant tokens often leads to the loss of crucial information. To address this issue, we propose a novel token compression scheme called TA-ASF. This scheme considers both the global role of low-importance tokens and the redundancy among similar tokens. TA-ASF employs novel approaches for token sampling and fusion, which are directly applicable to$V$iTs without introducing additional trainable parameters. A comprehensive evaluation against several edge devices demonstrates our method effectively reduces model complexity while preserving Top-1 accuracy. Experimental results show that on the ImageNet dataset, the proposed method reduces FLOPs by 37% and increases throughput by 1.48 times on the DeiT-S model, with only a 0.1% decrease in accuracy. Specifically, on the DeiT-B model, the proposed method decreases FLOPs by 35% and increases throughput by 1.52 times while maintaining the same accuracy. Junquan Chen, Xingzhou Zhang, Wei Zhou 0011, Weisong Shi |
SEC | 3 |
| 2024 | Generative Steganography Based on Dual-Branch Flow
Bingbing Song, Wei Zhou 0011, Chao Yi, Yunyun Dong |
PRCV (2) | 4 |
| 2024 | DTA: distribution transform-based attack for query-limited scenarioabstractAbstract In generating adversarial examples, the conventional black-box attack methods rely on sufficient feedback from the to-be-attacked models by repeatedly querying until the attack is successful, which usually results in thousands of trials during an attack. This may be unacceptable in real applications since Machine Learning as a Service Platform (MLaaS) usually only returns the final result (i.e., hard-label) to the client and a system equipped with certain defense mechanisms could easily detect malicious queries. By contrast, a feasible way is a hard-label attack that simulates an attacked action being permitted to conduct a limited number of queries. To implement this idea, in this paper, we bypass the dependency on the to-be-attacked model and benefit from the characteristics of the distributions of adversarial examples to reformulate the attack problem in a distribution transform manner and propose a distribution transform-based attack (DTA). DTA builds a statistical mapping from the benign example to its adversarial counterparts by tackling the conditional likelihood under the hard-label black-box settings. In this way, it is no longer necessary to query the target model frequently. A well-trained DTA model can directly and efficiently generate a batch of adversarial examples for a certain input, which can be used to attack un-seen models based on the assumed transferability. Furthermore, we surprisingly find that the well-trained DTA model is not sensitive to the semantic spaces of the training dataset, meaning that the model yields acceptable attack performance on other datasets. Extensive experiments validate the effectiveness of the proposed idea and the superiority of DTA over the state-of-the-art. Renyang Liu 0001, Wei Zhou 0011, Xin Jin 0005, Yuanyu Wang, Ruxin Wang 0002 |
Cybersecur. | 2 |
| 2024 | A survey on Deep-Learning-based image steganography
Bingbing Song, Sixing Wu, Wei Zhou 0011 |
Expert Syst. Appl. | 5 |
| 2024 | SIHNet: A safe image hiding method with less information leakingabstractAbstract Image hiding is a task that hides secret images into cover images. The purposes of image hiding are to ensure the secret images are invisible to the human and the secret images can be recovered. The current state‐of‐the‐art steganography methods run the risk of secret information leakage. A safe image hiding network (SIHNet) is presented to reduce the leakage of secret information. Based on some phenomena of image hiding methods which use invertible neural network, a reversible secret image processing (SIP) module is proposed to make the secret images suitable for hiding and make the stego images leak less secret information. Besides, a reversible lost information hiding (LIH) module is used to hide the lost information into the cover images, thus the method can recover the secret images better than the method that uses random noise to replace the lost information. Experimental results show that SIHNet outperforms other state‐of‐the‐art methods on the PSNR and SSIM values of the recovered secret images and the stego images. Besides, residual images of other state‐of‐the‐art methods all contain information about secret images while residual images of SIHNet leak almost no secret information. Thus the method can prevent the listener of transmission channel from obtaining the information of the secret image through the residual image, which means SIHNet performs better in security than other state‐of‐the‐art methods. Zien Cheng, Xin Jin 0005, Liwen Wu, Yunyun Dong, Wei Zhou 0011 |
IET Image Process. | 6 |
| 2024 | Hiding image with inception transformerabstractAbstract Image steganography aims to hide secret data in the cover media for covert communication. Though many deep‐learning‐based image steganography methods have been presented, these approaches suffer from the inefficiency of building long‐distance connections between the cover and secret images, leading to noticeable modification traces and poor steganalysis resistance. To improve the visual imperceptibility of generated stego images, it is essential to establish a global correlation between the cover and secret images. In this way, the secret image can be dispersed throughout the cover image globally. To bridge this gap, a novel image steganography framework called HiiT is proposed, which takes advantage of CNN and Transformer to learn both the local and global pixel correlation in image hiding. Specifically, a new Transformer structure called Inception Transformer is proposed, which incorporates the Inception Net in the attention‐based Transformer architecture. The Inception Net can learn different scaled image features using multiple convolution kernels, while the attention mechanism can learn the global pixel correlation. By this, the proposed Inception Transformer learns the long‐distance pixel dependency between the cover and secret images. Furthermore, we propose a ‘Skip Connection’ mechanism in the proposed Inception Transformer, which merges the low‐level visual features and high‐level semantic features and achieves better model performance. In detail, The HiiT generates higher‐quality stego images with 45.46 PSNR and 0.9915 SSIM. Besides, accurately restored secret images achieve 47.27 PSNR and 0.9952 SSIM. Extensive experimental results show the proposed HiiT significantly improves the image‐hiding performance compared with state‐of‐the‐art methods. Yunyun Dong, Ruxin Wang 0002, Bingbing Song, Tingchu Wei, Wei Zhou 0011 |
IET Image Process. | 6 |
| 2024 | MCDC-Net: Multi-scale forgery image detection network based on central difference convolutionabstractAbstract Generative Adversarial Networks (GANs) emerged thanks to the development of deep neural networks. Forgery images generated by various variants of GANs are widely spread on the Internet, which may be damage personal credibility and cause huge property losses. Thus, numerous methods are proposed to detect forgery images, but most of them are designed to detect forgery faces. Therefore, a method to detect forgery images of various scenes is proposed. In this work, central difference convolution and vanilla convolution (CDC‐Mix) are mixed after considering the depth and width features of neural networks and analyzing the influence of attention on network performance. Based on CDC‐Mix, a separable convolution (SeparableCDC‐Mix) is proposed. The proposed method consists of three parts: (1) CDC‐Mix and SeparableCDC‐Mix are used to extract the gradient information and texture features; (2) CDCM is used to extract the multi‐scale information of the image; (3) multi‐scale fusion module (MS‐Fusion) is used to fuse the multi‐scale information from different locations of the network. A large number of experiments have been carried out on several datasets generated by GAN, and the experimental results show that the proposed method has a great improvement compared with the existing advanced methods. Defen He, Xin Jin 0005, Zien Cheng, Shuai Liu 0009, Shaowen Yao 0001, Wei Zhou 0011 |
IET Image Process. | 7 |
| 2024 | E3-UAV: An Edge-Based Energy-Efficient Object Detection System for Unmanned Aerial VehiclesabstractMotivated by the advances in deep learning techniques, the application of unmanned aerial vehicle (UAV)-based object detection has proliferated across a range of fields, including vehicle counting, fire detection, and city monitoring. While most existing research studies only a subset of the challenges inherent to UAV-based object detection, there are few studies that balance various aspects to design a practical system for energy consumption reduction. In response, we present the E3-UAV, an edge-based energy-efficient object detection system for UAVs. The system is designed to dynamically support various UAV devices, edge devices, and detection algorithms, with the aim of minimizing energy consumption by deciding the most energy-efficient flight parameters (including flight altitude, flight speed, detection algorithm, and sampling rate) required to fulfill the detection requirements of the task. We first present an effective evaluation metric for actual tasks and construct a transparent energy consumption model based on hundreds of actual flight data to formalize the relationship between energy consumption and flight parameters. Then, we present a lightweight energy-efficient priority decision algorithm based on a large quantity of actual flight data to assist the system in deciding flight parameters. Finally, we evaluate the performance of the system, and our experimental results demonstrate that it can significantly decrease energy consumption in real-world scenarios. Additionally, we provide four insights that can assist researchers and engineers in their efforts to study UAV-based object detection further. Jiashun Suo, Xingzhou Zhang, Weisong Shi, Wei Zhou 0011 |
IEEE Internet Things J. | 4 |
| 2024 | Generative commonsense knowledge subgraph retrieval for open-domain dialogue response generation
Sixing Wu, Wei Zhou 0011 |
Neural Networks | 4 |
| 2024 | An efficient training-from-scratch framework with BN-based structural compressor
Fuyi Hu, Wei Zhou 0011, Ruxin Wang 0002 |
Pattern Recognit. | 5 |
| 2024 | Can LSH (locality-sensitive hashing) be replaced by neural network?
Renyang Liu 0001, Jun Zhao 0007, Xing Chu, Wei Zhou 0011, Jing He 0012 |
Soft Comput. | 5 |
| 2024 | Boosting Black-Box Attack to Deep Neural Networks With Conditional Diffusion ModelsabstractExisting black-box attacks have demonstrated promising potential in creating adversarial examples (AE) to deceive deep learning models. Most of these attacks need to handle a vast optimization space and require a large number of queries, hence exhibiting limited practical impacts in real-world scenarios. In this paper, we propose a novel black-box attack strategy, Conditional Diffusion Model Attack (CDMA), to improve the query efficiency of generating AEs under query-limited situations. The key insight of CDMA is to formulate the task of AE synthesis as a distribution transformation problem, i.e., benign examples and their corresponding AEs can be regarded as coming from two distinctive distributions and can transform from each other with a particular converter. Unlike the conventionalquery-and-optimizationapproach, we generate eligible AEs with direct conditional transform using the aforementioned data converter, which can significantly reduce the number of queries needed. CDMA adopts the conditional Denoising Diffusion Probabilistic Model as the converter, which can learn the transformation from clean samples to AEs, and ensure the smooth development of perturbed noise resistant to various defense strategies. We demonstrate the effectiveness and efficiency of CDMA by comparing it with nine state-of-the-art black-box attacks across three benchmark datasets. On average, CDMA can reduce the query count to a handful of times; in most cases, the query count is only ONE. We also show that CDMA can obtain > 99% attack success rate for untargeted attacks over all datasets and targeted attack over CIFAR-10 with the noise budget of ϵ = 16. Renyang Liu 0001, Wei Zhou 0011, Tianwei Zhang 0004, Kangjie Chen, Jun Zhao 0007, Kwok-Yan Lam |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Business Process Retrieval From Large Model Repositories for Industry 4.0abstractThe process model repository has demonstrated unprecedented success in a variety of industrial and process as a service scenarios. With the rapid increase of massive business process-related data under Industry 4.0, effectively retrieval of process models from large process model repositories becomes a critical challenge for process mining, process deployment and process model acquisition. To accelerate the retrieval of process models from a large process repository, existing retrieval methods rely solely on building single dimension process model indices. In this article we show that this single dimension indexing approach is not only inefficient but also cumbersome for supporting high performance retrieval services over large process model repositories. We propose a new business process model indexing and retrieval with structure and behavior fusion. In the indexing stage, we propose a process model index generation paradigm method with two novel features. First, our index algorithm can transform thetrace equivalent process model(TEPM) with complex structures into a process tree, which can better capture process sequence semantics than the existing approach based on block structured process model. Second, we improve the method for computing the process tree edit distance for measuring process model similarity by introducing the process tree similarity method, which can distinguish leaf nodes and non-leaf nodes and improve the limitations of the traditional edit distance algorithm. Extensive experiments using real world process repositories demonstrate that the proposed methods are under polynomial time in both the model index generation and model querying stages, and offer superior retrieval performance compared to existing process model retrieval methods in terms of efficiency, search capability and scope. Rui Zhu 0009, Ling Liu 0001, Wei Zhou 0011, Xuan Zhang 0002, Yeting Chen |
IEEE Trans. Serv. Comput. | 4 |
| 2023 | TIA: Token Importance Transferable Attack on Vision Transformers
Tingchao Fu, Fanxiao Li, Yuanyu Wang, Wei Zhou 0011 |
Inscrypt (2) | 6 |
| 2023 | Rewriting-Stego: Generating Natural and Controllable Steganographic Text with Pre-trained Language Model
Fanxiao Li, Sixing Wu, Shuoxin Wang, Bingbing Song, Renyang Liu 0001, Haoseng Lai, Wei Zhou 0011 |
DASFAA (1) | 8 |
| 2023 | AFLOW: Developing Adversarial Examples Under Extremely Noise-Limited Settings
Renyang Liu 0001, Haoran Li 0023, Yuanyu Wang, Wei Zhou 0011 |
ICICS | 6 |
| 2023 | Multi-scale Features Destructive Universal Adversarial Perturbations
Huangxinyue Wu, Haoran Li 0023, Wei Zhou 0011, Yunyun Dong |
ICICS | 4 |
| 2023 | SCME: A Self-contrastive Method for Data-Free and Query-Limited Model Extraction Attack
Renyang Liu 0001, Kwok-Yan Lam, Jun Zhao 0007, Wei Zhou 0011 |
ICONIP (5) | 5 |
| 2023 | Enhancing MOBA Game Commentary Generation with Fine-Grained Prototype Retrieval
Haoseng Lai, Shuoxin Wang, Sixing Wu, Wei Zhou 0011 |
NLPCC (2) | 6 |
| 2023 | Enhancing Semantic Consistency in Linguistic Steganography via Denosing Auto-Encoder and Semantic-Constrained Huffman Coding
Shuoxin Wang, Fanxiao Li, Haoseng Lai, Sixing Wu, Wei Zhou 0011 |
NLPCC (2) | 6 |
| 2023 | Dial-QP: A Multi-tasking and Keyword-Guided Approach for Enhancing Conversational Query Production
Sixing Wu, Shuoxin Wang, Haoseng Lai, Wei Zhou 0011 |
NLPCC (1) | 5 |
| 2023 | Improving the Adversarial Robustness of Object Detection with Contrastive Learning
Weiwei Zeng, Wei Zhou 0011, Yunyun Dong, Ruxin Wang 0002 |
PRCV (9) | 3 |
| 2023 | Model Inversion Attacks on Homogeneous and Heterogeneous Graph Neural Networks
Renyang Liu 0001, Wei Zhou 0011, Xiaoyuan Liu 0002, Peiyuan Si, Haoran Li 0023 |
SecureComm (1) | 2 |
| 2023 | Adversarial attacks on multi-focus image fusion models
Xin Jin 0005, Xin Jin 0021, Ruxin Wang 0002, Shin-Jye Lee, Shaowen Yao 0001, Wei Zhou 0011 |
Comput. Secur. | 7 |
| 2023 | A theoretical analysis of continuous firing condition for pulse-coupled neural networks with its applications
Xin Jin 0005, Pingfan Zhang, Youwei He, Puming Wang, Jingyu Hou 0001, Wei Zhou 0011, Shaowen Yao 0001 |
Eng. Appl. Artif. Intell. | 7 |
| 2023 | DBCT-Net:A dual branch hybrid CNN-transformer network for remote sensing image fusion
Quanli Wang, Xin Jin 0005, Liwen Wu, Yunchun Zhang, Wei Zhou 0011 |
Expert Syst. Appl. | 6 |
| 2023 | Energy-efficient computation offloading strategy with task priority in cloud assisted multi-access edge computing
Zhenli He, Di Liu 0002, Wei Zhou 0011, Keqin Li 0001 |
Future Gener. Comput. Syst. | 4 |
| 2023 | Improving robustness of convolutional neural networks using element-wise activation scaling
Zhi-Yuan Zhang, Zhenli He, Wei Zhou 0011, Di Liu 0002 |
Future Gener. Comput. Syst. | 4 |
| 2023 | OCAP: On-device Class-Aware Pruning for personalized edge DNN models
Ye-Da Ma, Zhi-chao Zhao, Di Liu 0002, Zhenli He, Wei Zhou 0011 |
J. Syst. Archit. | 5 |
| 2023 | Image colorization using deep convolutional auto-encoder with multi-skip connections
Xin Jin 0005, Yide Di, Xing Chu, Qing Duan, Shaowen Yao 0001, Wei Zhou 0011 |
Soft Comput. | 7 |
| 2023 | Detecting Adversarial Examples on Deep Neural Networks With Mutual Information Neural EstimationabstractDespite achieving exceptional performance, deep neural networks (DNNs) suffer from the harassment caused by adversarial examples, which are produced by corrupting clean examples with tiny perturbations. Many powerful defense methods have been presented such as training data augmentation and input reconstruction which, however, usually rely on the prior knowledge of the targeted models or attacks. A clean example and its adversarial version are very similar but have different high-level representations in a victim model. If we can obtain a space in which the representations of similar examples are also similar, then adversarial examples can be picked out by comparing the representations of input examples in this space and the high-level space of the victim model. Inspired by this, we propose a novel approach for detecting adversarial images, which can protect any pre-trained DNN classifiers and resist an endless stream of new attacks. Specifically, we first adopt a dual autoencoder to project images to a latent space. The dual autoencoder uses the self-supervised learning to ensure that small modifications to samples do not significantly alter their latent representations. Next, the mutual information neural estimation is utilized to enhance the discrimination of the latent representations. We then leverage the prior distribution matching to regularize the latent representations. To easily compare the representations of examples in the two spaces, and not rely on the prior knowledge of the targeted model, a simple fully connected neural network is used to embed the learned representations into an eigenspace, which is consistent with the output eigenspace of the targeted model. Through the distribution similarity of an input example in the two eigenspaces, we can judge whether the input example is adversarial or not. Extensive experiments on MNIST, CIFAR-10, and ImageNet show that the proposed method has superior defense performance and transferability than state-of-the-arts. Ruxin Wang 0002, Shui Yu 0001, Yunyun Dong, Shaowen Yao 0001, Wei Zhou 0011 |
IEEE Trans. Dependable Secur. Comput. | 7 |
| 2023 | Type-I Generative Adversarial AttackabstractDeep neural networks are vulnerable to adversarial attacks either by examples with indistinguishable perturbations which produce incorrect predictions, or by examples with noticeable transformations that are still predicted as the original label. The latter case is known as the Type I attack which, however, has achieved limited attention in literature. We advocate that the vulnerability comes from the ambiguous distributions among different classes in the resultant feature space of the model, which is saying that the examples with different appearances may present similar features. Inspired by this, we propose a novel Type I attack method called generative adversarial attack (GAA). Specifically, GAA aims at exploiting the distribution mapping from the source domain of multiple classes to the target domain of a single class by using generative adversarial networks. A novel loss and a U-net architecture with latent modification are elaborated to ensure the stable transformation between the two domains. In this way, the generated adversarial examples have similar appearances with examples of the target domain, yet obtaining the original prediction by the model being attacked. Extensive experiments on multiple benchmarks demonstrate that the proposed method generates adversarial images that are more visually similar to the target images than the competitors, and the state-of-the-art performance is achieved. Shenghong He, Ruxin Wang 0002, Tongliang Liu, Chao Yi, Xin Jin 0005, Renyang Liu 0001, Wei Zhou 0011 |
IEEE Trans. Dependable Secur. Comput. | 7 |
| 2023 | The Mining of Urban Hotspots Based on Multi-Source Location Data FusionabstractUrban hotspots reflect the degree of residents' travel gathering. The study of urban hotspots has important values for urban infrastructure planning, public security and other aspects. In existing researches, single-source location data and density-based clustering algorithms are used to mine hotspots. Due to the one-sidedness of using the single-source data, the mining of hotspots based on multi-source location data fusion has become a hot topic. Multi-source location data fusion requires a quantity balance between the data sets to be fused, because several famous clustering algorithms cannot handle multi-source imbalanced data sets. To solve this problem, we propose a novel framework to mine urban hotspots. First, we construct a data imputation model for the sparse data set so that reducing the difference in quantity between two types of data sets. Then, a clustering algorithm for imbalanced data sets is proposed, and a novel evaluation metric is designed to verify the effectiveness of clustering results. The experiment uses real data sets including POI data, check-in data and GPS trajectory data. The results show that the proposed method discovers all urban hotspots formed by fused imbalanced data sets, and it is more accurate and efficient than the state-of-the-art algorithms. Cong Sha, Wei Zhou 0011 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2023 | Priority-Based Offloading Optimization in Cloud-Edge Collaborative ComputingabstractAs an emerging computing paradigm, cloud-edge collaborative computing (CECC) combines computing resources at the back-end and the edge of the network to provide more flexible service delivery, thus striking a good balance between abundant computing resources and high responsiveness. However, mobile devices (MDs) must make strategic offloading decisions in such an environment. Although existing research has made remarkable progress in computation offloading strategies, most works ignore multi-priority settings in complex application scenarios. In this article, we focus on the impact of multi-priority settings and mixed queue disciplines on offloading decisions in CECC. First, we utilize queueing models to characterize all computing nodes in the environment and establish mathematical models to describe the considered scenario. Second, we formulate offloading decisions of the target MD into three multi-variable optimization problems to investigate the cost-performance tradeoff. Third, we propose numerical algorithms based on the Karush-Kuhn-Tucke conditions to address these problems. Finally, we construct numerical examples, a comparative experiment, and a simulation experiment to demonstrate the effectiveness of our methods. Our work provides important insights into the optimization of computation offloading for MDs in complex application scenarios, which can help achieve a better cost-performance tradeoff in CECC. Zhenli He, Mingxiong Zhao 0001, Wei Zhou 0011, Keqin Li 0001 |
IEEE Trans. Serv. Comput. | 4 |
| 2023 | Securing Deep Learning as a Service Against Adaptive High Frequency Attacks With MMCATabstractMost cloud providers offer Deep Learning as a Service (DLaaS) for different business, science and engineering domains. However, it is known that deep neural networks (DNNs) are vulnerable to adversarial examples, which can cause well-trained DNN models to misbehave by injecting human-imperceptible perturbations to the query input data. Securing deep learning as a service becomes a critical challenge in mitigating such adversarial input perturbations, and enhancing the robustness of DNNs. In this article, we report two important facts: First, most adversarial perturbations are high frequency signals or are added to high frequency signals. Second, due to Frequency Principle that neural networks overly pay attention to fit the low frequency signals during training, the models could be easily misled by the high frequency signals of adversarial examples. These facts consequently contribute to the vulnerability of DNNs service in the Cloud. We conjecture that the more robust the neural networks are in learning from high frequency signals, the more resilient these neural networks are against adversarial perturbed examples. We propose a novel method for generating high-frequency-enhanced adversarial examples, which is achieved by a high-pass filter in the frequency domain via Fourier Transform. This method enhances the learning ability for high frequency signals and ameliorates to over-fit useless low frequency signals. In order to improve the robustness of DNNs service under such signal frequency attacks, we propose a multi-modal collaborative adversarial training framework, named as MMCAT, which uses the multi-modal information of the input images for cross-modal collaborative training, delivering excellent extension for effectively learning of multi-modal image information. Extensive experiments show that under strong adaptive frequency attacks, the DNNs service trained with the proposed MMCAT method achieve superior performance and robustness over the state-of-the-art adversarial training approaches. Bingbing Song, Ruxin Wang 0002, Yunyun Dong, Ling Liu 0001, Wei Zhou 0011 |
IEEE Trans. Serv. Comput. | 6 |
| 2022 | RIA: A Reversible Network-based Imperceptible Adversarial AttackabstractThe robustness and security of deep neural network (DNN) models have received much attention in recent years. In-depth research on adversarial example generation methods that make DNN models make wrong judgments and decisions will facilitate further research on more comprehensive and practical adversarial defense methods. Most existing adversarial example generation methods focus too much on attack performance and design adversarial noise at the pixel level, resulting in the generated adversarial examples with redundant noise and evident perturbations. In this paper, we try to find the well-designed perturbations at the feature-level and propose a novel deep reversible network-based imperceptible adversarial examples generation method called RIA. Experimental results show that RIA can obtain more natural adversarial examples without losing attack performance and reducing redundant noise based on well-designed feature maps. To the best of our knowledge, in the white-box attack method research, this work is the first attempt to directly add perturbations to feature maps and use an reversible network to generate adversarial examples based on the perturbed feature maps. Fanxiao Li, Renyang Liu 0001, Zhenli He, Yunyun Dong, Wei Zhou 0011 |
ICTAI | 6 |
| 2022 | Towards Query-limited Adversarial Attacks on Graph Neural NetworksabstractGraph Neural Network (GNN) is a graph representation learning approach for graph-structured data, which has witnessed a remarkable progress in the past few years. As a counterpart, the robustness of such a model has also received considerable attention. Previous studies show that the performance of a well-trained GNN can be faded by black-box adversarial examples significantly. In practice, the attacker can only query the target model with very limited counts, yet the existing methods require hundreds of thousand queries to extend attacks, leading the attacker to be exposed easily. To perform a step forward in addressing this issue, in this paper, we propose a novel attack methods, namely Graph Query-limited Attack (GQA), in which we generate adversarial examples on the surrogate model to fool the target model. Specifically, in GQA, we use contrastive learning to fit the feature extraction layers of the surrogate model in a query-free manner, which can reduce the need of queries. Furthermore, in order to utilize query results sufficiently, we obtain a series of queries with rich information by changing the input iteratively, and storing them in a buffer for recycling usage. Experiments show that GQA can decrease the accuracy of the target model by 4.8%, with only 1% edges modified and 100 queries performed. Haoran Li 0023, Liwen Wu, Wei Zhou 0011, Ruxin Wang 0002 |
ICTAI | 5 |
| 2022 | Multiple Feature Mining Based on Local Correlation and Frequency Information for Face Forgery DetectionabstractAs facial image manipulation techniques developed, deep fake detection attracted extensive attentions. Although researchers have made remarkable progresses in deepfake detection recently, which is still suffering from two limitations: a) current detectors achieve high accuracy in the high-quality videos and images, but it is hard to capture local and subtle artifacts in the low-quality and high-compression media; b) few of deep fake detection methods gain satisfying performance under cross-database scenario, because detector overfit to specific color textures producing by same manipulation algorithm. Inspired the above issues, this paper proposes a novel framework fusing local related features and frequency information to mine the forgery patterns. Firstly, we design multi-feature enhancement module, which amplifies implicit local disc repancies and capture spatial correlation from three shallow feature layers and high-level semantic layer guided by attention maps. Secondly, dual frequency decomposition module is proposed for disassembling high-frequency and low-frequency features, the forgery artifacts are exposed after dual cross attention block processing in the frequency spectrum. Features from the two streams are fused to the classification for the final result. Comprehensive experiments demonstrate the superior performance of our proposed approach in the low-quality benchmark database and cross-dataset sce-nario. Shuai Liu 0009, Xin Jin 0005, Zhenli He, Wei Zhou 0011, Shaowen Yao 0001, Qiannian Wang |
ICTAI | 5 |
| 2022 | Deepfake Detection Using Multiple Feature Fusion
Xin Jin 0005, Yunyun Dong, Shaowen Yao 0001, Wei Zhou 0011 |
IFIP Int. Conf. Digital Forensics | 7 |
| 2022 | Low-power Robustness Learning Framework for Adversarial Attack on EdgesabstractRecent works on adversarial attacks uncover the intrinsic vulnerability of neural networks, which reveal a critical issue that the neural networks are easily misled by adversarial attacks. As the development of edge computing, more and more real-time tasks are deployed on edge devices. The safety of these neural network-based applications is threatened by adversarial attack. Therefore, the defense technique against adversarial attack has very important application value for edges. Especially, the defense technique should consider the deployment condition on edges, such as low power and low time consumption. Unfortunately, until now, very limited research considers the security problem under adversarial attack on edges. In this paper, we propose a low-power robust learning framework to deal with the adversarial attacks at resource-constrained edge devices. In this framework, we make a rough categorization of approaches on defending against adversarial attacks, and reveal how this edge device-based framework can be used to resist adversarial attacks. Furthermore, we propose a staged ensemble defense strategy in the framework, which achieves better defensive performance than a single defense algorithm. To verify our framework on real application, we build a Drone Search and Rescue System (DSRS) which is employed to examine the performance of the proposed framework. The results indicate that our framework achieves outstanding performance in all aspects, such as robustness, time and power consumption. Multiple evaluations of the low-power robust learning framework provide the advice that help to choose the optimal security configuration on power-constrained and performance-expected environments. Bingbing Song, Haiyang Chen 0003, Jiashun Suo, Wei Zhou 0011 |
MSN | 4 |
| 2022 | Multi-pipeline HotStuff: A High Performance Consensus for Permissioned BlockchainabstractThe state-of-the-art HotStuff operates an efficient pipeline in which a stable leader drives decisions with linear communication. However, with the unifying proposing-voting pattern, it takes two rounds of messages to produce a certified proposal, which severely limits the performance of the consensus protocol and makes it difficult for the blockchain to exert the bandwidth and concurrency of modern operating systems. Thus, this paper developed a new consensus protocol, called Multi-pipeline HotStuff, for permissioned blockchain. To the best of the authors’ knowledge, this is the first protocol that combines multiple pipelines of HotStuff to propose batches in order, such that proposals are built optimistically when a correct replica realizes that the current proposal is valid and will be certified by quorum votes in the near future. Simultaneous proposing and voting allow the protocol to produce more proposals in every two rounds of messages, it further boosts the throughput at a comparable latency with that of HotStuff. The evaluation experiment confirmed that the throughput of the proposed protocol outperformed HotStuff by approximately 60% without significantly increasing end-to-end latency under varying system sizes. Even if the protocol frequently performs view-change phase due to network asynchrony, its optimization continues to demonstrate better performance. Taining Cheng, Wei Zhou 0011, Shaowen Yao 0001, Libo Feng, Jing He 0012 |
TrustCom | 2 |
| 2022 | A novel deviation density peaks clustering algorithm and its applications of medical image segmentationabstractAbstract The density peaks clustering (DPC) algorithm can identify clusters with various shapes and densities in the underlying dataset. However, the DPC algorithm cannot exactly find the true quantity of clustering centers when computing the local density, and it is difficult to handle non‐convex datasets. Moreover, the DPC algorithm is difficult to identify boundary points and outliers without a reasonable allocation strategy when dealing with low‐density points. To solve these limitations, a novel deviation density peaks clustering (DeDPC) algorithm is proposed. First, the local deviation of the spatial distance of datasets with different structures is utilised to replace the local density to generate a more reasonable clustering center decision graph. Second, a threshold is defined to further divide and process low‐density points. Finally, outliers in low‐density points can be accurately found to accurately cluster the dataset. To evaluate the performance of the DeDPC algorithm, experiments are conducted on synthetic and real‐world datasets and the DeDPC is compared with other clustering methods. The DeDPC is also applied to medical image segmentation to further demonstrate its capability for medical image processing. The simulation results show that the DeDPC method has good validity and utility for both non‐convex datasets and medical image segmentation. Wei Zhou 0011, Limin Wang 0011, Xuming Han |
IET Image Process. | 1 |
| 2022 | Higher-Order Proximity-Based MiRNA-Disease Associations PredictionabstractMiRNA-disease association prediction plays an important role in identifying human disease-related miRNAs. This approach is helpful not only to formulate individualized diagnosis schemes, but also to understand the pathogenesis of diseases. Many studies have focused on enhancing the prediction performance using explicit side information, such as miRNA functional similarity and disease semantic similarity. The existing approaches, however, often ignore the higher-order implicit proximity among miRNAs and diseases. To this end, in this paper, we first propose a novel approach HOP_MDA (Higher-Order Proximity based MiRNA and Disease Association Prediction) for predicting potential association between miRNA and disease. Both explicit interaction information and implicit higher-order proximity information between miRNA and disease are encoded with different order proximity matrices which are weightily combined into a parameterized prediction matrix. A supervised learning approach based on the known miRNAs-disease associations is proposed to determine the optimal weight parameters. The prediction matrix is then used to achieve effective prediction. Additionally, a higher-order proximity approximation technique (HOPA_MDA) is presented to make more efficient predictions. 5-fold cross validation is used to evaluate the performance of our proposed method. The average AUC values of HOPA_MDA for two real datasets are 0.921+/-0.002 and 0.944+/-0.0015, respectively. Our method can also predict potential miRNAs specific to new diseases with no known related miRNAs. Jin Li 0007, Wei Zhou 0011, Tong Li 0004 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2022 | How to Analyze the Neurodynamic Characteristics of Pulse-Coupled Neural Networks? A Theoretical Analysis and Case Study of Intersecting Cortical ModelabstractThe intersecting cortical model (ICM), initially designed for image processing, is a special case of the biologically inspired pulse-coupled neural-network (PCNN) models. Although the ICM has been widely used, few studies concern the internal activities and firing conditions of the neuron, which may lead to an invalid model in the application. Furthermore, the lack of theoretical analysis has led to inappropriate parameter settings and consequent limitations on ICM applications. To address this deficiency, we first study the continuous firing condition of ICM neurons to determine the restrictions that exist between network parameters and the input signal. Second, we investigate the neuron pulse period to understand the neural firing mechanism. Third, we derive the relationship between the continuous firing condition and the neural pulse period, and the relationship can prove the validity of the continuous firing condition and the neural pulse period as well. A solid understanding of the neural firing mechanism is helpful in setting appropriate parameters and in providing a theoretical basis for widespread applications to use the ICM model effectively. Extensive experiments of numerical tests with a common image reveal the rationality of our theoretical results. Xin Jin 0005, Dongming Zhou 0001, Xing Chu, Shaowen Yao 0001, Keqin Li 0001, Wei Zhou 0011 |
IEEE Trans. Cybern. | 7 |
| 2022 | A Deep Multitask Convolutional Neural Network for Remote Sensing Image Super-Resolution and ColorizationabstractRemote sensing data have become increasingly vital in target detection, disaster monitoring, and military surveillance. Abundant pan-sharpening and super-resolution (SR) methods based on deep learning have been proposed and have achieved remarkable performance. However, pan-sharpening requires paired panchromatic (PAN) and multispectral (MS) images, and SR cannot increase the spectral resolution of PAN. Thus, we introduce a computational imaging-based method to recover or produce the incomplete data of single PAN or MS. This work also explores the integration of multiple tasks by a single neural network. We start with SR and colorization, study the feasibility of simultaneously finishing SR colorization, and use a model trained in SR colorization to finish pan-sharpening without MS. A generic neural network, remote sensing image improvement network (RSI-Net), is designed for remote sensing image SR, colorization, simultaneous SR colorization, and pan-sharpening. To verify its performance, RSI-Net is compared with the state-of-the-art SR and colorization methods. Experiments show that RSI-Net can be competitive in visual effects and evaluation indexes, and it performs well at simultaneous SR colorization, and RSI-Net finishes pan-sharpening and only needs to input PAN. Our experiments confirm the effect of integrating multiple tasks. Jianan Feng, Ching-Hsun Tseng, Xin Jin 0005, Ling Liu 0010, Wei Zhou 0011, Shaowen Yao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Distributed Bayesian Matrix Decomposition for Big Data Mining and ClusteringabstractMatrix decomposition is one of the fundamental tools to discover knowledge from big data generated by modern applications. However, it is still inefficient or infeasible to process very big data using such a method in a single machine. Moreover, big data are often distributedly collected and stored on different machines. Thus, such data generally bear strong heterogeneous noise. It is essential and useful to develop distributed matrix decomposition for big data analytics. Such a method should scale up well, model the heterogeneous noise, and address the communication issue in a distributed system. To this end, we propose a distributed Bayesian matrix decomposition model (DBMD) for big data mining and clustering. Specifically, we adopt three strategies to implement the distributed computing including 1) the accelerated gradient descent, 2) the alternating direction method of multipliers (ADMM), and 3) the statistical inference. We investigate the theoretical convergence behaviors of these algorithms. To address the heterogeneity of the noise, we propose an optimal plug-in weighted average that reduces the variance of the estimation. Synthetic experiments validate our theoretical results, and real-world experiments show that our algorithms scale up well to big data and achieves superior or competing performance compared to two typical distributed methods including Scalable-NMF and scalable k-means++. Chihao Zhang 0002, Wei Zhou 0011 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2021 | DLB: Deep Learning Based Load BalancingabstractIn this paper, we introduce DLB, a Deep Learning based load Balancing mechanism, to effectively address the data skew problem. The key idea of DLB is to replace hash functions in the load balancing mechanisms with deep learning models, which are trained to be able to map different distributions of workloads and data to the servers in a uniformed manner. We implemented DLB and deployed it on a practical Cloud environment using CloudSim. Experimental results using both synthetic and real-world data sets show that compared with traditional hash function based load balancing methods, DLB is able to achieve more balanced mappings, especially when the workload is highly skewed. Xiaoke Zhu, Qi Zhang 0009, Taining Cheng, Ling Liu 0001, Wei Zhou 0011, Jing He 0012 |
CLOUD | 5 |
| 2021 | A novel adaptive density-based spatial clustering of application with noise based on bird swarm optimization algorithm
Limin Wang 0011, Honghuan Wang, Xuming Han, Wei Zhou 0011 |
Comput. Commun. | 4 |
| 2021 | EnsembleFool: A method to generate adversarial examples based on model fusion strategy
Wenyu Peng, Renyang Liu 0001, Ruxin Wang 0002, Taining Cheng, Zifeng Wu, Wei Zhou 0011 |
Comput. Secur. | 7 |
| 2021 | Color-UNet++: A resolution for colorization of grayscale images using improved UNet++
Yide Di, Xiaoke Zhu, Xin Jin 0005, Qiwei Dou, Wei Zhou 0011, Qing Duan |
Multim. Tools Appl. | 5 |
| 2021 | Server configuration optimization in mobile edge computing: A cost-performance tradeoff perspectiveabstractAbstract Before service providers build up an mobile edge computing (MEC) platform, an important issue that needs to be considered is the configuration of computing resources on edge servers. Since the computing resources on an edge server are limited compared with a cloud server and the service provider's deployment budget is limited, it would be unrealistic to equip all edge servers with abundant computing resources. In addition, the edge servers have different computation demands due to their different geographies. Therefore, this article investigates the problem of server configuration optimization in an MEC environment based on a given computation demand statistics of the selected deployment locations. Our strategy is to treat each edge server as an M/M/m queueing model, and then establish the performance and cost models for the system. Two optimization problems, including cost constrained performance optimization, and performance constrained cost optimization are formulated based on our models and solved by a series of fast numerical algorithms. We also conduct extensive numerical simulation examples to show the effectiveness of the proposed algorithms. MEC service providers can use our strategy to get the appropriate type of processor and obtain the optimal processor number for each edge server to achieve two different goals: (1) deliver the highest‐quality services with a given cost constraint; (2) minimize the investment cost with a service‐quality guarantee. Our research is of great significance for service providers to control the tradeoff between investment cost and service quality. Zhenli He, Kenli Li 0001, Keqin Li 0001, Wei Zhou 0011 |
Softw. Pract. Exp. | 4 |
| 2020 | Neural inductive matrix completion with graph convolutional networks for miRNA-disease association predictionabstractMOTIVATION: Predicting the association between microRNAs (miRNAs) and diseases plays an import role in identifying human disease-related miRNAs. As identification of miRNA-disease associations via biological experiments is time-consuming and expensive, computational methods are currently used as effective complements to determine the potential associations between disease and miRNA. RESULTS: We present a novel method of neural inductive matrix completion with graph convolutional network (NIMCGCN) for predicting miRNA-disease association. NIMCGCN first uses graph convolutional networks to learn miRNA and disease latent feature representations from the miRNA and disease similarity networks. Then, learned features were input into a novel neural inductive matrix completion (NIMC) model to generate an association matrix completion. The parameters of NIMCGCN were learned based on the known miRNA-disease association data in a supervised end-to-end way. We compared the proposed method with other state-of-the-art methods. The area under the receiver operating characteristic curve results showed that our method is significantly superior to existing methods. Furthermore, 50, 47 and 48 of the top 50 predicted miRNAs for three high-risk human diseases, namely, colon cancer, lymphoma and kidney cancer, were verified using experimental literature. Finally, 100% prediction accuracy was achieved when breast cancer was used as a case study to evaluate the ability of NIMCGCN for predicting a new disease without any known related miRNAs. AVAILABILITY AND IMPLEMENTATION: https://github.com/ljatynu/NIMCGCN/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jin Li 0007, Chenxi Ning, Zhuoxuan Zhang, Wei Zhou 0011 |
Bioinform. | 6 |
| 2019 | HaloDPC: An Improved Recognition Method on Halo Node for Density Peak Clustering AlgorithmabstractThe density peaks clustering (DPC) is known as an excellent approach to detect some complicated-shaped clusters with high-dimensionality. However, it is not able to detect outliers, hub nodes and boundary nodes, or form low-density clusters. Therefore, halo is adopted to improve the performance of DPC in processing low-density nodes. This paper explores the potential reasons for adopting halos instead of low-density nodes, and proposes an improved recognition method on Halo node for Density Peak Clustering algorithm (HaloDPC). The proposed HaloDPC has improved the ability to deal with varying densities, irregular shapes, the number of clusters, outlier and hub node detection. This paper presents the advantages of the HaloDPC algorithm on several test cases. Jianhua Jiang, Wei Zhou 0011, Limin Wang 0011, Keqin Li 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2018 | A Comparative Study of Containers and Virtual Machines in Big Data EnvironmentabstractContainer technique is gaining increasing attention in recent years and has become an alternative to traditional virtual machines. Some of the primary motivations for the enterprise to adopt the container technology include its conveniency to encapsulate and deploy applications, lightweight operations, as well as efficiency and flexibility in resources sharing. However, there still lacks an in-depth and systematic comparison study on how big data applications, such as Spark jobs, perform between a container environment and a virtual machine environment. In this paper, by running various Spark applications with different configurations, we evaluate the two environments from many interesting aspects, such as how convenient the execution environment can be set up, what are makespans of different workloads running in each setup, how efficient the hardware resources, such as CPU and memory, are utilized, and how well each environment can scale. The results show that compared with virtual machines, containers provide a more easy-to-deploy and scalable environment for big data workloads. The research work in this paper can help practitioners and researchers to make more informed decisions on tuning their cloud environment and configuring the big data applications, so as to achieve better performance and higher resources utilization. Qi Zhang 0009, Ling Liu 0001, Calton Pu, Qiwei Dou, Liren Wu, Wei Zhou 0011 |
IEEE CLOUD | 6 |
| 2018 | Brain Tumor Segmention Based on Dilated Convolution Refine NetworksabstractA brain tumor is a growth of abnormal cells in the tissues of the brain, which is difficult for treatment and severely affects patients' cognitive ability. Recent year magnetic resonance imaging (MRI) has been widely used imaging technique to assess brain tumors. However manual segmentation and artificial extracting features block MRI's practice when facing with the huge amount of data produced by MRI. An efficient and automatic image segmentation of brain tumor is still needed. In this paper, a novel automatic segmentation framework of brain tumors, which have 5 parts and resnet-50 use as a backbone, is proposed based on convolutional neural network. A dilated convolution refine (DCR) structure is introduced to extract the local features and global features. After investigating different parameters of our framework, it is proved that DCR is an efficient and robust method in Brain Tumor Segmentation. The experiments are evaluated by Multimodal Brain Tumor Image Segmentation (BRATS 2015) dataset. The results show that our framework in complete tumor segmentation achieved excellent results with a DEC score of 0.87 and a PPV score of 0.92. (GitHub: https://github.com/wei-lab/DCR) Di Liu 0002, Xiaojuan Yu, Shaowen Yao 0001, Wei Zhou 0011 |
SERA | 6 |
| 2018 | Quadboost: A Scalable Concurrent QuadtreeabstractBuilding concurrent spatial trees is more complicated than binary search trees since a space hierarchy should be preserved during modifications. We present a non-blocking quadtree (quadboost) that supports concurrent insert, remove, move, and contain operations, in which the move operation combines the searches for different keys together and modifies different positions atomically. To increase its concurrency, a decoupling approach is proposed to separate physical adjustment from logical removal within the remove operation. In addition, we design a continuous find mechanism to reduce the search cost. Experimental results show that quadboost scales well on a multi-core system with 32 hardware threads. It outperforms existing concurrent trees in retrieving two-dimensional keys with up to 109 percent improvement when the number of threads is large. Furthermore, the move operation achieves better performance than the best-known algorithm with up to 47 percent. Keren Zhou 0001, Guangming Tan, Wei Zhou 0011 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2017 | An Experimental Study of a Biosequence Big Data Analysis ServiceabstractWith the development of next-generation sequencing (NGS), DNA/RNA sequencing has become cheaper and more efficient. Today, a whole human genome can be sequenced under $1,000, providing opportunities for large-scale bioinformatic analysis on big datasets. However, most of existing bioinformatic analysis tools are programmed for single server based computing platform and not suitable to process such big datasets. As Hadoop MapReduce and Spark are gaining popularity as cluster computing based big data processing platform, more and more bioinformatic applications start to explore cluster computing platform for large scale data analysis. In this paper we present an in-depth experimental study on deploying Spark clusters for high performance bioinformatic short sequence reconstruction. Our experimental results enable us to answer a number of challenging and yet most frequently asked questions regarding efficient management of bioinformatic data analysis services on Spark systems. Example questions include how to best split big dataset into multiple partitions, and how to distribute data partitions and bioinformatic analysis tasks on a Spark cluster for carrying out a high performance distributed analysis job? What types of memory models are effective for bioinformatic data analysis services on a Spark cluster? Why do different bioinformatic data analysis operations exhibit different throughput performance on the same Spark cluster? We conjecture that this experimental study not only demonstrates the feasibility of high performance bioinformatic data analysis on Spark platform, but also will help bioinformatic application developers to make more informed decisions on both design and configuration of Spark Cluster, managing and tuning parameters of Spark runtime system for enhancing the performance of large scale big data analytics. Wei Zhou 0011, Ling Liu 0001, Calton Pu, Qingyang Wang 0001, Wenkun Xiang, Shaowen Yao 0001 |
ICWS | 1 |
| 2017 | MetaSpark: a spark-based distributed processing tool to recruit metagenomic reads to reference genomesabstractSummary: With the advent of next-generation sequencing, traditional bioinformatics tools are challenged by massive raw metagenomic datasets. One of the bottlenecks of metagenomic studies is lack of large-scale and cloud computing suitable data analysis tools. In this paper, we proposed a Spark based tool, called MetaSpark, to recruit metagenomic reads to reference genomes. MetaSpark benefits from the distributed data set (RDD) of Spark, which makes it able to cache data set in memory across cluster nodes and scale well with the datasets. Compared with previous metagenomics recruitment tools, MetaSpark recruited significantly more reads than many programs such as SOAP2, BWA and LAST and increased recruited reads by ∼4% compared with FR-HIT when there were 1 million reads and 0.75 GB references. Different test cases demonstrate MetaSpark's scalability and overall high performance. Availability: https://github.com/zhouweiyg/metaspark. Contact: [email protected] , [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Wei Zhou 0011, Shaowen Yao 0001 |
Bioinform. | 1 |
| 2016 | HDCache: A Distributed Cache System for Real-Time Cloud Services
Jing Zhang 0015, Qianmu Li, Wei Zhou 0011 |
J. Grid Comput. | 3 |
| 2015 | Coupled Interdependent Attribute Analysis on Mixed DataabstractIn the real-world applications, heterogeneous interdependent attributes that consist of both discrete and numerical variables can be observed ubiquitously. The usual representation of these data sets is an information table, assuming the independence of attributes. However, very often, they are actually interdependent on one another, either explicitly or implicitly. Limited research has been conducted in analyzing such attribute interactions, which causes the analysis results to be more local than global. This paper proposes the coupled heterogeneous attribute analysis to capture the interdependence among mixed data by addressing coupling context and coupling weights in unsupervised learning. Such global couplings integrate the interactions within discrete attributes, within numerical attributes and across them to form the coupled representation for mixed type objects based on dimension conversion and feature selection. This work makes one step forward towards explicitly modeling the interdependence of heterogeneous attributes among mixed data, verified by the applications in data structure analysis, data clustering evaluation, and density comparison. Substantial experiments on 12 UCI data sets show that our approach can effectively capture the global couplings of heterogeneous attributes and outperforms the state-of-the-art methods, supported by statistical analysis. Can Wang 0004, Chihung Chi, Wei Zhou 0011, Raymond K. Wong 0001 |
AAAI | 3 |
| 2006 | A New Architecture of Data Access Middleware under Grid EnvironmentabstractData sharing is one of the most important research areas in data grid. Distributed data resource and heterogeneous data schema bring difficulties to data access and sharing. This article mainly focuses on how to deal with the heterogeneity of data schema and put forward a blueprint to solve the data access and sharing problem of heterogeneous physical data resources in the grid. To solve this problem we propose the SDB resource model which extracts three data layers from the physical data resource in order to facilitate the access and sharing of data resource Qingyang Wang 0001, Jingshu Chen, Xibin Gao, Wei Zhou 0011, Baoping Yan |
APSCC | 4 |
| 2006 | Truthful application-layer multicast in mesh-based selfish overlaysabstractBuilding a multicast tree in an overlay network is the major problem in application layer multicast. In a selfish overlay network, the end hosts may not be cooperative. They are willing to maximize their own benefits instead of being loyal to the protocol or altruistically contributing to the benefits of the whole overlay network. In this paper, we design a strategyproof mechanism in building a truthful minimum cost multicast tree in the mesh-tree way, which could simplify the overlay construction and maintenance. We model the network in both simple and general cost model. We conduct simulation to study the relationship between utility of cheating node and other nodes Wei Zhou 0011, Chihung Chi |
IPCCC | 1 |