Chuang Li 0004

dblp:10/4825-4 · DBLP profile ↗
← Back
20ranked-venue papers
11as first author
17since 2021 · last 2026
0000-0002-3139-2583ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 6 first-author · 4 since 2021Computer networks · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 ST-GCN and Reinforcement Learning-Assisted Dynamic Multistrategy Task Offloading in Edge-IoT Vehicular Networks
abstract
The Internet of Things (IoT) enables intelligent transportation services by connecting vehicles with roadside infrastructure and generating time-sensitive data. To support low-latency processing, edge-IoT vehicular networks deploy distributed edge servers near mobile users. However, high vehicular mobility and heterogeneous edge resources make it difficult for existing approaches to effectively exploit spatio-temporal mobility patterns and to support real-time offloading decisions. To address these challenges, this paper proposes TPADO, a Trajectory Prediction-Aware Dynamic Offloading framework that integrates a Spatio-Temporal Graph Convolutional Network (ST-GCN) with a multi-agent decision mechanism based on Proximal Policy Optimization (PPO). TPADO employs ST-GCN to perform high-fidelity trajectory prediction by explicitly modeling the graph structure of vehicular networks, thereby enabling proactive candidate-node selection and mobility-aware result delivery. Based on the predicted mobility information, we further design a hierarchical multi-strategy offloading framework, where a DRL-based policy layer adaptively selects offloading strategies, and a rule layer performs fine-grained node assignment and task partitioning. Extensive simulation results demonstrate that TPADO achieves the best overall performance among the compared methods. Compared with the centralized DQN baseline, it reduces global average latency by 6.4% and system saturation by 3.01 percentage points, while also delivering higher throughput and task success rate. These results validate the effectiveness and generalizability of the proposed framework.
Chuang Li 0004, Gang Liu 0038, Yanhua Wen, Junyan Hu, Qingyu Shi 0001, Zhao Tong 0001
IEEE Internet Things J.1
2026 Privacy-Preserving Federated Multimodal Agriproduct Anomaly Detection in AIoT via Modality-Under-Optimized Knowledge Distillation
abstract
As a key application of the Agricultural Internet of Things (AIoT), multimodal agriproduct anomaly detection faces severe privacy and security challenges. Existing federated learning (FL) methods struggle to capture fine-grained cross-modal correlations and to address the modality under-optimization problem, thereby limiting both detection accuracy and privacy levels. To this end, this paper proposes a privacy-preserving federated multimodal agriproduct anomaly detection scheme in AIoT based on modality-under-optimized knowledge distillation (PAMAD), achieving high-utility anomaly detection with enhanced privacy protection. Specifically, we develop a hierarchical privacy protection method for multimodal fine-grained alignment fusion based on meta-learning (HPPMF), which effectively captures cross-modal semantic correlations and protects the privacy of fused features. In addition, we propose a multi-task pre-training algorithm based on modality-under-optimized knowledge distillation (MTLMKD) to alleviate modal imbalance. We further design a pre-training dynamic protection algorithm based on adaptive gradient quantization (PDAGC) to ensure model security. Subsequently, a multimodal agriproduct anomaly detection method with a self-supervised denoising encoder (MAPADSE) is introduced to improve detection accuracy under noisy conditions. Rigorous security analysis demonstrates that the PAMAD scheme satisfies differential privacy. Experimental results show that, compared with existing state-of-the-art methods, our PAMAD scheme improves AUROC and accuracy by 7.56% and 8.71%, respectively, achieving a desirable balance between privacy protection and anomaly detection accuracy in AIoT services.
Chuang Li 0004, Yanhua Wen, Limei Liu, Qingyu Shi 0001
IEEE Internet Things J.3
2026 MC-ORAM: A Concurrent ORAM Scheme for Multi-User Shared Storage
abstract
The expansion of cloud-based shared storage increases data privacy concerns. While data encryption technologies can safeguard data content, they cannot prevent the leakage of data access patterns. By re-encrypting data and changing storage location after each access, Oblivious Random Access Machine (ORAM) can effectively avoid information leakage from memory access patterns. While ORAM was initially designed for single-user applications, most existing multi-user ORAM solutions have drawbacks, such as dependence on a trusted proxy, high client storage overhead, and low throughput. To address the issues of multi-user ORAM systems, this paper explores the design of proxyless ORAM solutions in shared storage scenarios and proposes MC-ORAM, a new multi-user oblivious data storage framework. It ensures data consistency through client collaboration in a proxyless architecture, achieves higher throughput by differentiating request processing based on privacy protection requirements, and uses recursion for optimization. We implemented MC-ORAM and analyzed its performance using a variety of indicators, as well as conducting comparative evaluations against alternative schemes. The results show that, on average, MC-ORAM reduces response latency by 18.1% and improves throughput by 23.8% compared to TaoStore.
Chuang Li 0004, Duo Hu, Gang Liu 0038, Yanhua Wen, Zhuo Tang
IEEE Trans. Computers1
2026 VM-ORAM: A Novel High-Performance ORAM Architecture for Efficient Data Integrity Verification in Industrial Cloud
abstract
With the rapid surge in industrial data, cloud computing has been integrated into Industrial Internet of Things (IIoT) systems to store, compute, and share massive data. In this process, privacy and data integrity are core concerns. Oblivious RAM (i.e., ORAM) is a technology widely applied to defend against cloud storage access pattern attacks. However, most existing ORAM systems do not consider the integration of data integrity verification technology. Although there are integrity verification systems integrated into conventional Path ORAM and Ring ORAM, they are not suitable for existing new ORAM systems. And the existing data integrity verification ORAM system still has the problem of excessive performance overhead. To address these challenges, this article proposes a novel high-performance data integrity ORAM system, VM-ORAM. Optimizes the ORAM integrity verification process by integrating dynamic scheduling and multipath eviction strategies, thereby minimizing performance loss. The comprehensive analysis and experimental results of this article show that VM-ORAM system not only defends against data tampering attacks, but also maintains high performance of the system.
Chuang Li 0004, Gang Liu 0038, Changyao Tan, Limei Liu, Wenhua Ye, Anthony T. Chronopoulos
IEEE Trans. Ind. Informatics1
2026 In-Network Load Balancing With Fast Congestion Flow Detection for Lossless Data Center Networks
abstract
To meet the high performance requirement of real time and critical applications in industrial Internet of Things, modern lossless ethernet data center networks (DCNs) deployed with remote direct memory access and priority-based flow control (PFC) are dedicated to delivering low latency and high bandwidth. However, existing load balancing schemes either lack sub-round-trip-time congestion sensing or fail to accurately detect and reroute flows that cause congestion in PFC-enabled lossless DCNs. Therefore, we propose LBoDSN, an in-network load balancing for lossless DCNs using direct switch notification (DSN) for fast congestion flow detection, to address above challenges. LBoDSN tracks the evolution of ingress queue lengths at destination switches to anticipate the initiation of PFC pause, precisely identifies congested flows before PFC pause, and then sends DSNs to source switches to perform rerouting. After rerouting, the congestion notification packet associated with the previous path is selectively discarded to improve transmission performance. Experiments under realistic workloads reveal that, LBoDSN outperforms CONGA by 13%–65%, and 25%–80% in average and tail Flow Completion Times (FCTs), respectively. Compared to ConWeave, LBoDSN achieves approximately 9% improvement in both average and tail FCTs, while reducing switch queue consumption for reordering.
Qingyu Shi 0001, Fangxue Jiang, Chuang Li 0004, Xiaocui Li 0001, Wenzhi Cao, Limei Liu
IEEE Trans. Ind. Informatics3
2026 Personalized Privacy-Preserving Task Allocation in Spatial Crowdsourcing
abstract
As a popular service management system, the spatial crowdsourcing (SC) server is responsible for allocating nearby workers to perform tasks based on outsourced locations. However, protecting the sensitive information contained in these outsourced locations is crucial. Traditional differential privacy (DP) methods suffer from two limitations: 1) they usually rely on a trusted third party, failing to protect both worker and task location privacy simultaneously, thus risking privacy breaches; 2) they ignore the personalized privacy demands of different users. In this paper, we propose a personalized local DP-based location obfuscation (PLDPLO) scheme, thereby providing personalized privacy-preserving both worker and task locations locally while allocating high-quality tasks. To achieve this, we introduce a personalized location indistinguishability (PLI) model, a new personalized Laplace mechanism achieving local DP, to jointly provide the protection of worker locations and different privacy levels for different workers. To address task privacy, we present a spatial mapping indistinguishability (SMI) algorithm to obfuscate task locations based on a random response mechanism, thereby ensuring data utility. Additionally, we propose a Zipf-Poisson model-based task allocation graph (ZPTAG) algorithm to perform one-task-multiple-workers allocation and achieve a high competitive ratio, which reduces the move distance of workers. Our PLDPLO scheme guarantees ϵ-LDP. Extensive experiments over real datasets demonstrate that our scheme achieves over 89% data utility for task allocation and outperforms state-of-the-art methods while providing personalized privacy levels.
Xiaolong Li 0004, Jun Cai 0001, Xin Yao 0002, Jin Zhang 0018, Yanhua Wen, Chuang Li 0004
IEEE Trans. Netw. Serv. Manag.9
2025 CST-ViT: Cascaded Spatio-Temporal Redundancy Elimination for Efficient Vision Transformers on Edge IoT Devices
abstract
Transformer-based models have demonstrated outstanding performance in video understanding tasks due to their capacity to capture long-range dependencies. However, their high computational cost, along with the massive volume of streaming video data, presents significant challenges for real-time deployment on resource-constrained edge devices integrated into internet of things (IoT) systems. Existing approaches typically eliminate spatial or temporal redundancy in isolation, failing to fully exploit the inherent spatio-temporal similarity in video data. To address this limitation, we propose CST-ViT, a cascaded spatio-temporal redundancy elimination framework that jointly reduces dynamic temporal and intra-frame spatial redundancy. CST-ViT incorporates three gating modules: the direct temporal gate for matching unchanged backgrounds, the offset temporal gate for capturing motion-related changes, and the spatial gate for intra-frame similarity matching. Together with a spatiotemporal caching and token reuse mechanism, CST-ViT enables efficient token filtering and computation reuse. Experimental results show that CST-ViT reduces computation by 55.88% with no loss in accuracy, and achieves up to a 74.75% reduction in computation with less than 1% accuracy degradation, outperforming state-of-the-art methods in terms of accuracy–efficiency trade-off for video transformers.
Qinyu Wang 0002, Xiaofeng Zou, Chuang Li 0004, Yujie Peng, Heshi Wang, Yanhua Wen, Minaer Yeerlan, Cen Chen 0002
IEEE Internet Things J.3
2025 Privacy-Preserving Sparse Traffic Flow Prediction in IIoT: A Three-Tier Federated Learning Framework
abstract
Traffic flow prediction, as a typical application of Industrial Internet of Things (IIoT) in urban infrastructure, faces critical security challenges. Existing privacy-preserving methods in two-tier federated learning (FL) frameworks primarily focus on dense data while neglecting privacy vulnerabilities in massive sparse traffic flow collected by clients, failing to effectively protect both high-sparsity traffic flow and federated pretrained models against privacy leakage risks. Therefore, this article proposes a novel three-tier FL framework-based privacy-preserving sparse traffic flow prediction (TFLST) scheme, achieving dual protection of sparse traffic flow and model parameters with high-precision prediction. Specifically, we innovatively design a spatiotemporal self-attention transformer-based Gestalt sparse key cell selection (STGSC) method to efficiently extract sparse key cells with high spatiotemporal correlations. Additionally, an adaptive truncated Gaussian mechanism-based local sparse traffic flow protection (ATLSP) algorithm is proposed, which dynamically allocates privacy budgets according to sparse correlations to achieve high-utility sparse data protection. A dynamic spatiotemporal matrix completion-based GCN pretraining protection (DSMGP) method is adopted to enhance the spatiotemporal features of sparse data efficiently, protect model parameter privacy, and improve FL training accuracy. Subsequently, we introduce a spatiotemporal self-supervised learning-based multiobjective weighted traffic flow prediction (SMWTP) method to achieve high-accuracy traffic flow prediction. Rigorous security analysis proves that our scheme satisfies differential privacy requirements. Experimental results on four real-world datasets show that our TFLST scheme reduces prediction errors by 6.21% compared to state-of-the-art methods, effectively balancing data privacy and utility.
Tingsen Zhou, Chuang Li 0004, Xin Yao 0002, Limei Liu, Yanhua Wen
IEEE Internet Things J.3
2025 STVAI: Exploring spatio-temporal similarity for scalable and efficient intelligent video inference
Chuang Li 0004, Heshi Wang, Yanhua Wen, Qingyu Shi 0001, Qinyu Wang 0002, Dongchen Wu
J. Parallel Distributed Comput.1
2025 DC-ORAM: An ORAM Scheme Based on Dynamic Compression of Data Blocks and Position Map
abstract
Oblivious RAM (ORAM) is an efficient cryptographic primitive that prevents leakage of memory access patterns. It has been referenced by modern secure processors and plays an important role in memory security protection. Although the most advanced ORAM has made great progress in performance optimization, the access overhead (i.e., data blocks) and on-chip (i.e., PosMap) storage overhead is still too high, which will lead to problems such as low system performance. To overcome the above challenges, in this paper, we propose a DC-ORAM system, which reduces the data access overhead and on-chip PosMap storage overhead by using dynamic compression technology. Specifically, we use byte stream redundancy compression technology to compress data blocks on the ORAM tree. And in PosMap, a high-bit multiplexing strategy is used to achieve data compression for binary high-bit repeated data of leaf labels (or path labels). By introducing the above compression technology, in this work, compared with conventional Path ORAM, the compression rate of the ORAM tree is$52.9\%$, and the compression rate of PosMap is$40.0\%$. In terms of performance, compared to conventional Path ORAM, our proposed DC-ORAM system reduces the average latency by$33.6\%$. In addition, we apply the compression technology proposed in this work to the Ring ORAM system. By comparison, it is found that with the same compression ratio as Path ORAM, our design can still reduce latency by an average of$21.5\%$.
Chuang Li 0004, Changyao Tan, Gang Liu 0038, Yanhua Wen, Yan Wang 0022, Kenli Li 0001
IEEE Trans. Computers1
2024 Adaptive Network Load Balancing at the End Host for Traffic Bursts in Data Centers
abstract
The network load balancing mechanism plays a pivotal role in enhancing transmission performance in modern cloud data centers. Conventional flowlet-based approaches at host side offer a balance between performance and deployment simplicity. However, their passive load balancing strategy restricts rerouting opportunities, and lacks precision in congestion detection as it necessitates at least one round-trip time (RTT) to acquire end-to-end congestion feedback. To overcome the performance loss caused by the above limitations, we propose BurstLoader, an enhanced flowlet-based mechanism that adapts to varying traffic burst intensities and improves congestion detection accuracy. BurstLoader proactively reroutes congested flows when no new flowlets are detected, while simultaneously avoiding the rerouting of flowlets that are in good transmission states. Furthermore, BurstLoader incorporates delay and its gradient for a more nuanced and precise congestion detection. The extensive experiments demonstrate that BurstLoader achieves a significant reduction in flow completion time (FCT) by up to 48% compared to other flowlet-based solutions deployed at the end host, while maintaining competitive performance even against schemes that require custom switches under realistic workloads.
Qingyu Shi 0001, Xiaocui Li 0001, Chuang Li 0004, Wenzhi Cao, Limei Liu
HPCC4
2024 LBoDSN: An In-Network Load Balancing Mechanism for Lossless Data Center Networks Based on Direct Switch Notification
Qingyu Shi 0001, Fangxue Jiang, Xiaocui Li 0001, Chuang Li 0004, Wenzhi Cao, Limei Liu
NPC (1)5
2024 GenoM7GNet: An Efficient N7-Methylguanosine Site Prediction Approach Based on a Nucleotide Language Model
abstract
N-methylguanosine (m7G), one of the mainstream post-transcriptional RNA modifications, occupies an exceedingly significant place in medical treatments. However, classic approaches for identifying m7G sites are costly both in time and equipment. Meanwhile, the existing machine learning methods extract limited hidden information from RNA sequences, thus making it difficult to improve the accuracy. Therefore, we put forward to a deep learning network, called "GenoM7GNet," for m7G site identification. This model utilizes a Bidirectional Encoder Representation from Transformers (BERT) and is pretrained on nucleotide sequences data to capture hidden patterns from RNA sequences for m7G site prediction. Moreover, through detailed comparative experiments with various deep learning models, we discovered that the one-dimensional convolutional neural network (CNN) exhibits outstanding performance in sequence feature learning and classification. The proposed GenoM7GNet model achieved 0.953in accuracy, 0.932in sensitivity, 0.976in specificity, 0.907in Matthews Correlation Coefficient and 0.984in Area Under the receiver operating characteristic Curve on performance evaluation. Extensive experimental results further prove that our GenoM7GNet model markedly surpasses other state-of-the-art models in predicting m7G sites, exhibiting high computing performance.
Chuang Li 0004, Heshi Wang, Yanhua Wen, Rui Yin 0002, Xiangxiang Zeng, Keqin Li 0001
IEEE ACM Trans. Comput. Biol. Bioinform.1
2024 RFL-APIA: A Comprehensive Framework for Mitigating Poisoning Attacks and Promoting Model Aggregation in IIoT Federated Learning
abstract
With the development of industrial Internet of Things (IIoT), federated learning (FL) is important for protecting sensitive data from various Internet of Things devices (i.e., clients in FL). Despite FL's privacy benefits, attackers (e.g., untrusted clients) can still compromise the performance of the global model through model poisoning attacks. Unfortunately, two key challenges hinder effective detection and impact the performance of the global model in FL: first, accurately identifying malicious models to defend against attacks, and second, efficiently aggregating local models after detecting malicious clients. To address these challenges, we propose an improved FL system based on fuzzy rules, termed RFL-APIA. Compared to the conventional FL system, we have designed two novel components, federated learning generalized depth detection (FedGDD) and Fedsv-Weighted, to enhance performance and mitigate model poisoning attacks. Specifically, FedGDD introduces variance reduction by examining the relationship between local and global model gradients, thereby mitigating interference in nonindependent and identical distributed settings. It further implements an adaptive penalty factor-based scoring system, leveraging variations in local model updates for precise identification and mitigation of attacks. Based on FedGDD's output, the Fedsv-Weighted mechanism dynamically updates the global model's aggregation weights by considering local models' contributions, thus improving model aggregation. Extensive experiments demonstrate that RFL-APIA effectively prevents model poisoning attacks during training, ensuring model security, and guaranteeing a certain level of accuracy and convergence for the global model.
Chuang Li 0004, Aoli He, Gang Liu 0038, Yanhua Wen, Anthony T. Chronopoulos, Aristotelis Giannakos
IEEE Trans. Ind. Informatics1
2023 MSSIF-Net: an efficient CNN automatic detection method for freight train images
Longxin Zhang, Jingsheng Chen, Chuang Li 0004, Keqin Li 0001
Neural Comput. Appl.4
2023 Optimal Trading Mechanism Based on Differential Privacy Protection and Stackelberg Game in Big Data Market
abstract
Big data has become a fundamental resource and a commodity in economic activities, thus, it is necessary to build a market model capable of supporting efficient data trading. However, two major challenges remain. First, researches have considered constructing data trading mechanisms, while few of them are based on the method of measuring data value in multiple dimensions. Second, a data market involved an intermediary trading platform (i.e., a third party) which is honest but curious, results may be obtained due the the leakage of private information. In this article, we design TM-OUE, a data trading mechanism based on Optimized Unary Encoding that enables reasonable trading mechanism and protects the privacy of data trading. First of all, we combine qualitative and quantitative methods to measure the value of data in multiple dimensions and formulate data trading model between the data provider and data users. Then, we utilize an Optimized Unary Encoding (OUE) protocol to protect the privacy of the data trading mechanism. Based on the above steps, we develop a two-stage single leader multi-follower Stackelberg game to jointly maximize profits of the data provider and data users. Experimental results demonstrate that TM-OUE can offer appropriate price for data and maximize benefits both for data providers and data users, which guarantees fair data trades while protecting privacy.
Chuang Li 0004, Aoli He, Yanhua Wen, Gang Liu 0038, Anthony T. Chronopoulos
IEEE Trans. Serv. Comput.1
2022 Optimizing Anchor Node Deployment for Fingerprint Localization With Low-Cost and Coarse-Grained Communication Chips
abstract
A part of off-the-shelf wireless communication chips can only provide very coarse-grained information of the distance range between the transmitter and receiver. Exploiting such chips for precise wireless indoor positioning (WIP) becomes very valuable since their costs are comparatively low. For the WIP systems built on the coarse-grained communication chips, this article discusses the optimal anchor node deployment problem for fingerprint localization. The problem is very challenging since it implies three particularly important requirements of adaptive localization precision, unique fingerprint, and a minimum number of anchor nodes. To minimize the required number of anchor nodes while satisfying the three requirements, we propose two anchor node deployment algorithms, namely, ANDAs and WANDA, both of which are developed for mediate and large-area deployment, respectively. We first prove the effectiveness of the chip in maintaining robust fingerprints and formulate the deployment optimization problem as a minimum attribute reduction problem. Then, to deal with the high computational complexity involved, ANDA is proposed by significantly reducing the search space of the optimization problem. To practically address the combination explosion problem for large-area deployment, the idea of area partitioning is adopted. Investigating anchor nodes at different grid vertices in subareas that may have different sharing degrees, WANDA is proposed to achieve an overall anchor node deployment optimization. Simulation and experimental results demonstrate the validity and the superiority of ANDA and WANDA over counterparts. For large-area deployment, compared to ANDA, WANDA can further decrease the total amount of anchor nodes by 17%.
Xiaolong Li 0004, Jun Cai 0001, Rongyang Zhao, Chuang Li 0004, Chengwen He, Dian He
IEEE Internet Things J.4
2019 SW-Tandem: a highly efficient tool for large-scale peptide identification with parallel spectrum dot product on Sunway TaihuLight
abstract
SUMMARY: Tandem mass spectrometry based database searching is a widely acknowledged and adopted method that identifies peptide sequence in shotgun proteomics. However, database searching is extremely computationally expensive, which can take days even weeks to process a large spectra dataset. To address this critical issue, this paper presents SW-Tandem, a new tool for large-scale peptide sequencing. SW-Tandem parallelizes the spectrum dot product scoring algorithm and leverages the advantages of Sunway TaihuLight, the No. 1 supercomputer in the world in 2017. Sunway TaihuLight is powered by the brand new many-core SW26010 processors and provides a peak computation performance greater than 100PFlops. To fully utilize the Sunway TaihuLights capacity, SW-Tandem employs three mechanisms to accelerate large-scale peptide identification, memory-access optimizations, double buffering and vectorization. The results of experiments conducted on multiple datasets demonstrate the performance of SW-Tandem against three state-of-the-art tools for peptide identification, including X!! Tandem, MR-Tandem and MSFragger. In addition, it shows high scalability in the experiments on extremely large datasets sized up to 12 GB. AVAILABILITY AND IMPLEMENTATION: SW-Tandem is an open source software tool implemented in C++. The source code and the parameter settings are available at https://github.com/Logic09/SW-Tandem. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Chuang Li 0004, Kenli Li 0001, Tao Chen 0005, Qiang He 0001
Bioinform.1
2019 MCtandem: an efficient tool for large-scale peptide identification on many integrated core (MIC) architecture
abstract
BACKGROUND: Tandem mass spectrometry (MS/MS)-based database searching is a widely acknowledged and widely used method for peptide identification in shotgun proteomics. However, due to the rapid growth of spectra data produced by advanced mass spectrometry and the greatly increased number of modified and digested peptides identified in recent years, the current methods for peptide database searching cannot rapidly and thoroughly process large MS/MS spectra datasets. A breakthrough in efficient database search algorithms is crucial for peptide identification in computational proteomics. RESULTS: This paper presents MCtandem, an efficient tool for large-scale peptide identification on Intel Many Integrated Core (MIC) architecture. To support big data processing capability, a novel parallel match scoring algorithm, named MIC-SDP (spectrum dot product), and its two-level parallelization are presented in MCtandem's design. In addition, a series of optimization strategies on both the host CPU side and the MIC side, which includes pre-fetching, optimized communication overlapping scheme, multithreading and hyper-threading, are exploited to improve the execution performance. CONCLUSIONS: For fair comparisons, we first set up experiments and verified the 28 fold times speedup on a single MIC against the original CPU-based implementation. We then execute the MCtandem for a very large dataset on an MIC cluster (a component of the Tianhe-2 supercomputer) and achieved much higher scalability than in a benchmark MapReduce-based programs, MR-Tandem. MCtandem is an open-source software tool implemented in C++. The source code and the parameter settings are available at https://github.com/LogicZY/MCtandem .
Chuang Li 0004, Kenli Li 0001, Keqin Li 0001
BMC Bioinform.1
2017 MRUniNovo: an efficient tool for de novo peptide sequencing utilizing the hadoop distributed computing framework
abstract
Summary: Tandem mass spectrometry-based de novo peptide sequencing is a complex and time-consuming process. The current algorithms for de novo peptide sequencing cannot rapidly and thoroughly process large mass spectrometry datasets. In this paper, we propose MRUniNovo, a novel tool for parallel de novo peptide sequencing. MRUniNovo parallelizes UniNovo based on the Hadoop compute platform. Our experimental results demonstrate that MRUniNovo significantly reduces the computation time of de novo peptide sequencing without sacrificing the correctness and accuracy of the results, and thus can process very large datasets that UniNovo cannot. Availability and Implementation: MRUniNovo is an open source software tool implemented in java. The source code and the parameter settings are available at http://bioinfo.hupo.org.cn/MRUniNovo/index.php. Contact: [email protected] ; [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Chuang Li 0004, Tao Chen 0005, Qiang He 0001, Kenli Li 0001
Bioinform.1