Moming Duan

dblp:224/5731 · DBLP profile ↗
← Back
21ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0001-5402-6634ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 'They've Stolen My GPL-Licensed Model!': Toward Standardized and Transparent Model Licensing
abstract
As model parameter sizes scale into the billions and training consumes zettaFLOPs of computation, the reuse of Machine Learning (ML) assets and collaborative development have become increasingly prevalent in the ML community. These ML assets, including models, datasets, and software, may originate from various sources and be published under different licenses, which govern the use and distribution of licensed works and their derivatives. However, commonly chosen licenses, such as GPL and Apache, are software-specific and are not clearly defined or bounded in the context of model publishing. Meanwhile, the reused assets may also be under free-content licenses and model licenses, which pose a potential risk of license noncompliance and rights infringement within the model production workflow. In this paper, we address these challenges along two lines: 1) For ML workflow compliance, we propose ModelGo (MG) Analyzer, a tool that incorporates a vocabulary for ML workflow management and encoded license rules, enabling ontological reasoning to analyze rights granting and compliance issues. 2) For standardized model publishing, we introduce ModelGo Licenses, a set of modell-specific licenses that provide flexible options to meet the diverse needs of the ML community. MG Analyzer is built on Turtle language and Notation3 reasoning engine, envisioned as a first step toward Linked Open Data for ML workflow management. We have also encoded our proposed model licenses into rules and demonstrated the effects of GPL and other commonly used licenses in model publishing, along with the flexibility advantages of our licenses, through comparisons and experiments.
Moming Duan, Rui Zhao 0009, Linshan Jiang, Nigel Shadbolt, Bingsheng He
WWW1
2026 OpenDigger: A Practical Framework for Assessing Community Health and Sustainability in Open Source Collaboration Platforms
abstract
The rapid development and widespread adoption of open source software, facilitated and accelerated by the web, have fostered a vibrant ecosystem for collaborative development and innovation. GitHub, a leading platform for collaborative software development, currently hosts more than 100 million registered users, creating a substantial ecosystem for examining open source community behaviors. Existing tools for measuring open source communities primarily focus on metrics such as issue response time, pull request response time, or incremental stars to provide insights into community activity. However, these tools are limited in their ability to assess the influence of communities from the perspective of collaboration networks. Moreover, current data collection solutions offer fixed functionalities and lack the flexibility to support multi-source, fine-grained, and customizable data acquisition, which is essential for comprehensive analysis of Open Source Ecosystems (OSEs). In this paper, we present OpenDigger, a framework for multi-dimensional assessment of collaboration activities in OSEs. To enable scalable, modular, and continuous acquisition of OSE data, we developed OpenCrawler, a one-line service providing customizable, fine-grained control over data collection. Using the collected data, OpenDigger computes 20 statistical and 2 network-based metrics, and our empirical analysis further verifies their effectiveness in enabling a comprehensive assessment of trends in OSEs. By continuously collecting logs from GitHub and Gitee, OpenDigger has now accumulated over 9 billion records. Our framework has already been deployed across multiple industrial environments, including Alibaba Group, Ant Group, Apache Foundation, and Mulan Open Source Community.
Wei Wang 0033, Fanyu Han, Shengyu Zhao, Xuan Zhou 0001, Weining Qian, Aoying Zhou, Xiaoya Xia, Moming Duan
WWW11
2025 MPNAS: Multimodal Sentiment Analysis Pruning via Neural Architecture Search
abstract
With the rapid development of social media, sentiment analysis from multimodal posts has garnered significant attention in recent years. However, the substantial size of these models impedes their deployment on resource-constrained embedded devices. Although pruning has been extensively studied to reduce the size of unimodal models, specific challenges remain for Multimodal Sentiment Analysis (MSA) models. First, existing techniques prune fixed original models into sparse models, while our findings indicate that different model architectures of identical size yield varying performance outcomes. Second, prior studies fail to explore the unique characteristics of MSA models, resulting in suboptimal pruning performance. To address these challenges, we propose MPNAS, a unified pruning framework via Neural Architecture Search (NAS) for MSA models. Specifically, we formulate pruning as a NAS problem and analyze MSA model characteristics to guide the subnet search. We conduct an initial coarse-grained NAS on the original model, expanding the search space slightly to identify suitable subnets that enhance pruning rates and accuracy. Subsequently, we refine coarse-grained subnets in a fine-grained NAS stage, where MSA model characteristics guide the search process. Extensive experiments on three representative datasets demonstrate the superiority of our approach over existing methods.
Binyan Zhang, Ao Ren, Zihao Zhang 0002, Moming Duan, Duo Liu 0002, Yujuan Tan, Kan Zhong
ICASSP4
2025 ML-Asset Management: Curation, Discovery, and Utilization
abstract
Machine learning (ML) assets, such as models, datasets, and metadata—are central to modern ML workflows. Despite their explosive growth in practice, these assets are often underutilized due to fragmented documentation, siloed storage, inconsistent licensing, and lack of unified discovery mechanisms, making ML-asset management an urgent challenge. This tutorial offers a comprehensive overview of ML-asset management activities across its lifecycle, including curation, discovery, and utilization. We provide a categorization of ML assets, and major management issues, survey state-of-the-art techniques, and identify emerging opportunities at each stage. We further highlight system-level challenges related to scalability, lineage, and unified indexing. Through live demonstrations of systems, this tutorial equips both researchers and practitioners with actionable insights and practical tools for advancing ML-asset management in real-world and domain-specific settings.
Mengying Wang 0001, Moming Duan, Yicong Huang 0002, Chen Li 0001, Bingsheng He, Yinghui Wu 0001
Proc. VLDB Endow.2
2024 Rethinking Literary Plagiarism in LLMs through the Lens of Copyright Laws
Huachen Tan, Moming Duan, Duo Liu 0002, Haojie Lu, Yuexin Mu, Longyi Zhou, Ao Ren, Yujuan Tan, Kan Zhong
ACML2
2024 ModelGo: A Practical Tool for Machine Learning License Analysis
abstract
Productionizing machine learning projects is inherently complex, involving a multitude of interconnected components that are assembled like LEGO blocks and evolve throughout development lifecycle. These components encompass software, databases, and models, each subject to various licenses governing their reuse and redistribution. However, existing license analysis approaches for Open Source Software (OSS) are not well-suited for this context. For instance, some projects are licensed without explicitly granting sublicensing rights, or the granted rights can be revoked, potentially exposing their derivatives to legal risks. Indeed, the analysis of licenses in machine learning projects grows significantly more intricate as it involves interactions among diverse types of licenses and licensed materials. To the best of our knowledge, no prior research has delved into the exploration of license conflicts within this domain. In this paper, we introduce ModelGo, a practical tool for auditing potential legal risks in machine learning projects to enhance compliance and fairness. With ModelGo, we present license assessment reports based on five use cases with diverse model-reusing scenarios, rendered by real-world machine learning components. Finally, we summarize the reasons behind license conflicts and provide guidelines for minimizing them. Our code is publicly available at https://github.com/Xtra-Computing/ModelGo.
Moming Duan, Qinbin Li, Bingsheng He
WWW1
2024 OFL-W3: A One-shot Federated Learning System on Web 3.0
abstract
Federated Learning (FL) addresses the challenges posed by data silos, which arise from privacy, security regulations, and ownership concerns. Despite these barriers, FL enables these isolated data repositories to participate in collaborative learning without compromising privacy or security. Concurrently, the advancement of blockchain technology and decentralized applications (DApps) within Web 3.0 heralds a new era of transformative possibilities in web development. As such, incorporating FL into Web 3.0 paves the path for overcoming the limitations of data silos through collaborative learning. However, given the transaction speed constraints of core blockchains such as Ethereum (ETH) and the latency in smart contracts, employing one-shot FL, which minimizes client-server interactions in traditional FL to a single exchange, is considered more apt for Web 3.0 environments. This paper presents a practical one-shot FL system for Web 3.0, termed OFL-W3. OFL-W3 capitalizes on blockchain technology by utilizing smart contracts for managing transactions. Meanwhile, OFL-W3 utilizes the Inter-Planetary File System (IPFS) coupled with Flask communication, to facilitate backend server operations to use existing one-shot FL algorithms. With the integration of the incentive mechanism, OFL-W3 showcases an effective implementation of one-shot FL on Web 3.0, offering valuable insights and future directions for AI combined with Web 3.0 studies.
Linshan Jiang, Moming Duan, Bingsheng He, Peishen Yan, Yang Hua 0001, Tao Song 0003
Proc. VLDB Endow.2
2023 LFPR: A Lazy Fast Predictive Repair Strategy for Mobile Distributed Erasure Coded Cluster
abstract
Mobile distributed erasure coded Internet of Things (IoT) clusters store popular data, reducing communication latency, and ensuring data reliability while requiring low storage overhead. However, it suffers a high repair overhead to ensure data reliability and availability due to mobile device failures or leaving the cluster. Predictive repair is an effective strategy for reducing repair overhead that has gained attention with the development in accurate failure and mobile node movement trajectory prediction technologies in recent years. We propose LFPR, a hybrid lazy fast predictive repair strategy that combines two baseline predictive repair approaches (reconstruction and migration), including LFPRH and LFPRC for a hot and cold data distributed cluster, respectively. LFPRC and LFPRH adopt different mechanisms to determine whether a block should perform predictive repair immediately. The predictive repair mechanisms of LFPR couples migration and reconstruction in parallel to reduce average repair time per block. LFPR significantly reduces average repair time per block via large-scale simulation and local cluster experiments, compared with existing predictive repair solutions, such as FastPR and the two baselines.
Yu Wu 0016, Duo Liu 0002, Yujuan Tan, Moming Duan, Longpan Luo, Weilve Wang, Xianzhang Chen
IEEE Internet Things J.4
2023 FedMDS: An Efficient Model Discrepancy-Aware Semi-Asynchronous Clustered Federated Learning Framework
abstract
Federated learning (FL) is an emerging distributed machine learning paradigm that protects privacy and tackles the problem of isolated data islands. At present, there are two main communication strategies of FL: synchronous FL and asynchronous FL. The advantages of synchronous FL are the high precision and easy convergence of the model. However, this synchronous communication strategy has the risk of the straggler effect. Asynchronous FL has a natural advantage in mitigating the straggler effect, but there are threats of model quality degradation and server crash. In this paper, we propose a model discrepancy-aware semi-asynchronous clustered FL framework,FedMDS, which alleviates the straggler effect by 1) a clustered strategy based on the delay and direction of the model update and 2) a synchronous trigger mechanism that limits the model staleness.FedMDSleverages the clustered algorithm to reschedule the clients. Each group of clients performs asynchronous updates until the synchronous update mechanism based on the model discrepancy is triggered. We evaluateFedMDSbased on four typical federated datasets in a non-IID setting and compareFedMDSto the baselines. The experimental results show thatFedMDSsignificantly improves average test accuracy by more than$+9.2\%$on the four datasets compared toTA-FedAvg. In particular,FedMDSimproves absolute Top-1 test accuracy by$+37.6\%$on FEMNIST compared toTA-FedAvg. The frequency of the average synchronization waiting time ofFedMDSis significantly lower than that ofTA-FedAvgon all datasets. Moreover,FedMDScan improve the accuracy and alleviate the straggler effect.
Yu Zhang 0184, Duo Liu 0002, Moming Duan, Xianzhang Chen, Ao Ren, Yujuan Tan, Chengliang Wang 0002
IEEE Trans. Parallel Distributed Syst.3
2022 Lazy repair with temporary redundancy(LRTR): reducing repair network traffic in erasure-coded storage
abstract
Erasure coding has gained popularity in today's storage systems as a low-storage overhead and high-reliability fault-tolerant method. However, it is hampered by the high repair costs. The temporary failures in storage systems amplify this drawback resulting in a lot of unnecessary repair traffic. It leads to a dilemma that traditional repair schemes can not optimize repair traffic and reliability at the same time.
Longpan Luo, Yujuan Tan, Duo Liu 0002, Moming Duan, Weilue Wang, Yu Wu 0016, Xianzhang Chen
CF4
2022 Federated learning with workload-aware client scheduling in heterogeneous systems
Duo Liu 0002, Moming Duan, Yu Zhang 0184, Ao Ren, Xianzhang Chen, Yujuan Tan, Chengliang Wang 0002
Neural Networks3
2022 Flexible Clustered Federated Learning for Client-Level Data Distribution Shift
abstract
Federated Learning (FL) enables the multiple participating devices to collaboratively contribute to a global neural network model while keeping the training data locally. Unlike the centralized training setting, the non-IID, imbalanced (statistical heterogeneity) and distribution shifted training data of FL is distributed in the federated network, which will increase the divergences between the local models and the global model, further degrading performance. In this paper, we propose a flexible clustered federated learning (CFL) framework named FlexCFL, in which we 1) group the training of clients based on the similarities between the clients’ optimization directions for lower training divergence; 2) implement an efficient newcomer device cold start mechanism for framework scalability and practicality; 3) flexibly migrate clients to meet the challenge of client-level data distribution shift. FlexCFL can achieve improvements by dividing joint optimization into groups of sub-optimization and can strike a balance between accuracy and communication efficiency in the distribution shift environment. The convergence and complexity are analyzed to demonstrate the efficiency of FlexCFL. We also evaluate FlexCFL on several open datasets and made comparisons with related CFL frameworks. The results show that FlexCFL can significantly improve absolute test accuracy by$+10.6\%$on FEMNIST compared withFedAvg,$+3.5\%$on FashionMNIST compared withFedProx,$+8.4\%$on MNIST compared withFeSEM,$+4.7\%$on Sentiment140 compare withIFCA. The experiment results show that FlexCFL is also communication efficient in the distribution shift environment.
Moming Duan, Duo Liu 0002, Xinyuan Ji, Yu Wu 0016, Liang Liang 0002, Xianzhang Chen, Yujuan Tan, Ao Ren
IEEE Trans. Parallel Distributed Syst.1
2021 Forseti: An Efficient Basic-block-level Sensitivity Analysis Framework Towards Multi-bit Faults
abstract
The per-instruction sensitivity analysis framework is developed to evaluate the resiliency of a program and identify the segments of the program needing protection. However, for multi-bit hardware faults, the per-instruction sensitivity analysis frameworks can cause large overhead for redundant analyses. In this paper, we propose a basic-block-level sensitivity analysis framework, Forseti, to reduce the analysis overhead in analyzing impacts of modern microprocessors' multi-bit faults on programs. We implement Forseti in LLVM and evaluate it with five typical workloads. Extensive experimental results show that Forseti can achieve more than 90% sensitivity classification accuracy and 6.16× speedup over instruction-level analysis.
Jinting Ren, Xianzhang Chen, Duo Liu 0002, Moming Duan, Renping Liu 0002, Chengliang Wang 0002
DATE4
2021 FedSAE: A Novel Self-Adaptive Federated Learning Framework in Heterogeneous Systems
abstract
Federated Learning (FL) is a novel distributed machine learning which allows thousands of edge devices to train model locally without uploading data concentrically to the server. But since real federated settings are resource-constrained, FL is encountered with systems heterogeneity which causes a lot of stragglers directly and then leads to significantly accuracy reduction indirectly. To solve the problems caused by systems heterogeneity, we introduce a novel self-adaptive federated framework FedSAE which adjusts the training task of devices automatically and selects participants actively to alleviate the performance degradation. In this work, we 1) propose FedSAE which leverages the complete information of devices' historical training tasks to predict the affordable training workloads for each device. In this way, FedSAE can estimate the reliability of each device and self-adaptively adjust the amount of training load per client in each round. 2)combine our framework with Active Learning to self-adaptively select participants. Then the framework accelerates the convergence of the global model. In our framework, the server evaluates devices' value of training based on their training loss. Then the server selects those clients with bigger value for the global model to reduce communication overhead. The experimental result indicates that in a highly heterogeneous system, FedSAE converges faster than FedAvg, the vanilla FL framework. Furthermore, FedSAE outperforms than FedAvg on several federated datasets - FedSAE improves test accuracy by 26.7% and reduces stragglers by 90.3% on average.
Moming Duan, Duo Liu 0002, Yu Zhang 0184, Ao Ren, Xianzhang Chen, Yujuan Tan, Chengliang Wang 0002
IJCNN2
2021 CSAFL: A Clustered Semi-Asynchronous Federated Learning Framework
abstract
Federated learning (FL) is an emerging distributed machine learning paradigm that protects privacy and tackles the problem of isolated data islands. At present, there are two main communication strategies of FL: synchronous FL and asynchronous FL. The advantages of synchronous FL are that the model has high precision and fast convergence speed. However, this synchronous communication strategy has the risk that the central server waits too long for the devices, namely, the straggler effect which has a negative impact on some time-critical applications. Asynchronous FL has a natural advantage in mitigating the straggler effect, but there are threats of model quality degradation and server crash. Therefore, we combine the advantages of these two strategies to propose a clustered semi-asynchronous federated learning (CSAFL) framework. We evaluate CSAFL based on four imbalanced federated datasets in a non-IID setting and compare CSAFL to the baseline methods. The experimental results show that CSAFL significantly improves test accuracy by more than +5% on the four datasets compared to TA-FedAvg. In particular, CSAFL improves absolute test accuracy by +34.4% on non-IID FEMNIST compared to TA-FedAvg.
Yu Zhang 0184, Moming Duan, Duo Liu 0002, Ao Ren, Xianzhang Chen, Yujuan Tan, Chengliang Wang 0002
IJCNN2
2021 A machine learning assisted data placement mechanism for hybrid storage systems
Jinting Ren, Xianzhang Chen, Duo Liu 0002, Yujuan Tan, Moming Duan, Ruolan Li, Liang Liang 0002
J. Syst. Archit.5
2021 Self-Balancing Federated Learning With Global Imbalanced Data in Mobile Systems
abstract
Federated learning (FL) is a distributed deep learning method that enables multiple participants, such as mobile and IoT devices, to contribute a neural network while their private training data remains in local devices. This distributed approach is promising in the mobile systems where have a large corpus of decentralized data and require high privacy. However, unlike the common datasets, the data distribution of the mobile systems is imbalanced which will increase the bias of model. In this article, we demonstrate that the imbalanced distributed training data will cause an accuracy degradation of FL applications. To counter this problem, we build a self-balancing FL framework named Astraea, which alleviates the imbalances by 1) Z-score-based data augmentation, and 2) Mediator-based multi-client rescheduling. The proposed framework relieves global imbalance by adaptive data augmentation and downsampling, and for averaging the local imbalance, it creates the mediator to reschedule the training of clients based on Kullback-Leibler divergence (KLD) of their data distribution. Compared with FedAvg, the vanilla FL algorithm, Astraea shows +4.39 and +6.51 percent improvement of top-1 accuracy on the imbalanced EMNIST and imbalanced CINIC-10 datasets, respectively. Meanwhile, the communication traffic of Astraea is reduced by 75 percent compared to FedAvg.
Moming Duan, Duo Liu 0002, Xianzhang Chen, Renping Liu 0002, Yujuan Tan, Liang Liang 0002
IEEE Trans. Parallel Distributed Syst.1
2019 Reducing Write Amplification for Inodes of Journaling File System using Persistent Memory
abstract
Conventional journaling file systems, such as Ext4, guarantee data consistency by writing in-memory dirty inodes to block devices twice. The write back of inodes may contain up to 80% clean inode that is unnecessary to be written back, which caused severe write amplification problem and largely reduce performance since the size of an inode is several times less than the size of a basic unit for updating the block device. Emerging persistent memories (PMs), such as phase change memory, provide the possibility for storing the offset of inodes in memory persistently. In this paper, we propose an efficient scheme, Updating Frequency based Inode Aggregation (UFIA), to reduce the write amplification of dirty inodes using PM. The main idea of UFIA is to identify the frequently-updated inodes and reorganize them in adjacent physical locations on block device. Firstly, UFIA adopts PM as an inode mapping table for remapping logical inodes to any physical inodes. Secondly, we design an efficient algorithm for UFIA to identify and reorganize the frequently-updated inodes. We implement UFIA and integrate it into Ext4 (denoted by UFIA-Ext4) in Linux kernel 4.4.4. The experiments are conducted with widely-used benchmark Filebench. Compared with original Ext4, the experimental results show that UFIA significantly reduces the write amplification of inodes and improves 54% of the performance on average.
Chaoshu Yang, Duo Liu 0002, Xianzhang Chen, Runyu Zhang 0002, Moming Duan, Yujuan Tan
DATE6
2019 Astraea: Self-Balancing Federated Learning for Improving Classification Accuracy of Mobile Deep Learning Applications
abstract
Federated learning (FL) is a distributed deep learning method which enables multiple participants, such as mobile phones and IoT devices, to contribute a neural network model while their private training data remains in local devices. This distributed approach is promising in the edge computing system where have a large corpus of decentralized data and require high privacy. However, unlike the common training dataset, the data distribution of the edge computing system is imbalanced which will introduce biases in the model training and cause a decrease in accuracy of federated learning applications. In this paper, we demonstrate that the imbalanced distributed training data will cause accuracy degradation in FL. To counter this problem, we build a self-balancing federated learning framework call Astraea, which alleviates the imbalances by 1) Global data distribution based data augmentation, and 2) Mediator based multi-client rescheduling. The proposed framework relieves global imbalance by runtime data augmentation, and for averaging the local imbalance, it creates the mediator to reschedule the training of clients based on Kullback-Leibler divergence (KLD) of their data distribution. Compared with FedAvg, the state-of-the-art FL algorithm, Astraea shows +5.59% and +5.89% improvement of top-1 accuracy on the imbalanced EMNIST and imbalanced CINIC-10 datasets, respectively. Meanwhile, the communication traffic of Astraea can be 92% lower than that of FedAvg.
Moming Duan, Duo Liu 0002, Xianzhang Chen, Yujuan Tan, Jinting Ren, Lei Qiao 0002, Liang Liang 0002
ICCD1
2019 Archivist: A Machine Learning Assisted Data Placement Mechanism for Hybrid Storage Systems
abstract
With the rapid growth of edge-cloud computing, emerging applications pose higher performance demand on the storage system for storing massive data that are generated from various sources. The multi-sourced data shows different properties in size, retention time, and read/write frequency. Hybrid storage system is promised to efficiently handle the data in edge-cloud computing environment satisfying different data demands. The key problem is how to place the data on the hybrid storage system according to the run-time status and the properties of both data and the storage systems. In this paper, we propose Archivist - a machine learning assisted data placement mechanism for hybrid storage systems to reduce file access latency. We first design a machine learning based approach for predicting the access patterns of the incoming data. Then, we present a data placement algorithm to optimize the data on the hybrid storage mediums by matching the properties of data and the features of storage mediums. Extensive experimental results show that Archivist can achieve up to 49% improvement of system performance for file accesses compared with baseline.
Jinting Ren, Xianzhang Chen, Yujuan Tan, Duo Liu 0002, Moming Duan, Liang Liang 0002, Lei Qiao 0002
ICCD5
2019 FitCNN: A cloud-assisted and low-cost framework for updating CNNs on IoT devices
Duo Liu 0002, Chaoshu Yang, Xianzhang Chen, Jinting Ren, Renping Liu 0002, Moming Duan, Yujuan Tan, Liang Liang 0002
Future Gener. Comput. Syst.7