Gang Mao

dblp:193/7144 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 2 since 2021Computer networks · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PD-FedOS: Prototype-driven federated open-set learning framework for collaborative intelligent fault diagnosis of aero-engine rotor systems
Gang Mao, Yongbo Li 0001, Zhiqiang Cai 0003, Teng Wang 0002, Khandaker Noman, Ran Zhang 0011
Expert Syst. Appl.1
2026 Federated Physics-Informed Graph Framework Guided by Multianchors for Heterogeneous Wheeled Robots Collaborative Fault Diagnosis
abstract
Wheeled robot fault diagnosis is indispensable for ensuring its reliable and safe operations. However, two challenges impede the application of prevalent intelligent diagnosis methods. 1) Multisensor fusion: The complexity of robot movements necessitates multisensor for comprehensive monitoring, generating strong-coupled, and high-dimensional data that complicate both intrinsic relationship mining and effective fusion; 2) Heterogeneous data silos: Dispersibility, heterogeneity and privacy constraints across different robots lead to non-independent and identically distributed (Non-IID) data silos, severely limiting the development of universal diagnostic models. To overcome these two problems, this article proposes a tailored federated physics-informed graph framework (FedMA-PIG). On the client side, the kinematics mathematical model is constructed for each robot, which explores the inter-sensor correlations and forms a physics-informed graph. It enables multisensor data fusion and assists the client in training a local graph neural network. On the federated framework side, a Non-IID federated framework based on a multianchor contrastive mechanism is devised. It employs multiple anchors to capture common knowledge from heterogeneous robot data, guiding feature representations toward corresponding anchors and away from others, thereby promoting consistency and mitigating inter-client data heterogeneity. Comprehensive experiments were conducted on three representative wheeled robots- Mecanum-wheeled, 4WD-wheeled, and Omni-wheeled- distributed across four federated clients. The results demonstrate that FedMA-PIG achieves generalized and superior diagnostic performance compared to state-of-the-art methods.
Gang Mao, Yongbo Li 0001, Teng Wang 0002, Khandaker Noman, Zhiqiang Cai 0003
IEEE Trans. Ind. Informatics1
2025 A survey on RFID anti-collision strategies: From avoidance to recovery
Xinning Xiong, Gang Mao
Comput. Networks4
2025 Dynamic Tsetlin Machine Accelerators for On-Chip Training Using FPGAs
abstract
The increased demand for data privacy and security in machine learning (ML) applications has put impetus on effective edge training on Internet-of-Things (IoT) nodes. Edge training aims to leverage speed, energy efficiency and adaptability within the resource constraints of the nodes. Deploying and training Deep Neural Networks (DNNs)-based models at the edge, although accurate, posit significant challenges from the back-propagation algorithm’s complexity, bit precision trade-offs, and heterogeneity of DNN layers. This paper presents a Dynamic Tsetlin Machine (DTM) training accelerator as an alternative to DNN implementations. DTM utilizes logic-based on-chip inference with finite-state automata-driven learning within the same Field Programmable Gate Array (FPGA) package. Underpinned on the Vanilla and Coalesced Tsetlin Machine algorithms, the dynamic aspect of the accelerator design allows for a run-time reconfiguration targeting different datasets, model architectures, and model sizes without resynthesis. This makes the DTM suitable for targeting multivariate sensor-based edge tasks. Compared to DNNs, DTM trains with fewer multiply-accumulates, devoid of derivative computation. It is a data-centric ML algorithm that learns by aligning Tsetlin automata with input data to form logical propositions enabling efficient Look-up-Table (LUT) mapping and frugal Block RAM usage in FPGA training implementations. The proposed accelerator offers 2.54x more Giga operations per second per Watt (GOP/s per W) and uses 6x less power than the next-best comparable design.
Gang Mao, Tousif Rahman, Sidharth Maheshwari, Bob Pattison, Rishad A. Shafik, Alexandre Yakovlev
IEEE Trans. Circuits Syst. I Regul. Pap.1
2024 MATADOR: Automated System-on-Chip Tsetlin Machine Design Generation for Edge Applications
abstract
System-on-Chip Field-Programmable Gate Arrays (SoC-FPGAs) offer significant throughput gains for machine learning (ML) edge inference applications via the design of co-processor accelerator systems. However, the design effort for training and translating ML models into SoC-FPGA solutions can be substantial and requires specialist knowledge aware trade-offs between model performance, power consumption, latency and resource utilization. Contrary to other ML algorithms, Tsetlin Machine (TM) performs classification by forming logic proposition between boolean actions from the Tsetlin Automata (the learning elements) and boolean input features. A trained TM model, usually, exhibits high sparsity and considerable overlapping of these logic propositions both within and among the classes. The model, thus, can be translated to RTL-level design using a miniscule number of AND and NOT gates. This paper presents MATADOR, an automated boolean-to-silicon tool with GUI interface capable of implementing optimized accelerator design of the TM model onto SoC-FPGA for inference at the edge. It offers automation of the full development pipeline: model training, system level design generation, design verification and deployment. It makes use of the logic sharing that ensues from propositional overlap and creates a compact design by effectively utilizing the TM model's sparsity. MATADOR accelerator designs are shown to be up to 13.4x faster, up to 7x more resource frugal and up to 2x more power efficient when compared to the state-of-the-art Quantized and Binary Deep Neural Network implementations.
Tousif Rahman, Gang Mao, Sidharth Maheshwari, Rishad A. Shafik, Alexandre Yakovlev
DATE2
2024 DISOC: A Trustworthy Decentralized Oracle Architecture With Strong Privacy Preservation for IIoT Data Sharing
abstract
Industrial blockchain is believed to be a promising technology for modern industries. Despite the capacity of blockchain to protect users’ privacy, its inherent transparency renders data provenance traceable, a situation that does not fully guarantee the confidentiality of users and data sources. Users’ preferences, habits, identity, and other sensitive privacy information can be analyzed by adversaries after collecting a series of users’ requests. This paper integrates existing oracle technologies and proposes a distributed oracle architecture DISOC based on TOTP and ring signature to overcome the privacy shortcomings of traditional oracles. DISOC can be used to implement trustworthy off-chain data sharing from different data source with strong privacy-preservation. It can effectively protect the privacy of data sources while protecting user privacy to mitigate potential data leakage issues. After analysis and testing, DISOC can perfectly meet the data sharing needs and privacy preservation of multiple data sources.
Chenxi Xiong, Yunwei Cao, Gang Mao
IEEE Internet Things J.5
2023 Asynchronous Control for Tsetlin Machine with Binary Memristor-Transistor Array
abstract
Tsetlin machines (TMs) are a novel machine learning paradigm based on learning automata and Boolean logic inference, with better energy-efficiency and explainability than neural networks. This work exploits non-volatile ReRAM-transistor memory arrays to perform efficient in-memory TM computing. To accommodate the large timing variability of ReRAM devices and enhance energy efficiency and speed, the control path is implemented with quasi delay-insensitive (QDI) asynchronous circuits. The design of these circuits are derived and synthesized from their signal-transition graph specifications using the Workcraft tool. The resulting circuits offer high event-driven controllability and high variation tolerance for the mixed-signal ReRAM data path. Compared to state of the art TM hardware, the new TM design uses less than 5% of the power to achieve better than$4\times$the performance.
Omar Ghazal, Gang Mao, Jesse Ojukwu, Fei Xia 0001, Alexandre Yakovlev, Rishad A. Shafik
ISCAS2
2023 Transferable dynamic enhanced cost-sensitive network for cross-domain intelligent diagnosis of rotating machinery under imbalanced datasets
Gang Mao, Yongbo Li 0001, Zhiqiang Cai 0003, Bin Qiao, Sixiang Jia
Eng. Appl. Artif. Intell.1
2023 Oscillatory Lempel-Ziv Complexity Calculation as a Nonlinear Measure for Continuous Monitoring of Bearing Health
abstract
As a nonlinear measure, Lempel–Ziv complexity (LZC) can be considered as a suitable parameter for characterizing bearing health status by measuring the complexity of vibration signals. However, in continuous monitoring scenario under noisy condition, all components of a multicomponent bearing signal are not equally sensitive toward a change of LZC value. As a result, a direct application of LZC for bearing health monitoring not only suffers from its inefficient early fault warning but also fails to infer the fault progression. In this article, instead of direct utilization of a whole vibration signal, its fundamental component (FC) sensitive to LZC calculation is separated with the help of continuously adjustable parameterized tunable$Q$factor wavelet transform (TQWT). In this context, a study based on sparsity indices has been done for$Q$factor selection of TQWT. Since TQWT uses an oscillation-based bearing FC separation scheme for LZC calculation, the proposed measure is termed as oscillatory Lempel–Ziv complexity (OLZC). Two experimental cases are used for validation. Performance of OLZC is compared with original LZC, representative sparsity indices and recently proposed multiscale symbolic Lempel–Ziv complexity. Results demonstrate that the proposed OLZC can not only overcome the limitations of the original LZC but also performs better than other indices in comparison to continuous monitoring of bearing health.
Khandaker Noman, Yongbo Li 0001, Shubin Si, Shun Wang 0003, Gang Mao
IEEE Trans. Reliab.5
2022 Editable asynchronous control logic for SAR ADCs
abstract
This paper presents a novel design method for asynchronous control logic targeting successive approximation register (SAR) analog-to-digital converters (ADCs). This work is based on modeling the control logic for SAR ADCs using signal transition graphs (STGs). Different from conventional synchronous controllers, the proposed method results in asynchronous controllers driven by the causality of signals rather than relying on clocks to control the conversion process. Moreover, the proposed asynchronous control logic can be modularized through the handshake protocol, making it possible to build ADCs of arbitrary precision based on single-bit control units. This work results in a formal, model-based asynchronous design flow for SAR ADC control, which is shown to produce resulting circuits of similar speeds but great power efficiency improvements.
Fei Xia 0001, Gang Mao, Shengqi Yu, Rishad A. Shafik, Alexandre Yakovlev
ISCAS3
2021 Forecasting air passenger traffic flow based on the two-phase learning model
Xinfang Wu, Gang Mao, Mingqian Du, Xiuqing Yang, Xinzhi Zhou
J. Supercomput.3
2020 Reconstruction of coverage hole model and cooperative repair optimization algorithm in heterogeneous wireless sensor networks
Pingzhang Gou, Gang Mao, Fen Zhang, Xiangdong Jia
Comput. Commun.2
2020 A High Generalizable Feature Extraction Method Using Ensemble Learning and Deep Auto-Encoders for Operational Reliability Assessment of Bearings
Xianguang Kong, Qibin Wang, Hongbo Ma, Gang Mao
Neural Process. Lett.6