Mohan Li

dblp:50/8279 · DBLP profile ↗
← Back
39ranked-venue papers
15as first author
34since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 12 first-author · 16 since 2021Artificial intelligence and machine learning · 15 · 5 first-author · 14 since 2021Computer networks · 5 · 1 first-author · 4 since 2021Security and privacy · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 AT-Field: Rethinking the Games in Adversarial Training
abstract
Adversarial training is often modeled as a two-player zero-sum game, relying on strong assumptions that limit its practical guidance. In this paper, we instead analyze the interactions between training samples and show that even the fundamental objective—minimizing training loss—may not converge. To address this, we propose AT-Field, an adversarial training framework guided by sample-wise game-theoretic relationships. Specifically, we prove that training samples across different batches can form a none-potential game, where gradient descent induces cyclic behaviors, preventing convergence. By strategically searching and grouping these samples within the same batch, AT-Field transforms none-potential games into exact potential games, which are more effectively optimized using gradient-based methods. Experiments demonstrate that AT-Field integrates seamlessly with existing adversarial training techniques, enhancing both accuracy and robustness.
Yixiao Xu, Mohan Li, Zhijie Shen, Yuan Liu 0002, Zhihong Tian 0001
AAAI2
2026 BDpackets: A Clean-label Backdoor Attack on Network Traffic Classifiers via Feature Fusion
Mengxia Zhang, Yixiao Xu, Mohan Li, Yanbin Sun, Zhihong Tian 0001
INFOCOM3
2026 Virtual-Mask Informed Prior for Sparse-View Dual-Energy CT Reconstruction
abstract
Sparse-view sampling in dual-energy computed tomography (DECT) significantly reduces radiation dose and increases imaging speed, yet is highly prone to artifacts. Although diffusion models have demonstrated potential in effectively handling incomplete data, most existing methods in this field focus on the image domain and lack global constraints, which consequently leads to insufficient reconstruction quality. In this study, we propose a dual-domain virtual-mask informed diffusion model for sparse-view reconstruction by leveraging the high inter-channel correlation in DECT. Specifically, the study designs a virtual mask and applies it to the high-energy and low-energy data to perform perturbation operations, thus constructing high-dimensional tensors that serve as the prior information of the diffusion model. In addition, a dual-domain collaboration strategy is adopted to integrate the information of the randomly selected high-frequency components in the wavelet domain with the information in the projection domain, for the purpose of optimizing the global structures and local details. The experimental results show that the method exhibits excellent performance on multiple datasets. Under 30-view sparse sampling conditions, VIP-DECT improves PSNR by at least 1.02 dB and enhances SSIM by 1.91% .
Zini Chen 0002, Mohan Li, Cunfeng Wei, Shaoyu Wang 0002, Liu Shi, Qiegen Liu
IEEE J. Biomed. Health Informatics4
2025 Mitigating Hallucinations in Large Vision-Language Models by Adaptively Constraining Information Flow
abstract
Large vision-language models show tremendous potential in understanding visual information through human languages. However, they are prone to suffer from object hallucination, i.e., the generated image descriptions contain objects that do not exist in the image. In this paper, we reveal that object hallucination can be attributed to overconfidence in irrelevant visual features when soft visual tokens map to the LLM's word embedding space. Specifically, by figuring out the semantic similarity between visual tokens and LLM's word embedding, we observe that the smoothness of similarity distribution strongly correlates with the emergence of object hallucinations. To mitigate hallucinations, we propose using the Variational Information Bottleneck (VIB) to alleviate overconfidence by introducing stochastic noise, facilitating the constraining of irrelevant information. Furthermore, we propose an entropy-based noise-controlling strategy to enable the injected noise to be adaptively constrained regarding the smoothness of the similarity distribution. We adapt the proposed AdaVIB across distinct model architectures. Experimental results demonstrate that the proposed AdaVIB mitigates object hallucinations by effectively alleviating the overconfidence in irrelevant visual features, with consistent improvements on two object hallucination benchmarks.
Jiaqi Bai 0001, Hongcheng Guo, Zhongyuan Peng, Jian Yang 0030, Zhoujun Li 0001, Mohan Li, Zhihong Tian
AAAI6
2025 A 16.4pJ/bit Fully Integrated IR-UWB Transmitter with PSK+PPM Modulation and On-Chip Antenna
abstract
This paper presents a fully integrated Impulse Radio Ultra-Wideband (IR-UWB) wireless transmitter (TX) chip with high transmission efficiency and on-chip antenna for brain-computer interface (BCI) application. A hybrid modulation method combining Pulse Phase Shift Keying (PSK) and Pulse Position Modulation (PPM) is proposed to boost the data transmission efficiency. An on-chip antenna is adopted to eliminate the external bulky UWB antenna thus minimizing the dimension of the whole BCI implant. This chip is designed and implemented using SMIC 180 nm standard CMOS process, with a core area of approximately 1.02 mm2. The measurement shows that a 30 Mb/s data stream can be transmitted over a distance of 2 cm through the proposed antenna, with 16.4 pJ/bit energy efficiency and 491 μW power consumption.
Mohan Li, Changhua You, Liu Yang 0003, Wenliang Yao
ISCAS1
2025 FLUX: Efficient Descriptor-Driven Clustered Federated Learning under Arbitrary Distribution Shifts
abstract
Federated Learning (FL) enables collaborative model training across multiple clients while preserving data privacy. Traditional FL methods often use a global model to fit all clients, assuming that clients' data are independent and identically distributed (IID). However, when this assumption does not hold, the global model accuracy may drop significantly, limiting FL applicability in real-world scenarios. To address this gap, we propose FLUX, a novel clustering-based FL (CFL) framework that addresses the four most common types of distribution shifts during both training and test time. To this end, FLUX leverages privacy-preserving client-side descriptor extraction and unsupervised clustering to ensure robust performance and scalability across varying levels and types of distribution shifts. Unlike existing CFL methods addressing non-IID client distribution shifts, FLUX i) does not require any prior knowledge of the types of distribution shifts or the number of client clusters, and ii) supports test-time adaptation, enabling unseen and unlabeled clients to benefit from the most suitable cluster-specific models. Extensive experiments across four standard benchmarks, two real-world datasets and ten state-of-the-art baselines show that FLUX improves performance and stability under diverse distribution shifts—achieving an average accuracy gain of up to 23 percentage points over the best-performing baselines—while maintaining computational and communication overhead comparable to FedAvg.
Dario Fenoglio, Mohan Li, Pietro Barbiero, Nicholas D. Lane, Marc Langheinrich, Martin Gjoreski
NeurIPS2
2025 Toward Robust Encrypted Traffic Detection via Graph Contrastive Learning
abstract
With the widespread adoption of network encryption, traditional Deep Packet Inspection (DPI) has become ineffective. Existing approaches that exploit side-channel information often suffer from limited representational capacity, poor robustness, and heavy reliance on large labeled datasets. To address these limitations, we present FGCL, a robust framework for malicious encrypted traffic detection based on graph contrastive learning. FGCL models each bidirectional flow as a Flow Graph (FG), capturing fine-grained interaction patterns between communicating entities. At its core is a novel two-stage augmentation strategy: at the traffic level, realistic obfuscation tactics are simulated to improve robustness, while at the graph level, structural transformations are applied to learn invariant representations. This enables the graph encoder to be pre-trained on large-scale unlabeled data, producing highly generalizable embeddings that can be effectively fine-tuned with only a few labeled samples. Extensive experiments on three real-world datasets show that FGCL consistently outperforms state-of-the-art methods in few-shot learning, adversarial robustness, and overall detection accuracy, achieving up to a 10% F1-score improvement in detecting obfuscated traffic. These results highlight FGCL as an effective and practical solution for encrypted traffic detection in scenarios characterized by label scarcity and adversarial conditions.
Mohan Li, Yanbin Sun
TrustCom2
2025 Invisible trigger image: A dynamic neural backdoor attack based on hidden feature
Mohan Li, Yanbin Sun, Zhihong Tian
Neurocomputing2
2025 Query-Efficient Model Inversion Attacks: An Information Flow View
abstract
Model Inversion Attacks (MIAs) pose a certain threat to the data privacy of learning-based systems, as they enable adversaries to reconstruct identifiable features of the training distribution with only query access to the victim model. In the context of deep learning, the primary challenges associated with MIAs are suboptimal attack success rates and the corresponding high computational costs. Prior efforts assumed that the expansive search space caused these limitations, employing generative models to constrain the dimensions of the search space. Despite the initial success of these generative-based solutions, recent experiments have cast doubt on this fundamental assumption, leaving two open questions about the influential factors determining MIA performance and how to manipulate these factors to improve MIAs. To answer these questions, we reframe MIAs from the perspective of information flow. This new formulation allows us to establish a lower bound for the error probability of MIAs, determined by two critical factors: (1) the size of the search space and (2) the mutual information between input and output random variables. Through a detailed analysis of generative-based MIAs within this theoretical framework, we uncover a trade-off between the size of the search space and the generation capability of generative models. Based on the theoretical conclusions, we introduce the Query-Efficient Model Inversion Approach (QE-MIA). By strategically selecting an appropriate search space and introducing additional mutual information, QE-MIA achieves a reduction of$60\%\sim 70\%$in query overhead while concurrently enhancing the attack success rate by$5\%\sim 25\%$.
Yixiao Xu, Binxing Fang, Mohan Li, Xiaolei Liu 0001, Zhihong Tian 0001
IEEE Trans. Inf. Forensics Secur.3
2025 Neural Honeypoint: An Active Defense Framework Against Model Inversion Attacks
abstract
Learning-based systems have been proved to be vulnerable against model inversion attacks (MIAs), where attackers steal private information of training data by querying the target model using synthetic samples. To alleviate the urgent threat introduced by MIAs, existing advancements are proposed to increase the attack overhead by limiting the information available. Although these methods successfully reduced the attack success rate (ASR) for a one-time inversion attempt, they usually compromise the usability of the protected model. More importantly, existing MIA defense methods fail to capture attack attempts, which can lead to persistent threats to data privacy. To bridge this gap, we propose Neural Honeypoint, an active defense framework against MIAs. The key insight is that MIA attackers will make a series of forward steps in the feature space while benign users will not. Motivated by the observation, defenders can deploy active defense devices (honeypoints) on critical paths to capture attack behaviors. Specifically, Neural Honeypoint first models the attackers' capabilities from the frequency domain and designs specialized honeypoints for protected classes in the training dataset. Subsequently, it deploys these honeypoints into the protected model via backdoor-like model fine-tuning. Then, defenders can distinguish model inversion examples by comparing the similarity of input features with deployed honeypoints. Experiments show that Neural Honeypoint reduces the ASRs of advanced MIAs to 0%~2%. Furthermore, it can effectively capture inversion queries, which helps defenders to detect and block attacks in time.
Yixiao Xu, Mohan Li, Binxing Fang, Yuan Liu 0002, Zhihong Tian 0001
IEEE Trans. Neural Networks Learn. Syst.2
2024 DiaLoc: An Iterative Approach to Embodied Dialog Localization
abstract
Multimodal learning has advanced the performance for many vision-language tasks. However, most existing works in embodied dialog research focus on navigation and leave the localization task understudied. The few existing dialogbased localization approaches assume the availability of entire dialog prior to Iocalizaiton, which is impractical for deployed dialog-based localization. In this paper, we propose DiaLoc, a new dialog-based localization framework which aligns with a real human operator behavior. Specifically, we produce an iterative refinement of location predictions which can visualize current pose believes after each dialog turn. DiaLoc effectively utilizes the multimodal data for multi-shot localization, where a fusion encoder fuses vision and dialog information iteratively. We achieve state-of-the-art results on embodied dialog-based localization task, in single-shot (+7.08% in Acc5@valUnseen) and multishot settings (+10.85% in Acc5@valUnseen). DiaLoc narrows the gap between simulation and real-world applications, opening doors for future research on collaborative localization and navigation.
Chao Zhang 0023, Mohan Li, Ignas Budvytis, Stephan Liwicki
CVPR2
2024 Multi-Frequency Federated Learning for Human Activity Recognition Using Head-Worn Sensors
abstract
Human Activity Recognition (HAR) benefits various application domains, including health and elderly care. Traditional HAR involves constructing pipelines reliant on centralized user data, which can pose privacy concerns as they necessitate the uploading of user data to a centralized server. This work proposes multi-frequency Federated Learning (FL) to enable: (1) privacy-aware ML; (2) joint ML model learning across devices with varying sampling frequency. We focus on head-worn devices (e.g., earbuds and smart glasses), a relatively unexplored domain compared to traditional smartwatch- or smartphone-based HAR. Results have shown improvements on two datasets against frequency-specific approaches, indicating a promising future in the multi-frequency FL-HAR task. The proposed network’s implementation is publicly available for further research and development.**
Dario Fenoglio, Mohan Li, Davide Casnici, Matías Laporte, Shkurta Gashi, Silvia Santini, Martin Gjoreski, Marc Langheinrich
IE2
2024 Prompting Whisper for QA-driven Zero-shot End-to-end Spoken Language Understanding
Mohan Li, Simon Keizer, Rama Sanand Doddipatla
INTERSPEECH1
2024 LT-Defense: Searching-free Backdoor Defense via Exploiting the Long-tailed Effect
abstract
Language models have shown vulnerability against backdoor attacks, threatening the security of services based on them. To mitigate the threat, existing solutions attempted to search for backdoor triggers, which can be time-consuming when handling a large search space. Looking into the attack process, we observe that poisoned data will create a long-tailed effect in the victim model, causing the decision boundary to shift towards the attack targets. Inspired by this observation, we introduce LT-Defense, the first searching-free backdoor defense via exploiting the long-tailed effect. Specifically, LT-Defense employs a small set of clean examples and two metrics to distinguish backdoor-related features in the target model. Upon detecting a backdoor model, LT-Defense additionally provides test-time backdoor freezing and attack target prediction. Extensive experiments demonstrate the effectiveness of LT-Defense in both detection accuracy and efficiency, e.g., in task-agnostic scenarios, LT-Defense achieves 98% accuracy across 1440 models with less than 1% of the time cost of state-of-the-art solutions.
Yixiao Xu, Binxing Fang, Mohan Li, Keke Tang, Zhihong Tian 0001
NeurIPS3
2024 WHISMA: A Speech-LLM to Perform Zero-Shot Spoken Language Understanding
abstract
Speech large language models (speech-LLMs) integrate speech and text-based foundation models to provide a unified framework for handling a wide range of downstream tasks. In this paper, we introduce WHISMA, a speech-LLM tailored for spoken language understanding (SLU) that demonstrates robust performance in various zero-shot settings. WHISMA combines the speech encoder from Whisper with the Llama-3 LLM, and is fine-tuned in a parameter-efficient manner on a comprehensive collection of SLU-related datasets. Our experiments show that WHISMA significantly improves the zero-shot slot filling performance on the SLURP benchmark, achieving a relative gain of 26.6% compared to the current state-of-the-art model. Furthermore, to evaluate WHISMA’s generalisation capabilities to unseen domains, we develop a new task-agnostic benchmark named SLU-GLUE. The evaluation results indicate that WHISMA outperforms an existing speech-LLM (Qwen-Audio) with a relative gain of 33.0%.
Mohan Li, Cong-Thanh Do, Simon Keizer, Youmna Farag, Svetlana Stoyanchev, Rama Sanand Doddipatla
SLT1
2024 An Access Control Method Against Unauthorized and Noncompliant Behaviors of Real-Time Data in Industrial IoT
abstract
There is a large amount of real-time data, e.g., measurement data and instructions, among controllers, sensors, and actuators in the Industrial IoT. These data are vulnerable to unauthorized access and tampering. In addition, once the controller is controlled by malicious code, it may send out dangerous instruction that is not compliant with the preset control process, which seriously interferes with the industrial control process. To achieve correct and undisturbed control based on real-time data, we propose an attribute-based access control (ABAC) method for real-time data in the Industrial IoT to mitigate unauthorized access and tampering, and noncompliant operation. First, we analyze the abnormal behaviors of real-time data interaction in the Industrial IoT. Second, we propose the multilevel hash identity authentication method to identify and block unauthorized access and tampering with real-time data. And, then we model the timing relationship and task logic relationship of the control process into the attribute fields of the ABAC method to identify and block noncompliant operations. Further, we design an access control module and display the lightweight deployment under the availability constraints of the control service. Finally, the proposed access control method is analyzed, proved, and experimented with. The results show that the proposed method can prevent unauthorized and noncompliant behaviors to field real-time data, meanwhile, it has a controllable delay and better scalability.
Mohan Li, Yanbin Sun, Zhihong Tian 0001
IEEE Internet Things J.3
2023 Towards a Unified End-to-End Language Understanding System for Speech and Text Inputs
abstract
End-to-end (E2E) spoken language understanding (SLU) systems facilitate mapping speech inputs directly to semantic outputs, eliminating the need for modular processing of speech-to-text and text-to-semantics sub-tasks using separate models. However, they are now limited to processing speech inputs only, and are not flexible to deal with plain texts. In this paper, we propose an E2E spoken and natural language understanding (SNLU) system that can handle both speech and text within a unified architecture. The system follows the Mask-CTC non-autoregressive approach, and the input flexibility is acquired by partially sharing the decoder between SLU and NLU tasks. Experiments on the SLURP dataset show that the proposed architecture achieves similar performance to using separate E2E SLU and NLU modules, but with relatively 43.7 % less model parameters. We also explore the use of pre-trained speech and language models into the SNLU system, and show that they further improve the performance.
Mohan Li, Catalin Zorila, Cong-Thanh Do, Rama Sanand Doddipatla
ASRU1
2023 Cumulative Attention Based Streaming Transformer ASR with Internal Language Model Joint Training and Rescoring
abstract
This paper presents an approach to improve the performance of streaming Transformer ASR by introducing an internal language model (ILM) as a part of the decoder layers. In the recently pro- posed cumulative attention (CA) based streaming ASR system, only the last or top few decoder layers are equipped with the CA module. Thus in this work, we propose to train the bottom (non-CA) layers as an ILM using an auxiliary LM loss jointly with the rest of the system. During inference, the outputs of the ILM are interpolated with those of the entire Transformer decoder as done in the conventional external language model (ELM) rescoring. The paper also proposes a refinement to the CA algorithm known as CTC look-ahead, in order to improve the precision of endpoint detection. Experiments conducted on AIShell-1, Aidatatang and Librispeech datasets show that the proposed ILM rescoring method achieves on par or better ASR performance when compared to the ELM rescoring baseline. Also, the CTC look-ahead strategy effectively alleviates the early end-of- speech (EOS) triggering issue suffered by the CA module, without bringing noticeable latency degradation.
Mohan Li, Cong-Thanh Do, Rama Sanand Doddipatla
ICASSP1
2023 Domain Adaptive Self-supervised Training of Automatic Speech Recognition
Cong-Thanh Do, Rama Sanand Doddipatla, Mohan Li, Thomas Hain
INTERSPEECH3
2023 Hierarchical Name-based Routing for Content Provider Mobility in ICN
abstract
ICN treats contents as the first citizens and faces serious scalability challenges. Especially for the content provider mobility scenario, the routing updates and routing efficiency should be cost-effective. This paper proposes a hierarchical name-based routing (HNR) for provider mobility. HNR focuses on how to find the mobile provider when given an immutable provider name. Based on the idea that the content provider generally moves within a certain topology area over a period of time, HNR first divides the topology into multiple partition topologies, then it adopts a hierarchical routing scheme that combines two suitable routing schemes for local mobility and global mobility. The experiments by simulation demonstrate that HNR achieves a good tradeoff between efficiency and scalability.
Yanbin Sun, Jianxun Zhou, Xiaoming Zhou, Mohan Li, Zhihong Tian 0001
IWCMC6
2023 Neighborhood Matching Entity Alignment Model for Vulnerability Knowledge Graphs
abstract
Entity alignment aims to match identical entities in different knowledge graphs (KGs). In recent years, entity alignment methods for encyclopedic KGs have achieved significant effectiveness. However, the characteristics of KGs of some specific domains differs from encyclopedic KGs, so that encyclopedic entity alignment methods do not perform well in domain KGs. Vulnerability KGs are a typical type of domain KG characterized by a large number of entities and limited structural variations, but strong heterogeneity across different graphs. The neighborhood matching-based entity alignment methods are effective on vulnerability KGs. However, previous neighborhood matching methods have primarily focused on aligning neighborhoods where both entities and relations are aligned simultaneously, neglecting the neighborhoods where only entities are aligned. Vulnerability KGs typically contain a large number of entities but have a limited types of relations. As a result, each relation often connects a significant number of entities. Incorrect matching of relations can potentially result in incorrect matching of all connected entities, leading to severe error propagation.In this paper, we propose a neighborhood matching based method VNM, for entity alignment in vulnerability KGs. VNM considers two layers of neighborhood, that is, the neighborhood that entities and relations both align and the neighborhood that only entities align. Our method not only mitigates the error propagation caused by incorrect relation matching but also leverages richer neighborhood information. Additionally, inspired by the phenomenon of semantic translation in word embeddings, we introduce a regularizer for semantic embedding of one-to-many and many-to-one relations in vulnerability KGs. Experimental results on four real-world vulnerability datasets demonstrate that our method outperforms existing methods in terms of performance.
Mohan Li, Yanbin Sun
TrustCom2
2023 SAT: sampling acceleration tree for adaptive database repartition
Xiaoxiao Xie, Shengfei Shi, Hongzhi Wang 0001, Mohan Li
World Wide Web (WWW)4
2022 Transformer-Based Streaming ASR with Cumulative Attention
abstract
In this paper, we propose an online attention mechanism, known as cumulative attention (CA), for streaming Transformer-based automatic speech recognition (ASR). Inspired by monotonic chunk-wise attention (MoChA) and head-synchronous decoder-end adaptive computation steps (HS-DACS) algorithms, CA triggers the ASR outputs based on the acoustic information accumulated at each encoding timestep, where the decisions are made using a trainable device, referred to as halting selector. In CA, all the attention heads of the same decoder layer are synchronised to have a unified halting position. This feature effectively alleviates the problem caused by the distinct behaviour of individual heads, which may otherwise give rise to severe latency issues as encountered by MoChA. The ASR experiments conducted on AIShell-1 and Librispeech datasets demonstrate that the proposed CA-based Transformer system can achieve on par or better performance with significant reduction in latency during inference, when compared to other streaming Transformer systems in literature.
Mohan Li, Shucong Zhang, Catalin Zorila, Rama Sanand Doddipatla
ICASSP1
2022 Multiple-hypothesis RNN-T Loss for Unsupervised Fine-tuning and Self-training of Neural Transducer
abstract
This paper proposes a new approach to perform unsupervised fine-tuning and self-training using unlabeled speech data for recurrent neural network (RNN)-Transducer (RNN-T) end-to-end (E2E) automatic speech recognition (ASR) systems.Conventional systems perform fine-tuning/self-training using ASR hypothesis as the targets when using unlabeled audio data and are susceptible to the ASR performance of the base model.Here in order to alleviate the influence of ASR errors while using unlabeled data, we propose a multiple-hypothesis RNN-T loss that incorporates multiple ASR 1-best hypotheses into the loss function.For the fine-tuning task, ASR experiments on Librispeech show that the multiple-hypothesis approach achieves a relative reduction of 14.2% word error rate (WER) when compared to the single-hypothesis approach, on the test other set.For the self-training task, ASR models are trained using supervised data from Wall Street Journal (WSJ), Aurora-4 along with CHiME-4 real noisy data as unlabeled data.The multiplehypothesis approach yields a relative reduction of 3.3% WER on the CHiME-4's single-channel real noisy evaluation set when compared with the single-hypothesis approach.
Cong-Thanh Do, Mohan Li, Rama Sanand Doddipatla
INTERSPEECH2
2022 Self-regularised Minimum Latency Training for Streaming Transformer-based Speech Recognition
Mohan Li, Rama Sanand Doddipatla, Catalin Zorila
INTERSPEECH1
2022 Multi-round Data Poisoning Attack and Defense against Truth Discovery in Crowdsensing Systems
abstract
Crowdsensing systems collect various types of data based on the personal sensing devices of ordinary users. These users are generally the workers of crowdsourcing tasks released by the systems. However, due to the lack of strict authentication for user identity in many crowdsensing services, malicious attackers can sneak into normal workers and submit malicious data to the system to launch data poisoning attacks. Truth discovery algorithms aim to calculate workers' trustworthiness and try to find the ground truth from inconsistent data submitted by different workers. It can help the systems filter out some data providers with poor quality, and thus can defend against some simple data poisoning attacks, such as random attacks, max-value attacks, etc. However, we found that in multi-rounds of data collaction tasks, the attackers can still attack successfully by slightly modifying the simple poisoning strategies. Attackers can deceive the truth discovery algorithm by alternately submitting real and fake data, thereby mislead the system to infer the wrong truth. Therefore, in this paper we study the attack and defense methods for multi-round data poisoning against truth discovery in crowdsensing systems. First, we verify the vulnerability of a class of commonly used truth discovery framework under a multi-rounds of data poisoning strategy named “Hide-AttackPois”. Experiments show that only using a simple hide attack strategy can cause great disturbance to the output of the algorithm. Second, we optimize this class of truth discovery framework to enhance its robustness and enable it to defend against “Hide-AttackPois”. Then, we further improve the data poisoning attack model, so that the model can learn better attack strategies and effectively attack the optimized truth discovery framework. We conduct experiments on real dataset to verify the effectiveness of the proposed method.
Hongniu Zhang, Mohan Li
MDM2
2022 Non-Autoregressive End-to-End Approaches for Joint Automatic Speech Recognition and Spoken Language Understanding
abstract
This paper presents the use of non-autoregressive (NAR) approaches for joint automatic speech recognition (ASR) and spoken language understanding (SLU) tasks. The proposed NAR systems employ a Conformer encoder that applies connectionist temporal classification (CTC) to transcribe the speech utterance into raw ASR hypotheses, which are further refined with a bidirectional encoder representations from Transformers (BERT)-like decoder. In the meantime, the intent and slot labels of the utterance are predicted simultaneously using the same decoder. Both Mask-CTC and self-conditioned CTC (SC-CTC) approaches are explored for this study. Experiments conducted on the SLURP dataset show that the proposed SC-Mask-CTC NAR system achieves 3.7% and 3.2% absolute gains in SLU metrics and a competitive level of ASR accuracy, when compared to a Conformer-Transformer based autoregressive (AR) model. Additionally, the NAR systems achieve 6× faster decoding speed than the AR baseline.
Mohan Li, Rama Sanand Doddipatla
SLT1
2022 Robust Truth Discovery Against Multi-round Data Poisoning Attacks
Hongniu Zhang, Mohan Li, Yanbin Sun, Guanqun Qu
WASA (1)2
2021 Improving HS-DACS Based Streaming Transformer ASR with Deep Reinforcement Learning
abstract
Transformer-based systems, though have achieved state-of-the-art performance on a wide range of automatic speech recognition (ASR) tasks, are subject to severe latency issues during inference that limit their deployment in real world applications. To enable online decoding, we recently proposed Decoder-end adaptive computation steps (DACS) and the head-synchronous version (HS-DACS) algorithms, which were shown to reduce the computational cost for decoding and close the gap in performance between offline and streaming Transformer ASR. In DACS/HS-DACS based systems, the halting position that triggers the ASR output was determined with an accumulation threshold that was set arbitrarily. In this paper, we propose a deep reinforcement learning approach to improve HS-DACS where an external agent is utilised to optimise the halting position. Experiments on AIShell-1, Tedlium-2 and Librispeech datasets show that the proposed method can further cut down the computation cost in inference with relative gains of 40.0%, 32.8% and 53.4% respectively, and still maintain similar ASR performance when compared with HS-DACS.
Mohan Li, Rama Sanand Doddipatla
ASRU1
2021 Head-Synchronous Decoding for Transformer-Based Streaming ASR
abstract
Online Transformer-based automatic speech recognition (ASR) systems have been extensively studied due to the increasing demand for streaming applications. Recently proposed Decoder-end Adaptive Computation Steps (DACS) algorithm for online Transformer ASR was shown to achieve state-of-the-art performance and outperform other existing methods. However, like any other online approach, the DACS-based attention heads in each of the Transformer decoder layers operate independently (or asynchronously) and lead to diverged attending positions. Since DACS employs a truncation threshold to determine the halting position, some of the attention weights are cut off untimely and might impact the stability and precision of decoding. To overcome these issues, here we propose a head-synchronous (HS) version of the DACS algorithm, where the boundary of attention is jointly detected by all the DACS heads in each decoder layer. ASR experiments on Wall Street Journal (WSJ), AIShell-1 and Lib- rispeech show that the proposed method consistently outperforms vanilla DACS and achieves state-of-the-art performance. We will also demonstrate that HS-DACS has reduced decoding cost when compared to vanilla DACS.
Mohan Li, Catalin Zorila, Rama Sanand Doddipatla
ICASSP1
2021 Transformer-Based Online Speech Recognition with Decoder-end Adaptive Computation Steps
abstract
Transformer-based end-to-end (E2E) automatic speech recognition (ASR) systems have recently gained wide popularity, and are shown to outperform E2E models based on recurrent structures on a number of ASR tasks. However, like other E2E models, Transformer ASR also requires the full input sequence for calculating the attentions on both encoder and decoder, leading to increased latency and posing a challenge for online ASR. The paper proposes Decoder-end Adaptive Computation Steps (DACS) algorithm to address the issue of latency and facilitate online ASR. The proposed algorithm streams the decoding of Transformer ASR by triggering an output after the confidence acquired from the encoder states reaches a certain threshold. Unlike other monotonic attention mechanisms that risk visiting the entire encoder states for each output step, the paper introduces a maximum look-ahead step into the DACS algorithm to prevent from reaching the end of speech too fast. A Chunkwise en-coder is adopted in our system to handle real-time speech inputs. The proposed online Transformer ASR system has been evaluated on Wall Street Journal (WSJ) and AIShell-1 datasets, yielding 5.5% word error rate (WER) and 7.1% character error rate (CER) respectively, with only a minor decay in performance when compared to the offline systems.
Mohan Li, Catalin Zorila, Rama Sanand Doddipatla
SLT1
2021 An Investigation into the Multi-channel Time Domain Speaker Extraction Network
abstract
This paper presents an investigation into the effectiveness of spatial features for improving time-domain speaker extraction systems. A two-dimensional Convolutional Neural Network (CNN) based encoder is proposed to capture the spatial information within the multichannel input, which are then combined with the spectral features of a single channel extraction network. Two variants of target speaker extraction methods were tested, one which employs a pre-trained i-vector system to compute a speaker embedding (System A), and one which employs a jointly trained neural network to extract the embeddings directly from time domain enrolment signals (System B). The evaluation was performed on the spatialized WSJ0-2mix dataset using the Signal-to-Distortion Ratio (SDR) metric, and ASR accuracy. In the anechoic condition, more than 10 dB and 7 dB absolute SDR gains were achieved when the 2-D CNN spatial encoder features were included with Systems A and B, respectively. The performance gains in reverberation were lower, however, we have demonstrated that retraining the systems by applying dereverberation preprocessing can significantly boost both the target speaker extraction and ASR performances.
Catalin Zorila, Mohan Li, Rama Sanand Doddipatla
SLT2
2021 Secure Data Sharing Framework via Hierarchical Greedy Embedding in Darknets
Yanbin Sun, Mohan Li, Shen Su, Zhihong Tian 0001, Wei Shi 0001
Mob. Networks Appl.2
2021 Honeypot Identification in Softwarized Industrial Cyber-Physical Systems
abstract
In softwarized industrial networking, honeypot identification is very important for both the attacker and the defender. Existing honeypot identification relies on simple features of honeypot. There exist two challenges: The simple feature is easily simulated, which causes inaccurate results, whereas the advanced feature relies on high interactions, which lead to security risks. To cope with these challenges, in this article, we propose a secure fuzzy testing approach for honeypot identification inspired by vulnerability mining. It utilizes error handling to distinguish honeypots and real devices. Specifically, we adopt a novel identification architecture with two steps. First, a multiobject fuzzy testing is proposed. It adopts mutation rules and security rules to generate effective and secure probe packets. Then, these probe packets are used for scanning and identification. Experiments show that the fuzzy testing is effective and corresponding probe packet can acquire more features than other packets. These features are helpful for honeypot identification.
Yanbin Sun, Zhihong Tian 0001, Mohan Li, Shen Su, Xiaojiang Du, Mohsen Guizani
IEEE Trans. Ind. Informatics3
2020 Deep Reinforcement Learning for Partially Observable Data Poisoning Attack in Crowdsensing Systems
abstract
Crowdsensing systems collect various types of data from sensors embedded on mobile devices owned by individuals. These individuals are commonly referred to as workers that complete tasks published by crowdsensing systems. Because of the relative lack of control over worker identities, crowdsensing systems are susceptible to data poisoning attacks which interfering with data analysis results by injecting fake data conflicting with ground truth. Frameworks like TruthFinder can resolve data conflicts by evaluating the trustworthiness of the data providers. These frameworks somehow make crowdsensing systems more robust since they can limit the impact of dirty data by reducing the value of unreliable workers. However, previous work has shown that TruthFinder may also be affected by the data poisoning attack when the malicious workers have access to global information. In this article, we focus on partially observable data poisoning attacks in crowdsensing systems. We show that even if the malicious workers only have access to local information, they can find effective data poisoning attack strategies to interfere with crowdsensing systems with TruthFinder. First, we formally model the problem of partially observable data poisoning attack against crowdsensing systems. Then, we propose a data poisoning attack method based on deep reinforcement learning, which helps malicious workers jeopardize with TruthFinder while hiding themselves. Based on the method, the malicious workers can learn from their attack attempts and evolve the poisoning strategies continuously. Finally, we conduct experiments on real-life data sets to verify the effectiveness of the proposed method.
Mohan Li, Yanbin Sun, Hui Lu 0005, Sabita Maharjan, Zhihong Tian 0001
IEEE Internet Things J.1
2019 End-to-end Speech Recognition with Adaptive Computation Steps
abstract
In this paper, we present Adaptive Computation Steps (ACS) algorithm, which enables end-to-end speech recognition models to dynamically decide how many frames should be processed to predict a linguistic output. The model that applies ACS algorithm follows the encoder-decoder framework, while unlike the attention-based models, it produces alignments independently at the encoder side using the correlation between adjacent frames. Thus, predictions can be made as soon as sufficient acoustic information is received, which makes the model applicable in online cases. Besides, a small change is made to the decoding stage of the encoder-decoder framework, which allows the prediction to exploit bidirectional contexts. We verify the ACS algorithm on a Mandarin speech corpus AIShell-1, and it achieves a 31.2% CER in the online occasion, compared to the 32.4% CER of the attention-based model. To fully demonstrate the advantage of ACS algorithm, offline experiments are conducted, in which our ACS model achieves an 18.7% CER, outperforming the attention-based counterpart with the CER of 22.0%.
Mohan Li, Masanori Hattori
ICASSP1
2019 Framewise Supervised Training Towards End-to-End Speech Recognition Models: First Results
Mohan Li, Yuanjiang Cao, Weicong Zhou
INTERSPEECH1
2019 Block-DEF: A secure digital evidence framework using blockchain
Zhihong Tian 0001, Mohan Li, Meikang Qiu, Yanbin Sun, Shen Su
Inf. Sci.2
2010 Efficient Duplicate Record Detection Based on Similarity Estimation
Mohan Li, Hongzhi Wang 0001, Jianzhong Li 0001, Hong Gao 0001
WAIM1