EDBT 2026 Demo / reviewers in the wild / expert
Mingye Li
dblp:283/6951
· DBLP profile ↗
17ranked-venue papers
6as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 5 first-author · 8 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate AttentionabstractYuxiang Huang, Mingye Li, Xu Han, Chaojun Xiao, Weilin Zhao, Ao Sun, Ziqi Yuan, Hao Zhou, Fandong Meng, Zhiyuan Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuxiang Huang 0001, Mingye Li, Xu Han 0007, Chaojun Xiao, Weilin Zhao, Hao Zhou 0012, Fandong Meng, Zhiyuan Liu 0001 |
ACL (1) | 2 |
| 2026 | Plug-and-play domain generalization: Empowering deployed RFFI systems with receiver agnosticism
Mingye Li, Xuelin Yang |
Ad Hoc Networks | 4 |
| 2026 | Flex-NTT: Design of a Flexible and Compact Number Theoretic Transform Architecture for Homomorphic Encryption ApplicationsabstractThis article presents Flex-NTT, a flexible (configurable) and area-efficient number theoretic transform (NTT) architecture featuring novel unified memory access patterns. The proposed design offers compile-time configurability (CTC) to support various parallel butterfly units (BUs) and run-time configurability (RTC) to accommodate diverse NTT operation sizes without recompilation. A single hardware instance is reconfigurable for NTT, inverse NTT (INTT), and point-wise multiplication, improving hardware utilization. The proposed coefficient and twiddle-factor memory access schemes achieve the theoretical minimum memory sizes, eliminate intermediate buffers, and enable natural-order outputs without additional reordering. FPGA evaluations demonstrate that Flex-NTT achieves up to$4.25 \times $BRAM savings and$2.82 \times $performance improvements over prior NTT/INTT designs. When applied to polynomial multiplication, it delivers up to$2.35 \times $performance and$2.93 \times $BRAM utilization improvements under the same metric. Zeming Cheng, Xiao Tuo, Mingye Li, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2025 | APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUsabstractWhile long-context inference is crucial for advancing large language model (LLM) applications, its prefill speed remains a significant bottleneck. Current approaches, including sequence parallelism strategies and compute reduction through approximate attention mechanisms, still fall short of delivering optimal inference efficiency. This hinders scaling the inputs to longer sequences and processing long-context queries in a timely manner. To address this, we introduce APB, an efficient long-context inference framework that leverages multi-host approximate attention to enhance prefill speed by reducing compute and enhancing parallelism simultaneously. APB introduces a communication mechanism for essential key-value pairs within a sequence parallelism framework, enabling a faster inference speed while maintaining task performance. We implement APB by incorporating a tailored FlashAttn kernel alongside optimized distribution strategies, supporting diverse models and parallelism configurations. APB achieves speedups of up to 9.2\times, 4.2\times, and 1.6\times compared with FlashAttn, RingAttn, and StarAttn, respectively, without any observable task performance degradation. Yuxiang Huang 0001, Mingye Li, Xu Han 0007, Chaojun Xiao, Weilin Zhao, Sun Ao, Hao Zhou 0012, Jie Zhou 0016, Zhiyuan Liu 0001, Maosong Sun 0001 |
ACL (1) | 2 |
| 2025 | Asymmetric Predictive Testing for Aging in SRAMsabstractTo avoid corruption of user data, predictive testing methods have been proposed to identify SRAMs likely to fail in the near future due to aging. These methods use aggressive operating conditions, e.g., adjustments to wordline voltages or the power supply voltage, that are calibrated to provide high coverage of SRAMs likely to fail in the near future, but end up with some over-testing, i.e., spuriously identifying some chips as likely to fail. We first present our study which discovered that a large fraction of over-tested chips fail due to read faults triggered during read-1 operations. Our analysis identifies asymmetric aging in SRAM cells, which are more likely to store zeros, as the root cause for this. We build on this discovery to propose an asymmetric predictive testing method which performs writes using normal voltages, read-0 at aggressive voltages, and read-1 at less aggressive voltages. We demonstrate that this method significantly reduces over-testing by over $3 \times$ to $5 \times$, for low limits on under-testing. We also propose and use a new statistical sampling and simulation method to enable fast convergence and accurate evaluation of asymmetric predictive testing. Yunkun Lin, Mingye Li |
DAC | 2 |
| 2025 | Radio Frequency Fingerprint Identification for Few-Shot Scenario via Grad-CAM Feature Augmentation and Meta-LearningabstractRadio Frequency Fingerprint Identification (RFFI), which leverages hardware-specific impairments in Internet of Things (IoT) devices, is widely used for device authentication and spoofing attack detection to enhance communication security. However, the existing RFFI methods heavily depend on large-scale training datasets in deep learning (DL), with severe overfitting issues if the training samples are scarce. This paper proposes a few-shot learning framework that combines feature augmentation and meta-learning to overcome these challenges. A novel data augmentation technique based on grad class activation maps (Grad-CAM) is introduced to address the scarcity of training samples, which generates augmented samples by adjusting the weights of receptive fields in feature maps, forming an auxiliary dataset for training. In meta-training, the auxiliary dataset is used to construct tasks comprising support and query sets. By extracting common features from the limited sample size, the framework trains a meta-model with robust generalization capabilities. In the deployment phase, a fine-tuning strategy further optimizes the classifier using a small labeled dataset from new IoT devices, allowing rapid adaptation with high accuracy. The proposed framework is evaluated on the large-scale open-source dataset, achieving an accuracy of 94.1% under 8-way 5-shot, with only 25 samples per device in meta-training. Meta-learning boosts performance by 15%-20%, with meta-training feature augmentation further increasing accuracy by 5%-6.6%. Compared to baseline methods, the proposed framework improves the accuracy by 30%, which outperforms the state-of-the-art algorithms by 4%-25%. Mingye Li, Yilin Qiu, Xuelin Yang |
IEEE Internet Things J. | 1 |
| 2025 | Efficient feature extraction of radio-frequency fingerprint using continuous wavelet transform
Mutala Mohammed, Xinyong Peng, Mingye Li, Rahel Abayneh, Xuelin Yang |
Wirel. Networks | 4 |
| 2024 | Challenges and Unexplored Frontiers in Electronic Design Automation for Superconducting Digital LogicabstractPositioned as a highly promising post-CMOS computing technology, superconductor electronics (SCE) offer the potential for unparalleled performance and energy efficiency gains compared to end-of-roadmap CMOS circuits. However, achieving very large-scale integration poses numerous challenges. These challenges span from the modeling and analysis of superconducting devices and logic gates to the intricate design of complex SCE circuits and systems. Addressing power and clock distribution issues, minimizing adverse effects of flux trappings, and mitigating stray electromagnetic fields in sensitive SCE circuitry are key challenges that need attention. Verification and testing of SCE circuits also remain open problems. Moreover, scaling the minimum feature sizes of SCE circuits, currently set at 150nm, presents critical scaling and physical design challenges that must be overcome. This review aims to delve into these issues, providing detailed insights while exploring existing or potential solutions to overcome them. Sasan Razmkhah, Robert Aviles, Mingye Li, Sandeep Gupta 0001, Peter A. Beerel, Massoud Pedram |
DATE | 3 |
| 2024 | Predictive Testing for Aging in SRAMs and MitigationabstractWe develop a method to estimate lifetime performance, yield, and power for Static Random-Access Memories (SRAMs) that captures the combination of process variations and aging. Using this method, we design and validate predictive tests to detect future aging failures. We use the results of predictive tests to reconfigure dynamic voltage and frequency scaling (DVFS) to reduce aging failures at minimal energy and latency overheads. Yunkun Lin, Mingye Li, Sandeep Gupta 0001 |
ITC | 2 |
| 2024 | Built in self test (BIST) for RSFQ circuitsabstractIn the era beyond the end of physical scaling of CMOS, growing attention is being paid to Superconducting electronics (SCE), especially Rapid Single Flux Quantum (RSFQ) logic due to its high-performance and low power consumption. In [1]–[3], static and delay fault models, corresponding automatic test pattern generator (ATPG) for testing these faults, and a scan architecture are developed. However, test pattern application involves moving patterns and responses via long wires from the test equipment at room temperature to the chip under test in liquid helium, which severely reduces the test clock frequency. At the same time, due to the high clock frequency of this technology, at-speed test is necessary for testing delay faults.In this paper, we present a scan-based BIST for RSFQ circuits which performs at-speed self-test including pseudo random pattern generation and response compression. We show that existing designs of pattern generators cannot be directly used for RSFQ and present new designs. Based on the scan architecture in [3], we design a new control strategy for at-speed self-test. We demonstrate that our new architecture supports testing at low overheads. Mingye Li, Yunkun Lin, Sandeep Gupta 0001 |
VTS | 1 |
| 2024 | Channel-Robust RF Fingerprint Identification Using Multi-Task Learning and Receiver CollaborationabstractRobust radio frequency fingerprint identification (RFFI) is crucial for physical layer authentication, while it suffers from channel effects and requires extra overhead to increase recognition accuracy (RA). To address this, an efficient channel-robust RFFI scheme is proposed, employing a specialized multi-task learning (MTL) framework to direct the neural network (NN) toward extracting channel-robust features. In addition, receiver collaboration (RC) is utilized for data augmentation and output calibration. Experimental results demonstrate that the RA is significantly increased from 51.72% to 99.97% when using the open-resource Wi-Fi signal datasets collected from different time periods. Meanwhile, the requirements for extra data transmission, NN structure, and feature crafting in the inferring stage are dramatically simplified. Xinyong Peng, Mingye Li, Xuelin Yang |
IEEE Signal Process. Lett. | 4 |
| 2023 | Design for testability (DFT) for RSFQ circuitsabstractSuperconducting electronics (SCE), especially Rapid Single Flux Quantum (RSFQ) logic, is being developed due to its high-performance and low power. In [1] –[3], we developed new static and delay fault models and an efficient automatic test pattern generator (ATPG) for testing both delay and static faults in RSFQ logic. However, test pattern application involves moving patterns and responses via long wires from the test equipment at room temperature to the chip under test in liquid helium. Due to the high cost associated with large numbers of such wires, testing is extremely expensive in absence of design for testability.We present a scan architecture for RSFQ circuits which enables the application of a large number of test patterns. Due to the unique characteristic of RSFQ, this scan architecture includes completely new scan cell design and a new scan control strategy. The on-chip test control logic enables scan chain to shift in test patterns from the test equipment at room temperature via a small number of wires, apply the pattern to the chip under test in parallel and at speed, and shift out the corresponding test response for checking. We demonstrate that our new scan architecture supports testing at low overheads. Mingye Li, Yunkun Lin, Sandeep Gupta 0001 |
VTS | 1 |
| 2023 | Efficient Data Offloading Using Markovian Decision on State Reward Action in Edge Computing
Mingye Li, Haiwei Lei, Riza Sulaiman, Wejdan Deebani, Meshal Shutaywi |
J. Grid Comput. | 1 |
| 2023 | Integration of Device Fingerprint Authentication and Physical-Layer Secret Key GenerationabstractAn integrated physical-layer security scheme is proposed, which combines dynamic physical-layer secret key generation (SKG) and static device fingerprint authentication (DFA) for secure transmission in frequency division duplex (FDD) systems. The proposed approach employs random forest (RF) algorithms to achieve high-speed SKG and utilizes autoencoders (AE) to realize high-accuracy DFA. Besides, channel interference to DFA is minimized during the integration. Experimental results demonstrated that the SKG achieved a key generation rate (KGR) of 57.71 Kbps, while the open-set DFA recognition accuracy reached 99.60%, with a miss alarm rate (MAR) of 0.93%. The proposed scheme successfully integrates physical-layer identity recognition and a fast key generation with excellent performance. Liuming Zhang, Mingye Li, Xuelin Yang |
IEEE Signal Process. Lett. | 4 |
| 2022 | Methods for testing path delay and static faults in RSFQ circuitsabstractSuperconducting electronics (SCE), especially Rapid Single Flux Quantum (RSFQ) logic, is being developed due to its high-performance and low power. In [1] [2], we developed new static and delay fault models and an automatic test pattern generator (ATPG) for path delay faults in RSFQ logic. Here we develop a method for selecting path delay faults by identifying the subset of paths for which the delay can exceed the clock period under the main cause of delay faults for RSFQ, namely extreme process variations. We show that this dramatically reduces the number of delay tests required due to the characteristics of gate-level pipelined design, a necessary requirement for RSFQ. We also extend our method to be the first ATPG to generate tests for RSFQ-specific static fault models derived in [1]. We demonstrate that our new ATPG achieves very high coverage of static and delay faults with small numbers of patterns. Mingye Li, Sandeep Gupta 0001 |
VTS | 1 |
| 2021 | Encoder-Decoder Couplet Generation Model Based on 'Trapezoidal Context' Character VectorabstractAbstract This paper studies the couplet generation model which automatically generates the second line of a couplet by giving the first line. Unlike other sequence generation problems, couplet generation not only considers the sequential context within a sentence line but also emphasizes the relationships between the corresponding words of first and second lines. Therefore, a trapezoidal context character embedding the vector model has been developed firstly, which considers the ‘sequence context’ and the ‘corresponding word context’ simultaneously. Afterwards, we chose the typical encoder–decoder framework to solve the sequence–sequence problems, of which the encoder and decoder are used by bi-directional GRU and GRU, respectively. In order to further increase the semantic consistency of the first and second lines of couplets, the pre-trained sentence vector of the first line is added to the attention mechanism in the model. To verify the effectiveness of the method, it is applied to the real data set. Experimental results show that our proposed model can compete with the up-to-date methods, and both adding sentence vectors to attention and using trapezoidal context character vectors can improve the effectiveness of the algorithm. Rui Gao 0004, Mingye Li, Shoufeng Li, Xiaohu Shi |
Comput. J. | 3 |
| 2020 | Data-driven fault model development for superconducting logicabstractSuperconducting technology is being seriously explored for certain applications. We propose a new clean-slate method to derive fault models from large numbers of simulation results. For this technology, our method identifies completely new fault models - overflow, pulse-escape, and pattern-sensitive - in addition to the well-known stuck-at faults. Mingye Li, Sandeep Gupta 0001 |
ITC | 1 |