EDBT 2026 Demo / reviewers in the wild / expert
Tong Qu
dblp:202/0895
· DBLP profile ↗
7ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0002-5300-8255ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | POFGSP: Priority-Based Out-of-Order Scheduling and Fine-Grain Status Polling for SSD Performance ImprovementabstractWith the development of flash technology, the increasing throughput gap betweennandflash memory (NFM) arrays and the I/O interface has become a performance bottleneck for NFM-based solid-state drives (SSDs). Multilevel parallelism techniques have been employed on modern SSDs to meet the challenge of increasing demands for bandwidth in I/O-intensive workloads. However, conventional parallel methods only monitor the status of ways, resulting in the “idle bubble”—idle time of the dies cannot execute subsequent operations until all the dies in the way complete command execution. This issue limits the resource utilization and performance of SSDs. To minimize the idle bubble, we propose priority-based out-of-order scheduling and fine-grain status polling (POFGSP). The priority-based out-of-order scheduling relaxes constraints on command execution order and schedules commands with the same execution time to be executed in parallel. Therefore, the scheduler reduces these idle bubbles caused by differences in command execution times. Moreover, the fine-grain status polling approach polls the die-level status during the interface’s idle time, reducing idle bubbles with accurate status. Compared to state-of-the-art schedulers, our POFGSP approach can reduce request response time by 35.6% under real-world cloud block storage workloads and improve the SSD system’s maximum bandwidth by 8.7%–74.9%. Wentian Wu, Qianhui Li, Tong Qu, Qi Wang 0041, Zongliang Huo, Tian-Chun Ye 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | NV-APP: Invalid Programming Performance Improved No-Verify and Adaptive Pulse Programming Scheme for 3-D QLC nand FlashabstractQuad-level cell (QLC) has received significant attention recently due to its extremely high storage capacity. However, because of its poor reliability, QLC-based solid-state drives (SSDs) require a two-step programming to reduce the layer interference. But during the interval between two programming steps on the same wordline (WL), data could be invalidated from update operations, leading to invalid programming and degraded performance. To mitigate the performance loss, we propose the NV-APP scheme to minimize the program and verify pulses during the second-step programming. NV-APP integrates the no-verify (NV) scheme and the adaptive pulse programming scheme (APP). The NV scheme omits verify pulses of invalid verify voltages. The APP scheme adaptively increases the programming step voltage$(V_{\mathrm { step}})$to accelerate cells’ threshold voltage shift, reducing the number of both program and verify pulses. Device-level simulation results show that the NV-APP scheme reduces the total number of program pulses by an average of 27.03% and verify pulses by an average of 48.70% across various invalid cases during the second-step programming. Based on a modified 3-D QLC SSD simulator with typical traces, the experiments demonstrate that our scheme reduces two-step programming time by an average of 17% on partially invalid WLs, close to the 19.8% reduction achieved by the ideal scheme with no performance loss. Qianqi Zhao, Jing He 0020, Tong Qu, Wentian Wu, Qianhui Li, Qi Wang 0041, Zongliang Huo, Tian-Chun Ye 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | Flexible Hotspot Detection Based on Fully Convolutional Network With Transfer LearningabstractLayout hotspot detection is one of the most important issues for the reliability enhancement of integrated circuits. Machine learning-based hotspot detectors have shown their advantages of efficiency and generalization compared with computationally intensive lithography process simulation. However, most machine learning-based hotspot detectors only accept layout clips of fixed size as input with the potential defect whose location is restricted at the center of each clip. Therefore, they cannot be used directly for multiple hotspots detection in a large area, which occurs frequently in real design cases. In this article, we build a new end-to-end hotspot detector based on a fully convolutional network, which has the flexibility of detecting a various number of hotspots in a layout of any size at one time. Moreover, we also develop a transfer learning scheme matching our proposed detector network, which can reduce the requirement of sample number when setting up a new model for a more advanced technology node. The experimental results demonstrate our proposed hotspot detector outstanding among state-of-the-art works and the transfer learning scheme is effective. Tianyang Gai, Tong Qu, Xiaojing Su, Renren Xu, Yajuan Su, Yayi Wei, Tian-Chun Ye 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | Asynchronous Reinforcement Learning Framework and Knowledge Transfer for Net-Order Exploration in Detailed RoutingabstractThe net orders in detailed routing are crucial to routing closure, especially in most modern routers following the sequential routing manner with the rip-up and reroute scheme. In advanced technology nodes, detailed routing has to deal with complicated design rules and large problem sizes, making its performance more sensitive to the order of nets to be routed. In the literature, the net orders are mostly determined by simple heuristic rules tuned for specific benchmarks. In this work, we propose an asynchronous reinforcement learning (RL) framework to automatically search for optimal ordering strategies and a transfer learning (TL) algorithm to improve performance. By asynchronous querying, the router, pretraining the RL agents, and finetuning with the TL algorithm, we can generate high-performance routing sequences to achieve a 26% reduction in the DRC violations and a 1.2% reduction in the total costs compared with the state-of-the-art detailed router. Yibo Lin, Tong Qu, Zongqing Lu 0002, Yajuan Su, Yayi Wei |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | Asynchronous Reinforcement Learning Framework for Net Order Exploration in Detailed RoutingabstractThe net orders in detailed routing are crucial to routing closure, especially in most modern routers following the sequential routing manner with the rip-up and reroute scheme. In advanced technology nodes, detailed routing has to deal with complicated design rules and large problem sizes, making its performance more sensitive to the order of nets to be routed. In literature, the net orders are mostly determined by simple heuristic rules tuned for specific benchmarks. In this work, we propose an asynchronous reinforcement learning (RL) framework to search for optimal ordering strategies automatically. By asynchronous querying the router and training the RL agents, we can generate highperformance routing sequences to achieve better solution quality. Tong Qu, Yibo Lin, Zongqing Lu 0002, Yajuan Su, Yayi Wei |
DATE | 1 |
| 2021 | Visual content-enhanced sequential recommendation with feature-level attention
Tong Qu, Shoujin Wang |
Neurocomputing | 1 |
| 2019 | Photoplethysmogram-based Cognitive Load Assessment Using Multi-Feature Fusion ModelabstractCognitive load assessment is crucial for user studies and human--computer interaction designs. As a noninvasive and easy-to-use category of measures, current photoplethysmogram- (PPG) based assessment methods rely on single or small-scale predefined features to recognize responses induced by people’s cognitive load, which are not stable in assessment accuracy. In this study, we propose a machine-learning method by using 46 kinds of PPG features together to improve the measurement accuracy for cognitive load. We test the method on 16 participants through the classical n-back tasks (0-back, 1-back, and 2-back). The accuracy of the machine-learning method in differentiating different levels of cognitive loads induced by task difficulties can reach 100% in 0-back vs. 2-back tasks, which outperformed the traditional HRV-based and single-PPG-feature-based methods by 12--55%. When using “leave-one-participant-out” subject-independent cross validation, 87.5% binary classification accuracy was reached, which is at the state-of-the-art level. The proposed method can also support real-time cognitive load assessment by beat-to-beat classifications with better performance than the traditional single-feature-based real-time evaluation method. Xiao Zhang 0008, Yongqiang Lyu 0001, Tong Qu, Pengfei Qiu, Xiaomin Luo, Shunjie Fan, Yuanchun Shi |
ACM Trans. Appl. Percept. | 3 |