EDBT 2026 Demo / reviewers in the wild / expert
Quan Fan
dblp:227/7157
· DBLP profile ↗
7ranked-venue papers
1as first author
6since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ECC-IDS: A Robust ECC-Based Identity Signcryption Scheme for UAV to Ground Station in Intelligent Search and Rescue
Yimin Wang 0004, Quan Fan, Hong Zhong 0001, Jie Cui 0004 |
IEEE Trans. Reliab. | 3 |
| 2025 | A fast recognition framework for identifying damage levels in rotating and small target solar greenhouse under complex scenarios
Danni Jia, Xinyue Ren, Cailong Cheng, Quan Fan |
Eng. Appl. Artif. Intell. | 7 |
| 2023 | OSMO: Enhanced Offloading for Data Stream Perception with Smoothness and OrderlinessabstractOn-device AI is taking over our daily lives by moving closer to mobile devices as perception applications. A data stream perception application generally has three essential requirements: timeliness, smoothness, and orderliness. Most researchers’ efforts to date have proposed various offloading approaches to accelerate compute-intensive AI algorithms in perception applications, thereby fulfilling the requirement of timeliness. However, the lack of concern about the smoothness and orderliness of the data stream will result in fluctuation and commotion anomalies that greatly impair the user experience. In this paper, we propose an enhanced Offloading System with sMoothness and Orderliness (OSMO) to guarantee perception applications’ smooth refresh rates while processing data streams in proper orders with low overhead. OSMO takes advantage of heterogeneous computing devices and data-level parallelism in the offloading process. A scheduling strategy is further devised that dynamically tunes a set of parameters to achieve the best trade-offs among the three requirements of perception applications. We implement a prototype system based on TensorFlow and its typical Android demos. Real-world evaluations demonstrate that our solution can effectively address the fluctuation and commotion issues while providing a high data processing rate with multi-device collaboration. Chaokun Zhang, Quan Fan, Jinlong E |
ICPADS | 2 |
| 2021 | FT-BLAS: a high performance BLAS implementation with online fault toleranceabstractBasic Linear Algebra Subprograms (BLAS) is a core library in scientific computing and machine learning. This paper presents FT-BLAS, a new implementation of BLAS routines that not only tolerates soft errors on the fly, but also provides comparable performance to modern state-of-the-art BLAS libraries on widely-used processors such as Intel Skylake and Cascade Lake. To accommodate the features of BLAS, which contains both memory-bound and computing-bound routines, we propose a hybrid strategy to incorporate fault tolerance into our brand-new BLAS implementation: duplicating computing instructions for memory-bound Level-1 and Level-2 BLAS routines and incorporating an Algorithm-Based Fault Tolerance mechanism for computing-bound Level-3 BLAS routines. Our high performance and low overhead are obtained from delicate assembly-level optimization and a kernel-fusion approach to the computing kernels. Experimental results demonstrate that FT-BLAS offers high reliability and high performance -- faster than Intel MKL, OpenBLAS, and BLIS by up to 3.50%, 22.14% and 21.70%, respectively, for routines spanning all three levels of BLAS we benchmarked, even under hundreds of errors injected per minute. Elisabeth Giem, Quan Fan, Kai Zhao 0008, Jinyang Liu 0003, Zizhong Chen |
ICS | 3 |
| 2021 | Locality-aware Thread Block Design in Single and Multi-GPU Graph ProcessingabstractGraphics Processing Unit (GPU) has been adopted to process graphs effectively. Recently, multi-GPU systems are also exploited for greater performance boost. To process graphs on multiple GPUs in parallel, input graphs should be partitioned into parts using partitioning schemes. The partitioning schemes can impact the communication overhead, locality of memory accesses, and further improve the overall performance. We found that both intra-GPU data sharing and inter-GPU communication can be summarized as inter-TB communication. Based on this key idea, we propose a new graph partitioning scheme by redefining the input graph as a TB Graph with calculated vertex and edge weights, and then partition it to reduce intra & inter-GPU communication overhead and improve the locality at the granularity of Thread Blocks (TB). We also propose to develop a partitioning and mapping scheme for heterogeneous architectures including physical links with different bandwidths. The experimental results on graph partitioning show that our scheme is effective to improve the overall performance of the Breadth First Search (BFS) by up to 33%. Quan Fan, Zizhong Chen |
NAS | 1 |
| 2021 | LocalityGuru: A PTX Analyzer for Extracting Thread Block-level Locality in GPGPUsabstractExploiting data locality in GPGPUs is critical for efficiently using the smaller data caches and handling the memory bottleneck problem. This paper proposes a thread block-centric locality analysis, which identifies the locality among the thread blocks (TBs) in terms of a number of common data references. In LocalityGuru, we seek to employ a detailed just-in-time (JIT) compilation analysis of the static memory accesses in the source code and derive the mapping between the threads and data indices at kernel-launch-time. Our locality analysis technique can be employed at multiple granularities such as threads, warps, and thread blocks in a GPU Kernel. This information can be leveraged to help make smarter decisions for locality-aware data-partition, memory page data placement, cache management, and scheduling in single-GPU and multi-GPU systems.The results of the LocalityGuru PTX analyzer are then validated by comparing with the Locality graph obtained through profiling. Since the entire analysis is carried out by the compiler before the kernel launch time, it does not introduce any timing overhead to the kernel execution time. Devashree Tripathy, AmirAli Abdolrashidi, Quan Fan, Daniel Wong 0001, Manoranjan Satpathy |
NAS | 3 |
| 2018 | Building Generic Scalable Middlebox Services Over Encrypted ProtocolsabstractThe trends of the increasing middleboxes make the middle network more and more complex. Today, many middleboxes work on application layer and offer significant network services by the plain-text traffic, such as firewalling, intrusion detecting and application layer gateways. At the same time, more and more network applications are encrypting their data transmission to protect security and privacy. It is becoming a critical task and hot topic to continue providing application-layer middlebox services in the encrypted Internet, however, the state of the art is far from being able to be deployed in the real network. In this paper, we propose a practical architecture, named PlainBox, to enable session key sharing between the communication client and the middleboxes in the network path. It employs Attribute-Based Encryption (ABE) in the key sharing protocol to support multiple chaining middleboxes efficiently and securely. We develop a prototype system and apply it to popular security protocols such as TLS and SSH. We have tested our prototype system in a lab testbed as well as real-world websites. Our result shows PlainBox introduces very little overhead and the performance is practically deployable. Cong Liu 0029, Yong Cui 0001, Kun Tan 0002, Quan Fan, Kui Ren 0001 |
INFOCOM | 4 |