EDBT 2026 Demo / reviewers in the wild / expert
Qiuming Luo
dblp:71/5289
· DBLP profile ↗
29ranked-venue papers
17as first author
15since 2021 · last 2026
0000-0003-3622-5386ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021Systems, architecture and hardware · 7 · 6 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MVE-Mamba: Multi-view Feature Enhanced Vision Mamba for Brain Tumor MRI Classification
Qiuming Luo |
ICIC (29) | 1 |
| 2026 | PRISM-Deblur: Parallel Routing of Implicit Spectral and Morphological Experts for High-Fidelity Deblurring in Non-uniform Dynamic Scenes
Qiuming Luo, Haigang Zhang, Chang Kong |
ICIC (10) | 1 |
| 2026 | Entropy-aware structural alignment for zero-shot Handwritten Chinese Character Recognition
Qiuming Luo, Heming Liu, Rui Mao 0001, Chang Kong |
Pattern Recognit. | 1 |
| 2025 | FPGA-Accelerated Error Diffusion Halftoning for High-Throughput Wide-Format Industrial Printing
Qiuming Luo, Yiwen Zhong, Chang Kong |
CGI (1) | 1 |
| 2025 | Radical Sequence Encoding with Fine-Tuned CLIP for Handwritten Chinese Character Recognition
Qiuming Luo, Chang Kong |
ICDAR (3) | 1 |
| 2025 | Challenges and Benchmarking of Object Detection Models on Edge AI SoCs
Chang Kong, Peng Mo, Qiuming Luo, Rui Mao 0001 |
ICIC (14) | 4 |
| 2025 | SoCDev: SoC Design Automation through TCL Code Generation via LLM-Based AgentsabstractThe design of System-on-Chip (SoC) is a complex task that requires considerable effort from designers. Recent advancements in Language Model (LLM)-based agents have brought opportunities to tackle this challenge. In this paper, we introduce SocDev, the first SoC design multi-agent system based on LLM, which can automatically generate SoC designs by creating TCL command scripts for the Vivado Design Suite. Additionally, we have developed SoCDevBench, a test bench to evaluate the capability of different models in generating SoC designs. We conducted experiments using four state-of-the-art LLMs with the test bench, assessing the pass rate, FPGA resource utilization, and logs throughout the process. The ablation study also verify the efficiency of each component. Furthermore, we successfully produced four functional SoCs with microprocessor and operating systems, and validated them using FPGAs. This study opens a new avenue for the generation of SoC designs leveraging LLMs. LinLin Shen, Qiuming Luo |
IJCNN | 4 |
| 2024 | Optimizing the DFCPP Dataflow Runtime Library for Resource Utilization in NUMA SystemsabstractIn parallel computing, Data Flow Graphs play a critical role by explicitly representing the dependencies between tasks, which is essential for task scheduling and resource utilization. In the process of scheduling optimization, it is crucial to address the dual challenges of balancing computational load and distributing data access pressure. The Dataflow for C++ (DFCPP) possesses the advantage of accurately sensing task data sizes, and based on this capability, this paper proposes a Primary-Secondary Core Selecting (PSCS) strategy specifically designed for Hyper-Threading-enabled Non-Uniform Memory Access systems (HT-NUMA). This strategy aims to mitigate interference between threads on the same physical core in a hyper-threading environment, thereby enhancing task execution stability and overall performance. Additionally, DFCPP integrates a task-stealing mechanism with a proactive allocation strategy for large-cache tasks, enabling the distribution of large tasks to low-load NUMA nodes, which reduces cache contention and improves cache hit rates. Experimental results demonstrate that, compared to common dataflow programming libraries such as Taskflow and TBB, DFCPP achieves over a 20% performance improvement when the system core count meets computational demands, and nearly a 25% improvement when handling large-dataset tasks. Furthermore, DFCPP exhibits high efficiency and adaptability in processing tasks with various directed acyclic graph topologies, fully leveraging the parallel computing advantages of NUMA systems, and demonstrates exceptional performance and broad applicability in dataflow task scheduling, resource optimization, and complex dependency management. Qiuming Luo, Zheng Du |
HPCC | 1 |
| 2024 | Leveraging Computer Vision for Automatic Modulation Classification: Insights from Spectrum and Constellation Diagram Analysis
Qiuming Luo |
ICPR (24) | 2 |
| 2024 | Enhancing Code Generation for Dataflow Programming: Fine-Tuning Large Language Models with the DFCPP DatasetabstractIn recent years, large language models (LLMs) based on the Transformer architecture have demonstrated excellent performance in code generation, but there have been fewer studies on data flow languages. This study proposes a scheme for fine-tuning large language models based on the DFCPP dataset. We demonstrate the model's ability to generate dataflow graph (DAG) topologies and achieve significant performance improvements. Experimental results show that the BLEU score of the fine-tuned model in the DFCPP code generation task reaches 0.193, which is an increase of 112.1% compared to the non-fine-tuned model (0.091). This demonstrates the effectiveness of fine-tuning techniques in domain-specific code generation. Qiuming Luo, Xi Ma |
ISPA | 1 |
| 2024 | M-DFCPP: A runtime library for multi-machine dataflow computingabstractSummary This article designs and implements a runtime library for general dataflow programming, DFCPP (Luo Q, Huang J, Li J, Du Z. Proceedings of the 52nd International Conference on Parallel Processing Workshops. ACM; 2023:145‐152.), and builds upon it to design and implement a multi‐machine C++ dataflow library, M‐DFCPP. In comparison to existing dataflow programming environments, DFCPP features a user‐friendly interface and richer expressive capabilities (Luo Q, Huang J, Li J, Du Z. Proceedings of the 52nd International Conference on Parallel Processing Workshops. ACM; 2023:145‐152.), enabling the representation of various types of dataflow actor tasks (static, dynamic and conditional task). Besides that, DFCPP addresses the memory management and task scheduling for non‐uniform memory access architectures, while other dataflow libraries lack attention to these issues. M‐DFCPP extends the capability of current dataflow runtime libraries (DFCPP, taskflow, openstream, etc.) and capable of multi‐machine computing, while maintains the API compatible with DFCPP. M‐DFCPP adopts the concepts of master and follower (Dean J, Ghemawat S. Commun ACM. 2008;51(1):107‐113; Ghemawat S, Gobioff H, Leung ST. ACM SIGOPS Operating Systems Review. ACM; 2003:29‐43.), which form a worksharing framework as many multi‐machine system. To shift to the M‐DFCPP framework, a filtering layer is inserted to the original DFCPP, transforming it into followers that can cooperate with each other. The master is made of modules for scheduling, data processing, graph partition, state management and so forth. In benchmark tests with workload with directed acyclic graph topology of binary trees and random graphs, DFCPP demonstrated performance improvements of 20% and 8%, respectively, compared to the second fastest library. M‐DFCPP consistently exhibits outstanding performance across varying levels of concurrency and task workloads, achieving a maximum speedup of more than 20 over DFCPP, when the task parallelism exceeds 5000 on 32 nodes. Moreover, M‐DFCPP, as a runtime library supporting multi‐node dataflow computation, is compared with MPI, a runtime library supporting multi‐node control flow computation. Qiuming Luo, Senhong Liu, Jinke Huang |
Concurr. Comput. Pract. Exp. | 1 |
| 2023 | Deep Multi-Input Multi-Stream Ordinal Model for age estimation: Based on spatial attention learning
Chang Kong, Qiuming Luo, Rui Mao 0001, Guoliang Chen 0005 |
Future Gener. Comput. Syst. | 3 |
| 2022 | Learning Deep Contrastive Network for Facial Age EstimationabstractAge estimation from a single facial image is an attractive and challenging research topic in the computer vision community. Most of previous works conventionally estimate absolute age from the input face images. However, telling someone's precise age at a glance without any reference information is essentially difficult even for humans. In this paper, we propose a novel Deep Contrastive Network (DCN) for age estimation, which can mine the variation information of samples. DCN model is trained end-to-end to learn the age distances between input unseen face and several reference images by contrasting deep feature maps. In the test phase, the input image is compared with a set of selected references to determine how many years younger or older than each of baselines which is more conform with human cognitive processes. We also propose a Cyclic Iterative Approximation Algorithm (CIAA) for post age voting which can further improve the accuracy. Therefore, the age estimation problem is cast as a metric learning task in our DCN model. By jointly learned with the cost sensitive loss and KL divergence loss, our DCN is easy to train and has very stable convergence. Extensive experiments show that the proposed approach significantly outperforms other state-of-the-art age estimation methods on MORPH II and FG-NET datasets. Chang Kong, Qiuming Luo, Guoliang Chen 0005 |
IJCNN | 2 |
| 2021 | RSFAD: A Large-Scale Real Scenario Face Age Dataset in the wildabstractAge estimation is a hot and challenging research topic in the computer vision community. Several facial datasets annotated with age and gender attributes became available in recent years. However, the statistical information of these datasets reveal the unbalanced label distribution which inevitably introduce bias during model training. In this work, we manually collect and label a large-scale age dataset called Real Scenario Face Age Dataset (RSFAD) which contains 85,044 facial images captured from surveillance cameras in the wild. Due to the COVID-19, we not only label the apparent age group and gender but also label the breathing mask, and the label distribution of RSFAD dataset is almost uniform which is the first age dataset to the best of our knowledge. In addition, we investigate the impact of age, gender and mask distribution on age group estimation by comparing GDEX CNN model trained on several different datasets. Our experiments show that the RSFAD dataset has good performance for age estimation task and also it is suitable for being an evaluation benchmark. Chang Kong, Qiuming Luo, Guoliang Chen 0005 |
FG | 2 |
| 2021 | A comparison study: the impact of age and gender distribution on age estimationabstractAge estimation from a single facial image is a challenging and attractive research area in the computer vision community. Several facial datasets annotated with age and gender attributes became available in the literature. However, one major drawback is that these datasets do not consider the label distribution during data collection. Therefore, the models training on these datasets inevitably have bias for the age having least number of images. In this work, we analyze the age and gender distribution of previous datasets and publish an Uniform Age and Gender Dataset (UAGD) which has almost equal number of female and male images in each age. In addition, we investigate the impact of age and gender distribution on age estimation by comparing DEX CNN model trained on several different datasets. Our experiments show that UAGD dataset has good performance for age estimation task and also it is suitable for being an evaluation benchmark. Chang Kong, Qiuming Luo, Guoliang Chen 0005 |
MMAsia | 2 |
| 2020 | The Compiler of DFC: A Source Code Converter that Transform the Dataflow Code to the Multi-threaded C Code
Zheng Du, Haixin Du, Jiwu Shu, Qiuming Luo |
PDCAT | 6 |
| 2020 | The Dataflow Runtime Environment of DFC
Zheng Du, Jiwu Shu, Qiuming Luo |
PDCAT | 5 |
| 2019 | Implementing the Matrix Multiplication with DFC on Kunlun Small Scale ComputerabstractIn this paper, we demonstrate a new dataflow platform of DFC, which can handle the successive dataflow computing passes with tagged data. By implementing the matrix multiplication in DFC, we show that DFC can exploit the parallelism automatically with a much simple dataflow graph constructed by DF functions of DFC. Different from the other dataflow execution platform, DFC support multiple worker threads for one dataflow node of DF functions. By running the matrix multiplication program of DFC on Kunlun system, it was verified that DFC get a reasonable speedup for large scale computing for thread number up to 512. Zheng Du, Shihao Sha, Qiuming Luo |
PDCAT | 4 |
| 2019 | The Library for Hadoop Deflate Compression Based on FPGA Accelerator with Load BalanceabstractHadoop application will produce lots of intermediate results in the map/reduce process that requires disk I/O and network transmission. By compressing the large-scale data of intermediate result, it will greatly improve disk access efficiently and reduce program run time. Hardware-accelerated solutions have become more desirable. This paper design a multi-FPGA compression accelerator on the Hadoop platform, and the system performance analysis compared with a software-only solution that mainly uses CPU to processing. The testing programs are zpipe, TestDFSIO and Terasort. In contrast with the software-only solution. The max speedup of zpipe is 6.55X (single FPGA) and 10.24X (dual FPGA), the max speedup of TestDFSIO is 6.28X (single FPGA) and 6.28X (dual FPGA), and the max speedup of Terasort application is up to 3.25X(single FPGA) and 3.35X(dual FPGA). Haixin Du, Jiankui Zhang, Shihao Sha, Cai Ye, Qiuming Luo |
PDCAT | 5 |
| 2019 | FPGA-Based Parallel Multi-Core GZIP Compressor in HDFSabstractWith the development of Big Data, data storage has been exposed to more challenges. Data compression which can save both storage and network bandwidth, is a very important technology to deal with the challenges. In this paper, we present an end-to-end, complete, high-throughput parallel multi-core GZIP compressor in FPGA for HDFS. The GZIP compressor is designed by the scalable architecture, which supports to increase throughput by expanding multiple compression cores based on systolic array architecture. We implemented and evaluated the hardware compressor in Alpha Data Adm-Pcie-KU3 FPGA board, utilizing RIFFA for data transfers over PCI Express. According to the evaluation results, up to 16-cores compressor can be implemented and the peak compression throughput exceeds 1.1 GB/s. It is 70X speedup compared with the software compression solution. When we load the hardware compressor into HDFS, the performance of HDFS is twice as much as that without loading the compressor. Haoxin Luo, Ye Cai 0001, Qiuming Luo, Rui Mao 0001 |
PDCAT | 3 |
| 2018 | Algorithms designed for compressed-gene-data transformation among gene banks with different referencesabstractBACKGROUND: With the reduction of gene sequencing cost and demand for emerging technologies such as precision medical treatment and deep learning in genome, it is an era of gene data outbreaks today. How to store, transmit and analyze these data has become a hotspot in the current research. Now the compression algorithm based on reference is widely used due to its high compression ratio. There exists a big problem that the data from different gene banks can't merge directly and share information efficiently, because these data are usually compressed with different references. The traditional workflow is decompression-and-recompression, which is too simple and time-consuming. We should improve it and speed it up. RESULTS: In this paper, we focus on this problem and propose a set of transformation algorithms to cope with it. We will 1) analyze some different compression algorithms to find the similarities and the differences among all of them, 2) come up with a naïve method named TDM for data transformation between difference gene banks and finally 3) optimize former method TDM and propose the method named TPI and the method named TGI. A number of experiment result proved that the three algorithms we proposed are an order of magnitude faster than traditional decompression-and-recompression workflow. CONCLUSIONS: Firstly, the three algorithms we proposed all have good performance in terms of time. Secondly, they have their own different advantages faced with different dataset or situations. TDM and TPI are more suitable for small-scale gene data transformation, while TGI is more suitable for large-scale gene data transformation. Qiuming Luo, Chao Guo 0004, Ye Cai 0001, Gang Liu 0028 |
BMC Bioinform. | 1 |
| 2014 | Optimization of Uncore Data Flow on NUMA Platform
Qiuming Luo, Chang Kong, Ye Cai 0001 |
NPC | 1 |
| 2014 | Understanding the Data Traffic of Uncore in Westmere NUMA ArchitectureabstractNon-Uniform Memory Access (NUMA) has become the main stream architecture of modern servers. In processors, Uncore part plays a very important role, especially in NUMA systems, because it is used to connect Cores, Last Level Caches (LLC), on-chip multiple Memory Controllers (MCs) and highspeed interconnections. Recent study shows that Uncore congestion plays a more important role than locality. It needs more understanding of Uncore behavior to alleviate the congestion and efficiently utilize certain architecture. Our work focuses on the unbalance and congestion of data traffic happened on processor's Uncore part. We choose an Intel NUMA architecture named "Westmere" and use hardware performance counters to investigate several benchmarks' data flow in Uncore. In our experiments we find that data unbalance of Global Queue (GQ) and QuickPath Home Logical (QHL)'s trackers is really serious, the biggest unbalance rate is more than 1000 times, new dynamic entries management algorithm is needed to improve entries' usage the congestion of GQ and QHL's trackers has different behaviors with threads number increases and also for a given memory access pattern the congestion of GQ and QHL's trackers grows linearly with the problem size increases. Qiuming Luo, Chang Kong, Chengjian Liu |
PDP | 1 |
| 2013 | Analyzing the Characteristics of Memory Subsystem on Two Different 8-Way NUMA Architectures
Qiuming Luo, Chang Kong, Ye Cai 0001, Xiaohui Lin 0001 |
NPC | 1 |
| 2012 | MAP-numa: Access Patterns Used to Characterize the NUMA Memory Access Optimization Techniques and Algorithms
Qiuming Luo, Chengjian Liu, Chang Kong, Ye Cai 0001 |
NPC | 1 |
| 2012 | Quantitatively Measuring the Memory Locality Leakage on NUMA Systems Based on Instruction-Based-SamplingabstractSustaining the memory locality is critical for obtaining high performance in NUMA system. But how to identify a locality leakage problem and how to measure the leakage is still open issue. This paper provides an algorithm to quantitatively measure the locality leakage based on the memory trace produced by IBS (Instruction-Based-Sampling). A """"perfect matrix"""" PM is generated from virtual memory address trace, which represents the highest locality pattern. A """"communication matrix"""" CM is obtained from physical memory address trace to describe the actual memory access pattern. The penalty factors are calculated from PM or CM with considering of the hardware NUMA factor. The leakage is measured by the difference between the penalty factors of PM and the penalty factors of CM, which can be used to estimate the performance decrease and guide the optimization. Some experiment results are show to testify the effectiveness and accuracy of our quantitative measurement. Qiuming Luo, Chengjian Liu, Chang Kong, Ye Cai 0001 |
PDCAT | 1 |
| 2011 | Performance Evaluation of OpenMP Constructs and Kernel Benchmarks on a Loongson-3A Quad-Core SMP SystemabstractAs a competitor and alternative to mainstream general-purpose CPU (Intel/AMD/etc.), Loongson is a family of general-purpose MIPS-compatible CPUs developed at the ICT of CAS in China. The quad-core Loongson 3A is evaluated in this paper. The performance of the basic OpenMP constructs on Loongson-3A quad-core SMP is obtained by applying the EPCC Micro benchmarks. And then the performance of NAS kernel codes is obtained by applying NAS Parallel Benchmarks (NPB). These benchmarking are carried out for three different OpenMP compilers (and the runtime system), which includes GCC, OMPipth (OMPi with pthread library) and OMPi-psth (OMPi with psthread library). The results show that OMPI-pth's performance is the best and OMPi-psth's performance is the worst. Those test results might help to program the OpenMP codes as well as to select the appropriate compiler and its runtime system. And an Intel core i5 quad-core platform is used for comparison purpose, by running NPB, which implies that Loongson 3A's performance is nearly one tenth of i5's. The NPB results can help to defining a Loongson system's scale when replacing an Intel i5 system for a given problem size. Qiuming Luo, Chang Kong, Ye Cai 0001, Gang Liu 0028 |
PDCAT | 1 |
| 2010 | A Novel Model and a Simulation Tool for Churn of P2P NetworkabstractThe prior studies setup the churn model by measuring the historical logs or records of a P2P network, and treat it as one whole black-box without understanding the inside of peer's population. The metrics used to characterize the churn is distributions of the node session lengths and arrival intervals. We investigate churn in a higher level point of view, and find that modeling it based on the global geographical distribution of peer nodes will result in a system which explain the fluctuation and cyclic phenomenon of network size. This model considers the user behavior pattern into account. Then we provide a Matlab tools that can provide churn events according to this model. From the output events of simulation, we do see some more future things than other models. We might expect or predict when and what nodes would return back, as well as when and what nodes would disappear at high possibility. So it is useful when designing a system optimized both to the pass and the future, which could reduce the overhead of the maintenance of underlying overlay network of DHTs and lower the redundant level of replications for P2P storage system. Qiuming Luo, Gang Liu 0028, Rui Mao 0001 |
PDCAT | 1 |
| 2003 | Stereo matching and occlusion detection with integrity and illusion sensitivity
Qiuming Luo, Jingli Zhou, Shengsheng Yu, Degui Xiao |
Pattern Recognit. Lett. | 1 |