EDBT 2026 Demo / reviewers in the wild / expert
Yonghua Hu
dblp:115/5897
· DBLP profile ↗
13ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Parallel implementation and optimization of LMS adaptive filtering algorithms based on vector DSP
Yonghua Hu, Linyun Deng, Zhezhuo Zhao |
CCF Trans. High Perform. Comput. | 1 |
| 2026 | Research on SIMD Instruction Sequence Generation Method for Vector DSP ProcessorabstractABSTRACT In the field of digital signal processing (DSP), the execution of vector operations depends on the optimization of Single Instruction Multiple Data (SIMD) technology. However, manual SIMD vectorization method is complex to develop, poorly portable, and costly to maintain. Therefore, we propose a method of SIMD instruction sequence generation based on LLVM. This method builds a hierarchical instruction generation framework, combines the characteristics of the target architecture, and uses LLVM automatic vectorization tool to gradually convert the vectorized intermediate representation into the target architecture instruction sequence containing SIMD instructions. Experiments on FT‐M7002 hardware platform show that, compared with the vectorization method of manually calling SIMD built‐in functions, the average execution performance of the instruction sequence generated by this method can be improved by up to 70%. Yonghua Hu, Fangjun Liu, Huifu Zhang, Ju Huang |
Concurr. Comput. Pract. Exp. | 1 |
| 2026 | GAS: A scheduling primitive dependency analysis-based cost model for tensor program optimizationabstractAutomatically generating high-performance tensor programs has become a promising approach for deploying deep neural networks. A key challenge lies in designing an effective cost model to navigate the vast scheduling search space. Existing approaches typically fall into two categories, each with limitations: offline learning cost models rely on large pre-collected datasets, which may be incomplete or device-specific, and online learning cost models depend on handcrafted features, requiring substantial manual effort and expertise. We propose GAS, a lightweight framework for generating tensor programs for deep learning applications. GAS reformulates feature extraction as a sequence-dependent analysis of scheduling primitives. Our cost model integrates three key factors to uncover performance-critical insights within scheduling sequences: (1) decision factors allocation, quantifying entropy and skewness of scheduling primitive factors to capture their dominance; (2) primitive contribution weights, measuring the relative impact of primitives on overall performance; and (3) structural semantic alignment, capturing correlations between scheduling primitive factors and hardware parallelism mechanisms. This approach reduces the complexity of handcrafted feature engineering and extensive pre-training datasets, significantly improving both efficiency and scalability. Experimental results on NVIDIA GPUs demonstrate that GAS achieves average speedups of 3.79 × over AMOS and 2.22 × over Ansor, while also consistently outperforming other state-of-the-art tensor compilers. Yonghua Hu, Anxing Xie, Zenghua Cheng, Junyang Tang |
J. Syst. Archit. | 1 |
| 2025 | GTA: Generating high-performance tensorized program with dual-task scheduling
Anxing Xie, Yonghua Hu, Yuxiang Gao, Zenghua Cheng |
J. Syst. Archit. | 2 |
| 2025 | Instruction selection optimization for VLIW architecture based on classification node merging
Fangjun Liu, Huifu Zhang, Yonghua Hu, Anxing Xie, Shangfeng Mo |
J. Supercomput. | 3 |
| 2024 | SORA: Rapid Software Pipelining Optimization after Register Allocation for Vector-DSPsabstractHeterogeneous multi-zone processors are key building blocks of high performance computing (HPC), and fully utilizing its hardware resources is essential to unlocking its full power. Software pipelining is a family of techniques that enhance program execution efficiency by accelerating the execution speed of loop programs through parallel execution of instructions from different loop bodies. However, existing software pipelining methods fail to account for the unique hardware architectures of digital signal processors (DSPs), resulting in suboptimal performance when applied directly.In this paper, we present SORA, a rapid Software pipelining Optimization architecture performing efficient instruction scheduling after Register Allocation for the Vector-DSPs. In addition, SORA conducts a comprehensive analysis encompassing instruction type, data flow, functional unit utilization, and register conflict, tailored to accommodate Vector-DSPs. Experimental results demonstrate that SORA can achieve 2.78∼24.1× speedup in the batch normalization, 2D convolution, relu and matrix multiplication, and 1.14× in the ResNet-18 neural network to the original program. Additionally, there is a 1.05∼6.59× improvement compared to loop unrolling optimization. Anxing Xie, Yonghua Hu |
ISPA | 2 |
| 2023 | Research on global register allocation for code containing array-unit dual-usage register namesabstractSummary Array‐unit dual‐usage register is a kind of register resource that can be read or written as a whole or individually. It is mainly configured in processors with SIMD processing units and provides register‐level speed data transfer between the scalar and vector processing units. To improve the efficiency of algorithms by using an array‐unit dual‐usage register, we investigate in this article the problem of adapting register allocation to code containing array‐unit dual‐usage register names. We propose a corresponding global register allocation method by combining the allocation of regular registers with array‐unit dual‐usage register, ensuring that the names of array‐unit dual‐usage register can be used in the input code of register allocation. Moreover, we present the processing framework of this method and the specific algorithms of some related vital aspects and demonstrate the working principles of the algorithms by an example. Experimental studies were conducted on a platform based on the FT‐M7002 DSP core, and showing that our register allocation method can effectively handle codes containing array‐unit dual‐usage register names and support relevant application algorithms to improve their data transfer scheme. For some typical algorithms with input matrix, substantial performance improvements of twofold or higher are achieved. Yonghua Hu, Xin Zhang 0141, Shuying Wang, Wei Liang 0005, Kuanching Li |
Concurr. Comput. Pract. Exp. | 1 |
| 2023 | Pulmonary Nodule Detection from 3D CT Image with a Two-Stage NetworkabstractEarly detection of lung nodules is an important means of reducing the lung cancer mortality rate. In this paper, we propose a three‐dimensional CT image lung nodule detection method based on parallel pooling and dense blocks, which includes two parts, i.e., candidate nodule extraction and false positive suppression. First, a dense U‐shaped backbone network with parallel pooling is proposed to obtain the candidate nodule probability map. The parallel pooling structure uses multiple pooling operations for downsampling to capture spatial information comprehensively and address the problem of information loss resulting from maximum and average pooling in the shallow layers. Then, a parasitic network with parallel pooling, dense blocks, and attention modules is designed to suppress false positive nodules. The parasitic network takes the multiscale feature maps of the backbone network as the input. The experimental results demonstrate that the proposed method significantly improves the accuracy of lung nodule detection, achieving a CPM score of 0.91, which outperforms many existing methods. Miao Liao, Zhiwei Chi, Huizhu Wu, Shuanhu Di, Yonghua Hu, Yunyi Li |
Int. J. Intell. Syst. | 5 |
| 2023 | PDPChain: A Consortium Blockchain-Based Privacy Protection Scheme for Personal DataabstractWith the advances and innovations in digital technologies, blockchain has empowered advancements in communications and networking, promising to build trust and establish secure decentralized communications networks. Unfortunately, current personal data privacy protection schemes still suffer from explicit storage, lack of data ownership and implementation of fine-grained access control by users, and lack of transparency and auditability of data. In this article, we propose a personal data privacy protection scheme based on consortium blockchain that stores original data encrypted with an improved Paillier homomorphic encryption mechanism, namely PDPChain, where users realize fine-grained access control based on ciphertext policy attribute-based encryption (CP-ABE) on blockchain. In this scheme, consortium blockchain combines distributed private clusters to store the encrypted data, improving data transmission efficiency, and guaranteeing user privacy and security through off-chain storage and on-chain transmission synergy. In addition, it is more lightweight encryption and demarcation, ultimately protecting personal data privacy and providing a secure and trusted way to obtain information for data mining. For the performance testing, data in the form of files are used as an example, and the scheme is designed and simulated on Hyperledger Fabric and InterPlanetary File System. Experimental results show that the improved Paillier encryption mechanism reduces the overall encryption and decryption elapsed time by 25% and encryption elapsed time by 48%. Furthermore, the proposed CP-ABE access control method is adaptive to storing and sharing a massive amount of data. With the increase in the number of access control policies, the overall time-consuming of the scheme does not increase, and the time-consuming of decryption can also be stabilized at about 2 s. Wei Liang 0005, Yang Yang 0197, Ce Yang 0007, Yonghua Hu, Songyou Xie, Kuanching Li, Jiannong Cao 0001 |
IEEE Trans. Reliab. | 4 |
| 2019 | Predicting Length of ICU Stay via Random Forest
Yonghua Hu, Guilan Kong |
AMIA | 4 |
| 2016 | Investigation on the Optimization for Storage Space in Register-Spilling
Yonghua Hu, Yaqiong Qiu, Wenti Huang |
CollaborateCom | 2 |
| 2016 | Belief rule-based inference for predicting trauma outcome
Guilan Kong, Dong-Ling Xu, Jian-Bo Yang, Xiaofeng Yin, Tianbing Wang, Baoguo Jiang, Yonghua Hu |
Knowl. Based Syst. | 7 |
| 2009 | Dependency Grammar Based English Subject-Verb Agreement Evaluation
Dongfeng Cai, Yonghua Hu, Xuelei Miao |
PACLIC | 2 |