Haobo Hua

dblp:378/3273 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0002-8015-6392ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Optimizing Standard Convolution for Diverse Precision on DCU
Haobo Hua, Chuangzheng Hou, Zhuxin Wen, Xiangkai Zhang, Jiandong Shang, Litao Zhang
CCF Trans. High Perform. Comput.1
2025 JOLT-SQL: Joint Loss Tuning of Text-to-SQL with Confusion-aware Noisy Schema Sampling
abstract
Text-to-SQL, which maps natural language to SQL queries, has benefited greatly from recent advances in Large Language Models (LLMs).While LLMs offer various paradigms for this task, including prompting and supervised fine-tuning (SFT), SFT approaches still face challenges such as complex multistage pipelines and poor robustness to noisy schema information.To address these limitations, we present JOLT-SQL, a streamlined single-stage SFT framework that jointly optimizes schema linking and SQL generation via a unified loss.JOLT-SQL employs discriminative schema linking, enhanced by local bidirectional attention, alongside a confusion-aware noisy schema sampling strategy with selective attention to improve robustness under noisy schema conditions.Experiments on the Spider and BIRD benchmarks demonstrate that JOLT-SQL achieves state-of-the-art execution accuracy among comparable-size open-source models, while significantly improving both training and inference efficiency.Our code is available at https://github.com/Songjw133/JOLT-SQL.
Jinwang Song, Hongying Zan, Kunli Zhang, Lingling Mu, Yingjie Han, Haobo Hua
EMNLP6
2025 Optimizing 2D convolution for DCUs
Wenlong Fan, Haobo Hua, Jiandong Shang, Zhuxin Wen, Hengliang Guo, Litao Zhang
CCF Trans. High Perform. Comput.2
2025 VBATS: an adaptive strategy for grouped GEMM on GPUs
Jiandong Shang, Zhuxin Wen, Haobo Hua, Hengliang Guo, Wenlong Fan, Guangsheng Qin
J. Supercomput.3
2024 A universal parallel simulation framework for energy pipeline networks on high-performance computers
abstract
Abstract Energy distribution networks represent crucial infrastructures for modern society, and various simulation tools have been widely used by energy suppliers to manage these intricate networks. However, simulation calculations include a large number of fluid control equations, and computational overhead limits the performance of simulation software. This paper proposes a universal parallel simulation framework for energy pipeline networks that takes advantages of data parallelism and computational independence between network elements. A non-pipe model of an energy supply network is optimized, and the input and output of the network model in the proposed framework are modified, which can reduce the development burden during the numerical computations of the pipeline network and weaken the computational correlation between different simulated components. In addition, independent computations can be performed concurrently through periodic data exchange procedures between component instances, improving the parallelism and efficiency of simulation computations. Further, a parallel water pipelines network simulation computing paradigm based on a heterogeneous computer hardware architecture is used to evaluate the proposed framework’s performance. A series of tests are conducted to verify the accuracy of the proposed framework, and simulation errors of less than 5% are achieved. The results of multi-threaded simulation experiments have demonstrated the feasibility of the proposed framework in a parallel computing approach. Moreover, an Advanced Micro Devices (AMD) Deep Computing Unit (DCU)-parallel program is implemented into a water supply network simulation system; the computational efficiency of this system is compared with that of its serial counterpart. The experimental results show that the proposed framework is appropriate for high-performance computer architectures, and the 18x speed-up ratio demonstrates that the parallel program based on the proposed universal framework outperforms the serial program. That provides the basis for the application of pipe network simulation on high-performance computers.
Pu Han, Haobo Hua, Changmao Wu, Jiandong Shang
J. Supercomput.2