Xiangbing Huang

dblp:305/2977 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
0009-0005-7761-9914ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Unseen Horizons: Unveiling the Real Capability of LLM Code Generation Beyond the Familiar
abstract
Recently, large language models (LLMs) have shown strong potential in code generation tasks. However, there are still gaps before they can be fully applied in actual software development processes. Accurately assessing the code generation capabilities of large language models has become an important basis for evaluating and improving the models. Some existing works have constructed datasets to evaluate the capabilities of these models. However, the current evaluation process may encounter the illusion of “Specialist in Familiarity”, primarily due to three gaps: the exposure of target code, case timeliness, and dependency availability. The fundamental reason for these gaps is that the code in current datasets may have been extensively exposed and exercised during the training phase, and due to the continuous training and development of LLM, their timeliness has been severely compromised. The key to solve the problem is to, as much as possible, evaluate the LLMs using code that they have not encountered before. Thus, the fundamental idea in this paper is to draw on the concept of code obfuscation, changing code at different levels while ensuring the functionality and output. To this end, we build a code-obfuscation based benchmark OBFusEvAL. We first collect 1,354 raw cases from five real-world projects, including function description and code. Then we use three-level strategy (symbol, structure and semantic) to obfuscate descriptions, code and context dependencies. We evaluate four LLMs on Obfu-sevaland compared the effectiveness of different obfuscation strategy. We use official test suites of these projects to evaluate the generated code. The results show that after obfuscation, the average decrease ratio of test pass rate can up to 62.5%.
Yuanliang Zhang, Shanshan Li 0001, Zhouyang Jia, Xiangbing Huang, Chaopeng Luo, Zhizheng Zheng, Rulin Xu, Si Zheng 0003, Xiangke Liao
ICSE7
2024 LatVision: Modeling and Predicting Persisting Tail Latency in SSDs
abstract
As Solid State Drives (SSDs) continue to evolve, the presence of tail latency within these devices remains a significant issue that can adversely affect overall performance. Various factors contribute to the emergence of tail latency spikes in SSDs. Current software-level management solutions primarily focus on the performance prediction of individual I/O operations, recognizing that persistent slow operations are prevalent in SSDs and tend to have a more pronounced impact. In this paper, we build a tool-LatVision to obtain I/O-related data directly from the kernel to predict persisting tail latency in SSDs by a neural network model. We conduct a comprehensive comparison and analysis of the input metrics and predictive models employed. Furthermore, we enhance LatVision’s performance through the application of heuristic algorithms. Through LatVision, we achieve real-time, lightweight, and high-accuracy performance prediction for low-latency SSDs.
Linxiao Bai, Zhijie Jiang, Yuanliang Zhang, Xiangbing Huang, Wang Li 0003, Bin Lin 0011
HPCC5
2023 Towards Better Multilingual Code Search through Cross-Lingual Contrastive Learning
abstract
Recent advances in deep learning have significantly improved the understanding of source code by leveraging large amounts of open-source software data. Thanks to the larger amount of data, code representation models trained with multilingual datasets show superior performance to monolingual models and attract much more attention. However, the entangled source code from various programming languages makes multilingual models hard to differentiate language-specific textual semantics or syntactic structures, which significantly increases the difficulty of model learning from multilingual datasets directly. On the other hand, for a given problem, developers are likely to choose similar identifiers, even if coding in different languages. However, the presence of similar identifiers in multilingual code snippets does not mean that they implement the same functionality, which may misdirect models to overemphasize these unreliable signals and ignore the semantic information of multilingual code. To tackle the above issues, we propose LAMCode, a language-aware multilingual code understanding model. Specifically, we propose a simple yet effective method to perceive linguistic information by injecting language-specific viewer into the language models. Furthermore, we introduce a cross-lingual contrastive learning method by generating more similar training instances but with fewer overlapping features. This method prevents the models from over-relying on similar identifiers across languages. We conduct extensive experiments to evaluate the effectiveness of our approach on a large-scale multilingual dataset. The experimental results show that our approach significantly outperforms the state-of-the-art methods.
Xiangbing Huang, Yingwei Ma, Haifang Zhou, Zhijie Jiang, Yuanliang Zhang, Teng Wang 0004, Shanshan Li 0001
Internetware1
2022 SEED: Semantic Graph Based Deep Detection for Type-4 Clone
Zhipeng Xue 0002, Zhijie Jiang, Chenlin Huang, Rulin Xu, Xiangbing Huang, Liumin Hu
ICSR5