VLDB 2026 Research / reviewers in the wild / expert
Mingzhe Hu
dblp:262/9460
· DBLP profile ↗
11ranked-venue papers
5as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 8 · 4 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RisConFix: LLM-Based Automated Repair of Risk-Prone Drone Configurations
Liping Han, Tingting Nie, Le Yu 0002, Mingzhe Hu, Tao Yue 0002 |
SANER | 4 |
| 2026 | NASchecker: Automatically Identifying the Performance, Security, and Privacy Issues of NAS DevicesabstractNetwork attached storage (NAS) devices are widely deployed for personal data storage. However, their distributed architecture and limited inspection interfaces pose significant challenges for comprehensive performance, security, and privacy analysis. In this paper, we first establish a threat model for NAS ecosystems. Then, we present a systematic framework NASchecker for discovering performance optimization mechanisms, security threats, and privacy leakage in NAS devices. By analyzing traffic generated during varied file operations on crafted files, NASchecker infers implemented optimizations and identifies security flaws within the traffic (e.g., susceptibility to passive sniffing and replay attacks). NASchecker also integrates NAS-specific protocol fuzzing and firmware reverse engineering to uncover deep-seated command injection, memory corruption, and improper access control vulnerabilities. NASchecker compares personally identifiable information (PII) leaked in traffic against declarations in privacy policies to detect privacy compliance issues. We evaluated NASchecker on twelve commercial NAS devices. Our results reveal that none of the tested devices employ file compression or deduplication. From a security standpoint, ten devices are vulnerable to passive sniffing and seven to replay attacks. Moreover, seven devices are affected by command injection, four by memory corruption, and eleven by improper access control. From a privacy perspective, four devices leaked PIIs that were not disclosed in their respective privacy policies. After reporting the findings to the manufacturers, we have been acknowledged by several manufacturers, resulting in the assignment of 20 CVEs and 6 NVDB entries (16 of them are rated as high severity). These findings validate NASchecker’s effectiveness and underscore the urgent need for improved design and testing practices of NAS. Guangyue Ren, Le Yu 0002, Liping Han, Mingzhe Hu, Wei Chen 0006, Tingting Liu 0005, Xiapu Luo, Guozi Sun |
IEEE Internet Things J. | 5 |
| 2026 | Clash: Enhancing context-sensitivity in data-flow analysis for mitigating the impact of indirect calls
Jinyan Xie, Yingzhou Zhang, Mingzhe Hu, Liping Han, Le Yu 0002, Qiuran Ding |
J. Syst. Softw. | 3 |
| 2025 | Wi-GPD Identification System Based on Gait Point DensityabstractA significant challenge currently facing Wi-Fi-based gait recognition technology is that changes in walking paths in a multipath environment can significantly interfere with the CSI gait signal collected via Wi-Fi, which greatly hinders the application of this technology in real life. To deal with this problem, most existing Wi-Fi gait recognition systems adopt the strategy of fixing walking paths or using multiple receivers, but these methods undoubtedly increase the complexity and cost of the system. This article proposes an innovative solution: an identification system independent of the walking path and requires only a pair of transceivers. The system is based on the identification of the IQ signal density characteristics. Specifically, the CSI signal is first decomposed and reconstructed using VMD technology, eliminating noise interference in a multipath environment. A unique point density feature is extracted from the IQ signal, which integrates both phase and amplitude information. This feature effectively distinguishes gait and is not influenced by changes in walking paths. At the same time, it can visually highlight the commonalities and differences when depicting the same person and distinguishing between different individuals' gaits, providing a strong basis for gait recognition classification. Finally, the deep learning model was introduced into the identification process, improving the system's accuracy. The experimental results show that in a dataset containing 5–20 testers, the Wi-GPD system achieves an identification accuracy of up to 84.75%–99.5%, thoroughly verifying its effectiveness and reliability. Ying Liu 0071, Zhiyang Cao, Jiaqi Cai, Mingzhe Hu |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2025 | A Survey of Multilanguage Interoperability and Its Program AnalysisabstractSince multilanguage programming has the strength of interoperating languages with different features and paradigms, and it also enables the reuse of existing libraries, developers often use multilanguage interoperability in software systems of different application domains. Program analysis is an effective way to maintain the reliability and safety of software. However, numerous challenges appear when using program analysis to analyze multilanguage interoperability. However, there still lacks a systematic overview of the multilanguage interoperability and the corresponding program analysis (e.g., what are the research trends, how existing works solve them). To bridge this gap, we conducted a comprehensive investigation of the 195 research works related to multilanguage interoperability and its program analysis in the past 38 years (1987–2024). In this article, we classify these works into three categories with 10 research perspectives, including foreign interface design, interface definition and generation, intermediate representation (IR), semantics, static analysis of memory management, type system, exception handling and concurrency, dynamic analysis, and others. Then, we evaluate research trends that affect multilanguage interoperability and key static/dynamic language futures of multilanguage program analysis. Finally, we discuss the open challenges and future research directions of multilanguage interoperability. Mingzhe Hu, Le Yu 0002, Yu Zhang 0086, Liping Han |
IEEE Trans. Reliab. | 1 |
| 2023 | CGORewritter: A better way to use C library in GabstractCGO is a foreign function interface mechanism that enables the creation of Go packages that call C code. It provides a way to reuse legacy C code or high-performance C libraries. However, it is tedious and error-prone to manually write the bindings of C libraries in Go. To make better use of CGO to realize the use of C libraries in Go, we propose CGORewritter in this paper. For a given Go library and corresponding C library, the internal code of the Go API function can be rewritten while keeping the Go API unchanged, and the C API function can be called through CGO in a semi-automatic way so that the application layer program can upgrade without any change. We use this method to rewrite the go/crypto with the OpenSSL library. Experimental results show that the rewritten code maintains full functionality and can be used by application code without modification. Besides, the rewritten code gains a 0.97-3.07× speedup. Boyao Ding, Yu Zhang 0086, Jinbao Chen, Mingzhe Hu, Qingwei Li |
SANER | 4 |
| 2023 | Cross-Language Call Graph Construction Supporting Different Host LanguagesabstractModern software systems are increasingly multi-lingual, which consist of components developed in different programming languages to reuse existing libraries and com-bine language features. Foreign function interface (FFI) is a mechanism that enables interoperation between a host language and a guest language. CFFI interoperating with external C is part of the language standard for almost all languages. For example, Python/C API is the CFFI between Python and C/C++. Python host with C guest can achieve both productivity and performance, and is widely used by many mainstream software systems in different application domains. The popularity of these software systems makes high demands on program analysis of multilingual codebases. A fundamental challenge is to construct call graphs that capture the connectivity between host and guest languages. In this work, we present a novel approach to call graph construction for calls from different host languages to C/C++ foreign functions. The semantics of the foreign function declaration interfaces are modeled to establish cross-language call relationships, taking into account the semantic abstraction to support different host languages. Graph transformation and function node fusions are defined to build the complete call graphs. We demonstrate experimentally that static call graphs for Python and JavaScript calling C/C++ can be constructed effectively and automatically. The call graph construction can further efficiently work as an IDE language server and integrate with other tools. Mingzhe Hu, Yu Zhang 0086, Yan Xiong 0001 |
SANER | 1 |
| 2023 | An empirical study of the Python/C API on evolution and bug patternsabstractAbstract Python is a popular programming language, and a large part of its appeal comes from diverse libraries and extension modules. In the bloom of data science and machine learning, Python frontend with C/C++ native implementation achieves both productivity and performance and has almost become the standard structure for many mainstream software systems. However, feature discrepancies between two languages such as exception handling, memory management, and type system can pose many safety hazards in the interface layer using the Python/C API. In this paper, we carry out an empirical study of the Python/C API on evolution and bug patterns. The evolution analysis includes Python/C API design in CPython compilers and its usage in mainstream software. By designing and applying a static analysis toolset, we reveal the evolution and usage statistics of the Python/C API and provide a summary and classification of 9 common bug patterns. In Pillow, a widely used Python imaging library, we find 48 bugs, 19 of which are undiscovered before. Our toolset can be easily extended to access different types of syntactic bug‐finding checkers, and our systematical taxonomy to classify bugs can guide the construction of more highly automated and high‐precision bug‐finding tools. Mingzhe Hu, Yu Zhang 0086 |
J. Softw. Evol. Process. | 1 |
| 2021 | Static Type Inference for Foreign Functions of PythonabstractStatic type inference is an effective way to maintain the safety of programs written in a dynamically typed language. However, foreign functions implemented in another programming language are often outside the inference range. Python, a popular dynamically typed language, has a lot of widely used packages which follow the multilingual structure with C/C++ extension modules. Existing deterministic Python static type inference tools which are not based on type annotations can do nothing about these foreign functions. In this paper, we propose a novel method to infer the type signature of foreign functions by analyzing implicit information in the layer of foreign function interface. We design a static type inference system, its evaluation on CPython, NumPy and Pillow shows that our method soundly infers the number and type of arguments for most foreign functions. Our results can further work as a complement to the state-of-the-art Python static type inference tool and enable it to analyze programs with foreign function calls. We catch 48 bugs of mismatch between foreign function declaration and its implementation, which make a parameter-free foreign function take argument of any type. 8 of the bugs we reported have been confirmed and fixed by communities. Mingzhe Hu, Yu Zhang 0086, Wenchao Huang 0001, Yan Xiong 0001 |
ISSRE | 1 |
| 2021 | An Empirical Study for Common Language Features Used in Python ProjectsabstractAs a dynamic programming language, Python is widely used in many fields. For developers, various language features affect programming experience. For researchers, they affect the difficulty of developing tasks such as bug finding and compilation optimization. Former research has shown that programs with Python dynamic features are more change-prone. However, we know little about the use and impact of Python language features in real-world Python projects. To resolve these issues, we systematically analyze Python language features and propose a tool named PYSCAN to automatically identify the use of 22 kinds of common Python language features in 6 categories in Python source code. We conduct an empirical study on 35 popular Python projects from eight application domains, covering over 4.3 million lines of code, to investigate the the usage of these language features in the project. We find that single inheritance, decorator, keyword argument, for loops and nested classes are top 5 used language features. Meanwhile different domains of projects may prefer some certain language features. For example, projects in DevOps use exception handling frequently. We also conduct in-depth manual analysis to dig extensive using patterns of frequently but differently used language features: exceptions, decorators and nested classes/functions. We find that developers care most about ImportError when handling exceptions. With the empirical results and in-depth analysis, we conclude with some suggestions and a discussion of implications for three groups of persons in Python community: Python designers, Python compiler designers and Python developers. Yu Zhang 0086, Mingzhe Hu |
SANER | 3 |
| 2020 | The Python/C API: Evolution, Usage Statistics, and Bug PatternsabstractPython has become one of the most popular programming languages in the era of data science and machine learning, especially for its diverse libraries and extension modules. Python front-end with C/C++ native implementation achieves both productivity and performance, almost becoming the standard structure for many mainstream software systems. However, feature discrepancies between two languages can pose many security hazards in the interface layer using the Python/C API. In this paper, we applied static analysis to reveal the evolution and usage statistics of the Python/C API, and provided a summary and classification of its 10 bug patterns with empirical bug instances from Pillow, a widely used Python imaging library. Our toolchain can be easily extended to access different types of syntactic bug-finding checkers. And our systematical taxonomy to classify bugs can guide the construction of more highly automated and high-precision bug-finding tools. Mingzhe Hu, Yu Zhang 0086 |
SANER | 1 |