VLDB 2026 Research / reviewers in the wild / expert
Huizhen Jiang
dblp:144/2871
· DBLP profile ↗
4ranked-venue papers
0as first author
3since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 3 · 3 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Automatically Deriving Developers' Technical Expertise from the GitHub Social NetworkabstractDevelopers’ technical expertise is crucial for numerous tasks within open-source communities, such as identifying suitable developers and maintainers. Despite its significance, GitHub, the world’s largest open-source code hosting platform, does not explicitly display developers’ technical expertise. Existing methods fall short in capturing the multi-faceted and dynamic nature of developers’ skills and knowledge. To address this gap, we propose a novel approach that leverages graph neural networks (GNNs) to express developers’ technical expertise. Our method constructs a comprehensive GitHub social network that integrates various social and development activities. We then employ a GNN model to learn a low-dimensional representation vector for each developer, encapsulating their technical expertise across different dimensions. We assess the effectiveness of our model by comparing it against five baselines on three GitHub social relationship recommendation tasks, including SimDeveloper, ContributionRepo, and RepoMaintainer. Our proposed method outperforms these baselines, achieving improvements of 5.6–9.5% on Hit Ratio@10 and 3.4–11.1% on F1 score. These results demonstrate promising performance in predicting technical preferences for both repositories and developers. This research contributes to a more nuanced understanding of developer expertise in open-source communities and has potential implications for improving collaboration and project management on platforms like GitHub. Yanchun Sun, Xiaohan Zhao, Haizhou Xu, Ye Zhu 0002, Zhenpeng Chen 0001, Huizhen Jiang, Gang Huang 0001 |
ACM Trans. Softw. Eng. Methodol. | 8 |
| 2025 | Leveraging BERT and Large Language Models for Mapping Heterogeneous Scientific and Technological Resources to Their IdentifiersabstractCurrently, there are various scientific and technological resource retrieval databases in the world. The resources stored in these databases may be identified by different identification systems. How to determine whether scientific and technological resources identified by different identification systems are the same resource is an urgent problem to be solved. This paper proposes a software service that leverages BERT and large language models to perform semantic analysis and similarity matching of scientific and technological resource content, and then maps the resources to their respective identifiers. The service effectively solves the problem of how to quickly retrieve the same resource from a large number of scientific and technological resources with diverse identification types, and improves the efficiency and quality of the retrieval. Implemented as a Chrome plugin, the service facilitates seamless mapping heterogeneous scientific and technological resources to their identifiers. We conduct a series of experiments. Their results demonstrate the effectiveness, scalability and stability of the service. To the best of our knowledge, we are the first to propose the service integrating BERT with large language models to extract and identify the key content of scientific and technological resources from web pages. Yanchun Sun, Xiaohan Zhao, Huizhen Jiang, Huaqian Cai, Changfa Lu, Gang Huang 0001 |
SSE | 4 |
| 2024 | Exploring GitHub Topics: Unveiling Their Content and PotentialabstractN owadays, software service design is increasingly oriented toward addressing human needs, aiming to extract users' needs and behavioral patterns from open-source data. GitHub's massive open-source repositories have emerged as a crucial data source for software service researchers seeking to extract valuable insights and develop software services tailored for developers. Both GitHub and researchers are making efforts to help researchers and developers better utilize GitHub data. In 2017, GitHub launched “topics”, enabling developers to assign keywords to repositories. This feature fosters linkages between repositories, aiding in their discovery by other developers. For software development, topics offer two significant values. First, topics provide researchers with new insights to better mine GitHub data and provide enhanced support for developers. Second, developers utilizing topics to annotate their repositories may enhance their visibility and engagement within the community, potentially bolstering their repository's popularity. Despite the increasing number of topics, no research has systematically analyzed their content and potential value. Therefore, we conduct the first empirical study on topics, providing valuable conclusions for future researchers and developers. We conduct a case study encompassing 900 repositories to analyze the information explicitly presented in the topic content, and three experiments to verify whether topics have the potential to be used as repository features and user features in GitHub-related studies. Furthermore, we delve into the correlation between topics and repository popularity, by analyzing the number of stars repositories received. Our findings cover the composition of topic content, the potential value of topics for GitHub-related research, and the impact of topics on repository popularity. Yanchun Sun, Huizhen Jiang, Gang Huang 0001 |
SSE | 5 |
| 2018 | PAVE: Personalized Academic Venue recommendation Exploiting co-publication networks
Shuo Yu 0001, Jiaying Liu 0006, Huizhen Jiang, Amr Tolba, Feng Xia 0001 |
J. Netw. Comput. Appl. | 5 |