Siyi Gong

dblp:296/1193 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 4 first-author · 4 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FlexNTN-Twin: A Flexible Hardware-in-the-Loop Emulation Platform for 5G-Advanced NTN
Qiuming Zhu, Siyi Gong, Hanpeng Li
INFOCOM5
2026 How challenging it is to identify real code authors: an empirical study
Siyi Gong, Hao Zhong 0001
Empir. Softw. Eng.1
2025 Communicating through Acting: The Role of Contextual Affordance in Intuitive Pantomimetic Gestural Communication
Siyi Gong, Jessica G. Li, Mireille Karadanaian, Ziyi Meng 0004, Tao Gao 0004
CogSci1
2025 Territorial Gestalt in the Strategy of Conflicts
Siyi Gong, Jifan Zhou, Mowei Shen, Tao Gao 0004
CogSci3
2025 Incremental learning of code authors over time
Siyi Gong, Hao Zhong 0001
J. Syst. Softw.1
2022 Perceptual Grouping for War and Peace
Siyi Gong, Jifan Zhou, Mowei Shen, Tao Gao 0004
CogSci4
2022 Exploring an Imagined "We" in Human Collective Hunting: Joint Commitment within Shared Intentionality
Siyi Gong, Minglu Zhao, Chenya Gu, Jifan Zhou, Mowei Shen, Tao Gao 0004
CogSci2
2022 A study on identifying code author from real development
abstract
Identifying code authors is important in many research topics, and various approaches have been proposed. Although these approaches achieve promising results on their datasets, their true effectiveness is still in question. To the best of our knowledge, only one large-scale study was conducted to explore the impacts of related factors (e.g., the temporal effect and the distribution of files per author). This study selected Google Code Jam programs as their subjects, but such programs are quite different from the source files that programmers write in daily development. To understand their effectiveness and challenges, we replicate their study and use their approach to analyze source files that are retrieved from real projects. The prior study claims that the temporal effect and the distribution of files per author have only minor impacts on their trained models. In the contrast, we find that in 85.48% pairs of training and testing sets, the accuracy of a trained model is less effective when the temporal effect is considered, and in total, the average accuracy decreases by 0.4298. In addition, when we use the real distribution of files as inputs, their approach can accurately identify only one or two core code authors, although a project can have more than ten authors. By revealing the limitations of the prior approach, our study sheds lights on where to make future improvements.
Siyi Gong, Hao Zhong 0001
ESEC/SIGSOFT FSE1
2021 Jointly Perceiving Physics and Mind: Motion, force and intention
Siyi Gong, Ziqian Liao, Haokui Xu, Jifan Zhou, Mowei Shen, Tao Gao 0004
CogSci2
2021 Code Authors Hidden in File Revision Histories: An Empirical Study
abstract
Although many programmers write their names in the comments of a source file, from such comments, it is unreliable to identify code authors, since the modifications of many programmers are not recorded. Even if they are recorded in a code repository, many authors are hidden in revision histories.The true authors of source files are important in many research topics. For example, when detecting plagiarism, if the authors of two source are overlapped, it becomes more challenging to determine plagiarism than the source files that are written by individual authors. As it is difficult to determine true authors of a source file, researchers typically use source files whose authors are already known (e.g., the source files from Google Code Jam), but such files are not many and less representative. Meanwhile, although some empirical studies touch code authors, to the best of our knowledge, no prior study has analyzed the characteristics of code authors that are hidden in revision histories. As a result, many research questions along with code authors are still open. For example, how many authors does a source file can have, and what are the proportions of contributions per source file, if they are written by more than one author?To answer the timely questions, in this paper, we conducted an empirical study on code authors that are hidden in revision histories. To support our study, we implemented a tool called CODA. By comparing the latest code lines with past commits, CODA identifies the true authors of all code lines. With its support, we analyzed 12,092 source files that were written by 506 programmers. Our study answers several interesting questions concerning code authors. For example, we find that 75.4% source files are written by multiple authors, and their contributions follow the famous 80/20 principle. These findings are useful to understand authors of source files in open source communities.
Siyi Gong, Hao Zhong 0001
ICPC1