VLDB 2026 Research / reviewers in the wild / expert
Siyi Gong
dblp:296/1193
· DBLP profile ↗
10ranked-venue papers
5as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 4 first-author · 4 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FlexNTN-Twin: A Flexible Hardware-in-the-Loop Emulation Platform for 5G-Advanced NTN
Qiuming Zhu, Siyi Gong, Hanpeng Li |
INFOCOM | 5 |
| 2026 | How challenging it is to identify real code authors: an empirical study
Siyi Gong, Hao Zhong 0001 |
Empir. Softw. Eng. | 1 |
| 2025 | Communicating through Acting: The Role of Contextual Affordance in Intuitive Pantomimetic Gestural Communication
Siyi Gong, Jessica G. Li, Mireille Karadanaian, Ziyi Meng 0004, Tao Gao 0004 |
CogSci | 1 |
| 2025 | Territorial Gestalt in the Strategy of Conflicts
Siyi Gong, Jifan Zhou, Mowei Shen, Tao Gao 0004 |
CogSci | 3 |
| 2025 | Incremental learning of code authors over time
Siyi Gong, Hao Zhong 0001 |
J. Syst. Softw. | 1 |
| 2022 | Perceptual Grouping for War and Peace
Siyi Gong, Jifan Zhou, Mowei Shen, Tao Gao 0004 |
CogSci | 4 |
| 2022 | Exploring an Imagined "We" in Human Collective Hunting: Joint Commitment within Shared Intentionality
Siyi Gong, Minglu Zhao, Chenya Gu, Jifan Zhou, Mowei Shen, Tao Gao 0004 |
CogSci | 2 |
| 2022 | A study on identifying code author from real developmentabstractIdentifying code authors is important in many research topics, and various approaches have been proposed. Although these approaches achieve promising results on their datasets, their true effectiveness is still in question. To the best of our knowledge, only one large-scale study was conducted to explore the impacts of related factors (e.g., the temporal effect and the distribution of files per author). This study selected Google Code Jam programs as their subjects, but such programs are quite different from the source files that programmers write in daily development. To understand their effectiveness and challenges, we replicate their study and use their approach to analyze source files that are retrieved from real projects. The prior study claims that the temporal effect and the distribution of files per author have only minor impacts on their trained models. In the contrast, we find that in 85.48% pairs of training and testing sets, the accuracy of a trained model is less effective when the temporal effect is considered, and in total, the average accuracy decreases by 0.4298. In addition, when we use the real distribution of files as inputs, their approach can accurately identify only one or two core code authors, although a project can have more than ten authors. By revealing the limitations of the prior approach, our study sheds lights on where to make future improvements. Siyi Gong, Hao Zhong 0001 |
ESEC/SIGSOFT FSE | 1 |
| 2021 | Jointly Perceiving Physics and Mind: Motion, force and intention
Siyi Gong, Ziqian Liao, Haokui Xu, Jifan Zhou, Mowei Shen, Tao Gao 0004 |
CogSci | 2 |
| 2021 | Code Authors Hidden in File Revision Histories: An Empirical StudyabstractAlthough many programmers write their names in the comments of a source file, from such comments, it is unreliable to identify code authors, since the modifications of many programmers are not recorded. Even if they are recorded in a code repository, many authors are hidden in revision histories.The true authors of source files are important in many research topics. For example, when detecting plagiarism, if the authors of two source are overlapped, it becomes more challenging to determine plagiarism than the source files that are written by individual authors. As it is difficult to determine true authors of a source file, researchers typically use source files whose authors are already known (e.g., the source files from Google Code Jam), but such files are not many and less representative. Meanwhile, although some empirical studies touch code authors, to the best of our knowledge, no prior study has analyzed the characteristics of code authors that are hidden in revision histories. As a result, many research questions along with code authors are still open. For example, how many authors does a source file can have, and what are the proportions of contributions per source file, if they are written by more than one author?To answer the timely questions, in this paper, we conducted an empirical study on code authors that are hidden in revision histories. To support our study, we implemented a tool called CODA. By comparing the latest code lines with past commits, CODA identifies the true authors of all code lines. With its support, we analyzed 12,092 source files that were written by 506 programmers. Our study answers several interesting questions concerning code authors. For example, we find that 75.4% source files are written by multiple authors, and their contributions follow the famous 80/20 principle. These findings are useful to understand authors of source files in open source communities. Siyi Gong, Hao Zhong 0001 |
ICPC | 1 |