Yi Bu 0001

dblp:181/4792-1 · DBLP profile ↗
← Back
31ranked-venue papers in the field
6as first author
23since 2021 · last 2026
0000-0003-2549-4580ORCID · conflict

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 29 (5 first)Data Mining & Knowledge Discovery · 2 (1 first)
YearPublicationVenuePosition
2026 IMGSL: Citation-based Innovation Motif Graph Structure Learning for scholar profiling
Mingshu Ai, Yi Bu 0001, Tao Jia 0001
Inf. Process. Manag.4
2025 The attention inequality of scientists: A core-periphery structure perspective
Win-Bin Huang, Yi Bu 0001
Inf. Process. Manag.3
2025 Knowledge integration and diffusion structures of interdisciplinary research: A large-scale analysis based on propensity score matching
abstract
Abstract While facilitating science, interdisciplinary research (IDR) has a heavier cognitive burden for researchers compared to unidisciplinary research (UDR). Yet, little has been known about patterns of knowledge integration and diffusion structures of IDR. Here we adopt a causal inference strategy, namely propensity score matching, with all journal publications in 2005 in Microsoft Academic Graph to better understand the IDR effect in various research fields. We use the diversity of reference fields of one paper as the proxy of the paper's interdisciplinarity and estimate the effect of a research article being IDR on its knowledge integration and diffusion measured by its high‐order citation/reference cascade. We find that, in disciplines where IDR articles are less popular, such as mathematics, physics, and chemistry, IDR needs a more extensive knowledge base than UDR to gain a similar number of citations. In disciplines where IDR articles are more popular, for example, psychology, geology, biology, and economics, a small knowledge base is enough for a high‐impact IDR article. As to knowledge diffusion, no matter whether IDR or UDR, a more extensive knowledge base leads to stronger knowledge diffusion ability. Findings imply potential drawbacks of pure interdisciplinarity‐oriented research policy; rather, the establishment of policies may vary across disciplines.
Jiawei Xu 0006, Zhihan Zheng, Win-bin Huang, Yi Bu 0001
J. Assoc. Inf. Sci. Technol.5
2025 Quantifying the dynamics of research teams' academic diversity
abstract
Abstract The growing complexity of modern scientific challenges demands research teams that integrate diverse perspectives, yet the role of academic status diversity—variation in team members' scholarly achievements—remains insufficiently understood. This study aims to bridge this gap by examining the dynamics of academic diversity within research teams and its association with innovation, analyzing more than 17 million articles across 292 fields. We introduce new metrics—academic entropy, academic standard deviation, and academic disparity—to capture the heterogeneity of team members' academic backgrounds. Using a network null model to account for temporal and disciplinary differences, we uncover significant increases in academic diversity, particularly within STEM fields and developed regions, where diversity levels are notably overrepresented. While we find a positive correlation between academic diversity and interdisciplinarity, higher diversity is associated with lower levels of scientific disruption. Teams with greater academic diversity tend to be associated with consolidating existing knowledge rather than producing disruptive innovations that challenge prevailing frameworks. This trend is especially evident in larger teams, where diversity is linked to incremental progress rather than transformative breakthroughs. These findings underscore the need for a balanced approach to promoting diversity in relation to scientific advancement.
Alex Jie Yang, Star X. Zhao, Sanhong Deng, Meijun Liu, Yi Bu 0001, Ying Ding 0001
J. Assoc. Inf. Sci. Technol.5
2025 Understanding discrepancies in the coverage of OpenAlex: The case of China
abstract
Abstract Citations indexes play a crucial role for understanding how science is produced, disseminated, and used. However, these databases often face a critical trade‐off: those offering extensive and high‐quality coverage are typically proprietary, whereas publicly accessible datasets frequently exhibit fragmented coverage and inconsistent data quality. OpenAlex was developed to address this challenge, providing a freely available database with broad open coverage, with a particular emphasis on non‐English speaking countries. Yet, few studies have assessed the quality of the OpenAlex dataset. This paper assesses the coverage by OpenAlex of China's papers, which shows an abnormal trend, and compares it with other countries that do not have English as their main language. Our analysis reveals that while OpenAlex increases the coverage of China's publications, primarily those disseminated by a national database, this coverage is incomplete and discontinuous when compared to other countries' records in the database. We observe similar issues in other non‐English‐speaking countries, with coverage varying across regions. These findings indicate that although OpenAlex expands coverage of research outputs, continuity issues persist and disproportionately affect certain countries. We emphasize the need for researchers to use OpenAlex data cautiously, being mindful of its potential limitations in cross‐national analyses.
Mengxue Zheng, Lili Miao, Yi Bu 0001, Vincent Larivière
J. Assoc. Inf. Sci. Technol.3
2024 The prominent and heterogeneous gender disparities in scientific novelty: Evidence from biomedical doctoral theses
Meijun Liu, Zihan Xie, Alex Jie Yang, Jian Xu 0003, Ying Ding 0001, Yi Bu 0001
Inf. Process. Manag.7
2024 The impact of heterogeneous shared leadership in scientific teams
Meijun Liu, Yi Bu 0001, Shujing Sun, Yi Zhang 0095, Daniel E. Acuna, Eric T. Meyer, Ying Ding 0001
Inf. Process. Manag.3
2024 Unveiling the loss of exceptional women in science
Yunhan Yang, Yi Bu 0001, Meijun Liu, Ying Ding 0001
Inf. Process. Manag.4
2024 Monodisciplinary collaboration disrupts science more than multidisciplinary collaboration
abstract
Abstract Collaboration across disciplines is a critical form of scientific collaboration to solve complex problems and make innovative contributions. This study focuses on the association between multidisciplinary collaboration measured by coauthorship in publications and the disruption of publications measured by the Disruption (D) index. We used authors' affiliations as a proxy of the disciplines to which they belong and categorized an article into multidisciplinary collaboration or monodisciplinary collaboration. The D index quantifies the extent to which a study disrupts its predecessors. We selected 13 journals that publish articles in six disciplines from the Microsoft Academic Graph (MAG) database and then constructed regression models with fixed effects and estimated the relationship between the variables. The findings show that articles with monodisciplinary collaboration are more disruptive than those with multidisciplinary collaboration. Furthermore, we uncovered the mechanism of how monodisciplinary collaboration disrupts science more than multidisciplinary collaboration by exploring the references of the sampled publications.
Xin Liu 0117, Yi Bu 0001, Jiang Li 0002
J. Assoc. Inf. Sci. Technol.2
2024 Understanding super-partnerships in scientific collaboration: Evidence from the field of economics
abstract
Abstract Super‐partnerships exist between scholars connected within densely‐knit collaboration networks. Understanding how such relationships affect scholars' careers is of great importance. In this paper, focusing on the longitudinal aspects of scientific collaboration, we analyze collaboration profiles from the egocentric perspective and use analytic extreme value thresholds to identify super‐partners. A total of 5722 pairs of super‐partners are found in the field of economics. The several interesting findings about super‐partners are summarized as follows. (1) The collaboration pattern of super‐partners can be divided into three types: the dual‐core, bridge, and triangle types. (2) Gender disparities are reflected in the collaboration among super‐partners, and the stability of super‐partnerships involving different combinations of genders displays different characteristics. The random‐effect model is constructed to explore the effect of a super‐partnership on both parties from the aspects of productivity and influence, which also shows gender disparities. (3) A super‐partnership contributes to above‐average productivity and citation impacts of the publication for three collaboration patterns, and the research improvement of the triangle type is the greatest among the three types. Overall, this paper explores the characteristics of super‐partners and the added value of a long‐term commitment, which provides quantitative insights into the effect on scientific collaboration associated with close collaboration.
Junwan Liu, Xiaofei Guo, Shuo Xu 0001, Yi Bu 0001, Cassidy R. Sugimoto, Vincent Larivière, Yinglu Song, Honghao Zhou
J. Assoc. Inf. Sci. Technol.4
2022 International Workshop on Data-driven Science of Science
abstract
Citation data, along with other bibliographic datasets, have long been adopted by the knowledge and data discovery community as an important direction for presenting the validity and effectiveness of proposed algorithms and strategies. Many top computer scientists are also excellent researchers in the science of science. The purpose of this workshop is to bridge the two communities (i.e., the knowledge discovery community and the science of science community) together as the scholarly activities become salient web and social activities that start to generate a ripple effect on broader knowledge discovery communities. This workshop will showcase the current data-driven science of science research by highlighting several studies and constructing a community of researchers to explore questions critical to the future of data-driven science of science, especially a community of data-driven science of science in Data Science so as to facilitate collaboration and inspire innovation. Through discussion on emerging and critical topics in the science of science, this workshop aims to help generate effective solutions for addressing environmental, societal, and technological problems in the scientific community.
Yi Bu 0001, Meijun Liu, Ying Ding 0001, Feng Xia 0001, Daniel E. Acuna, Yi Zhang 0095
KDD1
2022 Fine-grained citation count prediction via a transformer-based model with among-attention mechanism
Shengzhi Huang, Yong Huang 0008, Yi Bu 0001, Wei Lu 0019, Jiajia Qian
Inf. Process. Manag.3
2022 Revisiting the exploration-exploitation behavior of scholars' research topic selection: Evidence from a large-scale bibliographic database
Shengzhi Huang, Wei Lu 0019, Yi Bu 0001, Yong Huang 0008
Inf. Process. Manag.3
2022 Towards transdisciplinary impact of scientific publications: A longitudinal, comprehensive, and large-scale analysis on Microsoft Academic Graph
Yong Huang 0008, Wei Lu 0019, Qikai Cheng, Yi Bu 0001
Inf. Process. Manag.5
2022 Disclosing the relationship between citation structure and future impact of a publication
abstract
Abstract Each section header of an article has its distinct communicative function. Citations from distinct sections may be different regarding citing motivation. In this paper, we grouped section headers with similar functions as a structural function and defined the distribution of citations from structural functions for a paper as its citation structure. We aim to explore the relationship between citation structure and the future impact of a publication and disclose the relative importance among citations from different structural functions. Specifically, we proposed two citation counting methods and a citation life cycle identification method, by which the regression data were built. Subsequently, we employed a ridge regression model to predict the future impact of the paper and analyzed the relative weights of regressors. Based on documents collected from the Association for Computational Linguistics Anthology website, our empirical experiments disclosed that functional structure features improve the prediction accuracy of citation count prediction and that there exist differences among citations from different structural functions. Specifically, at the early stage of citation lifetime, citations from Introduction and Method are particularly important for perceiving future impact of papers, and citations from Result and Conclusion are also vital. However, early accumulation of citations from the Background seems less important.
Shengzhi Huang, Jiajia Qian, Yong Huang 0008, Wei Lu 0019, Yi Bu 0001, Jinqing Yang, Qikai Cheng
J. Assoc. Inf. Sci. Technol.5
2022 Pandemics are catalysts of scientific novelty: Evidence from COVID-19
abstract
Abstract Scientific novelty drives the efforts to invent new vaccines and solutions during the pandemic. First‐time collaboration and international collaboration are two pivotal channels to expand teams' search activities for a broader scope of resources required to address the global challenge, which might facilitate the generation of novel ideas. Our analysis of 98,981 coronavirus papers suggests that scientific novelty measured by the BioBERT model that is pretrained on 29 million PubMed articles, and first‐time collaboration increased after the outbreak of COVID‐19, and international collaboration witnessed a sudden decrease. During COVID‐19, papers with more first‐time collaboration were found to be more novel and international collaboration did not hamper novelty as it had done in the normal periods. The findings suggest the necessity of reaching out for distant resources and the importance of maintaining a collaborative scientific community beyond nationalism during a pandemic.
Meijun Liu, Yi Bu 0001, Chongyan Chen, Jian Xu 0003, Daifeng Li, Yan Leng, Richard B. Freeman 0002, Eric T. Meyer, Wonjin Yoon, Mujeen Sung, Minbyul Jeong, Jinhyuk Lee, Jaewoo Kang, Min Song 0001, Ying Ding 0001
J. Assoc. Inf. Sci. Technol.2
2022 Team power dynamics and team impact: New perspectives on scientific collaboration using career age as a proxy for team power
abstract
Abstract Power dynamics influence every aspect of scientific collaboration. Team power dynamics can be measured by team power level and team power hierarchy. Team power level is conceptualized as the average level of the possession of resources, expertise, or decision‐making authorities of a team. Team power hierarchy represents the vertical differences of the possessions of resources in a team. In Science of Science, few studies have looked at scientific collaboration from the perspective of team power dynamics. This research examines how team power dynamics affect team impact to fill the research gap. In this research, all coauthors of one publication are treated as one team. Team power level and team power hierarchy of one team are measured by the mean and Gini index of career age of coauthors in this team. Team impact is quantified by citations of a paper authored by this team. By analyzing over 7.7 million teams from Science (e.g., Computer Science, Physics), Social Sciences (e.g., Sociology, Library & Information Science), and Arts & Humanities (e.g., Art), we find that flat team structure is associated with higher team impact, especially when teams have high team power level. These findings have been repeated in all five disciplines except Art, and are consistent in various types of teams from Computer Science including teams from industry or academia, teams with different gender groups, teams with geographical contrast, and teams with distinct size.
Yi Bu 0001, Meijun Liu, Mengyi Sun, Yi Zhang 0095, Eric T. Meyer, Eduardo Salas, Ying Ding 0001
J. Assoc. Inf. Sci. Technol.2
2021 How wide is the citation impact of scientific publications? A cross-discipline and large-scale analysis
Yi Bu 0001, Wei Lu 0019, Hongkan Chen, Yong Huang 0008
Inf. Process. Manag.1
2021 Detecting research topic trends by author-defined keyword frequency
Wei Lu 0019, Shengzhi Huang, Jinqing Yang, Yi Bu 0001, Qikai Cheng, Yong Huang 0008
Inf. Process. Manag.4
2021 Identifying citation patterns of scientific breakthroughs: A perspective of dynamic citation process
Yi Bu 0001, Ding Wu, Ying Ding 0001, Yi Zhang 0095
Inf. Process. Manag.2
2021 Characterizing scientists leaving science before their time: Evidence from mathematics
Zhenyue Zhao, Yi Bu 0001, Jiang Li 0002
Inf. Process. Manag.2
2021 Topic diversity: A discipline scheme-free diversity measurement for journals
abstract
Abstract Scientometrics has many citation‐based measurements for characterizing diversity, but most of these measurements depend on human‐designed categories and the granularity of discipline classifications sometimes does not allow in‐depth analysis. As such, the current paper proposes a new measurement for quantifying journals' diversity by utilizing the abstracts of scientific publications in journals, namely topic diversity (TD). Specifically, we apply a topic detection method to extract fine‐grained topics, rather than disciplines, in journals and adapt certain diversity indicators to calculate TD. Since TD only needs as inputs abstracts of publications rather than citing relationships between publications, this measurement has the potential to be widely used in scientometrics.
Yi Bu 0001, Weiye Gu, Win-Bin Huang
J. Assoc. Inf. Sci. Technol.1
2021 Citation cascade and the evolution of topic relevance
abstract
Abstract Citation analysis, as a tool for quantitative studies of science, has long emphasized direct citation relations, leaving indirect or high‐order citations overlooked. However, a series of early and recent studies demonstrate the existence of indirect and continuous citation impact across generations. Adding to the literature on high‐order citations, we introduce the concept of a citation cascade: the constitution of a series of subsequent citing events initiated by a certain publication. We investigate this citation structure by analyzing more than 450,000 articles and over 6 million citation relations. We show that citation impact exists not only within the three generations documented in prior research but also in much further generations. Still, our experimental results indicate that two to four generations are generally adequate to trace a work's scientific impact. We also explore specific structural properties—such as depth, width, structural virality, and size—which account for differences among individual citation cascades. Finally, we find evidence that it is more important for a scientific work to inspire trans‐domain (or indirectly related domain) works than to receive only intradomain recognition in order to achieve high impact. Our methods and findings can serve as a new tool for scientific evaluation and the modeling of scientific history.
Erjia Yan, Yi Bu 0001
J. Assoc. Inf. Sci. Technol.4
2020 Considering author sequence in all-author co-citation analysis
Yi Bu 0001, Binglu Wang, Zaida Chinchilla-Rodríguez, Cassidy R. Sugimoto, Yong Huang 0008, Win-Bin Huang
Inf. Process. Manag.1
2020 The Gene of Scientific Success
abstract
This article elaborates how to identify and evaluate causal factors to improve scientific impact. Currently, analyzing scientific impact can be beneficial to various academic activities including funding application, mentor recommendation, discovering potential cooperators, and the like. It is universally acknowledged that high-impact scholars often have more opportunities to receive awards as an encouragement for their hard work. Therefore, scholars spend great efforts in making scientific achievements and improving scientific impact during their academic life. However, what are the determinate factors that control scholars’ academic success? The answer to this question can help scholars conduct their research more efficiently. Under this consideration, our article presents and analyzes the causal factors that are crucial for scholars’ academic success. We first propose five major factors including article-centered factors, author-centered factors, venue-centered factors, institution-centered factors, and temporal factors. Then, we apply recent advanced machine learning algorithms and jackknife method to assess the importance of each causal factor. Our empirical results show that author-centered and article-centered factors have the highest relevancy to scholars’ future success in the computer science area. Additionally, we discover an interesting phenomenon that the h -index of scholars within the same institution or university are actually very close to each other.
Xiangjie Kong 0001, Jun Zhang 0048, Da Zhang 0002, Yi Bu 0001, Ying Ding 0001, Feng Xia 0001
ACM Trans. Knowl. Discov. Data4
2019 From zero to one: A perspective on citing
abstract
This article investigates the lengths of time that publications with different numbers of citations take to receive their first citation (the beginning stage), and then compares the lengths of time to receive two or more citations after receiving the first citation (the accumulative stage) in the field of computer science. We find that in the beginning stage, that is, from zero to one citation, high‐, medium‐, and low‐cited publications do not obviously exhibit different lengths of time. However, in the accumulative stage, that is, from one to N citations, highly cited publications begin to receive citations much more rapidly than medium‐ and low‐cited publications. Moreover, as N increases, the difference in receiving new citations among high‐, medium‐, and low‐cited publications increases quite significantly.
Yong Huang 0008, Yi Bu 0001, Ying Ding 0001, Wei Lu 0019
J. Assoc. Inf. Sci. Technol.2
2019 Examining scientific writing styles from the perspective of linguistic complexity
abstract
Publishing articles in high‐impact English journals is difficult for scholars around the world, especially for non‐native English‐speaking scholars (NNESs), most of whom struggle with proficiency in English. To uncover the differences in English scientific writing between native English‐speaking scholars (NESs) and NNESs, we collected a large‐scale data set containing more than 150,000 full‐text articles published in PLoS between 2006 and 2015. We divided these articles into three groups according to the ethnic backgrounds of the first and corresponding authors, obtained by Ethnea, and examined the scientific writing styles in English from a two‐fold perspective of linguistic complexity: (a) syntactic complexity, including measurements of sentence length and sentence complexity; and (b) lexical complexity, including measurements of lexical diversity, lexical density, and lexical sophistication. The observations suggest marginal differences between groups in syntactical and lexical complexity.
Chao Lu 0010, Yi Bu 0001, Jie Wang 0044, Ying Ding 0001, Vetle I. Torvik, Matthew Schnaars
J. Assoc. Inf. Sci. Technol.2
2018 Understanding persistent scientific collaboration
abstract
Common sense suggests that persistence is key to success. In academia, successful researchers have been found more likely to be persistent in publishing, but little attention has been given to how persistence in maintaining collaborative relationships affects career success. This paper proposes a new bibliometric understanding of persistence that considers the prominent role of collaboration in contemporary science. Using this perspective, we analyze the relationship between persistent collaboration and publication quality along several dimensions: degree of transdisciplinarity, difference in coauthor's scientific age and their scientific impact, and research‐team size. Contrary to traditional wisdom, our results show that persistent scientific collaboration does not always result in high‐quality papers. We find that the most persistent transdisciplinary collaboration tends to output high‐impact publications, and that those coauthors with diverse scientific impact or scientific ages benefit from persistent collaboration more than homogeneous compositions. We also find that researchers persistently working in large groups tend to publish lower‐impact papers. These results contradict the colloquial understanding of collaboration in academia and paint a more nuanced picture of how persistent scientific collaboration relates to success, a picture that can provide valuable insights to researchers, funding agencies, policy makers, and mentor–mentee program directors. Moreover, the methodology in this study showcases a feasible approach to measure persistent collaboration.
Yi Bu 0001, Ying Ding 0001, Xingkun Liang, Dakota S. Murray
J. Assoc. Inf. Sci. Technol.1
2018 Understanding success through the diversity of collaborators and the milestone of career
abstract
Scientific collaboration is vital to many fields, and it is common to see scholars seek out experienced researchers or experts in a domain with whom they can share knowledge, experience, and resources. To explore the diversity of research collaborations, this article performs a temporal analysis on the scientific careers of researchers in the field of computer science. Specifically, we analyze collaborators using 2 indicators: the research topic diversity, measured by the Author‐Conference‐Topic model and cosine, and the impact diversity, measured by the normalized standard deviation of h‐indices. We find that the collaborators of high‐impact researchers tend to study diverse research topics and have diverse h‐indices. Moreover, by setting PhD graduation as an important milestone in researchers' careers, we examine several indicators related to scientific collaboration and their effects on a career. The results show that collaborating with authoritative authors plays an important role prior to a researcher's PhD graduation, but working with non‐authoritative authors carries more weight after PhD graduation.
Yi Bu 0001, Ying Ding 0001, Jian Xu 0003, Xingkun Liang, Gege Gao
J. Assoc. Inf. Sci. Technol.1
2018 Innovation or imitation: The diffusion of citations
abstract
Citations in scientific literature are important both for tracking the historical development of scientific ideas and for forecasting research trends. However, the diffusion mechanisms underlying the citation process remain poorly understood, despite the frequent and longstanding use of citation counts for assessment purposes within the scientific community. Here, we extend the study of citation dynamics to a more general diffusion process to understand how citation growth associates with different diffusion patterns. Using a classic diffusion model, we quantify and illustrate specific diffusion mechanisms which have been proven to exert a significant impact on the growth and decay of citation counts. Experiments reveal a positive relation between the “low p and low q” pattern and high scientific impact. A sharp citation peak produced by rapid change of citation counts, however, has a negative effect on future impact. In addition, we have suggested a simple indicator, saturation level, to roughly estimate an individual article's current stage in the life cycle and its potential to attract future attention. The proposed approach can also be extended to higher levels of aggregation (e.g., individual scientists, journals, institutions), providing further insights into the practice of scientific evaluation.
Ying Ding 0001, Jiang Li 0002, Yi Bu 0001
J. Assoc. Inf. Sci. Technol.4
2018 Understanding scientific collaboration: Homophily, transitivity, and preferential attachment
abstract
Scientific collaboration is essential in solving problems and breeding innovation. Coauthor network analysis has been utilized to study scholars' collaborations for a long time, but these studies have not simultaneously taken different collaboration features into consideration. In this paper, we present a systematic approach to analyze the differences in possibilities that two authors will cooperate as seen from the effects of homophily, transitivity, and preferential attachment. Exponential random graph models (ERGMs) are applied in this research. We find that different types of publications one author has written play diverse roles in his/her collaborations. An author's tendency to form new collaborations with her/his coauthors' collaborators is strong, where the more coauthors one author had before, the more new collaborators he/she will attract. We demonstrate that considering the authors' attributes and homophily effects as well as the transitivity and preferential attachment effects of the coauthorship network in which they are embedded helps us gain a comprehensive understanding of scientific collaboration.
Yi Bu 0001, Ying Ding 0001, Jian Xu 0003
J. Assoc. Inf. Sci. Technol.2