Ying Ding 0001

dblp:38/6013-1 · DBLP profile ↗
← Back
71ranked-venue papers in the field
13as first author
17since 2021 · last 2026
0000-0003-2567-2009ORCID · conflict

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 58 (10 first)Data Mining & Knowledge Discovery · 7 (2 first)Other / Interdisciplinary · 4Database Systems & Data Management · 1 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 1
YearPublicationVenuePosition
2026 Agent-Enhanced Heterogeneous Graph RAG for Academic Question Answering
abstract
Academic question answering requires reasoning over heterogeneous scholarly graphs, where queries range from simple attribute lookups to multi-hop inference across author--paper--venue structures. Existing retrieval-augmented generation (RAG) systems struggle in this setting due to three limitations: (1) fixed retrieval strategies that do not adapt to varying query complexity, (2) the absence of sufficiency evaluation leading to incomplete or misaligned evidence, and (3) a lack of structured verification against graph facts. To address these issues, we propose an agentic heterogeneous graph RAG method that transforms the three core stages of the RAG pipeline into explicit agentic decision steps. A query-aware retrieval agent analyzes query type and selects an appropriate graph traversal strategy; a sufficiency-aware reranking agent assesses evidence completeness and adaptively expands the retrieved subgraph; and a graph-grounded verification agent checks entity, relation, and attribute correctness before finalizing the answer. Experiments on heterogeneous graphs constructed from OpenAlex and DBLP suggest that our method consistently outperforms strong LLM, graph-augmented RAG, and agent-based baselines.
Runsong Jia, Mengjia Wu, Ying Ding 0001, Jie Lu 0001, Yi Zhang 0095
WWW3
2026 From Newborn to Impact: Bias-Aware Citation Prediction
abstract
As a key to accessing research impact, citation dynamics underpins research evaluation, scholarly recommendation, and the study of knowledge diffusion. Citation prediction is particularly critical for newborn papers, where early assessment must be performed without citation signals and under highly long-tailed distributions. We identify two key research gaps: (i) insufficient modeling of implicit factors of scientific impact, leading to reliance on coarse proxies; and (ii) a lack of bias-aware learning that can deliver stable predictions on lowly cited papers. We address these gaps by proposing a Bias-Aware Citation Prediction Framework, which combines multi-agent feature extraction with robust graph representation learning. First, a multi-agent x graph co-learning module derives fine-grained, interpretable signals, such as reproducibility, collaboration network, and text quality, from metadata and external resources, and fuses them with heterogeneous-network embeddings to provide rich supervision even in the absence of early citation signals. Second, we incorporate a set of robust mechanisms: a two-stage forward process that routes explicit factors through an intermediate exposure estimate, GroupDRO to optimize worst-case group risk across environments, and a regularization head that performs what-if analyses on controllable factors under monotonicity and smoothness constraints. Comprehensive experiments on two real-world datasets demonstrate the effectiveness of our proposed model. Specifically, our model achieves around a 13% reduction in error metrics (MALE and RMSLE) and a notable 5.5% improvement in the ranking metric (NDCG) over the baseline methods.
Mingfei Lu, Mengjia Wu, Jiawei Xu 0006, Weikai Li 0002, Feng Liu 0003, Ying Ding 0001, Yizhou Sun, Jie Lu 0001, Yi Zhang 0095
WWW6
2025 Quantifying the dynamics of research teams' academic diversity
abstract
Abstract The growing complexity of modern scientific challenges demands research teams that integrate diverse perspectives, yet the role of academic status diversity—variation in team members' scholarly achievements—remains insufficiently understood. This study aims to bridge this gap by examining the dynamics of academic diversity within research teams and its association with innovation, analyzing more than 17 million articles across 292 fields. We introduce new metrics—academic entropy, academic standard deviation, and academic disparity—to capture the heterogeneity of team members' academic backgrounds. Using a network null model to account for temporal and disciplinary differences, we uncover significant increases in academic diversity, particularly within STEM fields and developed regions, where diversity levels are notably overrepresented. While we find a positive correlation between academic diversity and interdisciplinarity, higher diversity is associated with lower levels of scientific disruption. Teams with greater academic diversity tend to be associated with consolidating existing knowledge rather than producing disruptive innovations that challenge prevailing frameworks. This trend is especially evident in larger teams, where diversity is linked to incremental progress rather than transformative breakthroughs. These findings underscore the need for a balanced approach to promoting diversity in relation to scientific advancement.
Alex Jie Yang, Star X. Zhao, Sanhong Deng, Meijun Liu, Yi Bu 0001, Ying Ding 0001
J. Assoc. Inf. Sci. Technol.6
2024 The prominent and heterogeneous gender disparities in scientific novelty: Evidence from biomedical doctoral theses
Meijun Liu, Zihan Xie, Alex Jie Yang, Jian Xu 0003, Ying Ding 0001, Yi Bu 0001
Inf. Process. Manag.6
2024 The impact of heterogeneous shared leadership in scientific teams
Meijun Liu, Yi Bu 0001, Shujing Sun, Yi Zhang 0095, Daniel E. Acuna, Eric T. Meyer, Ying Ding 0001
Inf. Process. Manag.10
2024 Unveiling the loss of exceptional women in science
Yunhan Yang, Yi Bu 0001, Meijun Liu, Ying Ding 0001
Inf. Process. Manag.6
2024 An editorial of "AI + informetrics": Robust models for large-scale analytics
Yi Zhang 0095, Philipp Mayr 0001, Arho Suominen, Ying Ding 0001
Inf. Process. Manag.5
2022 RoS-KD: A Robust Stochastic Knowledge Distillation Approach for Noisy Medical Imaging
abstract
AI-powered Medical Imaging has recently achieved enormous attention due to its ability to provide fast-paced healthcare diagnoses. However, it usually suffers from a lack of high-quality datasets due to high annotation cost, interobserver variability, human annotator error, and errors in computer-generated labels. Deep learning models trained on noisy labelled datasets are sensitive to the noise type and lead to less generalization on the unseen samples. To address this challenge, we propose a Robust Stochastic Knowledge Distillation (RoS-KD) framework which mimics the notion of learning a topic from multiple sources to ensure deterrence in learning noisy information. More specifically, RoS-KD learns a smooth, well-informed, and robust student manifold by distilling knowledge from multiple teachers trained on overlapping subsets of training data. Our extensive experiments on popular medical imaging classification tasks (cardiopulmonary disease and lesion classification) using real-world datasets, show the performance benefit of RoS-KD, its ability to distill knowledge from many popular large networks (ResNet-50, DenseNet-121, MobileNetV2) in a comparatively small network, and its robustness to adversarial attacks (PGD, FSGM). More specifically, RoS-KD achieves >2% and > 4% improvement on F1-score for lesion classification and cardiopulmonary disease classification tasks, respectively, when the underlying student is ResNet-18 against recent competitive knowledge distillation baseline. Additionally, on cardiopulmonary disease classification task, RoS-KD outperforms most of the SOTA baselines by ~1% gain in AUC score.
Ajay Jaiswal, Kumar Ashutosh, Justin F. Rousseau, Yifan Peng 0002, Zhangyang Wang, Ying Ding 0001
ICDM6
2022 International Workshop on Data-driven Science of Science
abstract
Citation data, along with other bibliographic datasets, have long been adopted by the knowledge and data discovery community as an important direction for presenting the validity and effectiveness of proposed algorithms and strategies. Many top computer scientists are also excellent researchers in the science of science. The purpose of this workshop is to bridge the two communities (i.e., the knowledge discovery community and the science of science community) together as the scholarly activities become salient web and social activities that start to generate a ripple effect on broader knowledge discovery communities. This workshop will showcase the current data-driven science of science research by highlighting several studies and constructing a community of researchers to explore questions critical to the future of data-driven science of science, especially a community of data-driven science of science in Data Science so as to facilitate collaboration and inspire innovation. Through discussion on emerging and critical topics in the science of science, this workshop aims to help generate effective solutions for addressing environmental, societal, and technological problems in the scientific community.
Yi Bu 0001, Meijun Liu, Ying Ding 0001, Feng Xia 0001, Daniel E. Acuna, Yi Zhang 0095
KDD4
2022 International Workshop on Knowledge Graphs: Open Knowledge Network
abstract
Knowledge networks/graphs provide a powerful approach for data discovery, integration, and reuse. The NSF's new Convergence Accelerator program, which focuses on transitioning research to practice and translational research, announced Track A on the Open Knowledge Network (OKN). The program calls for multidisciplinary and multi-sector teams to work together to build a cooperative and shared open knowledge network infrastructure to drive innovation across science, engineering, and humanities. This workshop aims to invite researchers, practitioners, and the general public to brainstorm the ideas related to OKN, collaboratively build KGs for different domains or applications, develop AI algorithms to provide intelligent services based on OKN, and discuss the social and economic implications related to OKN.
Ying Ding 0001, Amit P. Sheth, Krzysztof Janowicz, Sergio Baranzini, Sharat Israni, Ilkay Altintas, Lilit Yeghiazarian, Ellie Young, Sam Klein
KDD1
2022 Contributorship in scientific collaborations: The perspective of contribution-based byline orders
Chao Lu 0010, Chengrui Xiao, Ying Ding 0001
Inf. Process. Manag.4
2022 Pandemics are catalysts of scientific novelty: Evidence from COVID-19
abstract
Abstract Scientific novelty drives the efforts to invent new vaccines and solutions during the pandemic. First‐time collaboration and international collaboration are two pivotal channels to expand teams' search activities for a broader scope of resources required to address the global challenge, which might facilitate the generation of novel ideas. Our analysis of 98,981 coronavirus papers suggests that scientific novelty measured by the BioBERT model that is pretrained on 29 million PubMed articles, and first‐time collaboration increased after the outbreak of COVID‐19, and international collaboration witnessed a sudden decrease. During COVID‐19, papers with more first‐time collaboration were found to be more novel and international collaboration did not hamper novelty as it had done in the normal periods. The findings suggest the necessity of reaching out for distant resources and the importance of maintaining a collaborative scientific community beyond nationalism during a pandemic.
Meijun Liu, Yi Bu 0001, Chongyan Chen, Jian Xu 0003, Daifeng Li, Yan Leng, Richard B. Freeman 0002, Eric T. Meyer, Wonjin Yoon, Mujeen Sung, Minbyul Jeong, Jinhyuk Lee, Jaewoo Kang, Min Song 0001, Ying Ding 0001
J. Assoc. Inf. Sci. Technol.17
2022 Team power dynamics and team impact: New perspectives on scientific collaboration using career age as a proxy for team power
abstract
Abstract Power dynamics influence every aspect of scientific collaboration. Team power dynamics can be measured by team power level and team power hierarchy. Team power level is conceptualized as the average level of the possession of resources, expertise, or decision‐making authorities of a team. Team power hierarchy represents the vertical differences of the possessions of resources in a team. In Science of Science, few studies have looked at scientific collaboration from the perspective of team power dynamics. This research examines how team power dynamics affect team impact to fill the research gap. In this research, all coauthors of one publication are treated as one team. Team power level and team power hierarchy of one team are measured by the mean and Gini index of career age of coauthors in this team. Team impact is quantified by citations of a paper authored by this team. By analyzing over 7.7 million teams from Science (e.g., Computer Science, Physics), Social Sciences (e.g., Sociology, Library & Information Science), and Arts & Humanities (e.g., Art), we find that flat team structure is associated with higher team impact, especially when teams have high team power level. These findings have been repeated in all five disciplines except Art, and are consistent in various types of teams from Computer Science including teams from industry or academia, teams with different gender groups, teams with geographical contrast, and teams with distinct size.
Yi Bu 0001, Meijun Liu, Mengyi Sun, Yi Zhang 0095, Eric T. Meyer, Eduardo Salas, Ying Ding 0001
J. Assoc. Inf. Sci. Technol.9
2021 SCALP - Supervised Contrastive Learning for Cardiopulmonary Disease Classification and Localization in Chest X-rays using Patient Metadata
abstract
Computer-aided diagnosis plays a salient role in more accessible and accurate cardiopulmonary diseases classification and localization on chest radiography. Millions of people get affected and die due to these diseases without an accurate and timely diagnosis. Recently proposed contrastive learning heavily relies on data augmentation, especially positive data augmentation. However, generating clinically-accurate data augmentations for medical images is extremely difficult because the common data augmentation methods in computer vision, such as sharp, blur, and crop operations, can severely alter the clinical settings of medical images. In this paper, we proposed a novel and simple data augmentation method based on patient metadata and supervised knowledge to create clinically accurate positive and negative augmentations for chest X-rays. We introduce an end-to-end framework, SCALP, which extends the self-supervised contrastive approach to a supervised setting. Specifically, SCALP pulls together chest X-rays from the same patient (positive keys) and pushes apart chest X-rays from different patients (negative keys). In addition, it uses ResNet-50 along with the triplet-attention mechanism to identify cardiopulmonary diseases, and Grad-CAM++ to highlight the abnormal regions. Our extensive experiments demonstrate that SCALP outperforms existing baselines with significant margins in both classification and localization tasks. Specifically, the average classification AUCs improve from 82.8% (SOTA using DenseNet-121) to 83.9% (SCALP using ResNet-50), while the localization results improve on average by 3.7% over different IoU thresholds.
Ajay Jaiswal, Cyprian Zander, Yan Han 0001, Justin F. Rousseau, Yifan Peng 0002, Ying Ding 0001
ICDM7
2021 International Workshop on Knowledge Graph: Heterogenous Graph Deep Learning and Applications
abstract
Knowledge graph (KG) is the backbone to enable cognitive Artificial Intelligence (AI), which relies on cognitive computing and semantic reasoning. Knowledge graph is the connected data with the semantically enriched context. It is the necessary step for the next move of AI. Our daily activities have closely intermingled with various applications powered by knowledge graphs. It has even entered our healthcare system to facilitate clinical decision making and improve hospital efficiency. This workshop aims to bring researchers and practitioners to promote research and applications related to knowledge graph.
Ying Ding 0001, Bogdan G. Arsintescu, Ching-Hua Chen, Haoyun Feng, François Scharffe, Oshani Seneviratne, Juan F. Sequeda
KDD1
2021 Identifying citation patterns of scientific breakthroughs: A perspective of dynamic citation process
Yi Bu 0001, Ding Wu, Ying Ding 0001, Yi Zhang 0095
Inf. Process. Manag.4
2021 Innovation adoption: Broadcasting versus virality
abstract
Abstract Diffusion channels are critical to determining the adoption scale, which leads to the ultimate impact of an innovation. The aim of this study is to develop an integrative understanding of the impact of two diffusion channels (i.e., broadcasting vs. virality) on innovation adoption. Using citations of a series of classic algorithms and the time series of co‐authorship as the footprints of their diffusion trajectories, we propose a novel method to analyze the intertwining relationships between broadcasting and virality in the innovation diffusion process. Our findings show that broadcasting and virality have similar diffusion power, but play different roles across diffusion stages. Broadcasting is more powerful in the early stages but may be gradually caught up or even surpassed by virality in the later period. Meanwhile, diffusion speed in virality is significantly faster than broadcasting and members from virality channels tend to adopt the same innovation repetitively.
Ying Ding 0001, Hezhao Zhang
J. Assoc. Inf. Sci. Technol.2
2020 Co-contributorship network and division of labor in individual scientific collaborations
abstract
Abstract Collaborations are pervasive in current science. Collaborations have been studied and encouraged in many disciplines. However, little is known about how a team really functions from the detailed division of labor within. In this research, we investigate the patterns of scientific collaboration and division of labor within individual scholarly articles by analyzing their co‐contributorship networks. Co‐contributorship networks are constructed by performing the one‐mode projection of the author–task bipartite networks obtained from 138,787 articles published in PLoS journals. Given an article, we define 3 types of contributors: Specialists, Team‐players, and Versatiles. Specialists are those who contribute to all their tasks alone; team‐players are those who contribute to every task with other collaborators; and versatiles are those who do both. We find that team‐players are the majority and they tend to contribute to the 5 most common tasks as expected, such as “data analysis” and “performing experiments.” The specialists and versatiles are more prevalent than expected by our designed 2 null models. Versatiles tend to be senior authors associated with funding and supervision. Specialists are associated with 2 contrasting roles: the supervising role as team leaders or marginal and specialized contributors.
Chao Lu 0010, Yong-Yeol Ahn, Ying Ding 0001, Dandan Ma
J. Assoc. Inf. Sci. Technol.4
2020 The Gene of Scientific Success
abstract
This article elaborates how to identify and evaluate causal factors to improve scientific impact. Currently, analyzing scientific impact can be beneficial to various academic activities including funding application, mentor recommendation, discovering potential cooperators, and the like. It is universally acknowledged that high-impact scholars often have more opportunities to receive awards as an encouragement for their hard work. Therefore, scholars spend great efforts in making scientific achievements and improving scientific impact during their academic life. However, what are the determinate factors that control scholars’ academic success? The answer to this question can help scholars conduct their research more efficiently. Under this consideration, our article presents and analyzes the causal factors that are crucial for scholars’ academic success. We first propose five major factors including article-centered factors, author-centered factors, venue-centered factors, institution-centered factors, and temporal factors. Then, we apply recent advanced machine learning algorithms and jackknife method to assess the importance of each causal factor. Our empirical results show that author-centered and article-centered factors have the highest relevancy to scholars’ future success in the computer science area. Additionally, we discover an interesting phenomenon that the h -index of scholars within the same institution or university are actually very close to each other.
Xiangjie Kong 0001, Jun Zhang 0048, Da Zhang 0002, Yi Bu 0001, Ying Ding 0001, Feng Xia 0001
ACM Trans. Knowl. Discov. Data5
2019 From zero to one: A perspective on citing
abstract
This article investigates the lengths of time that publications with different numbers of citations take to receive their first citation (the beginning stage), and then compares the lengths of time to receive two or more citations after receiving the first citation (the accumulative stage) in the field of computer science. We find that in the beginning stage, that is, from zero to one citation, high‐, medium‐, and low‐cited publications do not obviously exhibit different lengths of time. However, in the accumulative stage, that is, from one to N citations, highly cited publications begin to receive citations much more rapidly than medium‐ and low‐cited publications. Moreover, as N increases, the difference in receiving new citations among high‐, medium‐, and low‐cited publications increases quite significantly.
Yong Huang 0008, Yi Bu 0001, Ying Ding 0001, Wei Lu 0019
J. Assoc. Inf. Sci. Technol.3
2019 Analyzing stock market trends using social media user moods and social influence
abstract
Information from microblogs is gaining increasing attention from researchers interested in analyzing fluctuations in stock markets. Behavioral financial theory draws on social psychology to explain some of the irrational behaviors associated with financial decisions to help explain some of the fluctuations. In this study we argue that social media users who demonstrate an interest in finance can offer insights into ways in which irrational behaviors may affect a stock market. To test this, we analyzed all the data collected over a 3‐month period in 2011 from Tencent Weibo (one of the largest microblogging websites in China). We designed a social influence (SI)‐based Tencent finance‐related moods model to simulate investors' irrational behaviors, and designed a Tencent Moods‐based Stock Trend Analysis (TM_STA) model to detect correlations between Tencent moods and the Hushen‐300 index (one of the most important financial indexes in China). Experimental results show that the proposed method can help explain the data fluctuation. The findings support the existing behavioral financial theory, and can help to understand short‐term rises and falls in a stock market. We use behavioral financial theory to further explain our findings, and to propose a trading model to verify the proposed model.
Daifeng Li, Yintian Wang, Andrew D. Madden, Ying Ding 0001, Jie Tang 0001, Gordon Guo-Zheng Sun, Ning Zhang 0041, Enguo Zhou
J. Assoc. Inf. Sci. Technol.4
2019 Examining scientific writing styles from the perspective of linguistic complexity
abstract
Publishing articles in high‐impact English journals is difficult for scholars around the world, especially for non‐native English‐speaking scholars (NNESs), most of whom struggle with proficiency in English. To uncover the differences in English scientific writing between native English‐speaking scholars (NESs) and NNESs, we collected a large‐scale data set containing more than 150,000 full‐text articles published in PLoS between 2006 and 2015. We divided these articles into three groups according to the ethnic backgrounds of the first and corresponding authors, obtained by Ethnea, and examined the scientific writing styles in English from a two‐fold perspective of linguistic complexity: (a) syntactic complexity, including measurements of sentence length and sentence complexity; and (b) lexical complexity, including measurements of lexical diversity, lexical density, and lexical sophistication. The observations suggest marginal differences between groups in syntactical and lexical complexity.
Chao Lu 0010, Yi Bu 0001, Jie Wang 0044, Ying Ding 0001, Vetle I. Torvik, Matthew Schnaars
J. Assoc. Inf. Sci. Technol.4
2018 Understanding persistent scientific collaboration
abstract
Common sense suggests that persistence is key to success. In academia, successful researchers have been found more likely to be persistent in publishing, but little attention has been given to how persistence in maintaining collaborative relationships affects career success. This paper proposes a new bibliometric understanding of persistence that considers the prominent role of collaboration in contemporary science. Using this perspective, we analyze the relationship between persistent collaboration and publication quality along several dimensions: degree of transdisciplinarity, difference in coauthor's scientific age and their scientific impact, and research‐team size. Contrary to traditional wisdom, our results show that persistent scientific collaboration does not always result in high‐quality papers. We find that the most persistent transdisciplinary collaboration tends to output high‐impact publications, and that those coauthors with diverse scientific impact or scientific ages benefit from persistent collaboration more than homogeneous compositions. We also find that researchers persistently working in large groups tend to publish lower‐impact papers. These results contradict the colloquial understanding of collaboration in academia and paint a more nuanced picture of how persistent scientific collaboration relates to success, a picture that can provide valuable insights to researchers, funding agencies, policy makers, and mentor–mentee program directors. Moreover, the methodology in this study showcases a feasible approach to measure persistent collaboration.
Yi Bu 0001, Ying Ding 0001, Xingkun Liang, Dakota S. Murray
J. Assoc. Inf. Sci. Technol.2
2018 Understanding success through the diversity of collaborators and the milestone of career
abstract
Scientific collaboration is vital to many fields, and it is common to see scholars seek out experienced researchers or experts in a domain with whom they can share knowledge, experience, and resources. To explore the diversity of research collaborations, this article performs a temporal analysis on the scientific careers of researchers in the field of computer science. Specifically, we analyze collaborators using 2 indicators: the research topic diversity, measured by the Author‐Conference‐Topic model and cosine, and the impact diversity, measured by the normalized standard deviation of h‐indices. We find that the collaborators of high‐impact researchers tend to study diverse research topics and have diverse h‐indices. Moreover, by setting PhD graduation as an important milestone in researchers' careers, we examine several indicators related to scientific collaboration and their effects on a career. The results show that collaborating with authoritative authors plays an important role prior to a researcher's PhD graduation, but working with non‐authoritative authors carries more weight after PhD graduation.
Yi Bu 0001, Ying Ding 0001, Jian Xu 0003, Xingkun Liang, Gege Gao
J. Assoc. Inf. Sci. Technol.2
2018 Innovation or imitation: The diffusion of citations
abstract
Citations in scientific literature are important both for tracking the historical development of scientific ideas and for forecasting research trends. However, the diffusion mechanisms underlying the citation process remain poorly understood, despite the frequent and longstanding use of citation counts for assessment purposes within the scientific community. Here, we extend the study of citation dynamics to a more general diffusion process to understand how citation growth associates with different diffusion patterns. Using a classic diffusion model, we quantify and illustrate specific diffusion mechanisms which have been proven to exert a significant impact on the growth and decay of citation counts. Experiments reveal a positive relation between the “low p and low q” pattern and high scientific impact. A sharp citation peak produced by rapid change of citation counts, however, has a negative effect on future impact. In addition, we have suggested a simple indicator, saturation level, to roughly estimate an individual article's current stage in the life cycle and its potential to attract future attention. The proposed approach can also be extended to higher levels of aggregation (e.g., individual scientists, journals, institutions), providing further insights into the practice of scientific evaluation.
Ying Ding 0001, Jiang Li 0002, Yi Bu 0001
J. Assoc. Inf. Sci. Technol.2
2018 Measuring the diffusion of an innovation: A citation analysis
abstract
Innovations transform our research traditions and become the driving force to advance individual, group, and social creativity. Meanwhile, interdisciplinary research is increasingly being promoted as a route to advance the complex challenges we face as a society. In this paper, we use Latent Dirichlet Allocation (LDA) citation as a proxy context for the diffusion of an innovation. With an analysis of topic evolution, we divide the diffusion process into five stages: testing and evaluation, implementation, improvement, extending, and fading. Through a correlation analysis of topic and subject, we show the application of LDA in different subjects. We also reveal the cross‐boundary diffusion between different subjects based on the analysis of the interdisciplinary studies. The results show that as LDA is transferred into different areas, the adoption of each subject is relatively adjacent to those with similar research interests. Our findings further support researchers' understanding of the impact formation of innovation.
Ying Ding 0001
J. Assoc. Inf. Sci. Technol.2
2018 Understanding scientific collaboration: Homophily, transitivity, and preferential attachment
abstract
Scientific collaboration is essential in solving problems and breeding innovation. Coauthor network analysis has been utilized to study scholars' collaborations for a long time, but these studies have not simultaneously taken different collaboration features into consideration. In this paper, we present a systematic approach to analyze the differences in possibilities that two authors will cooperate as seen from the effects of homophily, transitivity, and preferential attachment. Exponential random graph models (ERGMs) are applied in this research. We find that different types of publications one author has written play diverse roles in his/her collaborations. An author's tendency to form new collaborations with her/his coauthors' collaborators is strong, where the more coauthors one author had before, the more new collaborators he/she will attract. We demonstrate that considering the authors' attributes and homophily effects as well as the transitivity and preferential attachment effects of the coauthorship network in which they are embedded helps us gain a comprehensive understanding of scientific collaboration.
Yi Bu 0001, Ying Ding 0001, Jian Xu 0003
J. Assoc. Inf. Sci. Technol.3
2017 User-level microblogging recommendation incorporating social influence
abstract
With the information overload of user‐generated content in microblogging, users find it extremely challenging to browse and find valuable information in their first attempt. In this paper we propose a microblogging recommendation algorithm, TSI‐MR (Topic‐Level Social Influence‐based Microblogging Recommendation), which can significantly improve users' microblogging experiences. The main innovation of this proposed algorithm is that we consider social influences and their indirect structural relationships, which are largely based on social status theory, from the topic level. The primary advantage of this approach is that it can build an accurate description of latent relationships between two users with weak connections, which can improve the performance of the model; furthermore, it can solve sparsity problems of training data to a certain extent. The realization of the model is mainly based on Factor Graph. We also applied a distributed strategy to further improve the efficiency of the model. Finally, we use data from Tencent Weibo, one of the most popular microblogging services in China, to evaluate our methods. The results show that incorporating social influence can improve microblogging performance considerably, and outperform the baseline methods.
Daifeng Li, Ying Ding 0001, Jie Tang 0001, Gordon Guo-Zheng Sun, Xiaowen Dai, John Du, Shoubin Kong
J. Assoc. Inf. Sci. Technol.3
2016 Author credit-assignment schemas: A comparison and analysis
abstract
Credit assignment to multiple authors of a publication is a challenging task owing to the conventions followed within different areas of research. In this study, we present a review of different author credit‐assignment schemas, which are designed mainly based on author position and the total number of coauthors on the publication. We implemented, tested, and classified 15 author credit‐assignment schemas into 3 types: linear, curve, and “other” assignment schemas. Further investigation and analysis revealed that most of the methods provide reasonable credit‐assignment results, even though the credit‐assignment distribution approaches are quite different among different types. The evaluation of each schema based on P ub M ed articles published in 2013 shows that there exist positive correlations among different schemas and that the similarity of credit‐assignment distributions can be derived from the similar design principles that stress the number of coauthors or the author position, or consider both. We provide a summary about the features of each credit‐assignment schema to facilitate the selection of the appropriate one, depending on the different conditions required to meet diverse needs.
Jian Xu 0003, Ying Ding 0001, Min Song 0001, Tamy Chambers
J. Assoc. Inf. Sci. Technol.2
2015 A lead-lag analysis of the topic evolution patterns for preprints and publications
abstract
This study applied LDA (latent D irichlet allocation) and regression analysis to conduct a lead‐lag analysis to identify different topic evolution patterns between preprints and papers from arXiv and the W eb of S cience ( WoS ) in astrophysics over the last 20 years (1992–2011). Fifty topics in arXiv and WoS were generated using an LDA algorithm and then regression models were used to explain 4 types of topic growth patterns. Based on the slopes of the fitted equation curves, the paper redefines the topic trends and popularity. Results show that arXiv and WoS share similar topics in a given domain, but differ in evolution trends. Topics in WoS lose their popularity much earlier and their durations of popularity are shorter than those in arXiv . This work demonstrates that open access preprints have stronger growth tendency as compared to traditional printed publications.
Beibei Hu, Xianlei Dong, Timothy D. Bowman, Ying Ding 0001, Stasa Milojevic, Chaoqun Ni, Erjia Yan, Vincent Larivière
J. Assoc. Inf. Sci. Technol.5
2015 Topic-level opinion influence model (TOIM): An investigation using tencent microblogging
abstract
Text mining has been widely used in multiple types of user‐generated data to infer user opinion, but its application to microblogging is difficult because text messages are short and noisy, providing limited information about user opinion. Given that microblogging users communicate with each other to form a social network, we hypothesize that user opinion is influenced by its neighbors in the network. In this paper, we infer user opinion on a topic by combining two factors: the user's historical opinion about relevant topics and opinion influence from his/her neighbors. We thus build a topic‐level opinion influence model (TOIM) by integrating both topic factor and opinion influence factor into a unified probabilistic model. We evaluate our model in one of the largest microblogging sites in China, Tencent Weibo, and the experiments show that TOIM outperforms baseline methods in opinion inference accuracy. Moreover, incorporating indirect influence further improves inference recall and f1‐measure. Finally, we demonstrate some useful applications of TOIM in analyzing users' behaviors in Tencent Weibo.
Daifeng Li, Jie Tang 0001, Ying Ding 0001, Xin Shuai, Tamy Chambers, Gordon Guo-Zheng Sun
J. Assoc. Inf. Sci. Technol.3
2015 Influence diffusion detection using the influence style (INFUSE) model
abstract
Blogs are readily available sources of opinions and sentiments that in turn could influence the opinions of the blog readers. Previous studies have attempted to infer influence from blog features, but they have ignored the possible influence styles that describe the different ways in which influence is exerted. We propose a novel approach to analyzing bloggers' influence styles and using the influence styles as features to improve the performance of influence diffusion detection among linked bloggers. The proposed influence style (INFUSE) model describes bloggers' influence through their engagement style, persuasion style, and persona. Methods used include similarity analysis to detect the creating−sharing aspect of engagement style, subjectivity analysis to measure persuasion style, and sentiment analysis to identify persona style. We further extend the INFUSE model to detect influence diffusion among linked bloggers based on the bloggers' influence styles. The INFUSE model performed well with an average F1 score of 76% compared with the in‐degree and sentiment‐value baseline approaches. Previous studies have focused on the existence of influence among linked bloggers in detecting influence diffusion, but our INFUSE model is shown to provide a fine‐grained description of the manner in which influence is diffused based on the bloggers' influence styles.
Luke Kien-Weng Tan, Jin-Cheon Na, Ying Ding 0001
J. Assoc. Inf. Sci. Technol.3
2014 Modeling Paying Behavior in Game Social Networks
abstract
Online gaming is one of the largest industries on the Internet, generating tens of billions of dollars in revenues annually. One core problem in online game is to find and convert free users into paying customers, which is of great importance for the sustainable development of almost all online games. Although much research has been conducted, there are still several challenges that remain largely unsolved: What are the fundamental factors that trigger the users to pay? How does users? paying behavior influence each other in the game social network? How to design a prediction model to recognize those potential users who are likely to pay? In this paper, employing two large online games as the basis, we study how a user becomes a new paying user in the games. In particular, we examine how users' paying behavior influences each other in the game social network. We study this problem from various sociological perspectives including strong/weak ties, social structural diversity and social influence. Based on the discovered patterns, we propose a learning framework to predict potential new payers. The framework can learn a model using features associated with users and then use the social relationships between users to refine the learned model. We test the proposed framework using nearly 50 billion user activities from two real games. Our experiments show that the proposed framework significantly improves the prediction accuracy by up to 3-11% compared to several alternative methods. The study also unveils several intriguing social phenomena from the data. For example, influence indeed exists among users for the paying behavior. The likelihood of a user becoming a new paying user is 5 times higher than chance when he has 5 paying neighbors of strong tie. We have deployed the proposed algorithm into the game, and the Lift_Ratio has been improved up to 196% compared to the prior strategy.
Zhanpeng Fang, Jie Tang 0001, Longjun Sun, Ying Ding 0001, Jar-der Luo
CIKM7
2014 Content-based citation analysis: The next generation of citation analysis
abstract
Traditional citation analysis has been widely applied to detect patterns of scientific collaboration, map the landscapes of scholarly disciplines, assess the impact of research outputs, and observe knowledge transfer across domains. It is, however, limited, as it assumes all citations are of similar value and weights each equally. Content‐based citation analysis ( CCA ) addresses a citation's value by interpreting each one based on its context at both the syntactic and semantic levels. This paper provides a comprehensive overview of CAA research in terms of its theoretical foundations, methodical approaches, and example applications. In addition, we highlight how increased computational capabilities and publicly available full‐text resources have opened this area of research to vast possibilities, which enable deeper citation analysis, more accurate citation prediction, and increased knowledge discovery.
Ying Ding 0001, Guo Zhang 0007, Tamy Chambers, Min Song 0001, Xiaolong Wang 0009, ChengXiang Zhai
J. Assoc. Inf. Sci. Technol.1
2014 Patent citation analysis: Calculating science linkage based on citing motivation
abstract
Science linkage is a widely used patent bibliometric indicator to measure patent linkage to scientific research based on the frequency of citations to scientific papers within the patent. Science linkage is also regarded as noisy because the subject of patent citation behavior varies from inventors/applicants to examiners. In order to identify and ultimately reduce this noise, we analyzed the different citing motivations of examiners and inventors/applicants. We built 4 hypotheses based upon our study of patent law, the unique economic nature of a patent, and a patent citation's market effect. To test our hypotheses, we conducted an expert survey based on our science linkage calculation in the domain of catalyst from U.S. patent data (2006–2009) over 3 types of citations: self‐citation by inventor/applicant, non‐self‐citation by inventor/applicant, and citation by examiner. According to our results, evaluated by domain experts, we conclude that the non‐self‐citation by inventor/applicant is quite noisy and cannot indicate science linkage and that self‐citation by inventor/applicant, although limited, is more appropriate for understanding science linkage.
Rui Li 0097, Tamy Chambers, Ying Ding 0001, Guo Zhang 0007, Liansheng Meng
J. Assoc. Inf. Sci. Technol.3
2014 Productivity and influence in bioinformatics: A bibliometric analysis using PubMed central
abstract
Bioinformatics is a fast‐growing field based on the optimal use of “big data” gathered in genomic, proteomics, and functional genomics research. In this paper, we conduct a comprehensive and in‐depth bibliometric analysis of the field of bioinformatics by extracting citation data from PubMed Central full‐text. Citation data for the period 2000 to 2011, comprising 20,869 papers with 546,245 citations, was used to evaluate the productivity and influence of this emerging field. Four measures were used to identify productivity; most productive authors, most productive countries, most productive organizations, and most popular subject terms. Research impact was analyzed based on the measures of most cited papers, most cited authors, emerging stars, and leading organizations. Results show the overall trends between the periods 2000 to 2003 and 2004 to 2007 were dissimilar, while trends between the periods 2004 to 2007 and 2008 to 2011 were similar. In addition, the field of bioinformatics has undergone a significant shift, co‐evolving with other biomedical disciplines.
Min Song 0001, Su Yeon Kim, Guo Zhang 0007, Ying Ding 0001, Tamy Chambers
J. Assoc. Inf. Sci. Technol.4
2013 Workshop summary for the 2013 international workshop on mining unstructured big data using natural language processing
abstract
No abstract available.
Xiaozhong Liu 0001, Ying Ding 0001, Min Song 0001
CIKM3
2013 Journal impact and proximity: An assessment using bibliographic features
abstract
Journals in the Information Science & Library Science category of Journal Citation Reports (JCR) were compared using both bibliometric and bibliographic features. Data collected covered journal impact factor (JIF), number of issues per year, number of authors per article, longevity, editorial board membership, frequency of publication, number of databases indexing the journal, number of aggregators providing full‐text access, country of publication, JCR categories, Dewey decimal classification, and journal statement of scope. Three features significantly correlated with JIF: number of editorial board members and number of JCR categories in which a journal is listed correlated positively; journal longevity correlated negatively with JIF. Coword analysis of journal descriptions provided a proximity clustering of journals, which differed considerably from the clusters based on editorial board membership. Finally, a multiple linear regression model was built to predict the JIF based on all the collected bibliographic features.
Chaoqun Ni, Debora Shaw, Sean M. Lind, Ying Ding 0001
J. Assoc. Inf. Sci. Technol.4
2013 Citation content analysis (CCA): A framework for syntactic and semantic analysis of citation content
abstract
This study proposes a new framework for citation content analysis ( CCA ), for syntactic and semantic analysis of citation content that can be used to better analyze the rich sociocultural context of research behavior. This framework could be considered the next generation of citation analysis. The authors briefly review the history and features of content analysis in traditional social sciences and its previous application in library and information science (LIS). Based on critical discussion of the theoretical necessity of a new method as well as the limits of citation analysis, the nature and purposes of CCA are discussed, and potential procedures to conduct CCA , including principles to identify the reference scope, a two‐dimensional (citing and cited) and two‐module (syntactic and semantic) codebook, are provided and described. Future work and implications are also suggested.
Guo Zhang 0007, Ying Ding 0001, Stasa Milojevic
J. Assoc. Inf. Sci. Technol.2
2012 Meta path-based collective classification in heterogeneous information networks
abstract
Collective classification approaches exploit the dependencies of a group of linked objects whose class labels are correlated and need to be predicted simultaneously. In this paper, we focus on studying the collective classification problem in heterogeneous networks, which involves multiple types of data objects interconnected by multiple types of links. Intuitively, two objects are correlated if they are linked by many paths in the network. By considering different linkage paths in the network, one can capture the subtlety of different types of dependencies among objects. We introduce the concept of meta-path based dependencies among objects, where a meta path is a path consisting a certain sequence of linke types. We show that the quality of collective classification results strongly depends upon the meta paths used. To accommodate the large network size, a novel solution, called HCC (meta-path based Heterogenous Collective Classification), is developed to effectively assign labels to a group of instances that are interconnected through different meta-paths. The proposed HCC model can capture different types of dependencies among objects with respect to different meta paths. Empirical studies on real-world networks demonstrate that effectiveness of the proposed meta path-based collective classification approach.
Xiangnan Kong, Philip S. Yu, Ying Ding 0001, David J. Wild 0001
CIKM3
2012 Mining topic-level opinion influence in microblog
abstract
This paper proposes a Topic-Level Opinion Influence Model (TOIM) that simultaneously incorporates topic factor, user opinions and social influence in a unified probabilistic model with two stages learning processes. In the first stage, topic factor and user influence are integrated to generate users' influential relationship based on different topics; in the second stage, users' historical messages and social interaction records are leveraged by TOIM to construct their historical opinions and neighbors' opinion influence through a statistical learning process, which can be further utilized to predict users' future opinions on some specific topics. We evaluate our TOIM on a large-scaled dataset from Tencent Weibo, one of the largest microbloggings website in China. The experimental results show that TOIM can better predict users' opinion than other baseline methods.
Daifeng Li, Xin Shuai, Gordon Guo-Zheng Sun, Jie Tang 0001, Ying Ding 0001
CIKM5
2012 Mining competitive relationships by learning across heterogeneous networks
abstract
Detecting and monitoring competitors is fundamental to a company to stay ahead in the global market. Existing studies mainly focus on mining competitive relationships within a single data source, while competing information is usually distributed in multiple networks. How to discover the underlying patterns and utilize the heterogeneous knowledge to avoid biased aspects in this issue is a challenging problem. In this paper, we study the problem of mining competitive relationships by learning across heterogeneous networks. We use Twitter and patent records as our data sources and statistically study the patterns behind the competitive relationships. We find that the two networks exhibit different but complementary patterns of competitions. Our proposed model, Topical Factor Graph Model (TFGM), defines a latent topic layer to bridge the two networks and learns a semi-supervised learning model to classify the relationships between entities (e.g., companies or products). We test the proposed model on two real data sets and the experimental results validate the effectiveness of our model, with an average of +46\% improvement over alternative methods.
Yang Yang 0009, Jie Tang 0001, Jacklyne Keomany, Yanting Zhao, Juan-Zi Li, Ying Ding 0001
CIKM6
2012 Modeling Indirect Influence on Twitter
abstract
Social influence in social networks has been extensively researched. Most studies have focused on direct influence, while another interesting question can be raised as whether indirect influence exists between two users who’re not directly connected in the network and what affects such influence. In addition, the theory of complex contagion tells us that more spreaders will enhance the indirect influence between two users. The authors’ observation of intensity of indirect influence, propagated by n parallel spreaders and quantified by retweeting probability in two Twitter social networks, shows that complex contagion is validated globally but is violated locally. In other words, the retweeting probability increases non-monotonically with some local drops. A quantum cognition based probabilistic model is proposed to account for these local drops.
Xin Shuai, Ying Ding 0001, Jerome R. Busemeyer, Yuyin Sun, Jie Tang 0001
Int. J. Semantic Web Inf. Syst.2
2012 Information Concepts, Retrieval, and Services [Series]
Ying Ding 0001
J. Assoc. Inf. Sci. Technol.1
2012 The dynamic features of Delicious, Flickr, and YouTube
abstract
Abstract This article investigates the dynamic features of social tagging vocabularies in Delicious, Flickr, and YouTube from 2003 to 2008. Three algorithms are designed to study the macro‐ and micro‐tag growth as well as the dynamics of taggers' activities, respectively. Moreover, we propose a Tagger Tag Resource Latent Dirichlet Allocation (TTR‐LDA) model to explore the evolution of topics emerging from those social vocabularies. Our results show that (a) at the macro level, tag growth in all the three tagging systems obeys power law distribution with exponents lower than 1; at the micro level, the tag growth of popular resources in all three tagging systems follows a similar power law distribution; (b) the exponents of tag growth vary in different evolving stages of resources; (c) the growth of number of taggers associated with different popular resources presents a feature of convergence over time; (d) the active level of taggers has a positive correlation with the macro‐tag growth of different tagging systems; and (e) some topics evolve into several subtopics over time while others experience relatively stable stages in which their contents do not change much, and certain groups of taggers continue their interests in them.
Daifeng Li, Ying Ding 0001, Bing He 0003, Jie Tang 0001, Juan-Zi Li, Tianxi Dong
J. Assoc. Inf. Sci. Technol.3
2012 Scholarly network similarities: How bibliographic coupling networks, citation networks, cocitation networks, topical networks, coauthorship networks, and coword networks relate to each other
abstract
This study explores the similarity among six types of scholarly networks aggregated at the institution level, including bibliographic coupling networks, citation networks, cocitation networks, topical networks, coauthorship networks, and coword networks. Cosine distance is chosen to measure the similarities among the six networks. The authors found that topical networks and coauthorship networks have the lowest similarity; cocitation networks and citation networks have high similarity; bibliographic coupling networks and cocitation networks have high similarity; and coword networks and topical networks have high similarity. In addition, through multidimensional scaling, two dimensions can be identified among the six networks: Dimension 1 can be interpreted as citation‐based versus noncitation‐based, and Dimension 2 can be interpreted as social versus cognitive. The authors recommend the use of hybrid or heterogeneous networks to study research interaction and scholarly communications.
Erjia Yan, Ying Ding 0001
J. Assoc. Inf. Sci. Technol.2
2011 Popular and/or prestigious? Measures of scholarly esteem
Ying Ding 0001, Blaise Cronin
Inf. Process. Manag.1
2011 Discovering author impact: A PageRank perspective
Erjia Yan, Ying Ding 0001
Inf. Process. Manag.2
2011 Applying weighted PageRank to author citation networks
abstract
This article aims to identify whether different weighted PageRank algorithms can be applied to author citation networks to measure the popularity and prestige of a scholar from a citation perspective. Information retrieval (IR) was selected as a test field and data from 1956–2008 were collected from Web of Science. Weighted PageRank with citation and publication as weighted vectors were calculated on author citation networks. The results indicate that both popularity rank and prestige rank were highly correlated with the weighted PageRank. Principal component analysis was conducted to detect relationships among these different measures. For capturing prize winners within the IR field, prestige rank outperformed all the other measures.
Ying Ding 0001
J. Assoc. Inf. Sci. Technol.1
2011 Topic-based PageRank on author cocitation networks
abstract
Abstract Ranking authors is vital for identifying a researcher's impact and standing within a scientific field. There are many different ranking methods (e.g., citations, publications, h‐index, PageRank, and weighted PageRank), but most of them are topic‐independent. This paper proposes topic‐dependent ranks based on the combination of a topic model and a weighted PageRank algorithm. The author‐conference‐topic (ACT) model was used to extract topic distribution of individual authors. Two ways for combining the ACT model with the PageRank algorithm are proposed: simple combination (I_PR) or using a topic distribution as a weighted vector for PageRank (PR_t). Information retrieval was chosen as the test field and representative authors for different topics at different time phases were identified. Principal component analysis (PCA) was applied to analyze the ranking difference between I_PR and PR_t.
Ying Ding 0001
J. Assoc. Inf. Sci. Technol.1
2011 Mining enriched contextual information of scientific collaboration: A meso perspective
abstract
Studying scientific collaboration using coauthorship networks has attracted much attention in recent years. How and in what context two authors collaborate remain among the major questions. Previous studies, however, have focused on either exploring the global topology of coauthorship networks (macro perspective) or ranking the impact of individual authors (micro perspective). Neither of them has provided information on the context of the collaboration between two specific authors, which may potentially imply rich socioeconomic, disciplinary, and institutional information on collaboration. Different from the macro perspective and micro perspective, this article proposes a novel method (meso perspective) to analyze scientific collaboration, in which a contextual subgraph is extracted as the unit of analysis. A contextual subgraph is defined as a small subgraph of a large-scale coauthorship network that captures relationship and context between two coauthors. This method is applied to the field of library and information science. Topological properties of all the subgraphs in four time spans are investigated, including size, average degree, clustering coefficient, and network centralization. Results show that contextual subgprahs capture useful contextual information on two authors' collaboration.
Bing He 0003, Ying Ding 0001, Chaoqun Ni
J. Assoc. Inf. Sci. Technol.2
2011 Modeling topic and community structure in social tagging: The TTR-LDA-Community model
abstract
The presence of social networks in complex systems has made networks and community structure a focal point of study in many domains. Previous studies have focused on the structural emergence and growth of communities and on the topics displayed within the network. However, few scholars have closely examined the relationship between the thematic and structural properties of networks. Therefore, this article proposes the Tagger Tag Resource-Latent Dirichlet Allocation-Community model (TTR-LDA-Community model), which combines the Latent Dirichlet Allocation (LDA) model with the Girvan-Newman community detection algorithm through an inference mechanism. Using social tagging data from Delicious, this article demonstrates the clustering of active taggers into communities, the topic distributions within communities, and the ranking of taggers, tags, and resources within these communities. The data analysis evaluates patterns in community structure and topical affiliations diachronically. The article evaluates the effectiveness of community detection and the inference mechanism embedded in the model and finds that the TTR-LDA-Community model outperforms other traditional models in tag prediction. This has implications for scholars in domains interested in community detection, profiling, and recommender systems.
Daifeng Li, Ying Ding 0001, Cassidy R. Sugimoto, Bing He 0003, Jie Tang 0001, Erjia Yan, Tianxi Dong
J. Assoc. Inf. Sci. Technol.2
2011 The cognitive structure of Library and Information Science: Analysis of article title words
abstract
Abstract This study comprises a suite of analyses of words in article titles in order to reveal the cognitive structure of Library and Information Science (LIS). The use of title words to elucidate the cognitive structure of LIS has been relatively neglected. The present study addresses this gap by performing (a) co‐word analysis and hierarchical clustering, (b) multidimensional scaling, and (c) determination of trends in usage of terms. The study is based on 10,344 articles published between 1988 and 2007 in 16 LIS journals. Methodologically, novel aspects of this study are: (a) its large scale, (b) removal of non‐specific title words based on the “word concentration” measure (c) identification of the most frequent terms that include both single words and phrases, and (d) presentation of the relative frequencies of terms using “heatmaps”. Conceptually, our analysis reveals that LIS consists of three main branches: the traditionally recognized library‐related and information‐related branches, plus an equally distinct bibliometrics/scientometrics branch. The three branches focus on: libraries, information, and science, respectively. In addition, our study identifies substructures within each branch. We also tentatively identify “information seeking behavior” as a branch that is establishing itself separate from the three main branches. Furthermore, we find that cognitive concepts in LIS evolve continuously, with no stasis since 1992. The most rapid development occurred between 1998 and 2001, influenced by the increased focus on the Internet. The change in the cognitive landscape is found to be driven by the emergence of new information technologies, and the retirement of old ones.
Stasa Milojevic, Cassidy R. Sugimoto, Erjia Yan, Ying Ding 0001
J. Assoc. Inf. Sci. Technol.4
2011 The shifting sands of disciplinary development: Analyzing North American Library and Information Science dissertations using latent Dirichlet allocation
abstract
This work identifies changes in dominant topics in library and information science (LIS) over time, by analyzing the 3,121 doctoral dissertations completed between 1930 and 2009 at North American Library and Information Science programs. The authors utilize latent Dirichlet allocation (LDA) to identify latent topics diachronically and to identify representative dissertations of those topics. The findings indicate that the main topics in LIS have changed substantially from those in the initial period (1930–1969) to the present (2000–2009). However, some themes occurred in multiple periods, representing core areas of the field: library history occurred in the first two periods; citation analysis in the second and third periods; and information-seeking behavior in the fourth and last period. Two topics occurred in three of the five periods: information retrieval and information use. One of the notable changes in the topics was the diminishing use of the word library (and related terms). This has implications for the provision of doctoral education in LIS. This work is compared to other earlier analyses and provides validation for the use of LDA in topic analysis of a discipline.
Cassidy R. Sugimoto, Daifeng Li, Terrell G. Russell, S. Craig Finlay, Ying Ding 0001
J. Assoc. Inf. Sci. Technol.5
2011 P-Rank: An indicator measuring prestige in heterogeneous scholarly networks
abstract
Abstract Ranking scientific productivity and prestige are often limited to homogeneous networks. These networks are unable to account for the multiple factors that constitute the scholarly communication and reward system. This study proposes a new informetric indicator, P‐Rank, for measuring prestige in heterogeneous scholarly networks containing articles, authors, and journals. P‐Rank differentiates the weight of each citation based on its citing papers, citing journals, and citing authors. Articles from 16 representative library and information science journals are selected as the dataset. Principle Component Analysis is conducted to examine the relationship between P‐Rank and other bibliometric indicators. We also compare the correlation and rank variances between citation counts and P‐Rank scores. This work provides a new approach to examining prestige in scholarly communication networks in a more comprehensive and nuanced way.
Erjia Yan, Ying Ding 0001, Cassidy R. Sugimoto
J. Assoc. Inf. Sci. Technol.2
2010 Dynamic Features of Social Tagging Vocabulary: Delicious, Flickr and YouTube
abstract
This article investigates the dynamic features of social tagging vocabularies in Delicious, Flickr and YouTube from 2003 to 2008. It analyzes the evolution of the usage of the most popular tags in each of these three social networks. We find that for different tagging systems, the dynamic features reflect different cognitive processes. At the macro level, the tag growth obeys power-law distribution for all three tagging systems with exponents lower than one. At the micro level, the tag growth of popular resources in all three tagging systems follows a similar power-law distribution. Moreover, we find that the exponents of tag growth varied in different evolving stages of popular individual resources.
Daifeng Li, Ying Ding 0001, Stasa Milojevic, Bing He 0003, Erjia Yan, Tianxi Dong
ASONAM2
2010 Community-based topic modeling for social tagging
abstract
Exploring community is fundamental for uncovering the connections between structure and function of complex networks and for practical applications in many disciplines such as biology and sociology. In this paper, we propose a TTR-LDA-Community model which combines the Latent Dirichlet Allocation model (LDA) and the Girvan-Newman community detection algorithm with an inference mechanism. The model is then applied to data from Delicious, a popular social tagging system, over the time period of 2005-2008. Our results show that 1) users in the same community tend to be interested in similar set of topics in all time periods; and 2) topics may divide into several sub-topics and scatter into different communities over time. We evaluate the effectiveness of our model and show that the TTR-LDA-Community model is meaningful for understanding communities and outperforms TTR-LDA and LDA models in tag prediction.
Daifeng Li, Bing He 0003, Ying Ding 0001, Jie Tang 0001, Cassidy R. Sugimoto, Erjia Yan, Juan-Zi Li, Tianxi Dong
CIKM3
2010 Chem2Bio2RDF: A Linked Open Data Portal for Systems Chemical Biology
abstract
The Chem2Bio2RDF portal is a Linked Open Data (LOD) portal for systems chemical biology aiming for facilitating drug discovery. It converts around 25 different datasets on genes, compounds, drugs, pathways, side effects, diseases, and MEDLINE/PubMed documents into RDF triples and links them to other LOD bubbles, such as Bio2RDF, LODD and DBPedia. The portal is based on D2R server and provides a SPARQL endpoint, but adds on several unique features such as RDF faceted browser, user-friendly SPARQL query generator, MEDLINE/PubMed cross validation service, and Cytoscape visualization plugin. Three use cases demonstrate the functionality and usability of this portal. The portal is available at http://chem2bio2rdf.org.
Bin Chen 0002, Ying Ding 0001, David J. Wild 0001, Yuyin Sun, Qian Zhu 0003, Madhuvanthi Sankaranarayanan
Web Intelligence2
2010 Modeling Ontology of Folksonomy with Latent Semantics of Tags
abstract
Modeling ontology of folksonomy provides a way of learning light weight ontology's which is a hot topic investigated recently. Previous approaches for modeling ontology of folksonomy either ignores semantics (synonymy, hyponymy or polysemy) or do not simultaneously consider relationships between actors (users), concepts (tags) and instances(resources) or are based on the idea that title words are responsible for generating tags for resources. Latent semantics and user-tag dependencies instead of user-word dependencies however are extremely important. In this paper we address these problems by introducing a latent topic layer into the traditional tripartite Actor-Concept-Instance graph. We thus propose an Actor-Concept-Instance-Topic (ACIT) approach to model ontology from folksonomy in a unified way by directly using tags and users of resources. We illustrate on Bibsonomy dataset that our proposed approach ACIT outperforms title words based approaches Tag-Topic (TT) and (User-Word-Topic) UWT for modeling the ontology of folksonomy.
Ali Daud, Juan-Zi Li, Lizhu Zhou, Lei Zhang 0174, Ying Ding 0001, Faqir Muhammad
Web Intelligence5
2010 Muzk Mesh: Interlinking Semantic Music Data
abstract
The vision of the Semantic Web is to lift current Web into semantic repositories where heterogeneous data can be queried and different services can be mashed up. The Web becomes a platform for integrating data and services. The paper discusses the MuzkMesh music portal which mashups existing semantic music data from the Linked Open Data (LOD) bubbles and other common APIs. It aims to demo the power of semantic integration and useful use scenarios on music retrieval and entertainment.
Mayank Singhi, Ying Ding 0001, Yuyin Sun
Web Intelligence2
2010 Using Web Technologies for Integrative Drug Discovery
abstract
Recent years have seen a huge increase in the amount of publicly-available information relevant to drug discovery, including online databases of compound and bioassay information; scholarly publications linking compounds with genes, targets and diseases; and predictive models that can suggest new links between compounds, genes, targets and diseases. However, there is a lack of tools and methods to integrate this information, and in particular to look for pertinent knowledge and relationships across multiple sources. At Indiana University we are tackling this problem by applying aggregative data mining tools and semantic web technologies including using an extensive web service infrastructure, RDF networks and inference engines, ontologies, and automated extraction of information from scholarly literature.
Qian Zhu 0003, Sashikiran Challa, Prajakta Purohit, Yuyin Sun, Michael S. Lajiness, David J. Wild 0001, Ying Ding 0001
Web Intelligence7
2010 Upper tag ontology for integrating social tagging data
abstract
Abstract Data integration and mediation have become central concerns of information technology over the past few decades. With the advent of the Web and the rapid increases in the amount of data and the number of Web documents and users, researchers have focused on enhancing the interoperability of data through the development of metadata schemes. Other researchers have looked to the wealth of metadata generated by bookmarking sites on the Social Web. While several existing ontologies have capitalized on the semantics of metadata created by tagging activities, the Upper Tag Ontology (UTO) emphasizes the structure of tagging activities to facilitate modeling of tagging data and the integration of data from different bookmarking sites as well as the alignment of tagging ontologies. UTO is described and its utility in modeling, harvesting, integrating, searching, and analyzing data is demonstrated with metadata harvested from three major social tagging systems (Delicious, Flickr, and YouTube).
Ying Ding 0001, Elin K. Jacob, Michael A. H. Fried, Ioan Toma, Erjia Yan, Schubert Foo, Stasa Milojevic
J. Assoc. Inf. Sci. Technol.1
2010 Weighted citation: An indicator of an article's prestige
abstract
Abstract The authors propose using the technique of weighted citation to measure an article's prestige. The technique allocates a different weight to each reference by taking into account the impact of citing journals and citation time intervals. Weightedcitation captures prestige, whereas citation counts capture popularity. They compare the value variances for popularity and prestige for articles published in the Journal of the American Society for Information Science and Technology from 1998 to 2007, and find that the majority have comparable status.
Erjia Yan, Ying Ding 0001
J. Assoc. Inf. Sci. Technol.2
2009 A Categorical Model for Discovering Latent Structure in Social Annotations
Said Kashoob, James Caverlee, Ying Ding 0001
ICWSM3
2009 Perspectives on social tagging
abstract
Abstract Social tagging is one of the major phenomena transforming the World Wide Web from a static platform into an actively shared information space. This paper addresses various aspects of social tagging, including different views on the nature of social tagging, how to make use of social tags, and how to bridge social tagging with other Web functionalities; it discusses the use of facets to facilitate browsing and searching of tagging data; and it presents an analogy between bibliometrics and tagometrics, arguing that established bibliometric methodologies can be applied to analyze tagging behavior on the Web. Based on the Upper Tag Ontology (UTO), a Web crawler was built to harvest tag data from Delicious, Flickr, and YouTube in September 2007. In total, 1.8 million objects, including bookmarks, photos, and videos, 3.1 million taggers, and 12.1 million tags were collected and analyzed. Some tagging patterns and variations are identified and discussed.
Ying Ding 0001, Elin K. Jacob, Schubert Foo, Erjia Yan, Nicolas L. George, Lijiang Guo
J. Assoc. Inf. Sci. Technol.1
2009 PageRank for ranking authors in co-citation networks
abstract
Abstract This paper studies how varied damping factors in the PageRank algorithm influence the ranking of authors and proposes weighted PageRank algorithms. We selected the 108 most highly cited authors in the information retrieval (IR) area from the 1970s to 2008 to form the author co‐citation network. We calculated the ranks of these 108 authors based on PageRank with the damping factor ranging from 0.05 to 0.95. In order to test the relationship between different measures, we compared PageRank and weighted PageRank results with the citation ranking, h‐index, and centrality measures. We found that in our author co‐citation network, citation rank is highly correlated with PageRank with different damping factors and also with different weighted PageRank algorithms; citation rank and PageRank are not significantly correlated with centrality measures; and h‐index rank does not significantly correlate with centrality measures but does significantly correlate with other measures. The key factors that have impact on the PageRank of authors in the author co‐citation network are being co‐cited with important authors.
Ying Ding 0001, Erjia Yan, Arthur R. Frazho, James Caverlee
J. Assoc. Inf. Sci. Technol.1
2009 Applying centrality measures to impact analysis: A coauthorship network analysis
abstract
Abstract Many studies on coauthorship networks focus on network topology and network statistical mechanics. This article takes a different approach by studying micro‐level network properties with the aim of applying centrality measures to impact analysis. Using coauthorship data from 16 journals in the field of library and information science (LIS) with a time span of 20 years (1988–2007), we construct an evolving coauthorship network and calculate four centrality measures (closeness centrality, betweenness centrality, degree centrality, and PageRank) for authors in this network. We find that the four centrality measures are significantly correlated with citation counts. We also discuss the usability of centrality measures in author ranking and suggest that centrality measures can be useful indicators for impact analysis.
Erjia Yan, Ying Ding 0001
J. Assoc. Inf. Sci. Technol.2
2008 WWW 2008 workshop on social web search and mining: SWSM2008
abstract
No abstract available.
Juan-Zi Li, Gui-Rong Xue, Jie Tang 0001, Ying Ding 0001
WWW4
2002 The semantic web: yet another hip?
Ying Ding 0001, Dieter Fensel, Michel C. A. Klein, Borys Omelayenko
Data Knowl. Eng.1
2001 Bibliometric cartography of information retrieval research by using co-word analysis
Ying Ding 0001, Gobinda G. Chowdhury, Schubert Foo
Inf. Process. Manag.1
2000 Bibliometric information retrieval system (BIRS): A web search interface utilizing bibliometric research results
abstract
The aim of this article is to test whether the results obtained from a specific bibliographic research can be applied to a real search environment and enhance the level of utility of an information retrieval session for all levels of end users. In this respect, a Web-based Bibliometric Information Retrieval System (BIRS) has been designed and created, with facilities to assist the end users to get better understanding of their search domain, formulate and expand their search queries, and visualize the bibliographic research results. There are three specific features in the system design of the BIRS: the information visualization feature of the BIRS (cocitation maps) to guide the end users to identify the important research groups and capture the detailed information about the intellectual structure of the search domain; the multilevel browsing feature to allow the end users to go to different levels of interesting topics; and the common user interface feature to enable the end users to search all kinds of databases regardless of different searching systems, different working platforms, different database producer and supplier, such as different Web search engines, different library OPACs, or different on-line databases. A preliminary user evaluation study of BIRS revealed that users generally found it easy to form and expand their queries, and that BIRS helped them acquire useful background information about the search domain. They also pointed out aspects of information visualization, multilevel browsing, and common user interface as novel characteristics exhibited by BIRS.
Ying Ding 0001, Gobinda G. Chowdhury, Schubert Foo, Weizhong Qian
J. Am. Soc. Inf. Sci.1