Yong Huang 0008

dblp:35/198-8 · DBLP profile ↗
← Back
14ranked-venue papers in the field
2as first author
11since 2021 · last 2025
0000-0001-5953-6908ORCID · conflict

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 13 (2 first)Big Data, Cloud & Distributed Data Systems · 1
YearPublicationVenuePosition
2025 Are large language models qualified reviewers in originality evaluation?
Shengzhi Huang, Yong Huang 0008, Yinpeng Liu, Zhuoran Luo, Wei Lu 0019
Inf. Process. Manag.2
2025 Identifying potentially disruptive research via a comparative power-based large model
Shengzhi Huang, Wei Lu 0019, Zhenzhen Xu, Qikai Cheng, Jinqing Yang, Yong Huang 0008
Inf. Process. Manag.6
2025 PaperEval: A universal, quantitative, and explainable paper evaluation method powered by a multi-agent system
Shengzhi Huang, Qicong Wang, Wei Lu 0019, Lingyu Liu, Zhenzhen Xu, Yong Huang 0008
Inf. Process. Manag.6
2024 Evolutions of semantic consistency in research topic via contextualized word embedding
Shengzhi Huang, Wei Lu 0019, Qikai Cheng, Zhuoran Luo, Yong Huang 0008
Inf. Process. Manag.5
2023 Generating keyphrases for readers: A controllable keyphrase generation framework
abstract
Abstract With the wide application of keyphrases in many Information Retrieval (IR) and Natural Language Processing (NLP) tasks, automatic keyphrase prediction has been emerging. However, these statistically important phrases are contributing increasingly less to the related tasks because the end‐to‐end learning mechanism enables models to learn the important semantic information of the text directly. Similarly, keyphrases are of little help for readers to quickly grasp the paper's main idea because the relationship between the keyphrase and the paper is not explicit to readers. Therefore, we propose to generate keyphrases with specific functions for readers to bridge the semantic gap between them and the information producers, and verify the effectiveness of the keyphrase function for assisting users’ comprehension with a user experiment. A controllable keyphrase generation framework (the CKPG) that uses the keyphrase function as a control code to generate categorized keyphrases is proposed and implemented based on Transformer, BART, and T5, respectively. For the Computer Science domain, the Macro‐avgs of , , and on the Paper with Code dataset are up to 0.680, 0.535, and 0.558, respectively. Our experimental results indicate the effectiveness of the CKPG models.
Yong Huang 0008, Wei Lu 0019, Jiawei Liu 0002
J. Assoc. Inf. Sci. Technol.3
2022 Fine-grained citation count prediction via a transformer-based model with among-attention mechanism
Shengzhi Huang, Yong Huang 0008, Yi Bu 0001, Wei Lu 0019, Jiajia Qian
Inf. Process. Manag.2
2022 Revisiting the exploration-exploitation behavior of scholars' research topic selection: Evidence from a large-scale bibliographic database
Shengzhi Huang, Wei Lu 0019, Yi Bu 0001, Yong Huang 0008
Inf. Process. Manag.4
2022 Towards transdisciplinary impact of scientific publications: A longitudinal, comprehensive, and large-scale analysis on Microsoft Academic Graph
Yong Huang 0008, Wei Lu 0019, Qikai Cheng, Yi Bu 0001
Inf. Process. Manag.1
2022 Disclosing the relationship between citation structure and future impact of a publication
abstract
Abstract Each section header of an article has its distinct communicative function. Citations from distinct sections may be different regarding citing motivation. In this paper, we grouped section headers with similar functions as a structural function and defined the distribution of citations from structural functions for a paper as its citation structure. We aim to explore the relationship between citation structure and the future impact of a publication and disclose the relative importance among citations from different structural functions. Specifically, we proposed two citation counting methods and a citation life cycle identification method, by which the regression data were built. Subsequently, we employed a ridge regression model to predict the future impact of the paper and analyzed the relative weights of regressors. Based on documents collected from the Association for Computational Linguistics Anthology website, our empirical experiments disclosed that functional structure features improve the prediction accuracy of citation count prediction and that there exist differences among citations from different structural functions. Specifically, at the early stage of citation lifetime, citations from Introduction and Method are particularly important for perceiving future impact of papers, and citations from Result and Conclusion are also vital. However, early accumulation of citations from the Background seems less important.
Shengzhi Huang, Jiajia Qian, Yong Huang 0008, Wei Lu 0019, Yi Bu 0001, Jinqing Yang, Qikai Cheng
J. Assoc. Inf. Sci. Technol.3
2021 How wide is the citation impact of scientific publications? A cross-discipline and large-scale analysis
Yi Bu 0001, Wei Lu 0019, Hongkan Chen, Yong Huang 0008
Inf. Process. Manag.5
2021 Detecting research topic trends by author-defined keyword frequency
Wei Lu 0019, Shengzhi Huang, Jinqing Yang, Yi Bu 0001, Qikai Cheng, Yong Huang 0008
Inf. Process. Manag.6
2020 Considering author sequence in all-author co-citation analysis
Yi Bu 0001, Binglu Wang, Zaida Chinchilla-Rodríguez, Cassidy R. Sugimoto, Yong Huang 0008, Win-Bin Huang
Inf. Process. Manag.5
2019 From zero to one: A perspective on citing
abstract
This article investigates the lengths of time that publications with different numbers of citations take to receive their first citation (the beginning stage), and then compares the lengths of time to receive two or more citations after receiving the first citation (the accumulative stage) in the field of computer science. We find that in the beginning stage, that is, from zero to one citation, high‐, medium‐, and low‐cited publications do not obviously exhibit different lengths of time. However, in the accumulative stage, that is, from one to N citations, highly cited publications begin to receive citations much more rapidly than medium‐ and low‐cited publications. Moreover, as N increases, the difference in receiving new citations among high‐, medium‐, and low‐cited publications increases quite significantly.
Yong Huang 0008, Yi Bu 0001, Ying Ding 0001, Wei Lu 0019
J. Assoc. Inf. Sci. Technol.1
2016 A study of factuality, objectivity and relevance: three desiderata in large-scale information retrieval?
abstract
Much of the information processed by Information Retrieval (IR) systems is unreliable, biased, and generally untrust-worthy [15, 45, 48]. Yet, factuality & objectivity detection is not a standard component of IR systems, even though it has been possible in Natural Language Processing (NLP) in the last decade. Motivated by this, we ask if and how factuality & objectivity detection may benefit IR. We answer this in two parts. First, we use state-of-the-art NLP to compute the probability of document factuality & objectivity in two TREC collections, and analyse its relation to document relevance. We find that factuality is strongly and positively correlated to document relevance, but objectivity is not. Second, we study the impact of factuality & objectivity to retrieval effectiveness by treating them as query independent features that we combine with a competitive language modelling baseline. Experiments with 450 TREC queries show that factuality improves precision by more than 10% over strong baselines, especially for the type of uncurated data typically used in web search; objectivity gives mixed results. An overall clear trend is that document factuality & objectivity is much more beneficial to IR when searching uncurated (e.g. web) documents vs. curated (e.g. state documentation and newswire articles).
Christina Lioma, Birger Larsen, Wei Lu 0019, Yong Huang 0008
BDCAT4