Wei Lu 0019

dblp:98/6613-19 · DBLP profile ↗
← Back
32ranked-venue papers in the field
4as first author
23since 2021 · last 2025
0000-0002-0929-7416ORCID · conflict

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 29 (3 first)Knowledge Engineering, Semantic Web & Information Systems · 2 (1 first)Big Data, Cloud & Distributed Data Systems · 1
YearPublicationVenuePosition
2025 Towards Better Evaluating Multi-Query Sessions: A Measure Based on the Theory of Planned Behavior
abstract
To evaluate multi-query sessions, recent studies usually add a second ''session'' dimension to the query-level evaluation framework, deriving corresponding session-version evaluation metrics such as sDCG, sRBP, and sINST. However, these existing metrics do not sufficiently consider the different impacts of users' expectations of gains and costs on their behaviors such as query reformulation, nor the bounded rationality characteristic of users in expectation management. To address these issues and better explain user behavior in multi-query sessions, we design a user model based on the Theory of Planned Behavior (TPB), which links user expectations to user behaviors. Within the TPB framework, we propose sTPB, a new measure that adapts to users' expectation management modes by considering users' expectations of gains and costs. To demonstrate the effectiveness of sTPB in evaluating multi-query sessions, we compare it with existing session metrics on two publicly available user search behavior datasets. The results show that sTPB significantly outperforms other metrics in terms of both fitting user behavior and measuring user satisfaction. Additionally, we explore the differences between optimal parameters under different user characteristics and task types in session search evaluation. We find that different user characteristics and task types lead to various preferences in users' choices between continuing to examine results and reformulating queries. Our study not only validates the effectiveness of sTPB in evaluating multi-query sessions but also highlights the necessity of considering the influence of user characteristics and task types when designing metrics.
Fan Zhang 0053, Jia Chen 0003, Wei Lu 0019
SIGIR4
2025 Are large language models qualified reviewers in originality evaluation?
Shengzhi Huang, Yong Huang 0008, Yinpeng Liu, Zhuoran Luo, Wei Lu 0019
Inf. Process. Manag.5
2025 Identifying potentially disruptive research via a comparative power-based large model
Shengzhi Huang, Wei Lu 0019, Zhenzhen Xu, Qikai Cheng, Jinqing Yang, Yong Huang 0008
Inf. Process. Manag.2
2025 PaperEval: A universal, quantitative, and explainable paper evaluation method powered by a multi-agent system
Shengzhi Huang, Qicong Wang, Wei Lu 0019, Lingyu Liu, Zhenzhen Xu, Yong Huang 0008
Inf. Process. Manag.3
2024 Predicting Scientific Impact Through Diffusion, Conformity, and Contribution Disentanglement
abstract
The scientific impact of academic papers is influenced by intricate factors such as dynamic popularity and inherent contribution. Existing models typically rely on static graphs for citation count estimation, failing to differentiate among its sources. In contrast, we propose distinguishing effects derived from various factors and predicting citation increments as estimated potential impacts within the dynamic context. In this research, we introduce a novel model, DPPDCC, which Disentangles the Potential impacts of Papers into Diffusion, Conformity, and Contribution values. It encodes temporal and structural features within dynamic heterogeneous graphs derived from the citation networks and applies various auxiliary tasks for disentanglement. By emphasizing comparative and co-cited/citing information and aggregating snapshots evolutionarily, DPPDCC captures knowledge flow within the citation network. Afterwards, popularity is outlined by contrasting augmented graphs to extract the essence of citation diffusion and predicting citation accumulation bins for quantitative conformity modeling. Orthogonal constraints ensure distinct modeling of each perspective, preserving the contribution value. To gauge generalization across publication times and replicate the realistic dynamic context, we partition data based on specific time points and retain all samples without strict filtering. Extensive experiments on three datasets validate DPPDCC's superiority over baselines for papers published previously, freshly, and immediately, with further analyses confirming its robustness. Our codes and supplementary materials can be found at GitHub (https://github.com/ECNU-Text-Computing/DPPDCC).
Zhikai Xue, Guoxiu He, Zhuoren Jiang, Sichen Gu, Yangyang Kang, Star Zhao, Wei Lu 0019
CIKM7
2024 Evolutions of semantic consistency in research topic via contextualized word embedding
Shengzhi Huang, Wei Lu 0019, Qikai Cheng, Zhuoran Luo, Yong Huang 0008
Inf. Process. Manag.2
2023 H2CGL: Modeling dynamics of citation network for impact prediction
Guoxiu He, Zhikai Xue, Zhuoren Jiang, Yangyang Kang, Star Zhao, Wei Lu 0019
Inf. Process. Manag.6
2023 From "what" to "how": Extracting the Procedural Scientific Information Toward the Metric-optimization in AI
Jiawei Liu 0002, Wei Lu 0019, Qikai Cheng
Inf. Process. Manag.3
2023 A term function-aware keyword citation network method for science mapping analysis
Qikai Cheng, Wei Lu 0019, Yongxiang Dou, Pengcheng Li 0012
Inf. Process. Manag.3
2023 Re-examining lexical and semantic attention: Dual-view graph convolutions enhanced BERT for academic paper rating
Zhikai Xue, Guoxiu He, Jiawei Liu 0002, Zhuoren Jiang, Star Zhao, Wei Lu 0019
Inf. Process. Manag.6
2023 An effective method for figures and tables detection in academic literature
Fengchang Yu, Jiani Huang 0001, Zhuoran Luo, Li Zhang 0093, Wei Lu 0019
Inf. Process. Manag.5
2023 A multiple long short-term model for product sales forecasting based on stage future vision with prior knowledge
Daifeng Li, Xuting Li, Kaixin Lin, Jianbin Liao, Ruo Du, Wei Lu 0019, Andrew D. Madden
Inf. Sci.6
2023 Generating keyphrases for readers: A controllable keyphrase generation framework
abstract
Abstract With the wide application of keyphrases in many Information Retrieval (IR) and Natural Language Processing (NLP) tasks, automatic keyphrase prediction has been emerging. However, these statistically important phrases are contributing increasingly less to the related tasks because the end‐to‐end learning mechanism enables models to learn the important semantic information of the text directly. Similarly, keyphrases are of little help for readers to quickly grasp the paper's main idea because the relationship between the keyphrase and the paper is not explicit to readers. Therefore, we propose to generate keyphrases with specific functions for readers to bridge the semantic gap between them and the information producers, and verify the effectiveness of the keyphrase function for assisting users’ comprehension with a user experiment. A controllable keyphrase generation framework (the CKPG) that uses the keyphrase function as a control code to generate categorized keyphrases is proposed and implemented based on Transformer, BART, and T5, respectively. For the Computer Science domain, the Macro‐avgs of , , and on the Paper with Code dataset are up to 0.680, 0.535, and 0.558, respectively. Our experimental results indicate the effectiveness of the CKPG models.
Yong Huang 0008, Wei Lu 0019, Jiawei Liu 0002
J. Assoc. Inf. Sci. Technol.4
2023 LAGOS-AND: A large gold standard dataset for scholarly author name disambiguation
abstract
Abstract In this article, we present a method to automatically build large labeled datasets for the author ambiguity problem in the academic world by leveraging the authoritative academic resources, ORCID and DOI. Using the method, we built LAGOS‐AND, two large, gold‐standard sub‐datasets for author name disambiguation (AND), of which LAGOS‐AND‐BLOCK is created for clustering‐based AND research and LAGOS‐AND‐PAIRWISE is created for classification‐based AND research. Our LAGOS‐AND datasets are substantially different from the existing ones. The initial versions of the datasets (v1.0, released in February 2021) include 7.5 M citations authored by 798 K unique authors (LAGOS‐AND‐BLOCK) and close to 1 M instances (LAGOS‐AND‐PAIRWISE). And both datasets show close similarities to the whole Microsoft Academic Graph (MAG) across validations of six facets. In building the datasets, we reveal the variation degrees of last names in three literature databases, PubMed, MAG, and Semantic Scholar, by comparing author names hosted to the authors' official last names shown on the ORCID pages. Furthermore, we evaluate several baseline disambiguation methods as well as the MAG's author IDs system on our datasets, and the evaluation helps identify several interesting findings. We hope the datasets and findings will bring new insights for future studies. The code and datasets are publicly available.
Li Zhang 0093, Wei Lu 0019, Jinqing Yang
J. Assoc. Inf. Sci. Technol.2
2022 A comparative study of automated legal text classification using random forests and deep learning
Haihua Chen 0002, Jiangping Chen, Wei Lu 0019, Junhua Ding 0001
Inf. Process. Manag.4
2022 Fine-grained citation count prediction via a transformer-based model with among-attention mechanism
Shengzhi Huang, Yong Huang 0008, Yi Bu 0001, Wei Lu 0019, Jiajia Qian
Inf. Process. Manag.4
2022 Revisiting the exploration-exploitation behavior of scholars' research topic selection: Evidence from a large-scale bibliographic database
Shengzhi Huang, Wei Lu 0019, Yi Bu 0001, Yong Huang 0008
Inf. Process. Manag.2
2022 Towards transdisciplinary impact of scientific publications: A longitudinal, comprehensive, and large-scale analysis on Microsoft Academic Graph
Yong Huang 0008, Wei Lu 0019, Qikai Cheng, Yi Bu 0001
Inf. Process. Manag.2
2022 How humans obtain information from AI: Categorizing user messages in human-AI collaborative conversations
Yuhan Wei, Wei Lu 0019, Qikai Cheng, Tingting Jiang 0002, Shewei Liu
Inf. Process. Manag.2
2022 A novel emerging topic detection method: A knowledge ecology perspective
Jinqing Yang, Wei Lu 0019, Jiming Hu, Shengzhi Huang
Inf. Process. Manag.2
2022 Disclosing the relationship between citation structure and future impact of a publication
abstract
Abstract Each section header of an article has its distinct communicative function. Citations from distinct sections may be different regarding citing motivation. In this paper, we grouped section headers with similar functions as a structural function and defined the distribution of citations from structural functions for a paper as its citation structure. We aim to explore the relationship between citation structure and the future impact of a publication and disclose the relative importance among citations from different structural functions. Specifically, we proposed two citation counting methods and a citation life cycle identification method, by which the regression data were built. Subsequently, we employed a ridge regression model to predict the future impact of the paper and analyzed the relative weights of regressors. Based on documents collected from the Association for Computational Linguistics Anthology website, our empirical experiments disclosed that functional structure features improve the prediction accuracy of citation count prediction and that there exist differences among citations from different structural functions. Specifically, at the early stage of citation lifetime, citations from Introduction and Method are particularly important for perceiving future impact of papers, and citations from Result and Conclusion are also vital. However, early accumulation of citations from the Background seems less important.
Shengzhi Huang, Jiajia Qian, Yong Huang 0008, Wei Lu 0019, Yi Bu 0001, Jinqing Yang, Qikai Cheng
J. Assoc. Inf. Sci. Technol.4
2021 How wide is the citation impact of scientific publications? A cross-discipline and large-scale analysis
Yi Bu 0001, Wei Lu 0019, Hongkan Chen, Yong Huang 0008
Inf. Process. Manag.2
2021 Detecting research topic trends by author-defined keyword frequency
Wei Lu 0019, Shengzhi Huang, Jinqing Yang, Yi Bu 0001, Qikai Cheng, Yong Huang 0008
Inf. Process. Manag.1
2020 Think Beyond the Word: Understanding the Implied Textual Meaning by Digesting Context, Local, and Noise
abstract
Implied semantics is a complex language act that can appear everywhere on the Cyberspace. The prevalence of implied spam texts, such as implied pornography, sarcasm, and abuse hidden within the novel, tweet, microblog, or review, can be extremely harmful to the physical and mental health of teenagers. The non-literal interpretation of the implied text is hard to be understood by machine models due to its high context-sensitivity and heavy usage of figurative language. In this study, inspired by human reading comprehension, we propose a novel, simple, and effective deep neural framework, called Skim and Intensive Reading Model (SIRM), for figuring out implied textual meaning. The proposed SIRM consists of three main components, namely the skim reading component, intensive reading component, and adversarial training component. N-gram features are quickly extracted from the skim reading component, which is a combination of several convolutional neural networks, as skim (entire) information. An intensive reading component enables a hierarchical investigation for both sentence-level and paragraph-level representation, which encapsulates the current (local) embedding and the contextual information (context) with a dense connection. More specifically, the contextual information includes the near-neighbor information and the skim information mentioned above. Finally, besides the common training loss function, we employ an adversarial loss function as a penalty over the skim reading component to eliminate noisy information (noise) arisen from special figurative words in the training data. To verify the effectiveness, robustness, and efficiency of the proposed architecture, we conduct extensive comparative experiments on an industrial novel dataset involving implied pornography and three sarcasm benchmarks. Experimental results indicate that (1) the proposed model, which benefits from context and local modeling and consideration of figurative language (noise), outperforms existing state-of-the-art solutions, with comparable parameter scale and running speed; (2) the SIRM yields superior robustness in terms of parameter size sensitivity; (3) compared with ablation and addition variants of the SIRM, the final framework is efficient enough.
Guoxiu He, Zhuoren Jiang, Yangyang Kang, Changlong Sun, Xiaozhong Liu 0001, Wei Lu 0019
SIGIR7
2020 Creating a Children-Friendly Reading Environment via Joint Learning of Content and Human Attention
abstract
Technological advancements have led to increasing availability of erotic literature and pornography novels online, which can be alluring to adolescence and children. Unfortunately, because of the inherent complexity of these indecent contents and training data sparseness, it is a challenging task to detect these readings in the Cyberspace while children can easily access them. In this study, we propose a novel framework, Joint Learning of Content and Human Attention (GoodMan), to identify indecent readings by augmenting natural language understanding models with large scale human reading behaviors (dwell time per page) on portable devices. From the text modeling viewpoint, the innovative joint attention trained by joint learning is employed to orchestrate the content attention and human behavior attention via the BiGRU. From the data augmentation perspective, various users' reading behaviors on the same text can generate considerable training instances with joint attention, which can be effective to address the cold start problem. We conduct an extensive set of experiments on an online ebook dataset (with human reading behaviors on portable devices). The experimental results show insights into the task and demonstrate the superiority of the proposed model against alternative solutions.
Guoxiu He, Yangyang Kang, Zhuoren Jiang, Jiawei Liu 0002, Changlong Sun, Xiaozhong Liu 0001, Wei Lu 0019
SIGIR7
2020 Analyzing the topic distribution and evolution of foreign relations from parliamentary debates: A framework and case study
Wei Lu 0019, Jiming Hu
Inf. Process. Manag.1
2019 Finding Camouflaged Needle in a Haystack?: Pornographic Products Detection via Berrypicking Tree Model
abstract
It is an important and urgent research problem for decentralized eCommerce services, e.g., eBay, eBid, and Taobao, to detect illegal products, e.g., unclassified pornographic products. However, it is a challenging task as some sellers may utilize and change camouflaged text to deceive the current detection algorithms. In this study, we propose a novel task to dynamically locate the pornographic products from very large product collections. Unlike prior product classification efforts focusing on textual information, the proposed model, BerryPIcking TRee MoDel (BIRD), utilizes both product textual content and buyers' seeking behavior information as berrypicking trees. In particular, the BIRD encodes both semantic information with respect to all branches sequence and the overall latent buyer intent during the whole seeking process. An extensive set of experiments have been conducted to demonstrate the advantage of the proposed model against alternative solutions. To facilitate further research of this practical and important problem, the codes and buyers' seeking behavior data have been made publicly available1.
Guoxiu He, Yangyang Kang, Zhuoren Jiang, Changlong Sun, Xiaozhong Liu 0001, Wei Lu 0019, Luo Si
SIGIR7
2019 Result diversification in image retrieval based on semantic distance
Wei Lu 0019, Mengqi Luo, Guobiao Zhang, Heng Ding, Haihua Chen 0002, Jiangping Chen
Inf. Sci.1
2019 From zero to one: A perspective on citing
abstract
This article investigates the lengths of time that publications with different numbers of citations take to receive their first citation (the beginning stage), and then compares the lengths of time to receive two or more citations after receiving the first citation (the accumulative stage) in the field of computer science. We find that in the beginning stage, that is, from zero to one citation, high‐, medium‐, and low‐cited publications do not obviously exhibit different lengths of time. However, in the accumulative stage, that is, from one to N citations, highly cited publications begin to receive citations much more rapidly than medium‐ and low‐cited publications. Moreover, as N increases, the difference in receiving new citations among high‐, medium‐, and low‐cited publications increases quite significantly.
Yong Huang 0008, Yi Bu 0001, Ying Ding 0001, Wei Lu 0019
J. Assoc. Inf. Sci. Technol.4
2016 A study of factuality, objectivity and relevance: three desiderata in large-scale information retrieval?
abstract
Much of the information processed by Information Retrieval (IR) systems is unreliable, biased, and generally untrust-worthy [15, 45, 48]. Yet, factuality & objectivity detection is not a standard component of IR systems, even though it has been possible in Natural Language Processing (NLP) in the last decade. Motivated by this, we ask if and how factuality & objectivity detection may benefit IR. We answer this in two parts. First, we use state-of-the-art NLP to compute the probability of document factuality & objectivity in two TREC collections, and analyse its relation to document relevance. We find that factuality is strongly and positively correlated to document relevance, but objectivity is not. Second, we study the impact of factuality & objectivity to retrieval effectiveness by treating them as query independent features that we combine with a competitive language modelling baseline. Experiments with 450 TREC queries show that factuality improves precision by more than 10% over strong baselines, especially for the type of uncurated data typically used in web search; objectivity gives mixed results. An overall clear trend is that document factuality & objectivity is much more beneficial to IR when searching uncurated (e.g. web) documents vs. curated (e.g. state documentation and newswire articles).
Christina Lioma, Birger Larsen, Wei Lu 0019, Yong Huang 0008
BDCAT3
2012 Rhetorical relations for information retrieval
abstract
Typically, every part in most coherent text has some plausible reason for its presence, some function that it performs to the overall semantics of the text. Rhetorical relations, e.g. contrast, cause, explanation, describe how the parts of a text are linked to each other. Knowledge about this so-called discourse structure has been applied successfully to several natural language processing tasks. This work studies the use of rhetorical relations for Information Retrieval (IR): Is there a correlation between certain rhetorical relations and retrieval performance? Can knowledge about a document's rhetorical relations be useful to IR? We present a language model modification that considers rhetorical relations when estimating the relevance of a document to a query. Empirical evaluation of different versions of our model on TREC settings shows that certain rhetorical relations can benefit retrieval effectiveness notably (>10% in mean average precision over a state-of-the-art baseline).
Christina Lioma, Birger Larsen, Wei Lu 0019
SIGIR3
2012 Fixed versus dynamic co-occurrence windows in TextRank term weights for information retrieval
abstract
TextRank is a variant of PageRank typically used in graphs that represent documents, and where vertices denote terms and edges denote relations between terms. Quite often the relation between terms is simple term co-occurrence within a fixed window of k terms. The output of TextRank when applied iteratively is a score for each vertex, i.e. a term weight, that can be used for information retrieval (IR) just like conventional term frequency based term weights.
Wei Lu 0019, Qikai Cheng, Christina Lioma
SIGIR1