Weixue Lu

dblp:41/6582 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
3since 2021 · last 2023
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-authorDatabases, data management, data science and information retrieval · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1
YearPublicationVenuePosition
2023 Learning Discrete Document Representations in Web Search
abstract
Product quantization (PQ) has been usually applied to dense retrieval (DR) of documents thanks to its competitive time, memory efficiency and compatibility with other approximate nearest search (ANN) methods. Originally, PQ was learned to minimize the reconstruction loss, i.e., the distortions between the original dense embeddings and the reconstructed embeddings after quantization. Unfortunately, such an objective is inconsistent with the goal of selecting ground-truth documents for the input query, which may cause a severe loss of retrieval quality. Recent research has primarily concentrated on jointly training the biencoders and PQ to ensure consistency for improved performance. However, it is still difficult to design an approach that can cope with challenges like discrete representation collapse, mining informative negatives, and deploying effective embedding-based retrieval (EBR) systems in a real search engine.
Danfeng Zhang, Weixue Lu, Daiting Shi, Zhicong Cheng, Simiu Gu, Dawei Yin 0001
KDD3
2023 Pre-trained Language Model-based Retrieval and Ranking for Web Search
abstract
Pre-trained language representation models (PLMs) such as BERT and Enhanced Representation through kNowledge IntEgration (ERNIE) have been integral to achieving recent improvements on various downstream tasks, including information retrieval. However, it is nontrivial to directly utilize these models for the large-scale web search due to the following challenging issues: (1) the prohibitively expensive computations of massive neural PLMs, especially for long texts in the web document, prohibit their deployments in the web search system that demands extremely low latency; (2) the discrepancy between existing task-agnostic pre-training objectives and the ad hoc retrieval scenarios that demand comprehensive relevance modeling is another main barrier for improving the online retrieval and ranking effectiveness; and (3) to create a significant impact on real-world applications, it also calls for practical solutions to seamlessly interweave the resultant PLM and other components into a cooperative system to serve web-scale data. Accordingly, we contribute a series of successfully applied techniques in tackling these exposed issues in this work when deploying the state-of-the-art Chinese pre-trained language model, i.e., ERNIE, in the online search engine system. We first present novel practices to perform expressive PLM-based semantic retrieval with a flexible poly-interaction scheme and cost-efficiently contextualize and rank web documents with a cheap yet powerful Pyramid-ERNIE architecture. We then endow innovative pre-training and fine-tuning paradigms to explicitly incentivize the query-document relevance modeling in PLM-based retrieval and ranking with the large-scale noisy and biased post-click behavioral data. We also introduce a series of effective strategies to seamlessly interwoven the designed PLM-based models with other conventional components into a cooperative system. Extensive offline and online experimental results show that our proposed techniques are crucial to achieving more effective search performance. We also provide a thorough analysis of our methodology and experimental results.
Lixin Zou, Weixue Lu, Hengyi Cai, Xiaokai Chu, Dehong Ma, Daiting Shi, Yu Sun 0029, Zhicong Cheng, Simiu Gu, Shuaiqiang Wang, Dawei Yin 0001
ACM Trans. Web2
2021 Pre-trained Language Model for Web-scale Retrieval in Baidu Search
abstract
Retrieval is a crucial stage in web search that identifies a small set of query-relevant candidates from a billion-scale corpus. Discovering more semantically-related candidates in the retrieval stage is very promising to expose more high-quality results to the end users. However, it still remains non-trivial challenges of building and deploying effective retrieval models for semantic matching in real search engine. In this paper, we describe the retrieval system that we developed and deployed in Baidu Search. The system exploits the recent state-of-the-art Chinese pretrained language model, namely Enhanced Representation through kNowledge IntEgration (ERNIE), which facilitates the system with expressive semantic matching. In particular, we developed an ERNIE-based retrieval model, which is equipped with 1) expressive Transformer-based semantic encoders, and 2) a comprehensive multi-stage training paradigm. More importantly, we present a practical system workflow for deploying the model in web-scale retrieval. Eventually, the system is fully deployed into production, where rigorous offline and online experiments were conducted. The results show that the system can perform high-quality candidate retrieval, especially for those tail queries with uncommon demands. Overall, the new retrieval system facilitated by pretrained language model (i.e., ERNIE) can largely improve the usability and applicability of our search engine.
Weixue Lu, Suqi Cheng, Daiting Shi, Shuaiqiang Wang, Zhicong Cheng, Dawei Yin 0001
KDD2
2018 Recommendation with Multi-Source Heterogeneous Information
abstract
Network embedding has been recently used in social network recommendations by embedding low-dimensional representations of network items for recommendation. However, existing item recommendation models in social networks suffer from two limitations. First, these models partially use item information and mostly ignore important contextual information in social networks such as textual content and social tag information. Second, network embedding and item recommendations are learned in two independent steps without any interaction. To this end, we in this paper consider item recommendations based on heterogeneous information sources. Specifically, we combine item structure, textual content and tag information for recommendation. To model the multi-source heterogeneous information, we use two coupled neural networks to capture the deep network representations of items, based on which a new recommendation model Collaborative multi-source Deep Network Embedding (CDNE for short) is proposed to learn different latent representations. Experimental results on two real-world data sets demonstrate that CDNE can use network representation learning to boost the recommendation performance.
Hong Yang 0003, Jia Wu 0001, Chuan Zhou 0001, Weixue Lu, Yue Hu 0002
IJCAI5
2016 On the Minimum Differentially Resolving Set Problem for Diffusion Source Inference in Networks
abstract
In this paper we theoretically study the minimum Differentially Resolving Set (DRS) problem derived from the classical sensor placement optimization problem in network source locating. A DRS of a graph G = (V, E) is defined as a subset S ⊆ V where any two elements in V can be distinguished by their different differential characteristic sets defined on S. The minimum DRS problem aims to find a DRS S in the graph G with minimum total weight Σv∈S w(v). In this paper we establish a group of Integer Linear Programming (ILP) models as the solution. By the weighted set cover theory, we propose an approximation algorithm with the Θ(ln n) approximability for the minimum DRS problem on general graphs, where n is the graph size.
Chuan Zhou 0001, Weixue Lu, Peng Zhang 0001, Jia Wu 0001, Yue Hu 0002, Li Guo 0001
AAAI2
2016 Global and Local Influence-based Social Recommendation
abstract
Social recommendation has been widely studied in recent years. Existing social recommendation models use various explicit pieces of social information as regularization terms in recommendation, for instance, social links are considered as new constraints. However, social influence, an implicit source of information in social networks, is seldomly considered, even though it often drives recommendations in social networks. In this paper, we introduce a new global and local influence-based social recommendation model. Based on the observation that user purchase behaviour is influenced by both global influential nodes and the local influential nodes of the user, we formulate the global and local influence as an regularization terms, and incorporate them into a matrix factorization-based recommendation model. Experimental results on large data sets demonstrate the performance of the proposed method.
Qinzhe Zhang, Jia Wu 0001, Hong Yang 0003, Weixue Lu, Guodong Long, Chengqi Zhang
CIKM4
2016 A recursive method for big network influence estimation
abstract
Influence maximization aims to find a set of highly influential nodes in a social network to maximize the spread of influence. The most difficult part of the problem is to estimate the influence spread of any seed set, which has been proved to be #P-hard. There is no efficient method to estimate the influence spread of any seed set till now. Thus, the most common way to obtain the approximate influence spread is Monte Carlo simulation and two popular simulating strategies are applied: one is propagation strategy, the other is snapshot strategy. The former only fits for particular seed set and the latter incurs heavy memory cost. In this paper, we present a new algorithm to estimate the influence spread of any seed set. Our algorithm recursively estimates the influence spread using reachable probabilities from node to node. Accordingly, we provide three strategies to start the recursion by integrating the memory cost and computing efficiency. Experiments demonstrate high performance of our influence estimation.
Weixue Lu, Chuan Zhou 0001, Jia Wu 0001, Chun-Yi Liu 0003, Yue Hu 0002
IJCNN1
2016 Big social network influence maximization via recursively estimating influence spread
Weixue Lu, Chuan Zhou 0001, Jia Wu 0001
Knowl. Based Syst.1
2015 Influence Maximization in Big Networks: An Incremental Algorithm for Streaming Subgraph Influence Spread Estimation
Weixue Lu, Peng Zhang 0001, Chuan Zhou 0001, Chun-Yi Liu 0003
IJCAI1
1995 Adaptive image matching via spatial varying gray-level correction
abstract
This paper deals with the adaptive image matching and displacement estimation problems of dissimilar sequential images. Adaptive matching may be referred to a kind of functional optimization. Based on the additive measuring model for sequential images with gray-level deviation and the finite element technique, a class of gray-level correction model and robust matching criterion function are proposed, and an adaptive hierarchical searching algorithm is studied, which can automatically evaluate the temporal and spatial gray-level deviation existed in sequential images and filter out the interference thereof to increase the probability of correct estimation of displacement. Experimental results for the estimation of displacement fields with translation and/or small angle rotation in medical digital subtraction angiography (DSA) imaging and aerial remote sensing pictures are presented.>
Tianxu Zhang, Weixue Lu, Jiaxiong Peng
IEEE Trans. Commun.2
1992 Multiobjective optimization approach to image reconstruction from projections
Yuanmei Wang, Weixue Lu
Signal Process.2
1992 Multicriterion maximum entropy image reconstruction from projections
abstract
A solution algorithm for the image reconstruction problem with three criteria, maximum entropy, minimum nonuniformity and peakedness, and least square error between the original projection data and projection due to reconstruction is presented. Theoretical results of precedence properties which are respected by all noninferior solutions are first derived. These precedence properties are then incorporated into a multiple-criteria optimization framework to improve the computational efficiency. Comparisons of the new algorithm to the MART and MENT algorithms are carried out using computer-generated noise-free and Gaussian noisy projections. Results of the computational experiment and the efficiency of the multiobjective entropy optimization algorithm (MEOA) are reported.
Yuanmei Wang, Weixue Lu
IEEE Trans. Medical Imaging2
1992 Correspondence analysis for regional tracking in coronary arteriograms
abstract
A new method of correspondence analysis on the basis of geometrical similarity in tracking the region of interest (ROI) of coronary arteriograms is described. The method is insensitive to scaling, rotation, and gray-level variation. Experiments also show that the algorithm is valid even when large deformations exist. It is an important step toward a new approach to regional myocardial dysfunction analysis.
Weixin Xia, Weixue Lu
IEEE Trans. Medical Imaging2
1988 An optimized searching algorithm for image matching
abstract
It is shown that the statistical properties of the mean absolute difference mismatch measure surface is dependent on the autocorrelation function of image and signal-to-noise ratio. A method for finding the match point is presented, based on the idea of testing a set of hypotheses. Using the threshold estimation formulas and the optimized search scheme developed by the authors, a fast algorithm which may be adaptable to various expected operating conditions can be implemented.>
Tianxu Zhang, Jiaxiong Peng, Weixue Lu
ICPR3
1988 A new digital image registration algorithm based on the double spatial intensity gradients using pyramids
Zhixin Jiang, Weixue Lu, Dake Hu
Pattern Recognit. Lett.2