Dezhi Ye

dblp:251/9672 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
7since 2021 · last 2026
0000-0001-9776-2404ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 5 first-author · 6 since 2021Databases, data management, data science and information retrieval · 8 · 5 first-author · 6 since 2021Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TFRank: Think-Free Reasoning Enables Practical Pointwise LLM Ranking
abstract
Reasoning-intensive ranking models built on Large Language Models (LLMs) have made notable progress. However, existing approaches often rely on large-scale LLMs and explicit Chain-of-Thought (CoT) reasoning, resulting in high computational cost and latency that limit real-world use. To address this, we propose TFRank, an efficient pointwise reasoning ranker based on small-scale LLMs. To improve ranking performance, TFRank effectively integrates CoT data, fine-grained score supervision, and multi-task training. Furthermore, it achieves an efficient "Think-Free" reasoning capability by employing a "think-mode switch" and pointwise format constraints. Specifically, this allows the model to leverage explicit reasoning during training while delivering precise relevance scores for complex queries at inference without generating any reasoning chains. Experiments show that TFRank achieves performance comparable to models with four times more parameters on the BRIGHT benchmark, and demonstrates strong competitiveness on the BEIR benchmark. Further analysis shows that TFRank achieves an effective balance between performance and efficiency, providing a practical solution for integrating advanced reasoning into real-world systems.
Yongqi Fan, Xiaoyang Chen 0001, Dezhi Ye, Jie Liu 0075, Haijin Liang, Jin Ma 0003, Ben He 0001, Yingfei Sun, Tong Ruan
AAAI3
2025 Can LLMs Really Help Query Understanding In Web Search? A Practical Perspective
abstract
As a core module of web search, query understanding aims to bridge the semantic gap between user queries and web page documents, thereby enhancing the ability to deliver more relevant results.Recently, Large Language Models (LLMs) have achieved significant breakthroughs that have fundamentally altered the workflow of existing search ranking tasks.However, few researchers have explored the integration of LLMs into the field of query understanding.In this paper, we investigate the potential of LLMs in query understanding by conducting a comprehensive evaluation across three dimensions: term, structure, and topic.This evaluation includes several representative tasks such as segmentation, term weighting, error correction, query expansion, and intent recognition.The experimental results reveal that LLMs are particularly effective in query expansion and intent recognition but show limited improvement in other areas.This limitation may be attributed to LLMs' primary focus on modeling the semantic knowledge of entire queries, while lacking the capability to capture token-level information with finer granularity.Additionally, we explore potential practical applications of LLMs in query understanding, such as integrating the evaluation and training capabilities of smaller models with LLMs and constructing unsupervised samples.Based on comprehensive empirical results, collaborative training emerges as a promising approach to leverage LLMs for query understanding.We hope this research will advance the practical application of LLMs in query understanding and contribute to the development of this field.
Dezhi Ye, Ye Qin, Jiabin Fan, Jie Liu 0075, Haijin Liang, Jin Ma 0003
CIKM1
2025 Applying Large Language Model For Relevance Search In Tencent
abstract
Relevance plays a crucial role in commercial search engines by identifying documents related to user queries and fulfilling their search needs.Traditional approaches employ encoder-only models like BERT, which process concatenated query-document pairs to predict relevance scores.While autoregressive large language models (LLMs) have revolutionized numerous NLP domains, their direct application to web-scale search systems presents significant challenges.On one hand, the relevance modeling capabilities of LLMs have not been fully explored.On the other, the high computational costs and inference times make deploying LLMs in online search systems, which demand extremely low latency, nearly impossible.In this work, we address these challenges through two key contributions.First, we develop a comprehensive evaluation framework to systematically assess the effectiveness of LLMs in query-document relevance ranking.By conducting assessment experiments to LLMs in four perspectives: ranking objectives, model size, domain-specific continuous pre-training, and the integration of prior knowledge, we identify the best resource allocation strategy given a restricted budget and develop practical LLMs in a more efficient way.Second, we propose a novel framework to transfer the capabilities of LLMs in the ranking aspect to existing BERT models to avoid directly deploying LLMs.Finally, to fully leverage the improvements in relevance ranking brought by LLMs, we successfully nearline deploy LLMs in Tencent QQ Browser search engine using query-based ondemand computing and quantization.Experiments on real-world datasets and online A/B tests demonstrate that our approach significantly enhances search engine performance while maintaining practical operational efficiency.Our findings provide actionable insights for integrating LLMs into production search engines.
Dezhi Ye, Jie Liu 0075, Jiabin Fan, Haijin Liang, Jin Ma 0003
KDD (2)1
2024 Span Confusion is All You Need for Chinese Spelling Correction
Dezhi Ye, Haomei Jia, Jie Liu 0075, Haijin Liang, Jin Ma 0003, Wenmin Wang 0001
CIKM1
2024 Enhancing Asymmetric Web Search through Question-Answer Generation and Ranking
abstract
This paper addresses the challenge of the semantic gap between user queries and web content, commonly referred to as asymmetric text matching, within the domain of web search.By leveraging BERT for reading comprehension, current algorithms enable significant advancements in query understanding, but still encounter limitations in effectively resolving the asymmetrical ranking problem due to model comprehension and summarization constraints.To tackle this issue, we propose the QAGR (Question-Answer Generation and Ranking) method, comprising an offline module called QAGeneration and an online module called QARanking.The QAGeneration module utilizes large language models (LLMs) to generate high-quality question-answering pairs for each web page.This process involves two steps: generating question-answer pairs and performing verification to eliminate irrelevant questions, resulting in high-quality questions associated with their respective documents.The QARanking module combines and ranks the generated questions and web page content.To ensure efficient online inference, we design the QARanking model as a homogeneous dual-tower model, incorporating query intent to drive score fusion while balancing keyword matching and asymmetric matching.Additionally, we conduct a preliminary screening of questions for each document, selecting only the top-N relevant questions for further relevance calculation.Empirical results demonstrate the substantial performance improvement of our proposed method in web search.We achieve over
Dezhi Ye, Jie Liu 0075, Jiabin Fan, Tianhua Zhou, Jin Ma 0003
KDD1
2023 Improving Query Correction Using Pre-train Language Model In Search Engines
abstract
Query correction is a task that automatically detects and corrects errors in what users type into a search engine. Misspelled queries can lead to user dissatisfaction and churn. However, correcting a user query accurately is not an easy task. One major challenge is that a correction model must be capable of high-level language comprehension. Recently, pre-trained language models (PLMs) have been successfully applied to text correction tasks, but few works have been done on query correction. However, it is nontrivial to directly apply these PLMs to query correction in large-scale search systems due to the following challenging issues: 1) Expensive deployment. Deploying such a model requires expensive computations. 2) Lacking domain knowledge. A neural correction model needs massive training data to activate its power.
Dezhi Ye, Jiabin Fan, Jie Liu 0075, Tianhua Zhou, Jin Ma 0003
CIKM1
2022 CFFNN: Cross Feature Fusion Neural Network for Collaborative Filtering
abstract
Numerous state-of-the-art recommendation frameworks employ deep neural networks in Collaborative Filtering (CF). In this paper, we propose a cross feature fusion neural network (CFFNN) for the enhancement of CF. Existing studies overlook either user preferences for various item features or the relationship between item features and user features. To solve this problem, we construct a cross feature fusion network to enable the fusion of user features and item features as well as a self-attention network to determine users’ preferences for items. Specifically, we design a feature extraction layer with multiple MLP (Multilayer Perceptrons) modules to extract both user features and item features. Then, we introduce a cross feature fusion mechanism for an accurate determination of the relationship between different user-item interactions. The features of users and items are crossly embedded and then fed into a prediction network. The attention mechanism enables the model to focus on more effective features. The effectiveness of CFFNN model is demonstrated through extensive experiments on four real-world datasets. The experimental results indicate that CFFNN significantly outperforms the existing state-of-the-art models, with a relative improvement of 3.0 to 12.1 percent on hit ratio (HR) and normalized discounted cumulative gain (NDCG) compared with the baselines.
Ruiyun Yu, Dezhi Ye, Biyun Zhang, Ann Move Oguti, Jie Li 0008, Bo Jin 0001, Fadi J. Kurdahi
IEEE Trans. Knowl. Data Eng.2
2020 RePiDeM: A Refined POI Demand Modeling based on Multi-Source Data*
abstract
Point-of-Interest (POI) demand modeling in urban regions is critical for building smart cities with various applications, e.g., business location selection and urban planning. However, existing work does not fully utilize human mobility data and ignores the interactive-aware information. In this work, we design a refined POI demand modeling framework, named RePiDeM, to identify region POI demands based on multi-source data, including cellular data, POI data, satellite image, geographic data, etc. Specifically, we introduce a Cellular Data (CD) based visit inference algorithm to estimate the POI visit probability based on human mobility and POI data. Further, to address the data sparsity issue, we design a multi-source attention neural collaborative filtering (MANCF) model to output region POI demands considering various aspect attention. We conduct extensive experiments on real-world data collected in the Chinese city Shenyang, which show that RePiDeM is effective for modeling region POI demands.
Ruiyun Yu, Dezhi Ye, Jie Li 0008
INFOCOM2
2020 OptMatch: Optimized Matchmaking via Modeling the High-Order Interactions on the Arena
abstract
Matchmaking is a core problem for the e-sports and online games, which determines the player satisfaction and further influences the life cycle of the gaming products. Most of matchmaking systems take the form of grouping the queuing players into two opposing teams by following certain rules. The design and implementation of matchmaking systems are usually product-specific and labor-intensive.
Linxia Gong, Xiaochuan Feng, Dezhi Ye, Runze Wu 0001, Jianrong Tao, Changjie Fan, Peng Cui 0001
KDD3
2019 GMTL: A GART Based Multi-task Learning Model for Multi-Social-Temporal Prediction in Online Games
abstract
Multi-social-temporal (MST) data, which represent multi-attributed time series corresponding to the entities in multi-relational social network series, are ubiquitous in real-world and virtual-world dynamic systems, such as online games. Predictions over MST data such as social time series prediction and temporal link weight prediction are of great importance but challenging. They are affected by many complex factors, including temporal characteristics, social characteristics, collaborative characteristics, task characteristics and the intrinsic causality between them. In this paper, we propose a graph attention recurrent network (GART) based multi-task learning model (GMTL) to fuse information across multiple social-temporal prediction tasks. Experiments on an MMORPG dataset demonstrate that GMTL outperforms the state-of-the-art baselines and can significantly improve performances of specific social-temporal prediction task with additional information from others. Our work has been deployed to several MMORPGs in practice and can also expand to many related multi-social-temporal prediction tasks in real-world applications. Case studies on applications for multi-social-temporal prediction show that GMTL produces great value in the actual business in NetEase Games.
Jianrong Tao, Linxia Gong, Changjie Fan, Longbiao Chen, Dezhi Ye, Sha Zhao
CIKM5