Yukun Zheng

dblp:204/0083 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
3since 2021 · last 2022
0000-0003-0096-0979ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 10 · 4 first-author · 2 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 3 since 2021Security and privacy · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
8 papers
Information retrieval · 84% Recommender systems · 6% Web and social media mining · 4%
Artificial intelligence
1 paper
Graph learning · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 100%

Topics — the 26 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › user behavior › search behavior
click model
0.822020
Investigating Examination Behavior in Mobile Search · WSDM 2020
Constructing Click Model for Mobile Search with Viewport Time · ACM Trans. Inf. Syst. 2019
Information retrieval › evaluation
test collection
0.622018
Sogou-QCL: A New Dataset with Click Relevance Label · SIGIR 2018
SogouT-16: A New Web Corpus to Embrace IR Research · SIGIR 2017
Machine learning › Graph learning
graph neural network
0.612022
Joint Learning of E-commerce Search and Recommendation with a Unified Graph Neural Network · WSDM 2022
Machine learning › Graph learning › graph neural network
heterogeneous graph neural network
0.612022
Joint Learning of E-commerce Search and Recommendation with a Unified Graph Neural Network · WSDM 2022
Recommender systems
click-through rate prediction
0.612022
Joint Learning of E-commerce Search and Recommendation with a Unified Graph Neural Network · WSDM 2022
Information retrieval
search and recommendation
0.612022
Joint Learning of E-commerce Search and Recommendation with a Unified Graph Neural Network · WSDM 2022
Information retrieval › web search
mobile search
0.522020
Constructing Click Model for Mobile Search with Viewport Time · ACM Trans. Inf. Syst. 2019
Investigating Examination Behavior in Mobile Search · WSDM 2020
Information retrieval › evaluation › effectiveness metrics
cumulative gain
0.412020
Leveraging Passage-level Cumulative Gain for Document Ranking · WWW 2020
Information retrieval › ranking › text ranking
document ranking
0.412020
Leveraging Passage-level Cumulative Gain for Document Ranking · WWW 2020
Information retrieval › ranking › text ranking
passage ranking
0.412020
Leveraging Passage-level Cumulative Gain for Document Ranking · WWW 2020
Information retrieval
retrieval models
0.412020
Leveraging Passage-level Cumulative Gain for Document Ranking · WWW 2020
Information retrieval
user interaction
0.412020
Investigating Examination Behavior in Mobile Search · WSDM 2020
Performance modeling and evaluation
benchmarking
0.412020
Database Benchmarking for Supporting Real-Time Interactive Querying of Large Data · SIGMOD Conference 2020
Performance modeling and evaluation › benchmarking
database system benchmarking
0.412020
Database Benchmarking for Supporting Real-Time Interactive Querying of Large Data · SIGMOD Conference 2020
Information retrieval › user behavior
reading behavior
0.412019
Human Behavior Inspired Machine Reading Comprehension · SIGIR 2019
Information retrieval › ranking
relevance estimation
0.412019
Constructing Click Model for Mobile Search with Viewport Time · ACM Trans. Inf. Syst. 2019
Web and social media mining › user behavior analysis
user behavior modeling
0.412019
Constructing Click Model for Mobile Search with Viewport Time · ACM Trans. Inf. Syst. 2019
Information retrieval
evaluation
0.312018
Sogou-QCL: A New Dataset with Click Relevance Label · SIGIR 2018
Information retrieval › retrieval models › neural retrieval
neural ranking model
0.312018
Sogou-QCL: A New Dataset with Click Relevance Label · SIGIR 2018
Information retrieval
ranking
0.312018
Sogou-QCL: A New Dataset with Click Relevance Label · SIGIR 2018
Machine learning and data management
weak supervision
0.312018
Sogou-QCL: A New Dataset with Click Relevance Label · SIGIR 2018
Information retrieval
web corpus
0.312017
SogouT-16: A New Web Corpus to Embrace IR Research · SIGIR 2017
Information retrieval
e-commerce search
0.212022
Joint Learning of E-commerce Search and Recommendation with a Unified Graph Neural Network · WSDM 2022
Query processing and optimization
interactive query workload
0.112020
Database Benchmarking for Supporting Real-Time Interactive Querying of Large Data · SIGMOD Conference 2020
Visualization and visual analytics
interactive data exploration
0.112020
Database Benchmarking for Supporting Real-Time Interactive Querying of Large Data · SIGMOD Conference 2020
Information retrieval › retrieval models
ad-hoc retrieval
0.112017
SogouT-16: A New Web Corpus to Embrace IR Research · SIGIR 2017

Methods — techniques the papers use, named apart from their topics

graph neural network · 1.1attention aggregation · 1.1sequential model · 0.4lab-based user study · 0.4click model construction · 0.4BERT · 0.4feature extraction · 0.4eye tracking · 0.4click model · 0.4click-through data · 0.3
YearPublicationVenuePosition
2022 Query Rewriting in TaoBao Search
abstract
In e-commerce search engines, query rewriting (QR) is a crucial technique that improves shopping experience by reducing the vocabulary gap between user queries and product catalog. Recent works have mainly adopted the generative paradigm. However, they hardly ensure high-quality generated rewrites and do not consider personalization, which leads to degraded search relevance. In this work, we present Contrastive Learning Enhanced Query Rewriting (CLE-QR), the solution used in Taobao product search. It uses a novel contrastive learning enhanced architecture based on "query retrieval-semantic relevance ranking-online ranking". It finds the rewrites from hundreds of millions of historical queries while considering relevance and personalization. Specifically, we first alleviate the representation degeneration problem during the query retrieval stage by using an unsupervised contrastive loss, and then further propose an interaction-aware matching method to find the beneficial and incremental candidates, thus improving the quality and relevance of candidate queries. We then present a relevance-oriented contrastive pre-training paradigm on the noisy user feedback data to improve semantic ranking performance. Finally, we rank these candidates online with the user profile to model personalization for the retrieval of more relevant products. We evaluate CLE-QR on Taobao Product Search, one of the largest e-commerce platforms in China. Significant metrics gains are observed in online A/B tests. CLE-QR has been deployed to our large-scale commercial retrieval system and serviced hundreds of millions of users since December 2021. We also introduce its online deployment scheme, and share practical lessons and optimization tricks of our lexical match system.
Sen Li 0001, Fuyu Lv, Taiwei Jin, Guiyang Li, Yukun Zheng, Qingwen Liu 0002, Xiaoyi Zeng, James T. Kwok, Qianli Ma 0001
CIKM5
2022 Joint Learning of E-commerce Search and Recommendation with a Unified Graph Neural Network
abstract
Click-through rate (CTR) prediction plays an important role in search and recommendation, which are the two most prominent scenarios in e-commerce. A number of models have been proposed to predict CTR by mining user behaviors, especially users' interactions with items. But the sparseness of user behaviors is an obstacle to the improvement of CTR prediction. Previous works only focused on one scenario, either search or recommendation. However, on a practical e-commerce platform, search and recommendation share the same set of users and items, which means joint learning of both scenarios may alleviate the sparseness of user behaviors. In this paper, we propose a novel Search and Recommendation Joint Graph (SRJGraph) neural network to jointly learn a better CTR model for both scenarios. A key question of joint learning is how to effectively share information across search and recommendation, in spite of their differences. A notable difference between search and recommendation is that there are explicit queries in search, whereas no query exists in recommendation. We address this difference by constructing a unified graph to share representations of users and items across search and recommendation, as well as represent user-item interactions uniformly. In this graph, users and items are heterogeneous nodes, and search queries are incorporated into the user-item interaction edges as attributes. For recommendation where no query exists, a special attribute is attached on user-item interaction edges. We further propose an intention and upstream-aware aggregator to explore useful information from high-order connections among users and items. We conduct extensive experiments on a large-scale dataset collected from Taobao.com, the largest e-commerce platform in China. Empirical results show that SRJGraph significantly outperforms the state-of-the-art approaches of CTR prediction in both search and recommendation tasks.
Kai Zhao 0009, Yukun Zheng, Xiang Li 0107, Xiaoyi Zeng
WSDM2
2022 Adaptive neural control for mobile manipulator systems based on adaptive state observer
Yukun Zheng, Yixiang Liu, Rui Song 0002, Xin Ma 0001, Yibin Li 0001
Neurocomputing1
2020 Do People and Neural Nets Pay Attention to the Same Words: Studying Eye-tracking Data for Non-factoid QA Evaluation
abstract
We investigated how users evaluate passage-length answers for non-factoid questions. We conduct a study where answers were presented to users, sometimes shown with automatic word highlighting. Users were tasked with evaluating answer quality, correctness, completeness, and conciseness. Words in the answer were also annotated, both explicitly through user mark up and implicitly through user gaze data obtained from eye-tracking. Our results show that the correctness of an answer strongly depends on its completeness, conciseness is less important.
Valeria Bolotova-Baranova, Vladislav Blinov, Yukun Zheng, W. Bruce Croft, Falk Scholer, Mark Sanderson
CIKM3
2020 Database Benchmarking for Supporting Real-Time Interactive Querying of Large Data
abstract
In this paper, we present a new benchmark to validate the suitability of database systems for interactive visualization workloads. While there exist proposals for evaluating database systems on interactive data exploration workloads, none rely on real user traces for database benchmarking. To this end, our long term goal is to collect user traces that represent workloads with different exploration characteristics. In this paper, we present an initial benchmark that focuses on "crossfilter"-style applications, which are a popular interaction type for data exploration and a particularly demanding scenario for testing database system performance. We make our benchmark materials, including input datasets, interaction sequences, corresponding SQL queries, and analysis code, freely available as a community resource, to foster further research in this area: https://osf.io/9xerb/?view_only=81de1a3f99d04529b6b173a3bd5b4d23.
Leilani Battle, Philipp Eichmann, Marco Angelini, Tiziana Catarci, Giuseppe Santucci, Yukun Zheng, Carsten Binnig, Jean-Daniel Fekete, Dominik Moritz
SIGMOD Conference6
2020 Investigating Examination Behavior in Mobile Search
abstract
Examination is one of the most important user interactions in Web search. A number of works studied examination behavior in Web search and helped researchers better understand how users allocate their attention on search engine result pages (SERPs). Compared to desktop search, mobile search has a number of differences such as fewer results on the screen. These differences bring in mobile-specific factors affecting users' examination behavior. However, there still lacks research on users' attention allocation mechanism via viewports in mobile search. Therefore, we design a lab-based study to collect user's rich interaction behavior in mobile search. Based on the collected data, we first analyze how users examine SERPs and allocate their attention to heterogeneous results. Then we investigate the effect of mobile-specific factors and other common factors on users allocating attention. Finally, we apply the findings of user attention allocation from the user study into click model construction efforts, which significantly improves the state-of-the-art click model. Our work brings insights into a better understanding of users' interaction patterns in mobile search and may benefit other mobile search-related research.
Yukun Zheng, Jiaxin Mao, Yiqun Liu 0001, Mark Sanderson, Min Zhang 0006, Shaoping Ma
WSDM1
2020 Leveraging Passage-level Cumulative Gain for Document Ranking
abstract
Document ranking is one of the most studied but challenging problems in information retrieval (IR) research. A number of existing document ranking models capture relevance signals at the whole document level. Recently, more and more research has begun to address this problem from fine-grained document modeling. Several works leveraged fine-grained passage-level relevance signals in ranking models. However, most of these works focus on context-independent passage-level relevance signals and ignore the context information, which may lead to inaccurate estimation of passage-level relevance. In this paper, we investigate how information gain accumulates with passages when users sequentially read a document. We propose the context-aware Passage-level Cumulative Gain (PCG), which aggregates relevance scores of passages and avoids the need to formally split a document into independent passages. Next, we incorporate the patterns of PCG into a BERT-based sequential model called Passage-level Cumulative Gain Model (PCGM) to predict the PCG sequence. Finally, we apply PCGM to the document ranking task. Experimental results on two public ad hoc retrieval benchmark datasets show that PCGM outperforms most existing ranking models and also indicates the effectiveness of PCG signals. We believe that this work contributes to improving ranking performance and providing more explainability for document ranking.
Zhijing Wu 0001, Jiaxin Mao, Yiqun Liu 0001, Jingtao Zhan, Yukun Zheng, Min Zhang 0006, Shaoping Ma
WWW5
2019 Human Behavior Inspired Machine Reading Comprehension
abstract
Machine Reading Comprehension (MRC) is one of the most challenging tasks in both NLP and IR researches. Recently, a number of deep neural models have been successfully adopted to some simplified MRC task settings, whose performances were close to or even better than human beings. However, these models still have large performance gaps with human beings in more practical settings, such as MS MARCO and DuReader datasets. Although there are many works studying human reading behavior, the behavior patterns in complex reading comprehension scenarios remain under-investigated. We believe that a better understanding of how human reads and allocates their attention during reading comprehension processes can help improve the performance of MRC tasks. In this paper, we conduct a lab study to investigate human's reading behavior patterns during reading comprehension tasks, where 32 users are recruited to take 60 distinct tasks. By analyzing the collected eye-tracking data and answers from participants, we propose a two-stage reading behavior model, in which the first stage is to search for possible answer candidates and the second stage is to generate the final answer through a comparison and verification process. We also find that human's attention distribution is affected by both question-dependent factors (e.g., answer and soft matching signal with questions) and question-independent factors (e.g., position, IDF and Part-of-Speech tags of words). We extract features derived from the two-stage reading behavior model to predict human's attention signals during reading comprehension, which significantly improves performance in the MRC task. Findings in our work may bring insight into the understanding of human reading and information seeking processes, and help the machine to better meet users' information needs.
Yukun Zheng, Jiaxin Mao, Yiqun Liu 0001, Zixin Ye, Min Zhang 0006, Shaoping Ma
SIGIR1
2019 Constructing Click Model for Mobile Search with Viewport Time
abstract
A series of click models has been proposed to extract accurate and unbiased relevance feedback from valuable yet noisy click-through data in search logs. Previous works have shown that users search behavior in mobile and desktop scenarios are rather different in many aspects, therefore, the click models designed for desktop search may not be effective in the mobile context. To address this problem, we propose two novel click models for mobile search: (1) Mobile Click Model (MCM), which models click necessity bias and examination satisfaction bias; (2) Viewport Time Click Model (VTCM), which further extends MCM by utilizing the viewport time. Extensive experiments on large-scale real mobile search logs show that: (1) MCM and VTCM outperform existing models in predicting users’ clicks and estimating result relevance; (2) MCM and VTCM can extract richer information, such as the click necessity of search results and the probability of user satisfaction, from mobile click logs; (3) By modeling the viewport time distributions of heterogeneous results, VTCM can bring a significant improvement over MCM in click prediction and relevance estimation tasks. Our proposed click models can help better understand user behavior patterns in mobile search and improve the ranking performance of mobile search engines.
Yukun Zheng, Jiaxin Mao, Yiqun Liu 0001, Cheng Luo 0001, Min Zhang 0006, Shaoping Ma
ACM Trans. Inf. Syst.1
2018 Sogou-QCL: A New Dataset with Click Relevance Label
abstract
Data is of vital importance in the development of machine learning technologies. Recently, within the information retrieval field, a number of neural ranking frameworks have been proposed to address the ad-hoc search. These models usually need a large amount of query-document relevance judgments for training. However, obtaining this kind of relevance judgments needs a lot of money and manual effort. To shed light on this problem, researchers seek to use implicit feedback from users of search engines to improve the ranking performance. In this paper, we present a new dataset, Sogou-QCL, which contains 537,366 queries and five kinds of weak relevance labels for over 12 million query-document pairs. We apply Sogou-QCL dataset to train recent neural ranking models and show its potential to serve as weak supervision for ranking. We believe that Sogou-QCL will have a broad impact on corresponding areas.
Yukun Zheng, Zhen Fan 0003, Yiqun Liu 0001, Cheng Luo 0001, Min Zhang 0006, Shaoping Ma
SIGIR1
2017 A Quantitative Method for Evaluating Network Security Based on Attack Graph
Yukun Zheng, Changzhen Hu
NSS1
2017 SogouT-16: A New Web Corpus to Embrace IR Research
abstract
Web collection is essential for many Web based researches such as Web Information Retrieval (IR), Web data mining, Corpus linguistics and so on. However, it is usually expensive and time-consuming to collect a large scale of Web pages in lab-based environment and public-available collection becomes a necessity for these researches. In this study, we present a Chinese Web collection, SogouT-16, which is the largest free-of-charge public Chinese Web collection so far. We provide a variety of descriptive characteristics of SogouT-16 and discuss its adoption in a newly-designed ad-hoc retrieval task in NTCIR-13, We Want Web. SogouT-16 also provides online retrieval service and contains a number of auxiliary resources including hyperlink structure graph, query logs, word embedding, and etc. We believe that SogouT-16 will provide new opportunities for novel investigations and applications in IR and other related communities.
Cheng Luo 0001, Yukun Zheng, Yiqun Liu 0001, Jingfang Xu, Min Zhang 0006, Shaoping Ma
SIGIR2