Hongfei Yan

dblp:86/1423 · DBLP profile ↗
← Back
32ranked-venue papers
2as first author
3since 2021 · last 2026
0000-0001-5914-8585ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 1 since 2021Databases, data management, data science and information retrieval · 16Applied, interdisciplinary, general and emerging computing · 3Systems, architecture and hardware · 2 · 1 first-authorComputer networks · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Question answering and dialogue systems · 39% Knowledge representation and reasoning · 39% Information extraction and text analysis · 22%
Databases, data mining, and information retrieval
6 papers
Information retrieval · 54% Web and social media mining · 20% Data mining · 12%
Computer networks
1 paper
Content delivery and video streaming · 50% Transport protocols and congestion control · 50%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 100%

Topics — the 25 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Transport protocols and congestion control
cross-layer congestion control
0.812024
Context-Aware Cross-Layer Congestion Control for Large-Scale Live Streaming · IEEE/ACM Trans. Netw. 2024
Content delivery and video streaming
live streaming
0.812024
Context-Aware Cross-Layer Congestion Control for Large-Scale Live Streaming · IEEE/ACM Trans. Netw. 2024
Information retrieval › indexing
index compression
0.212015
A General SIMD-Based Approach to Accelerating Compression Algorithms · ACM Trans. Inf. Syst. 2015
Indexing and storage engines › data compression
integer compression
0.212015
A General SIMD-Based Approach to Accelerating Compression Algorithms · ACM Trans. Inf. Syst. 2015
Hardware accelerators and domain-specific architectures › data-parallel accelerator
SIMD accelerator
0.212015
A General SIMD-Based Approach to Accelerating Compression Algorithms · ACM Trans. Inf. Syst. 2015
Natural language and speech › Information extraction and text analysis
topic model
0.222012
SSHLDA: A Semi-Supervised Hierarchical Topic Model · EMNLP-CoNLL 2012
Jointly Modeling Aspects and Opinions with a MaxEnt-LDA Hybrid · EMNLP 2010
Data mining › structured data mining › graph mining
graph learning
0.212013
Mining New Business Opportunities: Identifying Trend related Products by Leveraging Commercial Intents from Microblogs · EMNLP 2013
Web and social media mining › social media analysis
microblog analysis
0.212013
Mining New Business Opportunities: Identifying Trend related Products by Leveraging Commercial Intents from Microblogs · EMNLP 2013
Natural language and speech › Information extraction and text analysis › topic model
hierarchical topic model
0.112012
SSHLDA: A Semi-Supervised Hierarchical Topic Model · EMNLP-CoNLL 2012
Natural language and speech › Information extraction and text analysis › topic model
semi-supervised topic modeling
0.112012
SSHLDA: A Semi-Supervised Hierarchical Topic Model · EMNLP-CoNLL 2012
Information retrieval › indexing › inverted index
block-max index
0.112012
Optimized top-k processing with global page scores on block-max indexes · WSDM 2012
Web and social media mining › event detection
burst detection
0.112012
EventSearch: a system for event discovery and retrieval on multi-type historical data · KDD 2012
Information retrieval › query processing
dynamic pruning
0.112012
Optimized top-k processing with global page scores on block-max indexes · WSDM 2012
Web and social media mining
event detection
0.112012
Identifying Event-related Bursts via Social Media Activities · EMNLP-CoNLL 2012
Information retrieval › document retrieval › temporal information retrieval
event retrieval
0.112012
EventSearch: a system for event discovery and retrieval on multi-type historical data · KDD 2012
Information retrieval
query processing
0.112012
Optimized top-k processing with global page scores on block-max indexes · WSDM 2012
Information retrieval
ranking
0.112012
Optimized top-k processing with global page scores on block-max indexes · WSDM 2012
Data mining › text mining
temporal text mining
0.112012
EventSearch: a system for event discovery and retrieval on multi-type historical data · KDD 2012
Query processing and optimization
top-k query processing
0.112012
Optimized top-k processing with global page scores on block-max indexes · WSDM 2012
Information retrieval › evaluation › online evaluation
click-based evaluation
0.112011
Efficiently collecting relevance information from clickthroughs for web retrieval system evaluation · SIGIR 2011
Information retrieval › user behavior › search behavior
click model
0.112011
Efficiently collecting relevance information from clickthroughs for web retrieval system evaluation · SIGIR 2011
Information retrieval
evaluation
0.112011
Efficiently collecting relevance information from clickthroughs for web retrieval system evaluation · SIGIR 2011
Natural language and speech › Information extraction and text analysis › sentiment analysis
aspect-based sentiment analysis
0.112010
Jointly Modeling Aspects and Opinions with a MaxEnt-LDA Hybrid · EMNLP 2010
Information retrieval › indexing › inverted index
document identifier assignment
0.012012
Optimized top-k processing with global page scores on block-max indexes · WSDM 2012
Web and social media mining
social media analysis
0.012012
Identifying Event-related Bursts via Social Media Activities · EMNLP-CoNLL 2012

Methods — techniques the papers use, named apart from their topics

large language model · 1.0state transition mechanism · 0.8cross-layer feedback · 0.8SIMD · 0.4graph-based method · 0.3latent dirichlet allocation · 0.3semi-supervised learning · 0.1maxscore · 0.1burst model · 0.1block-max index · 0.1WAND · 0.1reordering · 0.1optimization · 0.1maximum entropy · 0.1
YearPublicationVenuePosition
2026 LinkQA: Synthesizing Diverse QA from Multiple Seeds Strongly Linked by Knowledge Points
abstract
Xuemiao Zhang, Can Ren, Chengying Tu, Rongxiang Weng, Hongfei Yan, Jingang Wang, Xunliang Cai. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xuemiao Zhang, Can Ren, Chengying Tu, Rongxiang Weng, Hongfei Yan, Jingang Wang
ACL (1)5
2025 Understanding Operational CDN Live Streaming: A Measurement Study on Performance, Costs, and Enhancements
abstract
The escalating need for live video streaming has emerged as a significant catalyst for the business expansion of today’s content delivery networks (CDN). Selecting the right CDN live streaming architecture is fundamentally important in achieving the objective of enhancing users’ quality of experience (QoE) while reducing bandwidth costs. Regrettably, a limited number of studies have been conducted to systematically measure and compare the current typical solutions at production scale. Consequently, the performance and costs of different streaming architectures remain myths. This paper aims to address the existing research gap by undertaking a large-scale measurement study of three representative CDN live streaming architectures, defined by streaming protocol and overlay topology choices, currently running on Alibaba Cloud’s production video delivery network. By analyzing the results of over 500 million video plays over two months on a large live streaming platform hosted on Alibaba Cloud’s CDN, we reveal the impact of architectural compositions and operational factors on live streaming performance and bandwidth costs. In particular, our study reveals the trade-offs between QoE metrics and bandwidth costs for operational streaming architectures. Drawing upon the insights of this study, we further develop and deploy pragmatic strategies that yield remarkable real-world impact—our design saves over 17% bandwidth costs while maintaining the QoE.
Danfu Yuan, Weizhan Zhang, Haiyu Huang 0005, Xuan Zeng 0002, Hongfei Yan, Yubing Qiu, Jinghui Zhong
IEEE Trans. Circuits Syst. Video Technol.8
2024 Context-Aware Cross-Layer Congestion Control for Large-Scale Live Streaming
abstract
Live video streaming has come to dominate today’s Internet traffic. Content Delivery Network (CDN) providers, responsible for hosting outsourced live streaming services, are now striving to ensure an enhanced quality of experience (QoE) to meet the ever-increasing user expectations. Existing congestion control (CC) schemes in the kernel, however, suffer from unsatisfactory performance for live video delivery due to disparities in traffic characteristics and differentiated optimization goals between generic traffic and live video traffic. In this paper, we propose XCC, a streaming context-aware CC approach that helps achieve better QoE for the live streaming services from CDN provider. The core of XCC is to adaptively coordinate the transmission strategy and frame rate through a cross-layer feedback framework, responding to the fluctuating traffic dynamics and network conditions in the short term. Further, XCC matches the long-term traffic characteristics (i.e., two-stage delivery mode) by employing a task-specific state transition mechanism as the underlying TCP. XCC has been implemented in the Linux kernel’s TCP stack and media engine and has been fully deployed in Alibaba Cloud’s production service. Evaluation in experimental environments and A/B testing serving tens of millions of sessions demonstrate that XCC is competitive in streaming delay against the most prevalent TCP in today’s Operating Systems, while reducing startup delay by 9.9%, stall time by 36.4%, and stall frequency by 42.5% on average in deployment.
Danfu Yuan, Weizhan Zhang, Yubing Qiu, Haiyu Huang 0005, Hongfei Yan, Yaming He
IEEE/ACM Trans. Netw.8
2019 A Dynamic Financial Knowledge Graph Based on Reinforcement Learning and Transfer Learning
abstract
The knowledge graph is a means of visualizing data to aid information analysis and understanding. In this paper, we construct a novel dynamic financial knowledge graph, which utilizes time information to capture data changes and trends over time. Firstly, the basic dynamic financial knowledge graph is constructed through structured and semi-structured data related to A-share. Then, using the transfer learning algorithms, we train the financial entity recognition models based on BERT, BiLSTM, and CRF. Next, we train the financial entity linking models based on similarity features and prior knowledge. After that, to alleviate the noise brought by distant supervision, we explore to train the financial relation classification models with the help of reinforcement learning. Finally, we implement the dynamic knowledge graph based on these models and their predictions. Additionally, a display website is designed and implemented to dynamically display the structural changes of the knowledge graph over time. The financial knowledge graph constructed in this paper is practical and the construction pipeline provides insights for a professional dynamic knowledge graph as well.
Hongfei Yan, Chong Chen 0003
IEEE BigData3
2017 DroidForensics: Accurate Reconstruction of Android Attacks via Multi-layer Forensic Logging
abstract
The goal of cyber attack investigation is to fully reconstruct the details of an attack, so we can trace back to its origin, and recover the system from the damage caused by the attack. However, it is often difficult and requires tremendous manual efforts because attack events occurred days or even weeks before the investigation and detailed information we need is not available anymore. Consequently, forensic logging is significantly important for cyber attack investigation. In this paper, we present DroidForensics, a multi-layer forensic logging technique for Android. Our goal is to provide the user with detailed information about attack behaviors that can enable accurate post-mortem investigation of Android attacks. DroidForensics consists of three logging modules. API logger captures Android API calls that contain high-level semantics of an application. Binder logger records interactions between applications to identify causal relations between processes, and system call logger efficiently monitors low-level system events. We also provide the user interface that the user can compose SQL-like queries to inspect an attack. Our experiments show that Droid Forensics has low runtime overhead (2.9% on average) and low space overhead (105 ~ 169 MByte during 24 hours) on real Android devices. It is effective in the reconstruction of realworld Android attacks we have studied.
Xingzi Yuan, Omid Setayeshfar, Hongfei Yan, Pranav Panage, Xuetao Wei, Kyu Hyung Lee
AsiaCCS3
2017 Learning Graph-based Embedding For Time-Aware Product Recommendation
abstract
In this paper, we propose a novel Product Graph Embedding (PGE) model to investigate time-aware product recommendation by leveraging the network representation learning technique. Our model captures the sequential influences of products by transforming the historical purchase records into a product graph. Then the product can be transformed into a low dimensional vector by the network embedding model. Once products are projected into the latent space, we present a novel method to compute user's latest preferences, which projects users into the same latent space as products. This method is based on time-decay functions and the embedding of sequential products that the user purchased. Thus, relatedness between a product and a user can be measured by the similarity between the embedding vectors which represent the product and the user's preferences. The experimental results on purchase records crawled from JINGDONG, show the superiority of our proposed framework for personalized product recommendation.
Weizheng Chen, Hongfei Yan
CIKM3
2017 Deep joint discriminative learning for vehicle re-identification and retrieval
abstract
In this paper, we propose a novel vehicle re-identification method based on a Deep Joint Discriminative Learning (DJDL) model, which utilizes a deep convolutional network to effectively extract discriminative representations for vehicle images. To exploit properties and relationship among samples in different views, we design a unified framework to combine several different tasks efficiently, including identification, attribute recognition, verification and triplet tasks. The whole network is optimized jointly via a specific batch composition design. Extensive experiments are conducted on a large-scale VehicleID [1] dataset. Experimental results demonstrate the effectiveness of our method and show that it achieves the state-of-the-art performance on both vehicle re-identification and retrieval.
Yanghao Li, Hongfei Yan, Jiaying Liu 0001
ICIP3
2017 Mining E-commercial data: A text-rich heterogeneous network embedding approach
abstract
It is a great challenge to model and mine the e-commercial data, which is made up of multiple types of objects, such as products, users, comments and tags. To model the complicated interactive relationships in the the e-commercial data, we propose to transform the complex e-commercial data into a text-rich heterogeneous e-commercial network. Then three neural network based embedding algorithms named WTL (Weighted Text Learning), IBL (Identity Based Learning) and IBTSL (Identity Based Two Steps Learning) are proposed to consider both the network structure information and heterogeneous nodes attributes identity information to learn the embeddings. The key idea of our models is to map all objects in the e-commercial network to a same low-dimensional vector space, which is useful to produce meaningful features for many applications such as product classification, comment classification, product attributes forecasting, recommendation, and so on. Our algorithms are compared with other existing advanced methods on a real large-scale e-commercial dataset. Several applications are set to evaluate the effectivity of the learned embeddings. The experimental results show that the embeddings generated by our algorithms have superior performance in each application.
Weizheng Chen, Hongfei Yan, Yan Zhang 0004
IJCNN4
2016 Non-Linear Smoothed Transductive Network Embedding with Text Information
abstract
Network embedding is a classical task which aims to map the nodes of a network to low-dimensional vectors. Most of the previous network embedding methods are trained in an unsupervised scheme. Then the learned node embeddings can be used as inputs of many machine learning tasks such as node classification, attribute inference. However, the discrimination validity of the node embeddings maybe improved by considering the node label information and the node attribute information. Inspired by traditional semi-supervised learning techniques, we explore to train the node embeddings and the node classifiers simultaneously with the text attributes information in a flexible framework. We present NLSTNE (Non-Linear Smoothed Transductive Network Embedding), a transductive network embedding method, whose embeddings are enhanced by modeling the non-linear pairwise similarity between the nodes and the non-linear relationship between the nodes and the text attributes. We use the node classification task to evaluate the quality of node embeddings learned by different models on four real-world network datasets . The experimental results demonstrate that our model outperforms several state-of-the-art network embedding methods.
Weizheng Chen, Jinpeng Wang 0001, Yan Zhang 0004, Hongfei Yan, Xiaoming Li 0001
ACML5
2015 Educational Evaluation in the PKU SPOC Course "Data Structures and Algorithms"
abstract
In order to learn the impact of MOOCs, we conducted a SPOC experiment on the course of Data Structures and Algorithms in Peking University. In this paper, we analyze student online activities, test scores, and two surveys using statistical methods (t-test, analysis of variance, correlation analysis and OLS regression) to understand what factors will foster improvements in student learning. We find that the "SPOC + Flipped" is a helpful mode to teach algorithm, time spent on the course and students' confidence had a positive impact on learning effect, and SPOC resource should be made full use of.
Ming Zhang 0004, Jile Zhu, Yanzhen Zou, Hongfei Yan, Dan Hao 0001, Chuxiong Liu
L@S4
2015 A General SIMD-Based Approach to Accelerating Compression Algorithms
abstract
Compression algorithms are important for data-oriented tasks, especially in the era of “Big Data.” Modern processors equipped with powerful SIMD instruction sets provide us with an opportunity for achieving better compression performance. Previous research has shown that SIMD-based optimizations can multiply decoding speeds. Following these pioneering studies, we propose a general approach to accelerate compression algorithms. By instantiating the approach, we have developed several novel integer compression algorithms, called Group-Simple, Group-Scheme, Group-AFOR, and Group-PFD, and implemented their corresponding vectorized versions. We evaluate the proposed algorithms on two public TREC datasets, a Wikipedia dataset, and a Twitter dataset. With competitive compression ratios and encoding speeds, our SIMD-based algorithms outperform state-of-the-art nonvectorized algorithms with respect to decoding speeds.
Wayne Xin Zhao, Daniel Lemire, Dongdong Shan, Jian-Yun Nie, Hongfei Yan, Ji-Rong Wen
ACM Trans. Inf. Syst.6
2014 Group based Self Training for E-Commerce Product Record Linkage
Yuexin Wu, Hongfei Yan
COLING3
2013 Group-Scheme: SIMD-based compression algorithms for web text data
abstract
Compression algorithms have been quite important for data oriented tasks, especially in the era of Big Data. The rapid development of modern processors facilitates us with powerful SIMD instruction sets, which provides an opportunity for better performance. Although SIMD based optimization on compression have been explored in some studies [2, 7], these studies usually focus on modifying the existing algorithms to fit into the SIMD instruction. In this paper, we propose a compression framework with a novel storage layout format, which aims to improve instruction-level parallelizability of compression algorithms. By instantiating the framework, we design a novel compression algorithm family, called Group-Scheme, and present a parallelized version of Group-Scheme, called SIMD-Group-Scheme. We evaluate the proposed algorithms on two public TREC data sets. With very competitive performance on compression ratio and encoding speed, SIMD-Group-Scheme significantly outperforms the implementation without SIMD instructions and state-of-the-art algorithm (i.e. SIMD-G8IU [7]), w.r.t decoding speed.
Wayne Xin Zhao, Dongdong Shan, Hongfei Yan
IEEE BigData4
2013 Mining New Business Opportunities: Identifying Trend related Products by Leveraging Commercial Intents from Microblogs
abstract
Hot trends are likely to bring new business opportunities.For example, "Air Pollution" might lead to a significant increase of the sales of related products, e.g., mouth mask.For ecommerce companies, it is very important to make rapid and correct response to these hot trends in order to improve product sales.In this paper, we take the initiative to study the task of how to identify trend related products.The major novelty of our work is that we automatically learn commercial intents revealed from microblogs.We carefully construct a data collection for this task and present quite a few insightful findings.In order to solve this problem, we further propose a graph based method, which jointly models relevance and associativity.We perform extensive experiments and the results showed that our methods are very effective.
Jinpeng Wang 0001, Wayne Xin Zhao, Haitian Wei, Hongfei Yan, Xiaoming Li 0001
EMNLP4
2013 A Metric Learning Based Approach to Evaluate Task-Specific Time Series Similarity
Wayne Xin Zhao, Hongfei Yan, Xiaoming Li 0001
WAIM3
2012 Hierarchical topic integration through semi-supervised hierarchical topic modeling
abstract
Lots of document collections are well organized in hierarchical structure, and such structure can help users browse and understand these collections. Meanwhile, there are a large number of plain document collections loosely organized, and it is difficult for users to understand them effectively. In this paper we study how to automatically integrate latent topics in a plain collection with the topics in a hierarchical structured collection. We propose to use semi-supervised topic modeling to solve the problem in a principled way. The experiments show that the proposed method can generate both meaningful latent topics and expand high quality hierarchical topic structures.
Xianling Mao, Jing He 0010, Hongfei Yan, Xiaoming Li 0001
CIKM3
2012 Automatic labeling hierarchical topics
abstract
Recently, statistical topic modeling has been widely applied in text mining and knowledge management due to its powerful ability. A topic, as a probability distribution over words, is usually difficult to be understood. A common, major challenge in applying such topic models to other knowledge management problem is to accurately interpret the meaning of each topic. Topic labeling, as a major interpreting method, has attracted significant attention recently. However, previous works simply treat topics individually without considering the hierarchical relation among topics, and less attention has been paid to creating a good hierarchical topic descriptors for a hierarchy of topics. In this paper, we propose two effective algorithms that automatically assign concise labels to each topic in a hierarchy by exploiting sibling and parent-child relations among topics. The experimental results show that the inter-topic relation is effective in boosting topic labeling accuracy and the proposed algorithms can generate meaningful topic labels that are useful for interpreting the hierarchical topics.
Xianling Mao, Zhaoyan Ming, Zhengjun Zha, Tat-Seng Chua, Hongfei Yan, Xiaoming Li 0001
CIKM5
2012 SSHLDA: A Semi-Supervised Hierarchical Topic Model
Xianling Mao, Zhaoyan Ming, Tat-Seng Chua, Si Li 0001, Hongfei Yan, Xiaoming Li 0001
EMNLP-CoNLL5
2012 Identifying Event-related Bursts via Social Media Activities
Wayne Xin Zhao, Baihan Shu, Jing Jiang 0001, Yang Song 0021, Hongfei Yan, Xiaoming Li 0001
EMNLP-CoNLL5
2012 EventSearch: a system for event discovery and retrieval on multi-type historical data
abstract
We present EventSearch, a system for event extraction and retrieval on four types of news-related historical data, i.e., Web news articles, newspapers, TV news program, and micro-blog short messages. The system incorporates over 11 million web pages extracted from "Web InfoMall", the Chinese Web Archive since 2001. The newspaper and TV news video clips also span from 2001 to 2011. The system, upon a user query, returns a list of event snippets from multiple data sources. A novel burst model is used to discover events from time-stamped texts. In addition to offline event extraction, our system also provides online event extraction to further meet the user needs. EventSearch provides meaningful analytics that synthesize an accurate description of events. Users interact with the system by ranking the identified events using different criteria (scale, recency and relevance) and submitting their own information needs in different input fields.
Dongdong Shan, Wayne Xin Zhao, Rishan Chen, Baihan Shu, Ziqi Wang 0002, Hongfei Yan, Xiaoming Li 0001
KDD7
2012 Optimized top-k processing with global page scores on block-max indexes
abstract
Large web search engines are facing formidable performance challenges because they have to process thousands of queries per second on tens of billions of documents, within interactive response time. Among many others, Top-k query processing (also called early termination or dynamic pruning) is an important class of optimization techniques that can improve the search efficiency and achieve faster query processing by avoiding the scoring of documents that are unlikely to be in the top results. One recent technique is using Block-Max index. In the Block-Max index, the posting lists are organized as blocks and the maximum score for each block is stored to improve the query efficiency. Although query processing speedup is achieved with Block-Max index, the ranking function for the Top-k results is the term-based approach. It is well known that documents' static scores are also important for a good ranking function. In this paper, we show that the performance of the state-of-the-art algorithms with the Block-Max index is degraded when the static score is added in the ranking function. Then we study efficient techniques for Top-k query processing in the case where a page's static score is given, such as PageRank, in addition to the term-based approach. In particular, we propose a set of new algorithms based on the WAND and MaxScore with Block-Max index using local score, which outperform the existing ones. Then we propose new techniques to estimate a better score upper bound for each block. We also study the search efficiency on different index structures where the document identifiers are assigned by URL sorting or by static document scores. Experiments on TREC GOV2 and ClueWeb09B show that considerable performance gains are achieved.
Dongdong Shan, Shuai Ding 0006, Jing He 0010, Hongfei Yan, Xiaoming Li 0001
WSDM4
2011 Recommending citations with translation model
abstract
Citation Recommendation is useful for an author to find out the papers or books that can support the materials she is writing about. It is a challengeable problem since the vocabulary used in the content of papers and in the citation contexts are usually quite different. To address this problem, we propose to use translation model, which can bridge the gap between two heterogeneous languages. We conduct an experiment and find the translation model can provide much better candidates of citations than the state-of-the-art methods.
Jing He 0010, Dongdong Shan, Hongfei Yan
CIKM4
2011 Efficient phrase querying with flat position index
abstract
A large proportion of search engine queries contain phrases,namely a sequence of adjacent words. In this paper, we propose to use flat position index (a.k.a schema-independent index) for phrase query evaluation. In the flat position index, the entire document collection is viewed as a huge sequence of tokens. Each token is represented by one flat position, which is a unique position offset from the beginning of the collection. Each indexed term is associated with a list of the flat positions about that term in the sequence. To recover DocID from flat positions efficiently, we propose a novel cache sensitive look-up table (CSLT), which is much faster than existing search algorithms. Experiments on TREC GOV2 data collection show that flat position index can reduce the index size and speed up phrase querying substantially, compared with traditional word-level index.
Dongdong Shan, Wayne Xin Zhao, Jing He 0010, Rui Yan 0001, Hongfei Yan, Xiaoming Li 0001
CIKM5
2011 Comparing Twitter and Traditional Media Using Topic Models
Wayne Xin Zhao, Jing Jiang 0001, Jianshu Weng, Jing He 0010, Ee-Peng Lim, Hongfei Yan, Xiaoming Li 0001
ECIR6
2011 SizeSpotSigs: An Effective Deduplicate Algorithm Considering the Size of Page Content
Xianling Mao, Nan Di, Xiaoming Li 0001, Hongfei Yan
PAKDD (1)5
2011 Efficiently collecting relevance information from clickthroughs for web retrieval system evaluation
abstract
Various click models have been recently proposed as a principled approach to infer the relevance of documents from the clickthrough data. The inferred document relevance is potentially useful in evaluating the Web retrieval systems. In practice, it generally requires to acquire the accurate evaluation results within minimal users' query submissions. This problem is important for speeding up search engine development and evaluation cycle and acquiring reliable evaluation results on tail queries. In this paper, we propose a reordering framework for efficient evaluation problem in the context of clickthrough based Web retrieval evaluation. The main idea is to move up the documents that contribute more for the evaluation task. In this framework, we propose four intuitions and formulate them as an optimization problem. Both user study and TREC data based experiments validate that the reordering framework results in much fewer query submissions to get accurate evaluation results with only a little harm to the users' utility.
Jing He 0010, Wayne Xin Zhao, Baihan Shu, Xiaoming Li 0001, Hongfei Yan
SIGIR5
2010 Context modeling for ranking and tagging bursty features in text streams
abstract
Bursty features in text streams are very useful in many text mining applications. Most existing studies detect bursty features based purely on term frequency changes without taking into account the semantic contexts of terms, and as a result the detected bursty features may not always be interesting or easy to interpret. In this paper we propose to model the contexts of bursty features using a language modeling approach. We then propose a novel topic diversity-based metric using the context models to find newsworthy bursty features. We also propose to use the context models to automatically assign meaningful tags to bursty features. Using a large corpus of a stream of news articles, we quantitatively show that the proposed context language models for bursty features can effectively help rank bursty features based on their newsworthiness and to assign meaningful tags to annotate bursty features.
Wayne Xin Zhao, Jing Jiang 0001, Jing He 0010, Dongdong Shan, Hongfei Yan, Xiaoming Li 0001
CIKM5
2010 Jointly Modeling Aspects and Opinions with a MaxEnt-LDA Hybrid
Wayne Xin Zhao, Jing Jiang 0001, Hongfei Yan, Xiaoming Li 0001
EMNLP3
2007 On the peninsula phenomenon in web graph and its implications on web search
Hongfei Yan
Comput. Networks2
2004 The Evolution of Link-Attributes for Pages and Its Implications on Web Crawling
abstract
It is important for an incremental crawler to know how web pages evolve and the relation between their changing frequencies and the link-attributes such as indegrees. This paper proposes a model for incremental crawling and performs an experiment to verify the correlation between them, by monitoring the evolution of all the link-attributes of the web pages within one website. Particularly, we look deeply into one special kind of page named Index-pages. From the experiment, we can make four conclusions: (1) Pages which have bigger indegrees, outdegrees or PageRank values change more often, and these link-attributes all approximately obey a power-law distribution. (2) The link-attributes of pages seldom change though the pages change themselves. (3) A small proportion of the pages link to most of the vertexes in the web graph. (4) The Index-pages link to sizeable new pages in a website. These conclusions can be used to greatly enhance the performance of an incremental crawler, which is the foremost component for general search engines and web information stores.
Tao Meng II, Hongfei Yan, Jimin Wang
Web Intelligence2
2002 Architectural design and evaluation of an efficient Web-crawling system
Hongfei Yan, Jianyong Wang 0001, Xiaoming Li 0001
J. Syst. Softw.1
2001 Architectural Design and Evaluation of an Efficient Web-crawling System
abstract
This paper presents an architectural design and evaluation result of an efficient Web-crawling system. The design involves a fully distributed architecture, a URL allocating algorithm, and a method to assure system scalability and dynamic reconfigurability. Simulation experiment shows that load balance, scalability and efficiency can be achieved in the system. Currently this distributed Web-crawling subsystem has been successfully integrated with WebGather, a well-known Chinese and English Web search engine, aimed at collecting all the Web pages in China and keeping pace with the rapid growth of Chinese Web information. In addition, we believe that the design can also be useful in other context such as digital library, etc. 2002 Elsevier Science Inc. All rights reserved.
Hongfei Yan, Jianyong Wang 0001, Xiaoming Li 0001
IPDPS1