VLDB 2026 Research / reviewers in the wild / expert
Hongyan Liu 0002
dblp:28/374-2
· DBLP profile ↗
59ranked-venue papers in the field
7as first author
13since 2021 · last 2026
0000-0002-4902-1078ORCID · conflict
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 23 (1 first)Data Mining & Knowledge Discovery · 21 (2 first)Database Systems & Data Management · 8 (3 first)Knowledge Engineering, Semantic Web & Information Systems · 5 (1 first)Big Data, Cloud & Distributed Data Systems · 1Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing Multi-Valued Treatment Uplift Modeling with Knowledge Sharing and RCT DataabstractUplift modeling estimates individual-level treatment effects for personalized decisions in applications such as marketing and healthcare. Beyond binary treatments, many applications involve multi-valued treatments. As data are split across more groups, sparsity and imbalance pose challenges. We propose CAMU, which combines (i) a Codebook-enhanced representation network that anchors representations across mini-batches to stabilize learning and facilitate distribution alignment, and (ii) a Cross-Treatment Self-Attention (CTSA) enhanced prediction network for treatment-aware knowledge sharing. We extend CAMU to CAMUER, which transfers unbiased signals from randomized controlled trial (RCT) data via attention alignment and prediction-consistency distillation. The proposed model is deployed in Kuaishou's ad-serving system, serving tens of millions of users. Offline experiments and a large-scale online A/B test demonstrate consistent gains over baselines. Luo He, Zhenghao Zeng, Deqiang Kong, Funan Mu, Zilong Lu, Peng Jiang 0002, Hongyan Liu 0002 |
SIGIR | 9 |
| 2026 | QChunker: Learning Question-Aware Text Chunking for Domain RAG via Multi-Agent DebateabstractThe effectiveness upper bound of retrieval-augmented generation (RAG) is fundamentally constrained by the semantic integrity and information granularity of text chunks in its knowledge base. Moreover, domain documents are characterized by dense terminology and strong contextual dependencies, which exacerbate the semantic fragmentation of text chunks, thereby making it difficult to efficiently utilize their key information. To address these challenges, this paper proposes QChunker, which restructures the RAG paradigm from retrieval-augmentation to understanding-retrieval-augmentation. Firstly, QChunker models the text chunking as a composite task of text segmentation and knowledge completion to ensure the logical coherence and integrity of text chunks. Drawing inspiration from Hal Gregersen's ''Questions Are the Answer'' theory, we design a multi-agent debate framework comprising four specialized components: a question outline generator, text segmenter, integrity reviewer, and knowledge completer. This framework operates on the principle that questions serve as catalysts for profound insights. Through this pipeline, we successfully construct a high-quality dataset of 45K entries and transfer this capability to small language models. Additionally, to handle long evaluation chains and low efficiency in existing chunking evaluation methods, which overly rely on downstream QA tasks, we introduce a novel direct evaluation metric, ChunkScore. Both theoretical and experimental validations demonstrate that ChunkScore can directly and efficiently discriminate the quality of text chunks. Furthermore, during the text segmentation phase, we utilize document outlines for multi-path sampling to generate multiple candidate chunks and select the optimal solution employing ChunkScore. Extensive experimental results across four heterogeneous domains exhibit that QChunker effectively resolves aforementioned issues by providing RAG with more logically coherent and information-rich text chunks. Notably, this study also establishes a small-domain QA dataset concerning hazardous chemical safety, which fully reveals the significant value of RAG in specialized domains and the generalization capability of the QChunker framework. Jihao Zhao, Daixuan Li, Shuaishuai Zu, Biao Qin, Hongyan Liu 0002 |
WWW | 6 |
| 2025 | Diffusion Alignment for Cross Domain RecommendationabstractCross-domain recommendation (CDR) is a critical solution to address the sparsity issue of conventional recommendation systems. In most scenarios, only a subset of users have interactions on both domains, which is referred to as partial user overlap CDR. Despite this, most CDR methods tend to design alignment mechanisms only for overlapping users, resulting in partial alignment issue. In this paper, we analyze the cause of partial alignment, including the limited expressive capacity of the models and the lack of alignment signal for non-overlapping users. To address this issue, we propose a new model DACDR (Diffusion Alignment Cross-domain recommendation) with two specially designed modules: a user latent diffusion module and a user alignment module. The user latent diffusion module integrates a diffusion model into the CDR process to enhance the expressive capacity of the model. The user alignment module introduces novel alignment mechanisms that consider both overlapping and non-overlapping users. We conduct extensive experiments on three CDR scenarios to evaluate the performance of DACDR. Our results demonstrate that DACDR outperforms the baselines. Fengxin Li, Hongyan Liu 0002, Jun He 0008 |
ICMR | 2 |
| 2024 | PoseRec: 3D Human Pose Driven Online Advertisement Recommendation for Micro-videosabstractIn this paper, we present PoseRec, an innovative approach aimed at enhancing online advertisement recommendations for micro-videos to boost click-through rates. Addressing the inherent background bias introduced via direct video content learning from image frames, we exploit rich data within the 3D human pose. PoseRec capitalizes on the merits of 3D human pose detection and multi-frame pose data, resulting in superior advertisement recommendation performance. Additionally, we introduce a unique item-aware implicit prototype learning module and a pose-aware transductive hard-negative mining module to tackle the issues of ambiguity and sparsity in advertisement recommendation. Upon evaluation on our novel dataset, Pose-OBE, our method exhibits robust performance surpassing strong baselines, corroborating its effectiveness in resolving the complex challenges of micro-video advertisement recommendation. Zhaoxin Fan, Fengxin Li, Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001 |
ICMR | 3 |
| 2024 | ACR-Pose: Adversarial Canonical Representation Reconstruction Network for Category Level 6D Object Pose EstimationabstractIn the realm of category-level 6D object pose estimation, canonical 3D representation reconstruction is pivotal, yet current methods show limitations in reconstruction quality, a key step in current pose estimation pipeline. To address this, we introduce an innovative Adversarial Canonical Representation Reconstruction Network (ACR-Pose) in this paper. In particular, ACR-Pose comprises a Reconstructor, with novel sub-modules: a Pose-Irrelevant Module (PIM) for robustness to rotation and translation, and a Relational Reconstruction Module (RRM) for extracting relational information between input modalities. A Discriminator is incorporated to guide the generation of realistic canonical representations through adversarial optimization. Evaluated on the prevalent NOCS-CAMERA and NOCS-REAL datasets, our method significantly improves the performance of baseline models and achieves comparable performance with existing state-of-the-art methods, representing a promising advancement in the field of category-level 6D object pose estimation. Zhaoxin Fan, Zhenbo Song, Zhicheng Wang 0007, Jian Xu 0027, Kejian Wu, Hongyan Liu 0002, Jun He 0008 |
ICMR | 6 |
| 2024 | CoDancers: Music-Driven Coherent Group Dance Generation with Choreographic UnitabstractDance and music are intimately interconnected, with group dance being a crucial part of dance artistry. Consequently, Music-Driven Group Dance Generation has been a fundamental and challenging task in various fields like education, art, and sports. However, existing methods fail to fully explore group dance coherence. Thus, we propose CoDancers, a novel and efficient retrieval-based music-driven group dance generation framework. CoDancers improves performance by decomposing group dance coherence into individual movement coherence and group interaction coherence for specialized design, incorporating a Spatial-Temporal Group Dance Blender block, a Acoustic-Semantic Music Miner block, and a Stereotype-Reducing Dance Generator block. Experimental results on the public dataset demonstrate the superiority of our method over existing baselines, achieving state-of-the-art performance. The code is available at https://github.com/XulongT/CoDancers. Kaixing Yang, Xulong Tang, Ran Diao, Hongyan Liu 0002, Jun He 0008, Zhaoxin Fan |
ICMR | 4 |
| 2024 | BeatDance: A Beat-Based Model-Agnostic Contrastive Learning Framework for Music-Dance RetrievalabstractDance and music are closely related forms of expression, with mutual retrieval between dance videos and music being a fundamental task in various fields like education, art, and sports. However, existing methods often suffer from unnatural generation effects or fail to fully explore the correlation between music and dance. To overcome these challenges, we propose BeatDance, a novel beat-based model-agnostic contrastive learning framework. BeatDance incorporates a Beat-Aware Music-Dance InfoExtractor, a Trans-Temporal Beat Blender, and a Beat-Enhanced Hubness Reducer to improve Music-Dance retrieval performance by utilizing the alignment between music beats and dance movements. We also introduce the Music-Dance (M-D) dataset, a large-scale collection of over 10,000 Music-Dance video pairs for training and testing. Experimental results on the M-D dataset demonstrate the superiority of our method over existing baselines, achieving state-of-the-art performance. The code and dataset are available at https://github.com/XulongT/BeatDance. Kaixing Yang, Xukun Zhou, Xulong Tang, Ran Diao, Hongyan Liu 0002, Jun He 0008, Zhaoxin Fan |
ICMR | 5 |
| 2024 | STDG: Semi-Teacher-Student Training Paradigm for Depth-guided One-stage Scene Graph GenerationabstractScene Graph Generation is a critical enabler of environmental comprehension for autonomous robotic systems.Most of existing methods, however, are often thwarted by the intricate dynamics of background complexity, which limits their ability to fully decode the inherent topological information of the environment.Additionally, the wealth of contextual information encapsulated within depth cues is often left untapped, rendering existing approaches less effective.To address these shortcomings, we present STDG, an avant-garde Depth-Guided One-Stage Scene Graph Generation methodology.The innovative architecture of STDG is a triad of custom-built modules: The Depth Guided HHA Representation Generation Module, the Depth Guided Semi-Teaching Network Learning Module, and the Depth Guided Scene Graph Generation Module.This trifecta of modules synergistically harnesses depth information, covering all aspects from depth signal generation and depth feature utilization, to the final scene graph prediction.Importantly, this is achieved without imposing additional computational burden during the inference phase.Experimental results confirm that our method significantly enhances the performance of onestage scene graph generation baselines. Xukun Zhou, Zhenbo Song, Jun He 0008, Hongyan Liu 0002, Zhaoxin Fan |
ICMR | 4 |
| 2024 | CausalCDR: Causal Embedding Learning for Cross-domain RecommendationabstractCross-domain recommendation (CDR) methods achieve success in disentangling user preferences into domain-specific and domain-shared parts. However, recent research has shown that isolated domain-specific preference limits performance improvements. In this paper, we propose a new CDR framework, called CausalCDR, which identifies the limitations of existing methods and addresses existing issues. CausalCDR consists of two views: the causal view and the generative view. The causal view incorporates causality of variables into the CDR scenario, while the generative view implements the causal view by modeling the joint distribution of user interaction via encoding, causal, and generation stage. To optimize CausalCDR, we re-derive the Evidence Lower Bound (ELBO) and introduce a mutual information regularizer and an adversarial classifier. We evaluate CausalCDR on four real-world CDR scenarios and demonstrate its effectiveness in improving CDR performance. Fengxin Li, Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001 |
SDM | 2 |
| 2024 | A profile similarity-based personalized federated learning method for wearable sensor-based human activity recognition
Yidong Chai, Haoxin Liu 0003, Hongyi Zhu 0001, Yue Pan 0019, Anqi Zhou, Hongyan Liu 0002, Yang Qian 0001 |
Inf. Manag. | 6 |
| 2024 | Tagging Items with Emerging Tags: A Neural Topic Model Based Few-Shot Learning ApproachabstractThe tagging system has become a primary tool to organize information resources on the Internet, which benefits both users and the platforms. To build a successful tagging system, automatic tagging methods are desired. With the development of society, new tags keep emerging. The problem of tagging items with emerging tags is an open challenge for an automatic tagging system, and it has not been well studied in the literature. We define this problem as a tag-centered cold-start problem in this study and propose a novel neural topic model based few-shot learning method named NTFSL to solve the problem. In our proposed method, we innovatively fuse the topic modeling task with the few-shot learning task, endowing the model with the capability to infer effective topics to solve the tag-centered cold-start problem with the property of interpretability. Meanwhile, we propose a novel neural topic model for the topic modeling task to improve the quality of inferred topics, which helps enhance the tagging performance. Furthermore, we develop a novel inference method based on the variational auto-encoding framework for model inference. We conducted extensive experiments on two real-world datasets, and the results demonstrate the superior performance of our proposed model compared with state-of-the-art machine learning methods. Case studies also show the interpretability of the model. Shangkun Che, Hongyan Liu 0002 |
ACM Trans. Inf. Syst. | 2 |
| 2021 | Glaucoma diagnosis in the Chinese context: An uncertainty information-centric Bayesian deep learning model
Yidong Chai, Yiyang Bian, Hongyan Liu 0002, Jie Xu 0010 |
Inf. Process. Manag. | 3 |
| 2021 | A deep bi-directional prediction model for live streaming recommendation
Hongyan Liu 0002, Jun He 0008, Sanpu Han, Xiaoyong Du 0001 |
Inf. Process. Manag. | 2 |
| 2020 | Combining Global and Sequential Patterns for Multivariate Time Series ForecastingabstractMultivariate time series forecasting is very important for many applications. Many studies have been conducted for accurate and interpretable prediction methods. However, existing methods either cannot take both times series and covariates into consideration, lacking of interpretability, or ignore global trends across multivariate time series. In this paper, we aim to solve these issues. To this end, we propose a new model named TEDGE for accurate and interpretable time series prediction. In this model, we extract global trends hidden across multivariate times series to improve prediction accuracy. Meanwhile, we utilize a deep recurrent model with attention mechanism to find long-and short-term sequential patterns hidden in individual time series with interpretability. We conduct experiments on several datasets to evaluate the proposed models performance. Results demonstrate the superior performance of our proposed model. Jun He 0008, Hongyan Liu 0002, Xiaoyong Du 0001 |
IEEE BigData | 3 |
| 2020 | A Graph Attentive Network Model for P2P Lending Fraud Detection
Qiyi Wang, Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001 |
KSEM (1) | 2 |
| 2020 | DAGC: Employing Dual Attention and Graph Convolution for Point Cloud based Place RecognitionabstractPoint cloud based retrieval for place recognition remains to be a problem demanding prompt solution due to its difficulty in efficiently encoding local features into adequate global descriptor in scenes. Existing studies solve this problem by generating a global descriptor for each point cloud, which is used to retrieve matched point cloud in database. But existing studies do not make effective use of the relationship between points and neglect different feature's discrimination power. In this paper, we propose to employ Dual Attention and Graph Convolution for point cloud based place recognition (DAGC) to solve these issues. Specifically, we employ two modules to help extract discriminative and generalizable features to describe a point cloud. We introduce a Dual Attention module to help distinguish task-relevant features and to utilize other points' different contributions to a point to generate representation. Meanwhile, we introduce a Residual Graph Convolution Network (ResGCN) module to aggregate local features of each point's multi-level neighbor points to further improve the representation. In this way, we improve the descriptor generation by considering the importance of both point and feature and leveraging point relationship. Experiments conducted on different datasets show that our work outperforms current approaches on all evaluation metrics. Hongyan Liu 0002, Jun He 0008, Zhaoxin Fan, Xiaoyong Du 0001 |
ICMR | 2 |
| 2020 | Bi-Labeled LDA: Inferring Interest Tags for Non-famous Users in Social NetworkabstractAbstract User tags in social network are valuable information for many applications such as Web search, recommender systems and online advertising. Thus, extracting high quality tags to capture user interest has attracted many researchers’ study in recent years. Most previous studies inferred users’ interest based on text posted in social network. In some cases, ordinary users usually only publish a small number of text posts and text information is not related to their interest very much. Compared with famous user, it is more challenging to find non-famous (ordinary) user’s interest. In this paper, we propose a probabilistic topic model,Bi-Labeled LDA,to automatically find interest tags for non-famous users in social network such as Twitter. Instead of extracting tags from text posts, tags of non-famous users are inferred from interest topics of famous users. With the proposed model, the formulation of social relationship between non-famous users and famous user is simulated and interest tags of famous users are exploited to supervise the training of the model and to make use of latent relation among famous users. Furthermore, the influence of popularity of famous user and popular tags are considered, and tags of non-famous users are ranked based on random walk model. Experiments were conducted on Twitter real datasets. Comparison with state-of-the-art methods shows that our method is more superior in terms of both ranking and quality of the tagging results. Jun He 0008, Hongyan Liu 0002, Yiqing Zheng, Shu Tang, Xiaoyong Du 0001 |
Data Sci. Eng. | 2 |
| 2018 | Exploiting Instance Relationship for Effective Extreme Multi-label Learning
Feifei Li 0002, Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001 |
DASFAA (2) | 2 |
| 2017 | Topic Analysis and Influential Paper Discovery on Scientific PublicationsabstractWith the development of scientific research, scientific publications are valuable resources for new-comers in the research field. But massive scientific publications make it a challenge for researchers diving into a new research field. As a good practice to this problem, topics are put forward to organize publications. In this paper, we propose two modified LDA topic models as solutions to topic analysis and influential paper discovery on scientific publications, cc-LDA and cp-LDA. Compared to state-of-the-art researches on LDA, we incorporate citation information including its occurrence times and occurrence position into our models. Model cc-LDA integrates paper content and citation occurrence into LDA model, while cp-LDA considers both occurrence and position of citations. Both models can not only find topics in the form of citation distribution, but also help discover influential papers under certain topics. Furthermore, both models can extract more representative vectors for papers, which achieve good performance in subsequent clustering. Jun He 0008, Hongyan Liu 0002 |
WISA | 3 |
| 2017 | Mining Exploratory Behavior to Improve Mobile App RecommendationsabstractWith the widespread usage of smart phones, more and more mobile apps are developed every day, playing an increasingly important role in changing our lifestyles and business models. In this trend, it becomes a hot research topic for developing effective mobile app recommender systems in both industry and academia. Compared with existing studies about mobile app recommendations, our research aims to improve the recommendation effectiveness based on analyzing a psychological trait of human beings, exploratory behavior, which refers to a type of variety-seeking behavior in unfamiliar domains. To this end, we propose a novel probabilistic model named Goal-oriented Exploratory Model (GEM), integrating exploratory behavior identification with personalized item recommendation. An algorithm combining collapsed Gibbs sampling and Expectation Maximization is developed for model learning and inference. Through extensive experiments conducted on a real dataset, the proposed model demonstrates superior recommendation performances and good interpretability compared with state-of-art recommendation methods. Moreover, empirical analyses on exploratory behavior find that individuals with a strong exploratory tendency exhibit behavioral patterns of variety seeking, risk taking, and higher involvement. Besides, mobile apps that are less popular or in the long tail possess greater potential of arousing exploratory behavior in individuals. Jiangning He, Hongyan Liu 0002 |
ACM Trans. Inf. Syst. | 2 |
| 2016 | Finding Latest Influential Research Papers Through Modeling Two Views of Citation Links
Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001 |
APWeb (1) | 2 |
| 2016 | SocoTraveler: Travel-package recommendations leveraging social influence of different relationship types
Jiangning He, Hongyan Liu 0002, Hui Xiong 0001 |
Inf. Manag. | 2 |
| 2015 | Extracting Interest Tags for Non-famous Users in Social NetworkabstractInferring interests of users in social network is important for many applications such as personalized search, recommender systems and online advertising. Most previous studies inferred users' interests based on text posted in social network, which is usually not related to their interests. In this paper, we propose a modified topic model, Bi-Labeled LDA with a term weighting scheme, to extract interest tags for users in social network. The proposed model utilize only users' relationship information without requirement for text information, and incorporates supervision into traditional LDA. Specifically, we introduce method to extract tags for non-famous user through their relationship with famous users in Twitter, and study why a non-famous user follows famous users simultaneously. Comparison with state-of-the-art methods on real dataset shows that our method is far more superior in terms of precision and recall of the extracted tag set, and also more applicable for many personalized applications. Besides, we find that a reasonable term weighting scheme can actually improve the performance further. Hongyan Liu 0002, Jun He 0008, Shu Tang, Xiaoyong Du 0001 |
CIKM | 2 |
| 2014 | Constructing Decision Trees for Unstructured Data
Shucheng Gong, Hongyan Liu 0002 |
ADMA | 2 |
| 2014 | Selecting a Representative Set of Diverse Quality Reviews AutomaticallyabstractOnline user reviews are important information for both consumers and vendors. More and more people make their purchase decisions based on online reviews. Vendors also pay more and more attention to online reviews. However, as the number of reviews increases rapidly, the information overload problem prevents us making full use of online reviews. In this paper, we study how to find a representative set of high quality reviews to cover diversified aspects of user opinions. Existing work cannot solve this problem well. To overcome the drawbacks of existing methods, we define a new problem of finding the minimum set of reviews to cover all of features with different sentiment polarity and high quality without user-defined parameters. To solve the problem efficiently, we define potential objective function and develop greedy algorithm to find the solution in polynomial-time with approximation guarantee. We also propose two strategies to further reduce the number of reviews and to prune the search space respectively. Comprehensive experiments conducted on real review sets show that the proposed methods are effective and outperform existing methods. Nana Xu, Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001 |
SDM | 2 |
| 2014 | Assessing single-pair similarity over graphs by aggregating first-meeting probabilities
Jun He 0008, Hongyan Liu 0002, Jeffrey Xu Yu, Xiaoyong Du 0001 |
Inf. Syst. | 2 |
| 2013 | Detecting Event Rumors on Sina Weibo Automatically
Shengyun Sun, Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001 |
APWeb | 2 |
| 2013 | Mining User Interests from Information Sharing Behaviors in Social Media
Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001 |
PAKDD (2) | 2 |
| 2013 | Predicting Microblog User's Age Based on Text Information
Tao Liu 0001, Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001 |
WISE (1) | 3 |
| 2013 | Measuring Similarity Based on Link Information: A Comparative StudyabstractMeasuring similarity between objects is a fundamental task in domains such as data mining, information retrieval, and so on. Link-based similarity measures have attracted the attention of many researchers and have been widely applied in recent years. However, most previous works mainly focus on introducing new link-based measures, and seldom provide theoretical as well as experimental comparisons with other measures. Thus, selecting the suitable measure in different situations and applications is difficult. In this paper, a comprehensive analysis and critical comparison of various link-based similarity measures and algorithms are presented. Their strengths and weaknesses are discussed. Their actual runtime performances are also compared via experiments on benchmark data sets. Some novel and useful guidelines for users to choose the appropriate link-based measure for their applications are discovered. Hongyan Liu 0002, Jun He 0008, Charles Ling 0001, Xiaoyong Du 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2012 | Combining Spatial Cloaking and Dummy Generation for Location Privacy Preserving
Nana Xu, Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001, Tao Liu 0001 |
ADMA | 3 |
| 2012 | Bayesian Network Structure Learning from Attribute Uncertain Data
Wenting Song, Jeffrey Xu Yu, Hong Cheng 0001, Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001 |
WAIM | 4 |
| 2012 | Predicting Retweet Behavior in Weibo Social Network
Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001, Hong Chen 0001 |
WISE | 3 |
| 2012 | Detecting and Tracking Topics and Events from Web Search LogsabstractRecent years have witnessed increased efforts on detecting topics and events from Web search logs, since this kind of data not only capture web content but also reflect the users’ activities. However, the majority of existing work is focused on exploiting clustering techniques for topic and event detection. Due to the huge size and the evolving nature of Web data, existing clustering approaches are limited to meet the real-time demand. To that end, in this article, we propose a method called LETD to detect evolving topics in a timely manner. Also, we design the techniques to extract events from topics and to infer the evolving relationship among the events. For topic detection, we first provide a measurement to select the important URLs, which are most likely to describe a real-life topic. Then, starting from these selected URLs, we exploit the local expansion method to find other topic-related URLs. Moreover, in the LETD framework, we design algorithms based on Random Walk and Markov Random Fields (MRF), respectively. Because the LETD method exploits a divide-and-conquer strategy to process the data, it is more efficient than existing methods based on clustering techniques. To better illustrate the LETD framework, we develop a demo system StoryTeller which can discover hot topics and events, infer the evolving relationships among events, and visualize information in a storytelling way. This demo system can provide a global view of the topic development and help users target the interesting events more conveniently. Finally, experimental results on real-world Microsoft click-through data have shown that StoryTeller can find real-life hot topics and meaningful evolving relationships among events, and has also demonstrated the efficiency and effectiveness of the LETD method. Hongyan Liu 0002, Jun He 0008, Yingqin Gu, Hui Xiong 0001, Xiaoyong Du 0001 |
ACM Trans. Inf. Syst. | 1 |
| 2011 | Predicting New User's Behavior in Online Dating Systems
Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001 |
ADMA (2) | 2 |
| 2011 | Multi-view random walk framework for search task discovery from click-through logabstractSearch engine users often have clear search tasks hidden behind their queries. Inspired by this, the modern search engines are providing an increasing number of services to help users simplify their key tasks. However, the problem of what are the major user search tasks with high traffic for which search engines should design special services is still underexplored. In this paper, we propose a novel Multi-view Random Walk (MRW) algorithm to measure the search task oriented similarity between queries, and then group search queries with similar tasks so that the major search tasks of users can be identified from search engine click-through log. The proposed MRW, which is a general framework to combine knowledge from different views in a random walk process, allows the random surfer to walk across different views to integrate information for search task discovery. Experimental results on click-through log of a commonly used commercial search engine show that our proposed MRW algorithm can effectively discover user search tasks. Hongyan Liu 0002, Jun Yan 0001, Lei Ji 0001, Ruoming Jin, Jun He 0008, Yingqin Gu, Zheng Chen 0001, Xiaoyong Du 0001 |
CIKM | 2 |
| 2011 | Extract knowledge from semi-structured websites for search task simplificationabstractSimplifying the key tasks of search engine users by directly retrieving to them structured knowledge according to their queries is attracting much attention from both industry and academia. A bottleneck of this challenging problem is how to extract the structured knowledge from the noisy and complex Web scale websites automatically. In this paper, we propose an unsupervised automatic wrapper induction algorithm, named as Scalable Knowledge Extractor from webSites (SKES). SKES induces the wrapper in a divide and conquer mode, i.e., it divides the general wrapper into several sub-wrappers to learn from the data independently. Moreover, through employing techniques such as tag path representation of Web pages, SKES is verified to be efficient and noise-tolerant by the experimental results. Furthermore, based on our automatically extracted knowledge, we also built a prototype to serve structured knowledge to end users for simplifying their key search tasks. Very positive feedbacks were received on the prototype. Yingqin Gu, Jun Yan 0001, Hongyan Liu 0002, Jun He 0008, Lei Ji 0001, Ning Liu 0001, Zheng Chen 0001 |
CIKM | 3 |
| 2011 | Ranking Individuals and Groups by Influence Propagation
Jeffrey Xu Yu, Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001 |
PAKDD (2) | 3 |
| 2011 | TagClus: a random walk-based method for tag clustering
Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001, Puwei Wang |
Knowl. Inf. Syst. | 2 |
| 2011 | Methods for mining frequent items in data streams: an overview
Hongyan Liu 0002, Jiawei Han 0001 |
Knowl. Inf. Syst. | 1 |
| 2010 | Local Methods for Estimating SimRank ScoreabstractSimRank is a well known algorithm which conducts link analysis to measure similarity between each pair of nodes (nodepair). But it suffers from high computational cost, limiting its usage in large-scale datasets. Moreover, Links between nodes are changing over time. It may be desirable to quickly approximate the similarity score between certain nodepair without performing a large-scale computation on the entire graph. In our approach we propose a method to efficiently estimate the similarity score using only a small subgraph of the entire graph. We call this novel algorithm “Local-SimRank”. The experimental results conducted on real datasets and synthetic dataset show that our algorithm efficiently produces good approximations to the global SimRank scores. Meanwhile, we prove that the Local-SimRank score LS(a, b) is always less than original SimRank score S(a, b) mathematically. Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001, Yuanzhe Cai |
APWeb | 2 |
| 2010 | Mining Closed Episodes from Event Sequences Efficiently
Wenzhi Zhou, Hongyan Liu 0002, Hong Cheng 0001 |
PAKDD (1) | 2 |
| 2010 | Fast Single-Pair SimRank ComputationabstractSimRank is an intuitive and effective measure for link-based similarity that scores similarity between two nodes as the first-meeting probability of two random surfers, based on the random surfer model. However, when a user queries the similarity of a given node-pair based on SimRank, the existing approaches need to compute the similarities of other node-pairs beforehand, which we call an all-pair style. In this paper, we propose a Single-Pair SimRank approach. Without accuracy loss, this approach performs an iterative computation to obtain the similarity of a single node-pair. The time cost of our Single-Pair SimRank is always less than All-Pair SimRank and obviously efficient when we only need to assess similarity of one or a few node-pairs. We confirm the accuracy and efficiency of our approach in extensive experimental studies over synthetic and real datasets. Hongyan Liu 0002, Jeffrey Xu Yu, Jun He 0008, Xiaoyong Du 0001 |
SDM | 2 |
| 2010 | Detecting Hot Events from Web Search Logs
Yingqin Gu, Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001, Zhixu Li |
WAIM | 3 |
| 2010 | Comments on "an integrated efficient solution for computing frequent and top-k elements in data streams"abstractWe investigate a well-known algorithm, Space-Saving [Metwally et al. 2006], which has been proven efficient and effective at mining frequent elements in data streams. We discovered an error in one of the theorems in Metwally et al. [2006]. Experiments are conducted to illustrate the error. Hongyan Liu 0002, Yinghui Yang 0001 |
ACM Trans. Database Syst. | 1 |
| 2010 | Mining near-duplicate graph for cluster-based reranking of web video search resultsabstractRecently, video search reranking has been an effective mechanism to improve the initial text-based ranking list by incorporating visual consistency among the result videos. While existing methods attempt to rerank all the individual result videos, they suffer from several drawbacks. In this article, we propose a new video reranking paradigm called cluster-based video reranking (CVR). The idea is to first construct a video near-duplicate graph representing the visual similarity relationship among videos, followed by identifying the near-duplicate clusters from the video near-duplicate graph, then ranking the obtained near-duplicate clusters based on cluster properties and intercluster links, and finally for each ranked cluster, a representative video is selected and returned. Compared to existing methods, the new CVR ranks clusters and exhibits several advantages, including superior reranking by utilizing more reliable cluster properties, fast reranking on a small number of clusters, diverse and representative results. Particularly, we formulate the near-duplicate cluster identification as a novel maximally cohesive subgraph mining problem. By leveraging the designed cluster scoring properties indicating the cluster's importance and quality, random walk is applied over the near-duplicate cluster graph to rank clusters. An extensive evaluation study proves the novelty and superiority of our proposals over existing methods. Zi Huang, Hong Cheng 0001, Heng Tao Shen, Hongyan Liu 0002, Xiaofang Zhou 0001 |
ACM Trans. Inf. Syst. | 5 |
| 2009 | A Neighborhood Search Method for Link-Based Tag Clustering
Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001 |
ADMA | 3 |
| 2009 | Calculating Similarity Efficiently in a Small World
Yuanzhe Cai, Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001 |
ADMA | 3 |
| 2009 | Assessing the influence probability between objects: A random walker approachabstractInfluence between objects needs to be assessed in many applications. Lots of measures have been proposed, but a domain-independent method is still expected. In this paper, we give a probabilistic definition of influence based on the random walker model on graphs. Two approaches, linear systems method and Basic InfRank algorithm, are shown and return equal results, but Basic InfRank is more efficient by iterative computation. Two variants on bipartite graphs and star graphs are discussed. Experiments show InfRank algorithms have good accuracy, fast convergent rate and high performance. Zhixu Li, Jun He 0008, Xiaoyong Du 0001, Hongyan Liu 0002 |
CIDM | 5 |
| 2009 | An Adaptive Method for the Efficient Similarity Calculation
Yuanzhe Cai, Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001 |
DASFAA | 2 |
| 2009 | Efficient Algorithm for Computing Link-Based Similarity in Real World NetworksabstractSimilarity calculation has many applications, such as information retrieval, and collaborative filtering, among many others. It has been shown that link-based similarity measure, such as SimRank, is very effective in characterizing the object similarities in networks, such as the Web, by exploiting the object-to-object relationship. Unfortunately, it is prohibitively expensive to compute the link-based similarity in a relatively large graph. In this paper, based on the observation that link-based similarity scores of real world graphs follow the power-law distribution, we propose a new approximate algorithm, namely Power-SimRank, with guaranteed error bound to efficiently compute link-based similarity measure. We also prove the convergence of the proposed algorithm. Extensive experiments conducted on real world datasets and synthetic datasets show that the proposed algorithm outperforms SimRank by four-five times in terms of efficiency while the error generated by the approximation is small. Yuanzhe Cai, Gao Cong, Hongyan Liu 0002, Jun He 0008, Jiaheng Lu, Xiaoyong Du 0001 |
ICDM | 4 |
| 2009 | Exploiting the Block Structure of Link Graph for Efficient Similarity Computation
Yuanzhe Cai, Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001 |
PAKDD | 3 |
| 2009 | Top-down mining of frequent closed patterns from very high dimensional data
Hongyan Liu 0002, Jun He 0008, Jiawei Han 0001, Dong Xin, Zheng Shao |
Inf. Sci. | 1 |
| 2008 | S-SimRank: Combining Content and Link Information to Cluster Papers Effectively and Efficiently
Yuanzhe Cai, Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001 |
ADMA | 3 |
| 2008 | FARS: A Multi-relational Feature and Relation Selection Approach for Efficient Classification
Hongyan Liu 0002, Jun He 0008, Xiaoyong Du 0001 |
ADMA | 2 |
| 2008 | CRO: a system for online review structurizationabstractIn this paper, we present a system called CRO (Chinese Review Observer) for online product review structurization. By Structurization, we mean identifying, extracting and summarizing information from unstructured review text to a structured table. The core tasks include review collection, product feature and user opinion extraction, and polarity analysis of opinions. Existing research in this area is mainly English text oriented. To deal with Chinese effectively, we propose several novel approaches for fulfilling the core tasks. Then we integrated these approaches and implement the whole procedure of review structurization in the system CRO. Running results for reviews of real products show its performance is satisfactory. Hongyan Liu 0002, Hui Yang 0009, Jun He 0008, Xiaoyong Du 0001 |
KDD | 1 |
| 2007 | Separator: Sifting Hierarchical Heavy Hitters Accurately from Data Streams
Hongyan Liu 0002 |
ADMA | 2 |
| 2007 | DELAY : A Lazy Approach for Mining Frequent Patterns over High Speed Data Streams
Hui Yang 0009, Hongyan Liu 0002, Jun He 0008 |
ADMA | 2 |
| 2006 | Error-Adaptive and Time-Aware Maintenance of Frequency Counts over Data Streams
Hongyan Liu 0002, Ying Lu 0001, Jiawei Han 0001, Jun He 0008 |
WAIM | 1 |