VLDB 2026 Research / reviewers in the wild / expert
Yunhua Hu
dblp:34/3973
· DBLP profile ↗
16ranked-venue papers
4as first author
3since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 10 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 7 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Time series and sequential data · 97% Learning paradigms · 3% | |
| Databases, data mining, and information retrieval
5 papers |
Information retrieval · 100% |
Topics — the 17 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Time series and sequential data › time series analysis › time series forecasting
hierarchical reconciliation |
1.4 | 2 | 2024 | GMP-AR: Granularity Message Passing and Adaptive Reconciliation for Temporal Hierarchy Forecasting · AAAI 2024 SLOTH: Structured Learning and Task-Based Optimization for Time Series Forecasting on Hierarchies · AAAI 2023 |
Machine learning › Time series and sequential data › time series analysis › time series forecasting
hierarchical time series forecasting |
1.4 | 2 | 2024 | GMP-AR: Granularity Message Passing and Adaptive Reconciliation for Temporal Hierarchy Forecasting · AAAI 2024 SLOTH: Structured Learning and Task-Based Optimization for Time Series Forecasting on Hierarchies · AAAI 2023 |
Machine learning › Time series and sequential data › time series analysis
time series forecasting |
1.4 | 2 | 2024 | GMP-AR: Granularity Message Passing and Adaptive Reconciliation for Temporal Hierarchy Forecasting · AAAI 2024 SLOTH: Structured Learning and Task-Based Optimization for Time Series Forecasting on Hierarchies · AAAI 2023 |
Information retrieval
query understanding |
0.7 | 2 | 2022 | Building Multi-turn Query Interpreters for E-commercial Chatbots with Sparse-to-dense Attentive Modeling · WSDM 2022 Mining query subtopics from search log data · SIGIR 2012 |
Information retrieval
dialogue systems |
0.6 | 1 | 2022 | Building Multi-turn Query Interpreters for E-commercial Chatbots with Sparse-to-dense Attentive Modeling · WSDM 2022 |
Information retrieval › query understanding › query classification
query intent classification |
0.6 | 1 | 2022 | Building Multi-turn Query Interpreters for E-commercial Chatbots with Sparse-to-dense Attentive Modeling · WSDM 2022 |
Information retrieval › ranking › learning to rank
feature-based ranking |
0.1 | 1 | 2012 | Extracting search-focused key n-grams for relevance ranking in web search · WSDM 2012 |
Information retrieval › ranking
learning to rank |
0.1 | 1 | 2012 | Extracting search-focused key n-grams for relevance ranking in web search · WSDM 2012 |
Information retrieval › ranking › search ranking
relevance ranking |
0.1 | 1 | 2012 | Extracting search-focused key n-grams for relevance ranking in web search · WSDM 2012 |
Information retrieval › search engines
search result clustering |
0.1 | 1 | 2012 | Mining query subtopics from search log data · SIGIR 2012 |
Information retrieval › query understanding
subtopic mining |
0.1 | 1 | 2012 | Mining query subtopics from search log data · SIGIR 2012 |
Machine learning › Learning paradigms
multi-task learning |
0.1 | 1 | 2011 | Multi-Task Learning in Square Integrable Space · AAAI 2011 |
Information retrieval › text analysis
keyword extraction |
0.1 | 1 | 2009 | A ranking approach to keyphrase extraction · SIGIR 2009 |
Information retrieval
retrieval models |
0.1 | 1 | 2005 | Title extraction from bodies of HTML documents and its application to web page retrieval · SIGIR 2005 |
Information retrieval › web search
web information retrieval |
0.1 | 1 | 2005 | Title extraction from bodies of HTML documents and its application to web page retrieval · SIGIR 2005 |
Information retrieval
ranking |
0.0 | 1 | 2012 | Mining query subtopics from search log data · SIGIR 2012 |
Information retrieval
reranking |
0.0 | 1 | 2012 | Mining query subtopics from search log data · SIGIR 2012 |
Methods — techniques the papers use, named apart from their topics
task-based optimization · 0.8granularity message passing · 0.8top-down convolution · 0.7neural optimization · 0.7bottom-up attention · 0.7sparse-to-dense attentive modeling · 0.6pre-trained language model · 0.6hierarchical multi-grained classification · 0.6SAM-BERT · 0.6learning to rank · 0.2search log mining · 0.1search log analysis · 0.1clustering · 0.1ranking SVM · 0.1supervised machine learning · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | GMP-AR: Granularity Message Passing and Adaptive Reconciliation for Temporal Hierarchy ForecastingabstractTime series forecasts of different temporal granularity are widely used in real-world applications, e.g., sales prediction in days and weeks for making different inventory plans. However, these tasks are usually solved separately without ensuring coherence, which is crucial for aligning downstream decisions. Previous works mainly focus on ensuring coherence with some straightforward methods, e.g., aggregation from the forecasts of fine granularity to the coarse ones, and allocation from the coarse granularity to the fine ones. These methods merely take the temporal hierarchical structure to maintain coherence without improving the forecasting accuracy. In this paper, we propose a novel granularity message-passing mechanism (GMP) that leverages temporal hierarchy information to improve forecasting performance and also utilizes an adaptive reconciliation (AR) strategy to maintain coherence without performance loss. Furthermore, we introduce an optimization module to achieve task-based targets while adhering to more real-world constraints. Experiments on real-world datasets demonstrate that our framework (GMP-AR) achieves superior performances on temporal hierarchical forecasting tasks compared to state-of-the-art methods. In addition, our framework has been successfully applied to a real-world task of payment traffic management in Alipay by integrating with the task-based optimization module. Fan Zhou 0012, Lintao Ma, Yu Liu 0071, Siqiao Xue, James Y. Zhang, Jun Zhou 0011, Hongyuan Mei, Weitao Lin, Zi Zhuang, Wenxin Ning, Yunhua Hu |
AAAI | 12 |
| 2023 | SLOTH: Structured Learning and Task-Based Optimization for Time Series Forecasting on HierarchiesabstractMultivariate time series forecasting with hierarchical structure is widely used in real-world applications, e.g., sales predictions for the geographical hierarchy formed by cities, states, and countries. The hierarchical time series (HTS) forecasting includes two sub-tasks, i.e., forecasting and reconciliation. In the previous works, hierarchical information is only integrated in the reconciliation step to maintain coherency, but not in forecasting step for accuracy improvement. In this paper, we propose two novel tree-based feature integration mechanisms, i.e., top-down convolution and bottom-up attention to leverage the information of the hierarchical structure to improve the forecasting performance. Moreover, unlike most previous reconciliation methods which either rely on strong assumptions or focus on coherent constraints only, we utilize deep neural optimization networks, which not only achieve coherency without any assumptions, but also allow more flexible and realistic constraints to achieve task-based targets, e.g., lower under-estimation penalty and meaningful decision-making loss to facilitate the subsequent downstream tasks. Experiments on real-world datasets demonstrate that our tree-based feature integration mechanism achieves superior performances on hierarchical forecasting tasks compared to the state-of-the-art methods, and our neural optimization networks can be applied to real-world tasks effectively without any additional effort under coherence and task-based constraints. Fan Zhou 0012, Lintao Ma, Yu Liu 0071, Shiyu Wang 0001, James Zhang, Xuanwei Hu, Yunhua Hu, Yangfei Zheng, Lei Lei 0001, Hu Yun |
AAAI | 9 |
| 2022 | Building Multi-turn Query Interpreters for E-commercial Chatbots with Sparse-to-dense Attentive ModelingabstractPredicting query intents is crucial for understanding user demands in chatbots. In real-world applications, accurate query intent classification can be highly challenging as human-machine interactions are often conducted in multiple turns, which requires the models to capture related information from the entire contexts. In addition, query intents tend to be fine-grained (up to hundreds of classes), containing lots of casual chats without clear intents. Hence, it is difficult for standard transformer-based models to capture complicated language characteristics of dialogues to support these applications. In this demo, we present AliMeTerp, a multi-turn query interpretation system, which can be seamlessly integrated into e-commercial chatbots in order to generate appropriate responses. Specifically, in AliMeTerp, we introduce SAM-BERT, a pre-trained language model for fine-grained query intent understanding, based on Sparse-to-dense Attentive Modeling. For model pre-training, a stack of Sparse-to-dense Attentive Encoders are employed to model the complicated dialogue structures from different levels. We further design Hierarchical Multi-grained Classification tasks for model fine-tuning. Experiments show SAM-BERT consistently outperforms strong baselines over multiple multi-turn chatbot datasets. We further show how AliMeTerp is deployed in real-world e-commercial chatbots to support real-time customer service. Yan Fan 0004, Chengyu Wang 0001, Yunhua Hu |
WSDM | 4 |
| 2016 | Evaluation of geo-ecological environment bearing capacity along Dujiangyan-Wenchuan highwayabstractThis paper puts forward an evaluation index system with 20 indexes and quantitive criteria for geographical ecological environment bearing capacity based on national geographic conditions census data; the evaluation model is determined by comprehensive evaluation method and weighted by analytic hierarchy process; comprehensive suitability zone is introduced through the evaluation results of bearing capacity; taking Dujiangyan-Wenchuan Highway lines as research area, this paper conducts bearing capacity evaluation and construction suitable zoning. Guiyang Yu, Yunhua Hu |
IGARSS | 3 |
| 2015 | A new approach to query segmentation for relevance ranking in web search
Haocheng Wu, Yunhua Hu, Hang Li 0001, Enhong Chen |
Inf. Retr. J. | 2 |
| 2012 | Mining query subtopics from search log dataabstractMost queries in web search are ambiguous and multifaceted. Identifying the major senses and facets of queries from search log data, referred to as query subtopic mining in this paper, is a very important issue in web search. Through search log analysis, we show that there are two interesting phenomena of user behavior that can be leveraged to identify query subtopics, referred to as `one subtopic per search' and `subtopic clarification by keyword'. One subtopic per search means that if a user clicks multiple URLs in one query, then the clicked URLs tend to represent the same sense or facet. Subtopic clarification by keyword means that users often add an additional keyword or keywords to expand the query in order to clarify their search intent. Thus, the keywords tend to be indicative of the sense or facet. We propose a clustering algorithm that can effectively leverage the two phenomena to automatically mine the major subtopics of queries, where each subtopic is represented by a cluster containing a number of URLs and keywords. The mined subtopics of queries can be used in multiple tasks in web search and we evaluate them in aspects of the search result presentation such as clustering and re-ranking. We demonstrate that our clustering algorithm can effectively mine query subtopics with an F1 measure in the range of 0.896-0.956. Our experimental results show that the use of the subtopics mined by our approach can significantly improve the state-of-the-art methods used for search result clustering. Experimental results based on click data also show that the re-ranking of search result based on our method can significantly improve the efficiency of users' ability to find information. Yunhua Hu, Ya-nan Qian, Hang Li 0001, Daxin Jiang, Jian Pei 0001 |
SIGIR | 1 |
| 2012 | Extracting search-focused key n-grams for relevance ranking in web searchabstractIn web search, relevance ranking of popular pages is relatively easy, because of the inclusion of strong signals such as anchor text and search log data. In contrast, with less popular pages, relevance ranking becomes very challenging due to a lack of information. In this paper the former is referred to as head pages, and the latter tail pages. We address the challenge by learning a model that can extract search-focused key n-grams from web pages, and using the key n-grams for searches of the pages, particularly, the tail pages. To the best of our knowledge, this problem has not been previously studied. Our approach has four characteristics. First, key n-grams are search-focused in the sense that they are defined as those which can compose "good queries" for searching the page. Second, key n-grams are learned in a relative sense using learning to rank techniques. Third, key n-grams are learned using search log data, such that the characteristics of key n-grams in the search log data, particularly in the heads; can be applied to the other data, particularly to the tails. Fourth, the extracted key n-grams are used as features of the relevance ranking model also trained with learning to rank techniques. Experiments validate the effectiveness of the proposed approach with large-scale web search datasets. The results show that our approach can significantly improve relevance ranking performance on both heads and tails; and particularly tails, compared with baseline approaches. Characteristics of our approach have also been fully investigated through comprehensive experiments. Keping Bi, Yunhua Hu, Hang Li 0001, Guihong Cao |
WSDM | 3 |
| 2012 | A Kernel Approach to Multi-Task Learning with Task-Specific Kernels
Wei Wu 0014, Hang Li 0001, Yunhua Hu, Rong Jin 0001 |
J. Comput. Sci. Technol. | 3 |
| 2011 | Multi-Task Learning in Square Integrable Space
Wei Wu 0014, Hang Li 0001, Yunhua Hu, Rong Jin 0001 |
AAAI | 3 |
| 2011 | Combining machine learning and human judgment in author disambiguationabstractAuthor disambiguation in digital libraries becomes increasingly difficult as the number of publications and consequently the number of ambiguous author names keep growing. The fully automatic author disambiguation approach could not give satisfactory results due to the lack of signals in many cases. Furthermore, human judgment on the basis of automatic algorithms is also not suitable because the automatically disambiguated results are often mixed and not understandable for humans. In this paper, we propose a Labeling Oriented Author Disambiguation approach, called LOAD, to combine machine learning and human judgment together in author disambiguation. LOAD exploits a framework which consists of high precision clustering, high recall clustering, and top dissimilar clusters selection and ranking. In the framework, supervised learning algorithms are used to train the similarity functions between publications and a clustering algorithm is further applied to generate clusters. To validate the effectiveness and efficiency of the proposed LOAD approach, comprehensive experiments are conducted. Comparing to conventional author disambiguation algorithms, the LOAD yields much more accurate results to assist human labeling. Further experiments show that the LOAD approach can save labeling time dramatically. Ya-nan Qian, Yunhua Hu, Jianling Cui, Zaiqing Nie |
CIKM | 2 |
| 2009 | A ranking approach to keyphrase extractionabstractThis paper addresses the issue of automatically extracting keyphrases from a document. Previously, this problem was formalized as classification and learning methods for classification were utilized. This paper points out that it is more essential to cast the problem as ranking and employ a learning to rank method to perform the task. Specifically, it employs Ranking SVM, a state-of-art method of learning to rank, in keyphrase extraction. Experimental results on three datasets show that Ranking SVM significantly outperforms the baseline methods of SVM and Naive Bayes, indicating that it is better to exploit learning to rank techniques in keyphrase extraction. Xin Jiang 0002, Yunhua Hu, Hang Li 0001 |
SIGIR | 2 |
| 2007 | Web page title extraction and its application
Yewei Xue, Yunhua Hu, Guomao Xin, Ruihua Song, Shuming Shi 0001, Yunbo Cao, Chin-Yew Lin, Hang Li 0001 |
Inf. Process. Manag. | 2 |
| 2006 | Automatic extraction of titles from general documents using machine learning
Yunhua Hu, Hang Li 0001, Yunbo Cao, Li Teng 0002, Dmitriy Meyerzon |
Inf. Process. Manag. | 1 |
| 2005 | A new approach to intranet search based on information extractionabstractThis paper is concerned with 'intranet search'. By intranet search, we mean searching for information on an intranet within an organization. We have found that search needs on an intranet can be categorized into types, through an analysis of survey results and an analysis of search log data. The types include searching for definitions, persons, experts, and homepages. Traditional information retrieval only focuses on search of relevant documents, but not on search of special types of information. We propose a new approach to intranet search in which we search for information in each of the special types, in addition to the traditional relevance search. Information extraction technologies can play key roles in such kind of 'search by type' approach, because we must first extract from the documents the necessary information in each type. We have developed an intranet search system called 'Information Desk'. In the system, we try to address the most important types of search first - finding term definitions, homepages of groups or topics, employees' personal information and experts on topics. For each type of search, we use information extraction technologies to extract, fuse, and summarize information in advance. The system is in operation on the intranet of Microsoft and receives accesses from about 500 employees per month. Feedbacks from users and system logs show that users consider the approach useful and the system can really help people to find information. This paper describes the architecture, features, component technologies, and evaluation results of the system. Hang Li 0001, Yunbo Cao, Jun Xu 0001, Yunhua Hu, Shenjie Li, Dmitriy Meyerzon |
CIKM | 4 |
| 2005 | Taxonomy Building and Machine Learning Based Automatic Classification for Knowledge-Oriented Chinese Questions
Yunhua Hu, Huixian Bai, Haifeng Dang |
ICIC (1) | 1 |
| 2005 | Title extraction from bodies of HTML documents and its application to web page retrievalabstractThis paper is concerned with automatic extraction of titles from the bodies of HTML documents. Titles of HTML documents should be correctly defined in the title fields; however, in reality HTML titles are often bogus. It is desirable to conduct automatic extraction of titles from the bodies of HTML documents. This is an issue which does not seem to have been investigated previously. In this paper, we take a supervised machine learning approach to address the problem. We propose a specification on HTML titles. We utilize format information such as font size, position, and font weight as features in title extraction. Our method significantly outperforms the baseline method of using the lines in largest font size as title (20.9%-32.6% improvement in F1 score). As application, we consider web page retrieval. We use the TREC Web Track data for evaluation. We propose a new method for HTML documents retrieval using extracted titles. Experimental results indicate that the use of both extracted titles and title fields is almost always better than the use of title fields alone; the use of extracted titles is particularly helpful in the task of named page finding (23.1% -29.0% improvements). Yunhua Hu, Guomao Xin, Ruihua Song, Shuming Shi 0001, Yunbo Cao, Hang Li 0001 |
SIGIR | 1 |