VLDB 2026 Research / reviewers in the wild / expert
Huan Gui
dblp:153/5327
· DBLP profile ↗
12ranked-venue papers
6as first author
1since 2021 · last 2024
0000-0001-5621-1753ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 4 first-author · 1 since 2021Databases, data management, data science and information retrieval · 7 · 5 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Trustworthy machine learning · 39% Information extraction and text analysis · 15% Transfer learning and domain adaptation · 14% | |
| Databases, data mining, and information retrieval
3 papers |
Data mining · 89% Information retrieval · 11% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational social science and digital humanities · 100% |
Topics — the 26 heaviest of 28, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
robustness |
0.9 | 2 | 2024 | LEVI: Generalizable Fine-tuning via Layer-wise Ensemble of Different Views · ICML 2024 Robust Tensor Decomposition with Gross Corruption · NIPS 2014 |
Machine learning › Transfer learning and domain adaptation
fine-tuning |
0.8 | 1 | 2024 | LEVI: Generalizable Fine-tuning via Layer-wise Ensemble of Different Views · ICML 2024 |
Machine learning › Trustworthy machine learning
out-of-distribution generalization |
0.8 | 1 | 2024 | LEVI: Generalizable Fine-tuning via Layer-wise Ensemble of Different Views · ICML 2024 |
Machine learning › Trustworthy machine learning › robustness › distribution shift
robustness to distribution shift |
0.8 | 1 | 2024 | LEVI: Generalizable Fine-tuning via Layer-wise Ensemble of Different Views · ICML 2024 |
Data mining › structured data mining › graph mining
heterogeneous information network |
0.6 | 2 | 2017 | Embedding Learning with Events in Heterogeneous Information Networks · IEEE Trans. Knowl. Data Eng. 2017 PReP: Path-Based Relevance from a Probabilistic Perspective in Heterogeneous Information Networks · KDD 2017 |
Data mining › representation learning
hyperedge-based embedding |
0.5 | 2 | 2017 | Embedding Learning with Events in Heterogeneous Information Networks · IEEE Trans. Knowl. Data Eng. 2017 Large-Scale Embedding Learning in Heterogeneous Event Data · ICDM 2016 |
Natural language and speech › Language models and text generation › language modeling
character-level language modeling |
0.3 | 1 | 2018 | Empower Sequence Labeling with Task-Aware Neural Language Model · AAAI 2018 |
Natural language and speech › Information extraction and text analysis
named entity recognition |
0.3 | 1 | 2018 | Empower Sequence Labeling with Task-Aware Neural Language Model · AAAI 2018 |
Natural language and speech › Language models and text generation
neural language model |
0.3 | 1 | 2018 | Empower Sequence Labeling with Task-Aware Neural Language Model · AAAI 2018 |
Natural language and speech › Information extraction and text analysis
sequence labeling |
0.3 | 1 | 2018 | Empower Sequence Labeling with Task-Aware Neural Language Model · AAAI 2018 |
Natural language and speech › Information extraction and text analysis
relation extraction |
0.3 | 1 | 2017 | Heterogeneous Supervision for Relation Extraction: A Representation Learning Approach · EMNLP 2017 |
Data mining › structured data mining
graph mining |
0.3 | 1 | 2017 | PReP: Path-Based Relevance from a Probabilistic Perspective in Heterogeneous Information Networks · KDD 2017 |
Data mining › structured data mining › graph mining
network embedding |
0.3 | 1 | 2017 | Embedding Learning with Events in Heterogeneous Information Networks · IEEE Trans. Knowl. Data Eng. 2017 |
Data mining
pattern mining |
0.3 | 1 | 2017 | Embedding Learning with Events in Heterogeneous Information Networks · IEEE Trans. Knowl. Data Eng. 2017 |
Information retrieval › evaluation › effectiveness metrics
relevance measure |
0.3 | 1 | 2017 | PReP: Path-Based Relevance from a Probabilistic Perspective in Heterogeneous Information Networks · KDD 2017 |
Machine learning › Learning theory › high-dimensional statistics › matrix estimation
low-rank matrix estimation |
0.2 | 1 | 2016 | Towards Faster Rates and Oracle Property for Low-Rank Matrix Estimation · ICML 2016 |
Machine learning › Graph learning
network embedding |
0.2 | 1 | 2016 | Large-Scale Embedding Learning in Heterogeneous Event Data · ICDM 2016 |
Machine learning › Learning theory › statistical estimation
oracle property |
0.2 | 1 | 2016 | Towards Faster Rates and Oracle Property for Low-Rank Matrix Estimation · ICML 2016 |
Machine learning › Learning theory › statistical estimation › estimation error bounds
statistical rates |
0.2 | 1 | 2016 | Towards Faster Rates and Oracle Property for Low-Rank Matrix Estimation · ICML 2016 |
Data mining › structured data mining › graph mining › heterogeneous information network
heterogeneous information network mining |
0.2 | 1 | 2016 | Large-Scale Embedding Learning in Heterogeneous Event Data · ICDM 2016 |
Computational social science and digital humanities
causal inference |
0.2 | 1 | 2015 | Network A/B Testing: From Sampling to Estimation · WWW 2015 |
Computational social science and digital humanities › online controlled experiments › a/b testing
network a/b testing |
0.2 | 1 | 2015 | Network A/B Testing: From Sampling to Estimation · WWW 2015 |
Computational social science and digital humanities
online controlled experiments |
0.2 | 1 | 2015 | Network A/B Testing: From Sampling to Estimation · WWW 2015 |
Machine learning › Representation and self-supervised learning
tensor decomposition |
0.2 | 1 | 2014 | Robust Tensor Decomposition with Gross Corruption · NIPS 2014 |
Machine learning › Representation and self-supervised learning › representation learning
distributed representation learning |
0.1 | 1 | 2017 | Heterogeneous Supervision for Relation Extraction: A Representation Learning Approach · EMNLP 2017 |
Machine learning › Generative modeling
generative model |
0.1 | 1 | 2017 | PReP: Path-Based Relevance from a Probabilistic Perspective in Heterogeneous Information Networks · KDD 2017 |
Methods — techniques the papers use, named apart from their topics
layer-wise model ensembling · 0.8foundation model fine-tuning · 0.8maximum a posteriori estimation · 0.6generative modeling · 0.6embedding learning · 0.5word embeddings · 0.3transfer learning · 0.3neural language model · 0.3proximity prediction · 0.3hyperedge modeling · 0.3heterogeneous supervision · 0.3embedding · 0.3hyperedge prediction · 0.2spillover estimation · 0.2network sampling · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | LEVI: Generalizable Fine-tuning via Layer-wise Ensemble of Different ViewsabstractFine-tuning is becoming widely used for leveraging the power of pre-trained foundation models in new downstream tasks. While there are many successes of fine-tuning on various tasks, recent studies have observed challenges in the generalization of fine-tuned models to unseen distributions (i.e., out-of-distribution; OOD). To improve OOD generalization, some previous studies identify the limitations of fine-tuning data and regulate fine-tuning to preserve the general representation learned from pre-training data. However, potential limitations in the pre-training data and models are often ignored. In this paper, we contend that overly relying on the pre-trained representation may hinder fine-tuning from learning essential representations for downstream tasks and thus hurt its OOD generalization. It can be especially catastrophic when new tasks are from different (sub)domains compared to pre-training data. To address the issues in both pre-training and fine-tuning data, we propose a novel generalizable fine-tuning method LEVI (Layer-wise Ensemble of different VIews), where the pre-trained model is adaptively ensembled layer-wise with a small task-specific model, while preserving its efficiencies. By combining two complementing models, LEVI effectively suppresses problematic features in both the fine-tuning data and pre-trained model and preserves useful features for new tasks. Broad experiments with large language and vision models show that LEVI greatly improves fine-tuning generalization via emphasizing different views from fine-tuning data and pre-trained features. Yuji Roh, Qingyun Liu 0003, Huan Gui, Yujin Tang, Steven Euijong Whang, Liang Liu 0017, Shuchao Bi, Lichan Hong, Ed H. Chi, Zhe Zhao 0001 |
ICML | 3 |
| 2018 | Empower Sequence Labeling with Task-Aware Neural Language ModelabstractLinguistic sequence labeling is a general approach encompassing a variety of problems, such as part-of-speech tagging and named entity recognition. Recent advances in neural networks (NNs) make it possible to build reliable models without handcrafted features. However, in many cases, it is hard to obtain sufficient annotations to train these models. In this study, we develop a neural framework to extract knowledge from raw texts and empower the sequence labeling task. Besides word-level knowledge contained in pre-trained word embeddings, character-aware neural language models are incorporated to extract character-level knowledge. Transfer learning techniques are further adopted to mediate different components and guide the language model towards the key knowledge. Comparing to previous methods, these task-specific knowledge allows us to adopt a more concise model and conduct more efficient training. Different from most transfer learning methods, the proposed framework does not rely on any additional supervision. It extracts knowledge from self-contained order information of training sequences. Extensive experiments on benchmark datasets demonstrate the effectiveness of leveraging character-level knowledge and the efficiency of co-training. For example, on the CoNLL03 NER task, model training completes in about 6 hours on a single GPU, reaching F_1 score of 91.71+/-0.10 without using any extra annotations. Jingbo Shang, Xiang Ren 0001, Frank F. Xu, Huan Gui, Jian Peng 0001, Jiawei Han 0001 |
AAAI | 5 |
| 2018 | AspEm: Embedding Learning by Aspects in Heterogeneous Information NetworksabstractHeterogeneous information networks (HINs) are ubiquitous in real-world applications. Due to the heterogeneity in HINs, the typed edges may not fully align with each other. In order to capture the semantic subtlety, we propose the concept of aspects with each aspect being a unit representing one underlying semantic facet. Meanwhile, network embedding has emerged as a powerful method for learning network representation, where the learned embedding can be used as features in various downstream applications. Therefore, we are motivated to propose a novel embedding learning framework-ASPEM-to preserve the semantic information in HINs based on multiple aspects. Instead of preserving information of the network in one semantic space, ASPEM encapsulates information regarding each aspect individually. In order to select aspects for embedding purpose, we further devise a solution for ASPEM based on dataset-wide statistics. To corroborate the efficacy of ASPEM, we conducted experiments on two real-words datasets with two types of applications-classification and link prediction. Experiment results demonstrate that ASPEM can outperform baseline network embedding learning methods by considering multiple aspects, where the aspects can be selected from the given HIN in an unsupervised manner. Yu Shi 0002, Huan Gui, Qi Zhu 0008, Lance M. Kaplan, Jiawei Han 0001 |
SDM | 2 |
| 2017 | Heterogeneous Supervision for Relation Extraction: A Representation Learning ApproachabstractRelation extraction is a fundamental task in information extraction.Most existing methods have heavy reliance on annotations labeled by human experts, which are costly and time-consuming.To overcome this drawback, we propose a novel framework, REHESSION, to conduct relation extractor learning using annotations from heterogeneous information source, e.g., knowledge base and domain heuristics.These annotations, referred as heterogeneous supervision, often conflict with each other, which brings a new challenge to the original relation extraction task: how to infer the true label from noisy labels for a given instance.Identifying context information as the backbone of both relation extraction and true label discovery, we adopt embedding techniques to learn the distributed representations of context, which bridges all components with mutual enhancement in an iterative fashion.Extensive experimental results demonstrate the superiority of REHESSION over the state-of-the-art. Xiang Ren 0001, Qi Zhu 0008, Shi Zhi, Huan Gui, Heng Ji 0001, Jiawei Han 0001 |
EMNLP | 5 |
| 2017 | PReP: Path-Based Relevance from a Probabilistic Perspective in Heterogeneous Information NetworksabstractAs a powerful representation paradigm for networked and multi-typed data, the heterogeneous information network (HIN) is ubiquitous. Meanwhile, defining proper relevance measures has always been a fundamental problem and of great pragmatic importance for network mining tasks. Inspired by our probabilistic interpretation of existing path-based relevance measures, we propose to study HIN relevance from a probabilistic perspective. We also identify, from real-world data, and propose to model cross-meta-path synergy, which is a characteristic important for defining path-based HIN relevance and has not been modeled by existing methods. A generative model is established to derive a novel path-based relevance measure, which is data-driven and tailored for each HIN. We develop an inference algorithm to find the maximum a posteriori (MAP) estimate of the model parameters, which entails non-trivial tricks. Experiments on two real-world datasets demonstrate the effectiveness of the proposed model and relevance measure. Yu Shi 0002, Po-Wei Chan, Honglei Zhuang, Huan Gui, Jiawei Han 0001 |
KDD | 4 |
| 2017 | Embedding Learning with Events in Heterogeneous Information NetworksabstractIn real-world applications, objects of multiple types are interconnected, formingHeterogeneous Information Networks. In such heterogeneous information networks, we make the key observation that many interactions happen due to someeventand the objects in each event form a complete semantic unit. By taking advantage of such a property, we propose a generic framework calledHyperEdge-BasedEmbedding(Hebe) to learn object embeddings with events in heterogeneous information networks, where ahyperedgeencompasses the objects participating in one event. TheHebeframework models the proximity among objects in each event with two methods: (1) predicting a target object given other participating objects in the event, and (2) predicting if the event can be observed given all the participating objects. Since each hyperedge encapsulates more information of a given event,Hebeis robust to data sparseness and noise. In addition,Hebeis scalable when the data size spirals. Extensive experiments on large-scale real-world datasets show the efficacy and robustness of the proposed framework. Huan Gui, Fangbo Tao, Meng Jiang 0001, Brandon Norick, Lance M. Kaplan, Jiawei Han 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2016 | Downside management in recommender systemsabstractIn recommender systems, bad recommendations can lead to a net utility loss for both users and content providers. The downside (individual loss) management is a crucial and important problem, but has long been ignored. We propose a method to identify bad recommendations by modeling the users' latent preferences that are yet to be captured using a residual model, which can be applied independently on top of existing recommendation algorithms. We include two components in the residual utility: benefit and cost, which can be learned simultaneously from users' observed interactions with the recommender system. We further classify user behavior into fine-grained categories, based on which an efficient optimization algorithm to estimate the benefit and cost using Bayesian partial order is proposed. By accurately calculating the utility users obtained from recommendations based on the benefit-cost analysis, we can infer the optimal threshold to determine the downside portion of the recommender system. We validate the proposed method by experimenting with real-world datasets and demonstrate that it can help to prevent bad recommendations from showing. Huan Gui, Haishan Liu, Anmol Bhasin, Jiawei Han 0001 |
ASONAM | 1 |
| 2016 | Large-Scale Embedding Learning in Heterogeneous Event DataabstractHeterogeneous events, which are defined as events connecting strongly-typed objects, are ubiquitous in the real world. We propose a HyperEdge-Based Embedding (Hebe) framework for heterogeneous event data, where a hyperedge represents the interaction among a set of involving objects in an event. The Hebe framework models the proximity among objects in an event by predicting a target object given the other participating objects in the event (hyperedge). Since each hyperedge encapsulates more information on a given event, Hebe is robust to data sparseness. In addition, Hebe is scalable when the data size spirals. Extensive experiments on large-scale real-world datasets demonstrate the efficacy and robustness of Hebe. Huan Gui, Fangbo Tao, Meng Jiang 0001, Brandon Norick, Jiawei Han 0001 |
ICDM | 1 |
| 2016 | Towards Faster Rates and Oracle Property for Low-Rank Matrix EstimationabstractWe present a unified framework for low-rank matrix estimation with a nonconvex penalty. A proximal gradient homotopy algorithm is proposed to solve the proposed optimization problem. Theoretically, we first prove that the proposed estimator attains a faster statistical rate than the traditional low-rank matrix estimator with nuclear norm penalty. Moreover, we rigorously show that under a certain condition on the magnitude of the nonzero singular values, the proposed estimator enjoys oracle property (i.e., exactly recovers the true rank of the matrix), besides attaining a faster rate. Extensive numerical experiments on both synthetic and real world datasets corroborate our theoretical findings. Huan Gui, Jiawei Han 0001, Quanquan Gu |
ICML | 1 |
| 2015 | Network A/B Testing: From Sampling to EstimationabstractA/B testing, also known as bucket testing, split testing, or controlled experiment, is a standard way to evaluate user engagement or satisfaction from a new service, feature, or product. It is widely used in online websites, including social network sites such as Facebook, LinkedIn, and Twitter to make data-driven decisions. The goal of A/B testing is to estimate the treatment effect of a new change, which becomes intricate when users are interacting, i.e., the treatment effect of a user may spill over to other users via underlying social connections.When conducting these online controlled experiments, it is a common practice to make the Stable Unit Treatment Value Assumption (SUTVA) that each individual's response is affected by their own treatment only. Though this assumption simplifies the estimation of treatment effect, it does not hold when network interference is present, and may even lead to wrong conclusion. Huan Gui, Ya Xu, Anmol Bhasin, Jiawei Han 0001 |
WWW | 1 |
| 2014 | Modeling Topic Diffusion in Multi-Relational Bibliographic Information NetworksabstractInformation diffusion has been widely studied in networks, aiming to model the spread of information among objects when they are connected with each other. Most of the current research assumes the underlying network is homogeneous, i.e., objects are of the same type and they are connected by links with the same semantic meanings. However, in the real word, objects are connected via different types of relationships, forming multi-relational heterogeneous information networks. Huan Gui, Yizhou Sun, Jiawei Han 0001, George Brova |
CIKM | 1 |
| 2014 | Robust Tensor Decomposition with Gross Corruption
Quanquan Gu, Huan Gui, Jiawei Han 0001 |
NIPS | 2 |