VLDB 2026 Research / reviewers in the wild / expert
Weihua Hu
dblp:42/1232
· DBLP profile ↗
23ranked-venue papers
8as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
16 papers |
Graph learning · 28% Trustworthy machine learning · 26% Representation and self-supervised learning · 13% | |
| Databases, data mining, and information retrieval
7 papers |
Knowledge graphs · 24% Recommender systems · 22% Data mining · 20% | |
| Theoretical computer science
2 papers |
Coding theory · 79% Graph algorithms and graph theory · 21% |
Topics — the 30 heaviest of 46, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Graph learning
graph neural network |
2.6 | 5 | 2025 | RelBench: A Benchmark for Deep Learning on Relational Databases · NeurIPS 2024 Position: Relational Deep Learning - Graph Representation Learning on Relational Databases · ICML 2024 Strategies for Pre-training Graph Neural Networks · ICLR 2020 |
Machine learning › Trustworthy machine learning
robustness |
1.4 | 3 | 2022 | Extending the WILDS Benchmark for Unsupervised Adaptation · ICLR 2022 WILDS: A Benchmark of in-the-Wild Distribution Shifts · ICML 2021 Does Distributionally Robust Supervised Learning Give Robust Classifiers? · ICML 2018 |
Machine learning › Trustworthy machine learning › robustness
distribution shift |
0.9 | 2 | 2022 | Extending the WILDS Benchmark for Unsupervised Adaptation · ICLR 2022 Does Distributionally Robust Supervised Learning Give Robust Classifiers? · ICML 2018 |
Knowledge graphs
link prediction |
0.9 | 1 | 2025 | ContextGNN: Beyond Two-Tower Recommendation Systems · ICLR 2025 |
Recommender systems › neural recommendation
two-tower model |
0.9 | 1 | 2025 | ContextGNN: Beyond Two-Tower Recommendation Systems · ICLR 2025 |
Machine learning › Deep learning architectures and training › feature interaction
channel interaction |
0.8 | 1 | 2024 | From Similarity to Superiority: Channel Clustering for Time Series Forecasting · NeurIPS 2024 |
Machine learning › Time series and sequential data › time series analysis
time series forecasting |
0.8 | 1 | 2024 | From Similarity to Superiority: Channel Clustering for Time Series Forecasting · NeurIPS 2024 |
Machine learning › Transfer learning and domain adaptation › zero-shot learning
zero-shot forecasting |
0.8 | 1 | 2024 | From Similarity to Superiority: Channel Clustering for Time Series Forecasting · NeurIPS 2024 |
Data mining
predictive modeling |
0.8 | 1 | 2024 | RelBench: A Benchmark for Deep Learning on Relational Databases · NeurIPS 2024 |
Database system architecture and tuning
relational database system |
0.8 | 1 | 2024 | RelBench: A Benchmark for Deep Learning on Relational Databases · NeurIPS 2024 |
Machine learning and data management › relational machine learning
relational deep learning |
0.8 | 1 | 2024 | Position: Relational Deep Learning - Graph Representation Learning on Relational Databases · ICML 2024 |
Machine learning › Graph learning
dynamic graph learning |
0.7 | 1 | 2023 | Temporal Graph Benchmark for Machine Learning on Temporal Graphs · NeurIPS 2023 |
Machine learning › Representation and self-supervised learning › representation learning
embedding learning |
0.6 | 1 | 2022 | Learning Backward Compatible Embeddings · KDD 2022 |
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation |
0.6 | 1 | 2022 | Extending the WILDS Benchmark for Unsupervised Adaptation · ICLR 2022 |
Machine learning › Learning theory
generalization |
0.5 | 1 | 2021 | WILDS: A Benchmark of in-the-Wild Distribution Shifts · ICML 2021 |
Machine learning › Trustworthy machine learning
out-of-distribution generalization |
0.5 | 1 | 2021 | WILDS: A Benchmark of in-the-Wild Distribution Shifts · ICML 2021 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge graph reasoning |
0.4 | 1 | 2020 | Query2box: Reasoning over Knowledge Graphs in Vector Space Using Box Embeddings · ICLR 2020 |
Machine learning › Representation and self-supervised learning
pre-training |
0.4 | 1 | 2020 | Strategies for Pre-training Graph Neural Networks · ICLR 2020 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › semantic representation
query embedding |
0.4 | 1 | 2020 | Query2box: Reasoning over Knowledge Graphs in Vector Space Using Box Embeddings · ICLR 2020 |
Machine learning › Graph learning › graph neural network
scalable graph neural network |
0.4 | 1 | 2020 | Open Graph Benchmark: Datasets for Machine Learning on Graphs · NeurIPS 2020 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning |
0.4 | 1 | 2020 | Strategies for Pre-training Graph Neural Networks · ICLR 2020 |
Machine learning › Graph learning › graph neural network
expressive power |
0.4 | 1 | 2019 | How Powerful are Graph Neural Networks? · ICLR 2019 |
Graph algorithms and graph theory
graph isomorphism |
0.4 | 1 | 2019 | How Powerful are Graph Neural Networks? · ICLR 2019 |
Machine learning › Trustworthy machine learning › robustness › learning with noisy labels
co-teaching |
0.3 | 1 | 2018 | Co-teaching: Robust training of deep neural networks with extremely noisy labels · NeurIPS 2018 |
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels |
0.3 | 1 | 2018 | Co-teaching: Robust training of deep neural networks with extremely noisy labels · NeurIPS 2018 |
Machine learning › Trustworthy machine learning › learning from noisy data
noisy label training |
0.3 | 1 | 2018 | Co-teaching: Robust training of deep neural networks with extremely noisy labels · NeurIPS 2018 |
Machine learning › Trustworthy machine learning › robustness
robust learning |
0.3 | 1 | 2018 | Co-teaching: Robust training of deep neural networks with extremely noisy labels · NeurIPS 2018 |
Machine learning › Probabilistic and Bayesian machine learning
clustering |
0.3 | 1 | 2017 | Learning Discrete Representations via Information Maximizing Self-Augmented Training · ICML 2017 |
Machine learning › Representation and self-supervised learning › representation learning
discrete representation learning |
0.3 | 1 | 2017 | Learning Discrete Representations via Information Maximizing Self-Augmented Training · ICML 2017 |
Machine learning › Representation and self-supervised learning › representation learning › embedding learning
hash code learning |
0.3 | 1 | 2017 | Learning Discrete Representations via Information Maximizing Self-Augmented Training · ICML 2017 |
Methods — techniques the papers use, named apart from their topics
graph neural network · 6.2user study · 2.3manual feature engineering · 2.3two-tower architecture · 1.7pair-wise representation · 1.7representation learning · 1.5channel-independent strategy · 0.8channel-dependent strategy · 0.8channel clustering module · 0.8self-training · 0.6embedding model retraining · 0.6benchmark construction · 0.5vector space reasoning · 0.4box embedding · 0.4weisfeiler-lehman test · 0.4worst-case redundancy derivation · 0.3code tree analysis · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ContextGNN: Beyond Two-Tower Recommendation SystemsabstractRecommendation systems predominantly utilize two-tower architectures, which evaluate user-item rankings through the inner product of their respective embeddings. However, one key limitation of two-tower models is that they learn a pair-agnostic representation of users and items. In contrast, pair-wise representations either scale poorly due to their quadratic complexity or are too restrictive on the candidate pairs to rank. To address these issues, we introduce Context-based Graph Neural Networks (ContextGNNs), a novel deep learning architecture for link prediction in recommendation systems. The method employs a pair-wise representation technique for familiar items situated within a user's local subgraph, while leveraging two-tower representations to facilitate the recommendation of exploratory items. A final network then predicts how to fuse both pair-wise and two-tower recommendations into a single ranking of items. We demonstrate that ContextGNN is able to adapt to different data characteristics and outperforms existing methods, both traditional and GNN-based, on a diverse set of practical recommendation tasks, improving performance by 20\% on average. Yiwen Yuan, Zecheng Zhang, Akihiro Nitta, Weihua Hu, Manan Shah, Blaz Stojanovic, Shenyang Huang, Jan Eric Lenssen, Jure Leskovec, Matthias Fey |
ICLR | 5 |
| 2024 | Position: Relational Deep Learning - Graph Representation Learning on Relational DatabasesabstractMuch of the world’s most valued data is stored in relational databases and data warehouses, where the data is organized into tables connected by primary-foreign key relations. However, building machine learning models using this data is both challenging and time consuming because no ML algorithm can directly learn from multiple connected tables. Current approaches can only learn from a single table, so data must first be manually joined and aggregated into this format, the laborious process known as feature engineering. Feature engineering is slow, error prone and leads to suboptimal models. Here we introduce Relational Deep Learning (RDL), a blueprint for end-to-end learning on relational databases. The key is to represent relational databases as a temporal, heterogeneous graphs, with a node for each row in each table, and edges specified by primary-foreign key links. Graph Neural Networks then learn representations that leverage all input data, without any manual feature engineering. We also introduce RelBench, and benchmark and testing suite, demonstrating strong initial results. Overall, we define a new research area that generalizes graph machine learning and broadens its applicability. Matthias Fey, Weihua Hu, Jan Eric Lenssen, Rishabh Ranjan, Joshua Robinson 0001, Rex Ying, Jiaxuan You, Jure Leskovec |
ICML | 2 |
| 2024 | RelBench: A Benchmark for Deep Learning on Relational DatabasesabstractWe present RelBench, a public benchmark for solving predictive tasks in relational databases with deep learning. RelBench provides databases and tasks spanning diverse domains, scales, and database dimensions, and is intended to be a foundational infrastructure for future research in this direction. We use RelBench to conduct the first comprehensive empirical study of graph neural network (GNN) based predictive models on relational data, as recently proposed by Fey et al. 2024. End-to-end learned GNNs are capable fully exploiting the predictive signal encoded in links between entities, marking a significant shift away from the dominant paradigm of manual feature engineering combined with tabular machine learning. To thoroughly evaluate GNNs against the prior gold-standard we conduct a user study, where an experienced data scientist manually engineers features for each task. In this study, GNNs learn better models whilst reducing human work needed by more than an order of magnitude. This result demonstrates the power of GNNs for solving predictive tasks in relational databases, opening up new research opportunities. Joshua Robinson 0001, Rishabh Ranjan, Weihua Hu, Jiaqi Han 0001, Alejandro Dobles, Matthias Fey, Jan Eric Lenssen, Yiwen Yuan, Zecheng Zhang, Jure Leskovec |
NeurIPS | 3 |
| 2024 | From Similarity to Superiority: Channel Clustering for Time Series ForecastingabstractTime series forecasting has attracted significant attention in recent decades.
Previous studies have demonstrated that the Channel-Independent (CI) strategy improves forecasting performance by treating different channels individually, while it leads to poor generalization on unseen instances and ignores potentially necessary interactions between channels. Conversely, the Channel-Dependent (CD) strategy mixes all channels with even irrelevant and indiscriminate information, which, however, results in oversmoothing issues and limits forecasting accuracy.
There is a lack of channel strategy that effectively balances individual channel treatment for improved forecasting performance without overlooking essential interactions between channels. Motivated by our observation of a correlation between the time series model's performance boost against channel mixing and the intrinsic similarity on a pair of channels, we developed a novel and adaptable \textbf{C}hannel \textbf{C}lustering \textbf{M}odule (CCM). CCM dynamically groups channels characterized by intrinsic similarities and leverages cluster information instead of individual channel identities, combining the best of CD and CI worlds. Extensive experiments on real-world datasets demonstrate that CCM can (1) boost the performance of CI and CD models by an average margin of 2.4% and 7.2% on long-term and short-term forecasting, respectively; (2) enable zero-shot forecasting with mainstream time series forecasting models; (3) uncover intrinsic time series patterns among channels and improve interpretability of complex time series models. Jan Eric Lenssen, Aosong Feng, Weihua Hu, Matthias Fey, Leandros Tassiulas, Jure Leskovec, Rex Ying |
NeurIPS | 4 |
| 2024 | Utilizing ChatGPT as a scientific reasoning engine to differentiate conflicting evidence and summarize challenges in controversial clinical questionsabstractOBJECTIVE: Synthesizing and evaluating inconsistent medical evidence is essential in evidence-based medicine. This study aimed to employ ChatGPT as a sophisticated scientific reasoning engine to identify conflicting clinical evidence and summarize unresolved questions to inform further research. MATERIALS AND METHODS: We evaluated ChatGPT's effectiveness in identifying conflicting evidence and investigated its principles of logical reasoning. An automated framework was developed to generate a PubMed dataset focused on controversial clinical topics. ChatGPT analyzed this dataset to identify consensus and controversy, and to formulate unsolved research questions. Expert evaluations were conducted 1) on the consensus and controversy for factual consistency, comprehensiveness, and potential harm and, 2) on the research questions for relevance, innovation, clarity, and specificity. RESULTS: The gpt-4-1106-preview model achieved a 90% recall rate in detecting inconsistent claim pairs within a ternary assertions setup. Notably, without explicit reasoning prompts, ChatGPT provided sound reasoning for the assertions between claims and hypotheses, based on an analysis grounded in relevance, specificity, and certainty. ChatGPT's conclusions of consensus and controversies in clinical literature were comprehensive and factually consistent. The research questions proposed by ChatGPT received high expert ratings. DISCUSSION: Our experiment implies that, in evaluating the relationship between evidence and claims, ChatGPT considered more detailed information beyond a straightforward assessment of sentimental orientation. This ability to process intricate information and conduct scientific reasoning regarding sentiment is noteworthy, particularly as this pattern emerged without explicit guidance or directives in prompts, highlighting ChatGPT's inherent logical reasoning capabilities. CONCLUSION: This study demonstrated ChatGPT's capacity to evaluate and interpret scientific claims. Such proficiency can be generalized to broader clinical research literature. ChatGPT effectively aids in facilitating clinical studies by proposing unresolved challenges based on analysis of existing studies. However, caution is advised as ChatGPT's outputs are inferences drawn from the input literature and could be harmful to clinical practice. Shiyao Xie, Guanghui Deng, Guohua He, Zhenhua Lu, Weihua Hu |
J. Am. Medical Informatics Assoc. | 7 |
| 2023 | Temporal Graph Benchmark for Machine Learning on Temporal GraphsabstractWe present the Temporal Graph Benchmark (TGB), a collection of challenging and diverse benchmark datasets for realistic, reproducible, and robust evaluation of machine learning models on temporal graphs. TGB datasets are of large scale, spanning years in duration, incorporate both node and edge-level prediction tasks and cover a diverse set of domains including social, trade, transaction, and transportation networks. For both tasks, we design evaluation protocols based on realistic use-cases. We extensively benchmark each dataset and find that the performance of common models can vary drastically across datasets. In addition, on dynamic node property prediction tasks, we show that simple methods often achieve superior performance compared to existing temporal graph models. We believe that these findings open up opportunities for future research on temporal graphs. Finally, TGB provides an automated machine learning pipeline for reproducible and accessible temporal graph research, including data loading, experiment setup and performance evaluation. TGB will be maintained and updated on a regular basis and welcomes community feedback. TGB datasets, data loaders, example codes, evaluation setup, and leaderboards are publicly available at https://tgb.complexdatalab.com/. Shenyang Huang, Farimah Poursafaei, Jacob Danovitch, Matthias Fey, Weihua Hu, Emanuele Rossi 0001, Jure Leskovec, Michael M. Bronstein, Guillaume Rabusseau, Reihaneh Rabbany |
NeurIPS | 5 |
| 2022 | Extending the WILDS Benchmark for Unsupervised Adaptation
Shiori Sagawa, Pang Wei Koh, Irena Gao, Sang Michael Xie, Kendrick Shen, Ananya Kumar, Weihua Hu, Michihiro Yasunaga, Henrik Marklund, Sara Beery, Etienne David, Ian Stavness, Wei Guo 0002, Jure Leskovec, Kate Saenko, Tatsunori B. Hashimoto, Sergey Levine, Chelsea Finn, Percy Liang |
ICLR | 8 |
| 2022 | Learning Backward Compatible EmbeddingsabstractEmbeddings, low-dimensional vector representation of objects, are fundamental in building modern machine learning systems. In industrial settings, there is usually an embedding team that trains an embedding model to solve intended tasks (e.g., product recommendation). The produced embeddings are then widely consumed by consumer teams to solve their unintended tasks (e.g., fraud detection). However, as the embedding model gets updated and retrained to improve performance on the intended task, the newly-generated embeddings are no longer compatible with the existing consumer models. This means that historical versions of the embeddings can never be retired or all consumer teams have to retrain their models to make them compatible with the latest version of the embeddings, both of which are extremely costly in practice. Weihua Hu, Rajas Bansal, Kaidi Cao, Nikhil Rao 0001, Karthik Subbian, Jure Leskovec |
KDD | 1 |
| 2022 | LE-MSFE-DDNet: a defect detection network based on low-light enhancement and multi-scale feature extraction
Weihua Hu, Yangsai Wang, Guoheng Huang |
Vis. Comput. | 1 |
| 2021 | WILDS: A Benchmark of in-the-Wild Distribution ShiftsabstractDistribution shifts—where the training distribution differs from the test distribution—can substantially degrade the accuracy of machine learning (ML) systems deployed in the wild. Despite their ubiquity in the real-world deployments, these distribution shifts are under-represented in the datasets widely used in the ML community today. To address this gap, we present WILDS, a curated benchmark of 10 datasets reflecting a diverse range of distribution shifts that naturally arise in real-world applications, such as shifts across hospitals for tumor identification; across camera traps for wildlife monitoring; and across time and location in satellite imaging and poverty mapping. On each dataset, we show that standard training yields substantially lower out-of-distribution than in-distribution performance. This gap remains even with models trained by existing methods for tackling distribution shifts, underscoring the need for new methods for training models that are more robust to the types of distribution shifts that arise in practice. To facilitate method development, we provide an open-source package that automates dataset loading, contains default model architectures and hyperparameters, and standardizes evaluations. The full paper, code, and leaderboards are available at https://wilds.stanford.edu. Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard L. Phillips, Irena Gao, Etienne David, Ian Stavness, Wei Guo 0002, Berton Earnshaw, Imran S. Haque, Sara Beery, Jure Leskovec, Anshul Kundaje, Emma Pierson, Sergey Levine, Chelsea Finn, Percy Liang |
ICML | 7 |
| 2020 | Strategies for Pre-training Graph Neural Networks
Weihua Hu, Bowen Liu 0014, Joseph Gomes, Marinka Zitnik, Percy Liang, Vijay S. Pande, Jure Leskovec |
ICLR | 1 |
| 2020 | Query2box: Reasoning over Knowledge Graphs in Vector Space Using Box Embeddings
Hongyu Ren, Weihua Hu, Jure Leskovec |
ICLR | 2 |
| 2020 | Open Graph Benchmark: Datasets for Machine Learning on GraphsabstractWe present the Open Graph Benchmark (OGB), a diverse set of challenging and realistic benchmark datasets to facilitate scalable, robust, and reproducible graph machine learning (ML) research. OGB datasets are large-scale, encompass multiple important graph ML tasks, and cover a diverse range of domains, ranging from social and information networks to biological networks, molecular graphs, source code ASTs, and knowledge graphs. For each dataset, we provide a unified evaluation protocol using meaningful application-specific data splits and evaluation metrics. In addition to building the datasets, we also perform extensive benchmark experiments for each dataset. Our experiments suggest that OGB datasets present significant challenges of scalability to large-scale graphs and out-of-distribution generalization under realistic data splits, indicating fruitful opportunities for future research. Finally, OGB provides an automated end-to-end graph ML pipeline that simplifies and standardizes the process of graph data loading, experimental setup, and model evaluation. OGB will be regularly updated and welcomes inputs from the community. OGB datasets as well as data loaders, evaluation scripts, baseline code, and leaderboards are publicly available at https://ogb.stanford.edu . Weihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong, Hongyu Ren, Bowen Liu 0014, Michele Catasta, Jure Leskovec |
NeurIPS | 1 |
| 2019 | How Powerful are Graph Neural Networks?
Keyulu Xu, Weihua Hu, Jure Leskovec, Stefanie Jegelka |
ICLR | 2 |
| 2018 | Does Distributionally Robust Supervised Learning Give Robust Classifiers?abstractDistributionally Robust Supervised Learning (DRSL) is necessary for building reliable machine learning systems. When machine learning is deployed in the real world, its performance can be significantly degraded because test data may follow a different distribution from training data. DRSL with f-divergences explicitly considers the worst-case distribution shift by minimizing the adversarially reweighted training loss. In this paper, we analyze this DRSL, focusing on the classification scenario. Since the DRSL is explicitly formulated for a distribution shift scenario, we naturally expect it to give a robust classifier that can aggressively handle shifted distributions. However, surprisingly, we prove that the DRSL just ends up giving a classifier that exactly fits the given training distribution, which is too pessimistic. This pessimism comes from two sources: the particular losses used in classification and the fact that the variety of distributions to which the DRSL tries to be robust is too wide. Motivated by our analysis, we propose simple DRSL that overcomes this pessimism and empirically demonstrate its effectiveness. Weihua Hu, Gang Niu 0001, Issei Sato, Masashi Sugiyama |
ICML | 1 |
| 2018 | Co-teaching: Robust training of deep neural networks with extremely noisy labelsabstractDeep learning with noisy labels is practically challenging, as the capacity of deep models is so high that they can totally memorize these noisy labels sooner or later during training. Nonetheless, recent studies on the memorization effects of deep neural networks show that they would first memorize training data of clean labels and then those of noisy labels. Therefore in this paper, we propose a new deep learning paradigm called ''Co-teaching'' for combating with noisy labels. Namely, we train two deep neural networks simultaneously, and let them teach each other given every mini-batch: firstly, each network feeds forward all data and selects some data of possibly clean labels; secondly, two networks communicate with each other what data in this mini-batch should be used for training; finally, each network back propagates the data selected by its peer network and updates itself. Empirical results on noisy versions of MNIST, CIFAR-10 and CIFAR-100 demonstrate that Co-teaching is much superior to the state-of-the-art methods in the robustness of trained deep models. Bo Han 0003, Quanming Yao, Xingrui Yu, Gang Niu 0001, Miao Xu 0001, Weihua Hu, Ivor W. Tsang, Masashi Sugiyama |
NeurIPS | 6 |
| 2017 | Learning Discrete Representations via Information Maximizing Self-Augmented TrainingabstractLearning discrete representations of data is a central machine learning task because of the compactness of the representations and ease of interpretation. The task includes clustering and hash learning as special cases. Deep neural networks are promising to be used because they can model the non-linearity of data and scale to large datasets. However, their model complexity is huge, and therefore, we need to carefully regularize the networks in order to learn useful representations that exhibit intended invariance for applications of interest. To this end, we propose a method called Information Maximizing Self-Augmented Training (IMSAT). In IMSAT, we use data augmentation to impose the invariance on discrete representations. More specifically, we encourage the predicted representations of augmented data points to be close to those of the original data points in an end-to-end fashion. At the same time, we maximize the information-theoretic dependency between data and their predicted discrete representations. Extensive experiments on benchmark datasets show that IMSAT produces state-of-the-art results for both clustering and unsupervised hash learning. Weihua Hu, Takeru Miyato, Seiya Tokui, Eiichi Matsumoto, Masashi Sugiyama |
ICML | 1 |
| 2017 | Learning from Complementary LabelsabstractCollecting labeled data is costly and thus a critical bottleneck in real-world classification tasks. To mitigate this problem, we propose a novel setting, namely learning from complementary labels for multi-class classification. A complementary label specifies a class that a pattern does not belong to. Collecting complementary labels would be less laborious than collecting ordinary labels, since users do not have to carefully choose the correct class from a long list of candidate classes. However, complementary labels are less informative than ordinary labels and thus a suitable approach is needed to better learn from them. In this paper, we show that an unbiased estimator to the classification risk can be obtained only from complementarily labeled data, if a loss function satisfies a particular symmetric condition. We derive estimation error bounds for the proposed method and prove that the optimal parametric convergence rate is achieved. We further show that learning from complementary labels can be easily combined with learning from ordinary labels (i.e., ordinary supervised learning), providing a highly practical implementation of the proposed method. Finally, we experimentally demonstrate the usefulness of the proposed methods. Takashi Ishida 0001, Gang Niu 0001, Weihua Hu, Masashi Sugiyama |
NIPS | 3 |
| 2017 | Worst-case Redundancy of Optimal Binary AIFV Codes and Their Extended CodesabstractBinary almost instantaneous fixed-to-variable length (AIFV) codes are lossless codes that generalize the class of instantaneous fixed-to-variable length codes. The code uses two code trees and assigns source symbols to incomplete internal nodes as well as to leaves. AIFV codes are empirically shown to attain better compression ratio than Huffman codes. Nevertheless, an upper bound on the redundancy of optimal binary AIFV codes is only known to be 1, which is the same as the bound of Huffman codes. In this paper, the upper bound is improved to 1/2, which is shown to coincide with the worst-case redundancy of the codes. Along with this, the worst-case redundancy is derived for sources with pmax ≥1/2, where pmax is the probability of the most likely source symbol. In addition, we propose an extension of binary AIFV codes, which use m code trees and allow at most m-bit decoding delay. We show that the worst-case redundancy of the extended binary AIFV codes is 1/m for m ≤ 4. Weihua Hu, Hirosuke Yamamoto, Junya Honda |
IEEE Trans. Inf. Theory | 1 |
| 2016 | Tight upper bounds on the redundancy of optimal binary AIFV codesabstractAIFV codes are lossless codes that generalize the class of instantaneous FV codes. The code uses multiple code trees and assigns source symbols to incomplete internal nodes as well as to leaves. AIFV codes are empirically shown to attain better compression ratio than Huffman codes. Nevertheless, an upper bound on the redundancy of optimal binary AIFV codes is only known to be 1, the same as the bound of Huffman codes. In this paper, the upper bound is improved to 1/2, which is shown to be tight. Along with this, a tight upper bound on the redundancy of optimal binary AIFV codes is derived for the case pmax≥1/2, where pmaxis the probability of the most likely source symbol. This is the first theoretical work on the redundancy of optimal binary AIFV codes, suggesting superiority of the codes over Huffman codes. Weihua Hu, Hirosuke Yamamoto, Junya Honda |
ISIT | 1 |
| 2015 | Efficient Auto-Scaling Approach in the Telco Cloud Using Self-Learning AlgorithmabstractNetwork Function Virtualization (NFV) and Software Defined Network (SDN) technologies makes it possible for the Telco Operators to assign resource for virtual network functions (VNF) on demand. Provision and orchestration of physical and virtual resource is crucial for both Quality of Service (QoS) guarantee and cost management in cloud computing environment. Auto-scaling mechanism is essential in the lifecycle management of those VNFs. Threshold based policy is always applied in classic IT cloud environments which can not satisfy carrier grade requirements such as reliability and stability. In this paper, we present a novel SLA-aware and Resource-efficient Self-learning Approach (SRSA) for auto-scaling policy decision. The scenarios of the service volatility is categorized into daily busy-and-idle scenario and burst-traffic scenario. First, we formulate the workload of the VNF as discrete-time series and treat procedure of policy-making in auto-scaling as a Markov Decision Process (MDP). Second, parameters in the Reinforcement Learning process are tuned cautiously. Finally the experiments show that our solution outperforms threshold based policy and voting policy adopted by RightScale in oscillation suppression, QoS guarantee, and energy saving. Pengcheng Tang, Weihua Hu |
GLOBECOM | 4 |
| 2007 | Embedded education for Computer Rank ExaminationabstractEmbedded system has become one of the most important directions in computer education. Embedded system is at the intersection of control system, command and control, wireless data systems, real-time system and so on. The Computer Rank Examination (CRE) is designed to promote the certification for computer application. In this paper, we describe the efforts on embedded education for CRE including the curriculum design and practice. Tianzhou Chen, Weihua Hu, Qingsong Shi |
ICPADS | 2 |
| 2005 | Interactive learning of CG in networked virtual environments
Jiejie Zhu, Weihua Hu, Hung Pak Lun |
Comput. Graph. | 3 |