VLDB 2026 Research / reviewers in the wild / expert
Zhichao Huang 0001
dblp:153/7931-1
· DBLP profile ↗
20ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0002-7662-5184ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Spectral Disentanglement and Enhancement: A Dual-domain Contrastive Framework for Representation LearningabstractLarge-scale multimodal contrastive learning has recently achieved impressive success in learning rich and transferable representations, yet it remains fundamentally limited by the uniform treatment of feature dimensions and the neglect of the intrinsic spectral structure of the learned features. Empirical evidence indicates that high-dimensional embeddings tend to collapse into narrow cones, concentrating task-relevant semantics in a small subspace, while the majority of dimensions remain occupied by noise and spurious correlations. Such spectral imbalance and entanglement undermine model generalization. We propose Spectral Disentanglement and Enhancement (SDE), a novel framework that bridges the gap between the geometry of the embedded spaces and their spectral properties. Our approach leverages singular value decomposition to adaptively partition feature dimensions into strong signals that capture task-critical semantics, weak signals that reflect ancillary correlations, and noise representing irrelevant perturbations. A curriculum-based spectral enhancement strategy is then applied, selectively amplifying informative components with theoretical guarantees on training stability. Building upon the enhanced features, we further introduce a dual-domain contrastive loss that jointly optimizes alignment in both the feature and spectral spaces, effectively integrating spectral regularization into the training process and encouraging richer, more robust representations. Extensive experiments on large-scale multimodal benchmarks demonstrate that SDE consistently improves representation robustness and generalization, outperforming state-of-the-art methods. SDE integrates seamlessly with existing contrastive pipelines, offering an effective solution for multimodal representation learning. Jinjin Guo, Yexin Li, Zhichao Huang 0001, Pengzhang Liu, Qixia Jiang |
WWW | 3 |
| 2025 | Core Knowledge Learning Framework for GraphabstractGraph classification is a pivotal challenge in machine learning, especially within the realm of graph-based data, given its importance in numerous real-world applications such as social network analysis, recommendation systems, and bioinformatics. Despite its significance, graph classification faces several hurdles, including adapting to diverse prediction tasks, training across multiple target domains, and handling small-sample prediction scenarios. Current methods often tackle these challenges individually, leading to fragmented solutions that lack a holistic approach to the overarching problem. In this paper, we propose an algorithm aimed at addressing the aforementioned challenges. By incorporating insights from various types of tasks, our method aims to enhance adaptability, scalability, and generalizability in graph classification. Motivated by the recognition that the underlying subgraph plays a crucial role in GNN prediction, while the remainder is task-irrelevant, we introduce the Core Knowledge Learning (CKL) framework for graph adaptation and scalability learning. CKL comprises several key modules, including the core subgraph knowledge submodule, graph domain adaptation module, and few-shot learning module for downstream tasks. Each module is tailored to tackle specific challenges in graph classification, such as domain shift, label inconsistencies, and data scarcity. By learning the core subgraph of the entire graph, we focus on the most pertinent features for task relevance. Consequently, our method offers benefits such as improved model performance, increased domain adaptability, and enhanced robustness to domain variations. Experimental results demonstrate significant performance enhancements achieved by our method compared to state-of-the-art approaches. Specifically, our method achieves notable improvements in accuracy and generalization across various datasets and evaluation metrics, underscoring its effectiveness in addressing the challenges of graph classification. Bowen Zhang 0005, Zhichao Huang 0001, Guangning Xu, Xiaomao Fan, Mingyan Xiao, Genan Dai, Hu Huang 0009 |
AAAI | 2 |
| 2025 | Knowledge-Augmented Interpretable Network for Zero-Shot Stance Detection on Social MediaabstractStance detection on social media has become increasingly important for understanding public opinions on controversial issues. Existing methods often require large amounts of labeled data to learn target-independent transferable knowledge, which is infeasible under zero-shot settings where the target is unseen. Furthermore, most current stance detection models, primarily based on end-to-end deep learning architectures, lack transparency and may produce counter-intuitive and uninterpretable predictions. In this article, we propose a novel knowledge-augmented interpretable network (KAI) to enable zero-shot stance detection (ZSSD). First, we introduce an unsupervised approach based on large language models (LLMKE) to elicit analysis perspectives, which is target-independent knowledge shared across different targets. This transferable knowledge bridges connections between seen and unseen targets. Second, we develop a bidirectional knowledge-guided neural production system (Bi-KGNPS) that effectively integrates such transferable knowledge through an iterative knowledge-variable binding process to guide stance predictions. Extensive experiments on benchmark datasets demonstrate KAI achieves new state-of-the-art performance on ZSSD. Moreover, our approach also delivers strong results on conventional in-target and cross-target stance detection. With the dual benefits of knowledge-augmented accuracy and model interpretability, this work represents an important advance toward practical stance detection systems that can generalize to emerging topics of interest. The proposed KAI framework provides an interpretable approach to effectively transfer knowledge across domains for zero-shot learning. Bowen Zhang 0005, Daijun Ding, Zhichao Huang 0001, Ang Li 0047, Baoquan Zhang, Hu Huang 0009 |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2024 | EDDA: An Encoder-Decoder Data Augmentation Framework for Zero-Shot Stance DetectionabstractStance detection aims to determine the attitude expressed in text towards a given target. Zero-shot stance detection (ZSSD) has emerged to classify stances towards unseen targets during inference. Recent data augmentation techniques for ZSSD increase transferable knowledge between targets through text or target augmentation. However, these methods exhibit limitations. Target augmentation lacks logical connections between generated targets and source text, while text augmentation relies solely on training data, resulting in insufficient generalization. To address these issues, we propose an encoder-decoder data augmentation (EDDA) framework. The encoder leverages large language models and chain-of-thought prompting to summarize texts into target-specific if-then rationales, establishing logical relationships. The decoder generates new samples based on these expressions using a semantic correlation word replacement strategy to increase syntactic diversity. We also analyze the generated expressions to develop a rationale-enhanced network that fully utilizes the augmented data. Experiments on benchmark datasets demonstrate our approach substantially improves over state-of-the-art ZSSD techniques. The proposed EDDA framework increases semantic relevance and syntactic variety in augmented texts while enabling interpretable rationale-based learning. Daijun Ding, Li Dong 0011, Zhichao Huang 0001, Guangning Xu, Liwen Jing 0001, Bowen Zhang 0005 |
LREC/COLING | 3 |
| 2024 | An offline-to-online reinforcement learning approach based on multi-action evaluation with policy extension
Xuebo Cheng, Xiaohui Huang 0003, Zhichao Huang 0001, Nan Jiang 0013 |
Appl. Intell. | 3 |
| 2024 | TLS-MWP: A Tensor-Based Long- and Short-Range Convolution for Multiple Weather PredictionabstractWeather prediction plays a crucial role in human development. Recently, deep learning has demonstrated promising prospects in weather forecasting by integrating convolutional neural networks (CNNs) and recurrent neural networks (RNNs). However, two main challenges still exist in multiple weather condition prediction. The first challenge considers multiple weather condition correlations in predictions. The second challenge is how to model long- and short-range spatial dependencies under multiple weather conditions. A novel operator named as tensor-based long- and short-range convolution (TLS-Conv) is proposed to address these challenges. Within this operator, the node & relation attention is utilized to identify the contributions of spatial grid points and weather conditions for prediction. Additionally, the adaptive tensor graph convolution (ATGCN) is tailored to dynamically capture long-range spatial dependencies within multiple weather conditions. Finally, the traditional convolution is integrated with the ATGCN to model both long- and short-range spatial dependencies and weather condition correlations. Building upon the TLS-Conv, the tensor-based long- and short-range convolution for multiple weather prediction (TLS-MWP) model is proposed to predict multiple weather conditions. Extensive experiments are conducted under real-world weather conditions to evaluate its performance. These results unequivocally demonstrate that TLS-MWP surpasses previous methods. The code is available on GitHub at: https://github.com/xuguangning1218/TLS_MWP. Guangning Xu, Michael Kwok-Po Ng, Yunming Ye, Xutao Li 0003, Bowen Zhang 0005, Zhichao Huang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2023 | Knowledge-Aware Few Shot Learning for Event Detection from Short TextsabstractEvent detection in a city is crucial for the government to listen to the voice of the citizens, be aware of the real occurrences in a city, and then make wiser policies. However, in reality some important events with few samples are easily to be overwhelmed by the massive information and hard to be recognized, and additionally the limited word description from the short texts even makes the recognition harder. To address the problems, we propose a knowledge-aware event detector by incorporating the external knowledge to detect the events with few examples. The external knowledge incorporation with different semantic relations is capable to enrich the short texts. In addition, we leverage the representative few shot learning framework to formulate the event detection as the text classification problem. The proposed model is evaluated on two widely event-detection datasets. The experiments show a consistent accuracy improvement. The findings validates that our model with the knowledge infusion is effective to detect the few shot events from the short texts. Jinjin Guo, Zhichao Huang 0001, Guangning Xu, Bowen Zhang 0005, Chaoqun Duan |
ICASSP | 2 |
| 2023 | Int-GNN: A User Intention Aware Graph Neural Network for Session-Based RecommendationabstractSession-Based Recommendation (SBR) is a spotlight research problem. Although many efforts have been made, challenges still exist. The key to unlocking this shackle is the user intention, an intuitive but hard-to-model concept in the anonymous session. Unlike previous research, we suggest mining potential user intention by counting the number of item occurrences in a user session and considering the long interval between item re-interactions. Beyond these, we take user preference, a biased user intention, into account in the prediction stage. Forming these together, we propose a model named user Intention aware Graph Neural Network (Int-GNN) aiming at capturing user intention. Extensive experiments have been conducted on three real-world datasets, and the results show the superiority of our method. The code is available on GitHub: https://github.com/xuguangning1218/IntGNN_ICASSP2023 Guangning Xu, Jinyang Yang, Jinjin Guo, Zhichao Huang 0001, Bowen Zhang 0005 |
ICASSP | 4 |
| 2023 | Twitter Stance Detection via Neural Production SystemsabstractStance detection is an important task, which aims to classify the attitude of an opinionated text toward a given target. In this paper, we develop an interpretable neural production system for stance detection (NPS4SD). NPS4SD is an end-to-end deep learning model, which consists of a set of knowledge rules that are applied by binding with specific entities. NPS4SD consists of two main components: a pretrained model for learning the text representation and a variable binding network (VBN) to bind the knowledge rules with text entities. Extensive experiments are conducted to evaluate the effectiveness of the proposed NPS4SD model on three real-world datasets with in-domain, cross-target and zero-shot setups. Experimental results demonstrate that NPS4SD achieves substantially better performance than the strong competitors for the stance detection task. Bowen Zhang 0005, Daijun Ding, Guangning Xu, Jinjin Guo, Zhichao Huang 0001 |
ICASSP | 5 |
| 2023 | Multi-view knowledge graph fusion via knowledge-aware attentional graph neural network
Zhichao Huang 0001, Xutao Li 0003, Yunming Ye, Baoquan Zhang, Guangning Xu, Wensheng Gan |
Appl. Intell. | 1 |
| 2023 | TFG-Net: Tropical Cyclone Intensity Estimation from a Fine-grained perspective with the Graph convolution neural network
Guangning Xu, Yan Li 0040, Xutao Li 0003, Yunming Ye, Qingquan Lin, Zhichao Huang 0001, Shidong Chen |
Eng. Appl. Artif. Intell. | 7 |
| 2022 | Sentiment Interpretable Logic Tensor Network for Aspect-Term Sentiment AnalysisabstractAspect-term sentiment analysis (ATSA) is an important task that aims to infer the sentiment towards the given aspect-terms. It is often required in the industry that ATSA should be performed with interpretability, computational efficiency and high accuracy. However, such an ATSA method has not yet been developed. This study aims to develop an ATSA method that fulfills all these requirements. To achieve the goal, we propose a novel Sentiment Interpretable Logic Tensor Network (SILTN). SILTN is interpretable because it is a neurosymbolic formalism and a computational model that supports learning and reasoning about data with a differentiable first-order logic language (FOL). To realize SILTN with high inferring accuracy, we propose a novel learning strategy called the two-stage syntax knowledge distillation (TSynKD). Using widely used datasets, we experimentally demonstrate that the proposed TSynKD is effective for improving the accuracy of SILTN, and the SILTN has both high interpretability and computational efficiency. Bowen Zhang 0005, Zhichao Huang 0001, Hu Huang 0009, Baoquan Zhang, Xianghua Fu, Liwen Jing 0001 |
COLING | 3 |
| 2022 | LS-NTP: Unifying long- and short-range spatial correlations for near-surface temperature prediction
Guangning Xu, Xutao Li 0003, Shanshan Feng 0001, Yunming Ye, Zhihua Tu, Kenghong Lin, Zhichao Huang 0001 |
Neural Networks | 7 |
| 2022 | Multisensor Fusion and Explicit Semantic Preserving-Based Deep Hashing for Cross-Modal Remote Sensing Image RetrievalabstractCross-modal hashing is an important tool for retrieving useful information from very-high-resolution (VHR) optical images and synthetic aperture radar (SAR) images. Dealing with the intermodal discrepancies, including both spatial–spectral and visual semantic aspects, between VHR and SAR images is extremely vital to generate high-quality common hash codes in the Hamming space. However, existing cross-modal hashing methods ignore the spatial–spectral discrepancy when representing VHR and SAR images. Moreover, existing methods employ derived supervised signals, such as pairwise training images, to implicitly guide hashing learning, which fails to effectively deal with the visual semantic discrepancy, i.e., cannot adequately preserve the intraclass similarity and interclass discrimination between VHR and SAR images. To address these drawbacks, this article proposes a multisensor fusion and explicit semantic preserving-based deep Hashing method, termed as MsEspH, which can effectively deal with the discrepancies. Specifically, we design a novel cross-modal hashing network to eliminate the spatial–spectral discrepancies by fusing extra multispectral images (MSIs), which are generated in real time by a generative adversarial network. Then, we propose an explicit semantic preserving-based objective function by analyzing the connection between classification and hash learning. The objective function can preserve the intraclass similarity and interclass discrimination with class labels directly. Moreover, we theoretically verify that hash learning and classification can be unified into a learning framework under certain conditions. To evaluate our method, we construct and release a large-scale VHR-SAR image dataset. Extensive experiments on the dataset demonstrate that our method outperforms various state-of-the-art cross-modal hashing methods. Yuxi Sun 0002, Shanshan Feng 0001, Yunming Ye, Xutao Li 0003, Jian Kang 0005, Zhichao Huang 0001, Chuyao Luo |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2021 | Prototype Completion With Primitive Knowledge for Few-Shot LearningabstractFew-shot learning is a challenging task, which aims to learn a classifier for novel classes with few examples. Pre-training based meta-learning methods effectively tackle the problem by pre-training a feature extractor and then fine-tuning it through the nearest centroid based meta-learning. However, results show that the fine-tuning step makes very marginal improvements. In this paper, 1) we figure out the key reason, i.e., in the pre-trained feature space, the base classes already form compact clusters while novel classes spread as groups with large variances, which implies that fine-tuning the feature extractor is less meaningful; 2) instead of fine-tuning the feature extractor, we focus on estimating more representative prototypes during meta-learning. Consequently, we propose a novel prototype completion based meta-learning framework. This framework first introduces primitive knowledge (i.e., class-level part or attribute annotations) and extracts representative attribute features as priors. Then, we design a prototype completion network to learn to complete prototypes with these priors. To avoid the prototype completion error caused by primitive knowledge noises or class differences, we further develop a Gaussian based prototype fusion strategy that combines the mean-based and completed prototypes by exploiting the unlabeled samples. Extensive experiments show that our method: (i) can obtain more accurate prototypes; (ii) out-performs state-of-the-art techniques by 2%~9% in terms of classification accuracy. Our code is available online1. Baoquan Zhang, Xutao Li 0003, Yunming Ye, Zhichao Huang 0001, Lisai Zhang |
CVPR | 4 |
| 2020 | MR-GCN: Multi-Relational Graph Convolutional Networks based on Generalized Tensor ProductabstractGraph Convolutional Networks (GCNs) have been extensively studied in recent years. Most of existing GCN approaches are designed for the homogenous graphs with a single type of relation. However, heterogeneous graphs of multiple types of relations are also ubiquitous and there is a lack of methodologies to tackle such graphs. Some previous studies address the issue by performing conventional GCN on each single relation and then blending their results. However, as the convolutional kernels neglect the correlations across relations, the strategy is sub-optimal. In this paper, we propose the Multi-Relational Graph Convolutional Network (MR-GCN) framework by developing a novel convolution operator on multi-relational graphs. In particular, our multi-dimension convolution operator extends the graph spectral analysis into the eigen-decomposition of a Laplacian tensor. And the eigen-decomposition is formulated with a generalized tensor product, which can correspond to any unitary transform instead of limited merely to Fourier transform. We conduct comprehensive experiments on four real-world multi-relational graphs to solve the semi-supervised node classification task, and the results show the superiority of MR-GCN against the state-of-the-art competitors. Zhichao Huang 0001, Xutao Li 0003, Yunming Ye, Michael Kwok-Po Ng |
IJCAI | 1 |
| 2020 | TLVANE: a two-level variation model for attributed network embedding
Zhichao Huang 0001, Xutao Li 0003, Yunming Ye, Feng Li 0022, Feng Liu 0034, Yuan Yao 0016 |
Neural Comput. Appl. | 1 |
| 2019 | Low-resolution image categorization via heterogeneous domain adaptation
Yuan Yao 0016, Xutao Li 0003, Yunming Ye, Feng Liu 0034, Michael Kwok-Po Ng, Zhichao Huang 0001, Yu Zhang 0006 |
Knowl. Based Syst. | 6 |
| 2017 | Joint Weighted Nonnegative Matrix Factorization for Mining Attributed Graphs
Zhichao Huang 0001, Yunming Ye, Xutao Li 0003, Feng Liu 0034, Huajie Chen |
PAKDD (1) | 1 |
| 2016 | A Semi-supervised Clustering Method through Bottleneck Distance ExplorationabstractSemi-supervised clustering is one of the most active research area in machine learning and pattern recognition, which can improve the performance of unsupervised clustering efficiently. This paper focuses on exploiting both the label information of a few labeled samples and the spatial distribution information of large amount of unlabeled samples. We proposed a new semi-supervised clustering method, named Bottleneck Distance based Semi-supervised Clustering (BDSC), which is based on the idea of label propagation and can perform clustering with no parameters. BDSC works by firstly obtaining small amount of labeled samples for each class. Then, a minimum spanning tree is constructed from both labeled and unlabeled samples, where the distances between an unlabeled sample and labeled samples are computed to get the bottleneck distance for each unlabeled sample. Finally, labels are propagated by comparing the bottleneck distances. Experimental results demonstrate that the proposed technique outperforms classical clustering algorithms with respect to the precision and the capability of recognizing nonspherical-shaped clusters. Yuan Yao 0016, Yan Li 0040, Ke Wang 0068, Zhichao Huang 0001, Yunming Ye |
ICSS | 4 |