Zhiwei Hu

dblp:89/7644 · DBLP profile ↗
← Back
23ranked-venue papers
12as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 11 first-author · 14 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 FedRMamba: Federated Residual Mamba for Multivariate Time-Series Forecasting
abstract
Time series forecasting underpins many real-world services. Recent trends have focused on foundation models inspired by the paradigm of large language models, which rely on large volumes of centralized time-series data across diverse domains. However, such approaches raise significant concerns regarding data privacy. Federated learning (FL) has emerged as a promising paradigm for training unified time-series models using isolated datasets distributed across multiple clients. Nevertheless, existing FL methods face two critical challenges: heterogeneous variables and heterogeneous temporal correlations. To address these issues, we propose FedRMamba, a personalized federated forecasting framework built entirely from Mamba state-space blocks. Each client adopts a residual-coupled architecture, where a global frequency-aware Mamba module captures the common low-frequency structures shared across different variables, while a local patch-wise Mamba module learns personalized high-frequency patterns within the multivariate context. To clearly separate these responsibilities, we introduce a frequency-aware supervision that aligns the global path with low-frequency components and the local path with high-frequency residuals. Additionally, we design a gated fusion mechanism that dynamically combines the low-frequency and high-frequency components for improved prediction. We conduct extensive experiments to evaluate the performance of our proposed framework, demonstrating its effectiveness in handling heterogeneous data in federated settings.
Zhiwei Hu, Liang Zhang 0042, Guangxu Zhu
WWW1
2026 Plan Then Retrieve: Reinforcement Learning-Guided Complex Reasoning over Knowledge Graphs
abstract
Knowledge Graph Question Answering (KGQA) aims to answer natural language questions by reasoning over structured knowledge graphs (KGs). While large language models (LLMs) have advanced KGQA through their strong reasoning capabilities, existing methods continue to struggle to fully exploit both the rich knowledge encoded in KGs and the reasoning capabilities of LLMs, particularly in complex scenarios. They often assume complete KG coverage and lack mechanisms to judge when external information is needed, and their reasoning remains locally myopic, failing to maintain coherent multi-step planning, leading to reasoning failures even when relevant knowledge exists. We propose Graph-RFT, a novel two-stage reinforcement fine-tuning KGQA framework with a ''plan–KGsearch–and–Websearch–during–think'' paradigm, that enables LLMs to perform autonomous planning and adaptive retrieval scheduling across KG and web sources under incomplete knowledge conditions. Graph-RFT introduces a chain-of-thought (CoT) fine-tuning method with a customized plan–retrieval dataset activates structured reasoning and resolves the GRPO cold-start problem. It then introduces a novel plan–retrieval guided reinforcement learning process integrates explicit planning and retrieval actions with a multi-reward design, enabling coverage-aware retrieval scheduling. It employs a Cartesian-inspired planning module to decompose complex questions into ordered sub-questions, and logical expression to guide tool invocation for globally consistent multi-step reasoning. This reasoning–retrieval process is optimized with a multi-reward combining outcome and retrieval-specific signals, enabling the model to learn when and how to combine KG and web retrieval effectively. Experiments on multiple KGQA benchmarks demonstrate that Graph-RFT achieves superior performance over strong baselines, even with smaller LLM backbones, and substantially improves complex question decomposition, factual coverage, and tool coordination.
Yanlin Song, Ben Liu 0002, Víctor Gutiérrez-Basulto, Zhiwei Hu, Qianqian Xie, Min Peng 0002, Sophia Ananiadou, Jeff Z. Pan
WWW4
2026 Leveraging intra-modal and inter-modal interaction for multi-modal entity alignment
Zhiwei Hu, Víctor Gutiérrez-Basulto, Zhiliang Xiang, Ru Li 0001, Jeff Z. Pan
Neurocomputing1
2025 Multi-level Matching Network for Multimodal Entity Linking
abstract
Multimodal entity linking (MEL) aims to link ambiguous mentions within multimodal contexts to corresponding entities in a multimodal knowledge base. Most existing approaches to MEL are based on representation learning or vision-and-language pre-training mechanisms for exploring the complementary effect among multiple modalities. However, these methods suffer from two limitations. On the one hand, they overlook the possibility of considering negative samples from the same modality. On the other hand, they lack mechanisms to capture bidirectional cross-modal interaction. To address these issues, we propose a Multi-level Matching network for Multimodal Entity Linking(M3EL). Specifically, M3EL is composed of three different modules: (i) a Multimodal Feature Extraction module, which extracts modality-specific representations with a multimodal encoder and introduces an intra-modal contrastive learning sub-module to obtain better discriminative embeddings based on uni-modal differences; (ii) an Intra-modal Matching Network module, which contains two levels of matching granularity: Coarse-grained Global-to-Global and Fine-grained Global-to-Local, to achieve local and global level intra-modal interaction; (iii) a Cross-modal Matching Network module, which applies bidirectional strategies, Textual-to-Visual and Visual-to-Textual matching, to implement bidirectional cross-modal interaction. Extensive experiments conducted on WikiMEL, RichpediaMEL, and WikiDiverse datasets demonstrate the outstanding performance of M3EL when compared to the state-of-the-art baselines.
Zhiwei Hu, Víctor Gutiérrez-Basulto, Ru Li 0001, Jeff Z. Pan
KDD (1)1
2025 Multi-level Mixture of Experts for Multimodal Entity Linking
abstract
Multimodal Entity Linking (MEL) aims to link ambiguous mentions within multimodal contexts to associated entities in a multimodal knowledge base. Existing approaches to MEL introduce multimodal interaction and fusion mechanisms to bridge the modality gap and enable multi-grained semantic matching. However, they do not address two important problems: (i) mention ambiguity, i.e., the lack of semantic content caused by the brevity and omission of key information in the mention's textual context; (ii) dynamic selection of modal content, i.e., to dynamically distinguish the importance of different parts of modal information. To mitigate these issues, we propose a Multi-level Mixture of Experts (MMoE) model for MEL. MMoE has four components: (i) the description-aware mention enhancement module leverages large language models to identify the WikiData descriptions that best match a mention, considering the mention's textual context; (ii) the multimodal feature extraction module adopts multimodal feature encoders to obtain textual and visual embeddings for both mentions and entities; (iii)-(iv) the intra-level mixture of experts and inter-level mixture of experts modules apply a switch mixture of experts mechanism to dynamically and adaptively select features from relevant regions of information. Extensive experiments on WikiMEL, RichpediaMEL and WikiDiverse datasets demonstrate the outstanding performance of MMoE compared to the state-of-the-art. MMoE's code is available at: https://github.com/zhiweihu1103/MEL-MMoE.
Zhiwei Hu, Víctor Gutiérrez-Basulto, Zhiliang Xiang, Ru Li 0001, Jeff Z. Pan
KDD (2)1
2024 Learning From Box Annotations for Referring Image Segmentation
abstract
Referring image segmentation (RIS) has obtained an impressive achievement by fully convolutional networks (FCNs). However, previous RIS methods require a large number of pixel-level annotations. In this article, we present a weakly supervised RIS method by using bounding box (BB) annotations. In the first stage, we introduce an adversarial boundary loss to extract the object contour from the BB, which is then used to select appropriate region proposals for pseudoground-truth (PGT) generation. In the second stage, we design a co-training (Co-T) strategy to purify the pseudolabels. Specifically, we train two networks and interactively guide them to pick clean labels for each other's networks, which can weaken the effect of noisy labels on model training. Experiment results on four benchmark datasets demonstrate that the proposed method can produce high-quality masks with a speed of 63 frames/s.
Lihe Zhang, Zhiwei Hu, Huchuan Lu
IEEE Trans. Neural Networks Learn. Syst.3
2023 HyperFormer: Enhancing Entity and Relation Interaction for Hyper-Relational Knowledge Graph Completion
abstract
Hyper-relational knowledge graphs (HKGs) extend standard knowledge graphs by associating attribute-value qualifiers to triples, which effectively represent additional fine-grained information about its associated triple. Hyper-relational knowledge graph completion (HKGC) aims at inferring unknown triples while considering its qualifiers. Most existing approaches to HKGC exploit a global-level graph structure to encode hyper-relational knowledge into the graph convolution message passing process. However, the addition of multi-hop information might bring noise into the triple prediction process. To address this problem, we propose HyperFormer, a model that considers local-level sequential information, which encodes the content of the entities, relations and qualifiers of a triple. More precisely, HyperFormer is composed of three different modules: an entity neighbor aggregator module allowing to integrate the information of the neighbors of an entity to capture different perspectives of it; a relation qualifier aggregator module to integrate hyper-relational knowledge into the corresponding relation to refine the representation of relational content; a convolution-based bidirectional interaction module based on a convolutional operation, capturing pairwise bidirectional interactions of entity-relation, entity-qualifier, and relation-qualifier. Furthermore, we introduce a Mixture-of-Experts strategy into the feed-forward layers of HyperFormer to strengthen its representation capabilities while reducing the amount of model parameters and computation. Extensive experiments on three well-known datasets with four different conditions demonstrate HyperFormer's effectiveness. Datasets and code are available at https://github.com/zhiweihu1103/HKGC-HyperFormer.
Zhiwei Hu, Víctor Gutiérrez-Basulto, Zhiliang Xiang, Ru Li 0001, Jeff Z. Pan
CIKM1
2023 MLPST: MLP is All You Need for Spatio-Temporal Prediction
abstract
Traffic prediction is a typical spatio-temporal data mining task and has great significance to the public transportation system. Considering the demand for its grand application, we recognize key factors for an ideal spatio-temporal prediction method: efficient, lightweight, and effective. However, the current deep model-based spatio-temporal prediction solutions generally own intricate architectures with cumbersome optimization, which can hardly meet these expectations. To accomplish the above goals, we propose an intuitive and novel framework, MLPST, a pure multi-layer perceptron architecture for traffic prediction. Specifically, we first capture spatial relationships from both local and global receptive fields. Then, temporal dependencies in different intervals are comprehensively considered. Through compact and swift MLP processing, MLPST can well capture the spatial and temporal dependencies while requiring only linear computational complexity, as well as model parameters that are more than an order of magnitude lower than baselines. Extensive experiments validated the superior effectiveness and efficiency of MLPST against advanced baselines, and among models with optimal accuracy, MLPST achieves the best time and space efficiency.
Zijian Zhang 0009, Ze Huang, Zhiwei Hu, Xiangyu Zhao 0001, Zitao Liu 0001, Junbo Zhang 0004, S. Joe Qin
CIKM3
2023 Multi-view Contrastive Learning for Entity Typing over Knowledge Graphs
abstract
Knowledge graph entity typing (KGET) aims at inferring plausible types of entities in knowledge graphs.Existing approaches to KGET focus on how to better encode the knowledge provided by the neighbors and types of an entity into its representation.However, they ignore the semantic knowledge provided by the way in which types can be clustered together.In this paper, we propose a novel method called Multi-view Contrastive Learning for knowledge graph Entity Typing (MCLET), which effectively encodes the coarse-grained knowledge provided by clusters into entity and type embeddings.MCLET is composed of three modules: i) Multi-view Generation and Encoder module, which encodes structured information from entity-type, entity-cluster and cluster-type views; ii) Cross-view Contrastive Learning module, which encourages different views to collaboratively improve view-specific representations of entities and types; iii) Entity Typing Prediction module, which integrates multi-head attention and a Mixture-of-Experts strategy to infer missing entity types.Extensive experiments show the strong performance of MCLET compared to the state-of-the-art.
Zhiwei Hu, Víctor Gutiérrez-Basulto, Zhiliang Xiang, Ru Li 0001, Jeff Z. Pan
EMNLP1
2023 Only Classification Head Is Sufficient for Medical Image Segmentation
Hongbin Wei, Zhiwei Hu, Zhilong Ji, Hongpeng Jia, Lihe Zhang, Huchuan Lu
PRCV (13)2
2023 Referring Segmentation via Encoder-Fused Cross-Modal Attention Network
abstract
This paper focuses on referring segmentation, which aims to selectively segment the corresponding visual region in an image (or video) according to the referring expression. However, the existing methods usually consider the interaction between multi-modal features at the decoding end of the network. Specifically, they interact the visual features of each scale with language respectively, thus ignoring the correlation between multi-scale features. In this work, we present an encoder fusion network (EFN), which transfers the multi-modal feature learning process from the decoding end to the encoding end and realizes the gradual refinement of multi-modal features by the language. In EFN, we also adopt a co-attention mechanism to promote the mutual alignment of language and visual information in feature space. In the decoding stage, a boundary enhancement module (BEM) is proposed to enhance the network's attention to the details of the target. For video data, we introduce an asymmetric cross-frame attention module (ACFM) to effectively capture the temporal information from the video frames by computing the relationship between each pixel of the current frame and each pooled sub-region of the reference frames. Extensive experiments on referring image/video segmentation datasets show that our method outperforms the state-of-the-art performance.
Lihe Zhang, Zhiwei Hu, Huchuan Lu
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 A Span-based Target-aware Relation Model for Frame-semantic Parsing
abstract
Frame-semantic Parsing (FSP) is a challenging and critical task in Natural Language Processing (NLP). Most of the existing studies decompose the FSP task into frame identification (FI) and frame semantic role labeling (FSRL) subtasks, and adopt a pipeline model architecture that clearly causes error propagation problem. However, recent jointly learning models aim to address the above problem and generally treat FSP as a span-level structured prediction task, which, unfortunately, leads to cascading error propagation problem between roles and less-efficient solutions due to huge search space of roles. To address these problems, we reformulate the FSRL task into a target-aware relation classification task and propose a novel and lightweight jointly learning framework that simultaneously processes three subtasks of FSP, including frame identification, argument identification, and role classification. The novel task formulation and jointly learning with interaction mechanisms among subtasks can help improve the overall system performance and reduce the search space and time complexity, compared with existing methods. Extensive experimental results demonstrate that our proposed model significantly outperforms 10 state-of-the-art models in terms of F1 score across two benchmark datasets.
Xuefeng Su, Ru Li 0001, Xiaoli Li 0001, Baobao Chang, Zhiwei Hu, Xiaoqi Han, Zhichao Yan 0002
ACM Trans. Asian Low Resour. Lang. Inf. Process.5
2023 Bidirectional Relationship Inferring Network for Referring Image Localization and Segmentation
abstract
Recently, referring image localization and segmentation has aroused widespread interest. However, the existing methods lack a clear description of the interdependence between language and vision. To this end, we present a bidirectional relationship inferring network (BRINet) to effectively address the challenging tasks. Specifically, we first employ a vision-guided linguistic attention module to perceive the keywords corresponding to each image region. Then, language-guided visual attention adopts the learned adaptive language to guide the update of the visual features. Together, they form a bidirectional cross-modal attention module (BCAM) to achieve the mutual guidance between language and vision. They can help the network align the cross-modal features better. Based on the vanilla language-guided visual attention, we further design an asymmetric language-guided visual attention, which significantly reduces the computational cost by modeling the relationship between each pixel and each pooled subregion. In addition, a segmentation-guided bottom-up augmentation module (SBAM) is utilized to selectively combine multilevel information flow for object localization. Experiments show that our method outperforms other state-of-the-art methods on three referring image localization datasets and four referring image segmentation datasets.
Zhiwei Hu, Lihe Zhang, Huchuan Lu
IEEE Trans. Neural Networks Learn. Syst.2
2022 Transformer-based Entity Typing in Knowledge Graphs
abstract
We investigate the knowledge graph entity typing task which aims at inferring plausible entity types.In this paper, we propose a novel Transformer-based Entity Typing (TET) approach, effectively encoding the content of neighbors of an entity.More precisely, TET is composed of three different mechanisms: a local transformer allowing to infer missing types of an entity by independently encoding the information provided by each of its neighbors; a global transformer aggregating the information of all neighbors of an entity into a single long sequence to reason about more complex entity types; and a context transformer integrating neighbors content based on their contribution to the type inference through information exchange between neighbor pairs.Furthermore, TET uses information about class membership of types to semantically strengthen the representation of an entity.Experiments on two real-world datasets demonstrate the superior performance of TET compared to the state-of-the-art.
Zhiwei Hu, Víctor Gutiérrez-Basulto, Zhiliang Xiang, Ru Li 0001, Jeff Z. Pan
EMNLP1
2022 Type-aware Embeddings for Multi-Hop Reasoning over Knowledge Graphs
abstract
Multi-hop reasoning over real-life knowledge graphs (KGs) is a highly challenging problem as traditional subgraph matching methods are not capable to deal with noise and missing information. Recently, to address this problem a promising approach based on jointly embedding logical queries and KGs into a low-dimensional space to identify answer entities has emerged. However, existing proposals ignore critical semantic knowledge inherently available in KGs, such as type information. To leverage type information, we propose a novel type-aware model, TypE-aware Message Passing (TEMP), which enhances the entity and relation representation in queries, and simultaneously improves generalization, and deductive and inductive reasoning. Remarkably, TEMP is a plug-and-play model that can be easily incorporated into existing embedding-based models to improve their performance. Extensive experiments on three real-world datasets demonstrate TEMP’s effectiveness.
Zhiwei Hu, Víctor Gutiérrez-Basulto, Zhiliang Xiang, Xiaoli Li 0001, Ru Li 0001, Jeff Z. Pan
IJCAI1
2021 Encoder Fusion Network With Co-Attention Embedding for Referring Image Segmentation
abstract
Recently, referring image segmentation has aroused widespread interest. Previous methods perform the multi-modal fusion between language and vision at the decoding side of the network. And, linguistic feature interacts with visual feature of each scale separately, which ignores the continuous guidance of language to multi-scale visual features. In this work, we propose an encoder fusion network (EFN), which transforms the visual encoder into a multi-modal feature learning network, and uses language to refine the multi-modal features progressively. Moreover, a co-attention mechanism is embedded in the EFN to realize the parallel update of multi-modal features, which can promote the consistent of the cross-modal information representation in the semantic space. Finally, we propose a boundary enhancement module (BEM) to make the network pay more attention to the fine structure. The experiment results on four benchmark datasets demonstrate that the proposed approach achieves the state-of-the-art performance under different evaluation metrics without any post-processing.
Zhiwei Hu, Lihe Zhang, Huchuan Lu
CVPR2
2021 Infrared Target Tracking Based on Improved Particle Filtering
abstract
Infrared target tracking technology is one of the core technologies in infrared imaging guidance systems and is also a hot research topic. The problem of particle degradation could be always found in traditional particle filtering, and a large number of particles are additionally required for accurate estimation. It is difficult to meet the requirements of a modern infrared imaging guidance system for accurate target tracking. To solve the problem of particle degradation and improve the performance of infrared target tracking, the extended Kalman filter and genetic algorithm are introduced into particle filtering, and an improved algorithm for infrared target tracking is proposed in this paper. In the framework of a particle filter algorithm, the Gaussian distribution for each particle is generated and propagated by a separate extended Kalman filter to improve the sampling accuracy and effectiveness of the probability density function of particles. Genetic algorithm is used to perform a resampling process to solve particle degradation and ensure the diversity of particle states in particle swarm. Simulation results show that the improved tracking algorithm based on improved particle filtering proposed in this paper can effectively solve the phenomenon of particle degradation and track the infrared target.
Zhiwei Hu
Int. J. Pattern Recognit. Artif. Intell.1
2020 Bi-Directional Relationship Inferring Network for Referring Image Segmentation
abstract
Most existing methods do not explicitly formulate the mutual guidance between vision and language. In this work, we propose a bi-directional relationship inferring network (BRINet) to model the dependencies of cross-modal information. In detail, the vision-guided linguistic attention is used to learn the adaptive linguistic context corresponding to each visual region. Combining with the language-guided visual attention, a bi-directional cross-modal attention module (BCAM) is built to learn the relationship between multi-modal features. Thus, the ultimate semantic context of the target object and referring expression can be represented accurately and consistently. Moreover, a gated bi-directional fusion module (GBFM) is designed to integrate the multi-level features where a gate function is used to guide the bi-directional flow of multi-level information. Extensive experiments on four benchmark datasets demonstrate that the proposed method outperforms other state-of-the-art methods under different evaluation metrics.
Zhiwei Hu, Lihe Zhang, Huchuan Lu
CVPR1
2020 Infrared Small Target Detection Based on Morphology and SUSAN Algorithm
abstract
Infrared small target detection is one of the key techniques in infrared imaging guidance system. The technology of infrared small target detection still needs to be further studied to improve the detection performance. This paper combines the high-pass filtering characteristics of morphological top-hat transform with SUSAN algorithm, and proposes a small infrared target detection method based on morphology and SUSAN algorithm. This method uses top-hat transform to detect the high-frequency region in infrared image, and filters out the low-frequency region in the image to implement the preliminary background suppression of infrared image. Then the SUSAN algorithm is used to detect small targets in the image after background suppression. The proposed method is applied to the single infrared image which is acquired by the infrared guidance system in the process of detecting and tracking the target under specific conditions. The experimental results show that the method is effective and can detect infrared small targets under different background.
Zhiwei Hu
Int. J. Pattern Recognit. Artif. Intell.1
2020 Improving biterm topic model with word embeddings
Min Peng 0002, Pengwei Li, Zhiwei Hu
World Wide Web4
2019 Feature Selection Based on Graph Structure
Zhiwei Hu, Zhaogong Zhang, Zongchao Huang, Dayuan Zheng
COCOA1
2011 Robust object tracking with occlusion handle
Gang Yu 0002, Zhiwei Hu
Neural Comput. Appl.2
2009 Robust Incremental Subspace Learning for Object Tracking
Gang Yu 0002, Zhiwei Hu
ICONIP (1)2