VLDB 2026 Research / reviewers in the wild / expert
Tengfei Liu 0005
dblp:64/8133-5
· DBLP profile ↗
19ranked-venue papers
10as first author
18since 2021 · last 2026
0000-0002-2739-5220ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 6 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MARE: Multimodal Analogical Reasoning for Disease Evolution-Aware Radiology Report GenerationabstractRadiology report generation from longitudinal medical data is critical for assessing disease progression and automating diagnostic workflows. While recent methods incorporate longitudinal information, they primarily rely on multimodal feature fusion, with limited capacity for explicit disease evolution modeling and temporal reasoning. To address this, we propose MARE, an end-to-end framework that formulates longitudinal radiology report generation as a multimodal analogical reasoning task. Inspired by the Abduction–Mapping–Induction paradigm, MARE models latent relational structures underlying disease evolution by aligning lesion-level visual features across time and mapping them to the textual domain for temporally coherent and clinically meaningful report generation. To mitigate the spatial misalignment caused by patient positioning or imaging variation, we introduce an Adaptive Region Alignment (ARA) module for robust temporal correspondence. Additionally, we design Dual Evolution Consistency (DEC) losses to regularize analogical reasoning by enforcing temporal coherence in both visual and textual evolution paths. Extensive experiments on the Longitudinal-MIMIC dataset demonstrate that MARE significantly outperforms state-of-the-art baselines across both natural language generation and clinical effectiveness metrics, highlighting the value of structured analogical reasoning for disease evolution-aware report generation. Qingqing Gao, Tengfei Liu 0005, Xiaodan Zhang 0003, Zhongfan Sun, Boyue Wang |
AAAI | 2 |
| 2026 | Multi-agent role-playing by LLMs and LMMs: An explainable open-world multi-modal crisis tweet classification method
Tong Bie, Yongli Hu, Linjia Hao, Tengfei Liu 0005, Huajie Jiang, Junbin Gao |
Expert Syst. Appl. | 5 |
| 2026 | Distilling Object Detectors via Monte Carlo DropoutabstractKnowledge distillation (KD) has become a fundamental technique for model compression in object detection tasks. The data noise and training randomness may cause the knowledge of the teacher model to be unreliable, referred to as knowledge uncertainty. Existing methods neglect this uncertainty, potentially hindering the student's capacity to capture and understand latent "dark knowledge". In this work, we introduce a novel strategy that explicitly incorporates knowledge uncertainty, named Uncertainty-Driven Knowledge Extraction and Transfer (UET). Given the unknown, high-dimensional nature of the knowledge distribution, we employ Monte Carlo dropout to effectively estimate the teacher's uncertainty. Leveraging information theory, we combine uncertainty with deterministic knowledge, enabling the student to benefit from both precision and diversity. UET is a plug-and-play method that integrates seamlessly with existing distillation techniques. We validate our approach through comprehensive experiments across various distillation strategies, detectors, and backbones. Specifically, UET achieves state-of-the-art results, with a ResNet50-based GFL detector obtaining 44.1% mAP on the COCO dataset-surpassing baseline performance by 3.9%. Junfei Yi, Hui Zhang 0023, Jianxu Mao, Tengfei Liu 0005, Mingjie Li 0006, Sihao Lin, Hanyu Gu, Zhihui Li 0001, Xiaojun Chang, Yaonan Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | M3Former: Memory-Guided Multi-Modal Generation and Adaptive Mixture Reasoning for Incomplete-Modality Crisis Event DetectionabstractMulti-modal crisis event detection is critical for timely situational awareness across diverse real-world emergencies. In practice, however, multi-modal data—typically composed of images and textual reports—are often incomplete, with many instances providing only a single modality because of data loss, platform constraints, or real-time limitations. This modality-incomplete setting poses three key challenges: (1) how to reliably reconstruct missing modalities to restore cross-modal context, (2) how to extract deep semantic clues across original and completed data, and (3) how to bridge the distributional gap between reconstructed and fully-observed samples during model training. To tackle these issues, we propose M3Former, a unified framework tailored for modality-incomplete multi-modal crisis event detection. It consists of four dedicated modules: Memory-Guided Modality Completion builds a memory bank of paired image–text data to retrieve semantically related samples and keywords, guiding powerful pretrained generators—diffusion models for image synthesis and multi-modal large language models for text generation; Crisis-Aware Heterogeneous Dual-Stream Encoders jointly capture modality-specific cues and establish initial cross-modal semantic alignment between original and completed data; Hierarchical Attention Refinement Network progressively refines representations through hierarchical attention and guided cross-modal interaction to suppress noise and semantic drift; Modality-aware Routing Experts designs a gated mixture-of-experts architecture that dynamically selects both modality-specific and shared experts to mitigate distributional shifts and enhance the quality of fused representations. Extensive experiments on two real-world crisis datasets demonstrate that M3Former significantly outperforms existing baselines under a variety of modality-incomplete scenarios. The code is available: https://github.com/lcygky/M3former. Chenyang Lu 0014, Boyue Wang, Tengfei Liu 0005, Yongli Hu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | HC-LLM: Historical-Constrained Large Language Models for Radiology Report GenerationabstractRadiology report generation (RRG) models typically focus on individual exams, often overlooking the integration of historical visual or textual data, which is crucial for patient follow-ups. Traditional methods usually struggle with long sequence dependencies when incorporating historical information, but large language models (LLMs) excel at in-context learning, making them well-suited for analyzing longitudinal medical data. In light of this, we propose a novel Historical-Constrained Large Language Models (HC-LLM) framework for RRG, empowering LLMs with longitudinal report generation capabilities by constraining the consistency and differences between longitudinal images and their corresponding reports. Specifically, our approach extracts both time-shared and time-specific features from longitudinal chest X-rays and diagnostic reports to capture disease progression. Then, we ensure consistent representation by applying intra-modality similarity constraints and aligning various features across modalities with multimodal contrastive and structural constraints. These combined constraints effectively guide the LLMs in generating diagnostic reports that accurately reflect the progression of the disease, achieving state-of-the-art results on the Longitudinal-MIMIC dataset. Notably, our approach performs well even without historical data during testing and can be easily adapted to other multimodal large models, enhancing its versatility. Tengfei Liu 0005, Jiapu Wang, Yongli Hu, Mingjie Li 0006, Junfei Yi, Xiaojun Chang, Junbin Gao |
AAAI | 1 |
| 2025 | Tackling Real-World Complexity: Hierarchical Modeling and Dynamic Prompting for Multimodal Long Document ClassificationabstractWith the rapid growth of internet content, multimodal long document data has become increasingly prominent, drawing significant attention from researchers. However, most existing methods primarily focus on scenarios where all modalities are present, often overlooking more challenging and realistic cases involving missing image modality. To address this limitation, we propose a robust multimodal long document classification (MLDC) framework that integrates hierarchical modeling and dynamic prompting to handle complex multimodal long document data. Our approach begins by leveraging hierarchical modeling combined with an Adaptive Correlation Multimodal Transformer (ACMT) to effectively capture relationships between text and images at both section and sentence levels. We also introduce a Dynamic Prompt Generation (DPG) module at both levels to enhance the model’s robustness in handling missing image data. By evaluating sample uncertainty, the DPG module dynamically adjusts both the number of prompts and the prompts themselves, allowing the model to better adapt to the varying needs of different samples. Finally, a Hierarchical Heterogeneous Graph (HHG) is introduced to enhance feature interactions across levels, further improving the coherence and accuracy of the model. Extensive experiments on four multi-modal long document datasets demonstrate that our model shows superior performance compared to existing state-of-the-art MLDC classification methods in various conditions. Tengfei Liu 0005, Yongli Hu, Mingjie Li 0006, Junfei Yi, Xiaojun Chang, Junbin Gao |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Toward Efficient Power Scene Detection via Topology-Preserved Knowledge DistillationabstractThe power industry relies on efficient inspection systems to ensure stability and safety. While deep learning has advanced automated inspection, its reliance on custom modules for specific tasks can impact efficiency. Knowledge distillation (KD) offers a balanced solution, but the complex textures and structures of power equipment challenge conventional KD methods, which often fail to capture essential local semantic and topological relationships. To address this, we proposeTopNet, a novel topology-preserved KD framework for power scene detection tasks. Specifically, we model the teacher’s knowledge as a graph, where nodes encode local fine-grained features and edges capture global topological relationships. Based on this, we introduce node feature distillation and edge feature distillation to transfer local–global structural knowledge, which can enhance the student’s ability to perceive objects. Furthermore, we also introduce aggregated feature distillation to incorporate and transfer contextual semantic knowledge. Comprehensive experiments are conducted on two different benchmark datasets to demonstrate that TopNet achieves state-of-the-art detection performance with high efficiency, offering a robust solution for automated power equipment inspection. Junfei Yi, Tengfei Liu 0005, Jianxu Mao, Yaonan Wang 0001, Hui Zhang 0023, He Xie, Hang Zhong, Xiaojun Chang |
IEEE Trans. Ind. Informatics | 2 |
| 2025 | Balancing Accuracy and Efficiency With a Multiscale Uncertainty-Aware Knowledge-Based Network for Transmission Line InspectionabstractReal-world transmission line inspections (RTLIs) ensure power stability and safety. Deep learning (DL) models have become prevalent approaches for performing RTLI tasks. However, the high computational demands and substantial parameter requirements of DL models limit their real-world applicability. This article introduces a novel approach, a multiscale uncertainty-aware knowledge-based network, which is designed to balance the accuracy and efficiency in RTLI tasks. Specifically, we propose an uncertainty-aware knowledge distillation method that incorporates pixel-level uncertainty into the knowledge transfer process, mitigating the impact of noisy knowledge derived from extra background information contained in ground truths. In addition, our method integrates a multiscale relationship distillation technique, thus enhancing the transfer of multiscale information between the teacher and student models. Consequently, RTLI tasks can be efficiently accomplished using the well-learned lightweight student model. Comprehensive experiments conducted on a real-world dataset collected via uncrewedaerial vehicles demonstrate the efficacy of our proposed approach in terms of achieving high detection accuracy with reduced computational costs. Junfei Yi, Jianxu Mao, Hui Zhang 0023, Yurong Chen 0003, Tengfei Liu 0005, Kai Zeng 0010, He Xie, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2025 | Hierarchical Multi-Modal Transformer for Cross-Modal Long Document ClassificationabstractLong Document Classification (LDC) has gained significant attention recently. However, multi-modal data in long documents such as texts and images are not being effectively utilized. Prior studies in this area have attempted to integrate texts and images in document-related tasks, but they have only focused on short text sequences and images of pages. How to classify long documents with hierarchical structure texts and embedding images is a new problem and faces multi-modal representation difficulties. In this paper, we propose a novel approach called Hierarchical Multi-modal Transformer (HMT) for cross-modal long document classification. The HMT conducts multi-modal feature interaction and fusion between images and texts in a hierarchical manner. Our approach uses a multi-modal transformer and a dynamic multi-scale multi-modal transformer to model the complex relationships between image features, and the section and sentence features. Furthermore, we introduce a new interaction strategy called the dynamic mask transfer module to integrate these two transformers by propagating features between them. To validate our approach, we conduct cross-modal LDC experiments on two newly created and two publicly available multi-modal long document datasets, and the results show that the proposed HMT outperforms state-of-the-art single-modality and multi-modality methods. Tengfei Liu 0005, Yongli Hu, Junbin Gao |
IEEE Trans. Multim. | 1 |
| 2024 | Multi-modal long document classification based on Hierarchical Prompt and Multi-modal Transformer
Tengfei Liu 0005, Yongli Hu, Junbin Gao, Jiapu Wang |
Neural Networks | 1 |
| 2024 | Hierarchical Multi-Granularity Interaction Graph Convolutional Network for Long Document ClassificationabstractWith the growing demand for text analytics, long document classification (LDC) has received extensive attention, and great progress has been made. To reveal the complex structure and extract the intrinsic feature, the current approaches focus on modeling a long sequence with sparse attention or representing word-sentence or word-section relations partially. However, the thorough hierarchical structure from words, sentences to sections of long documents remains relatively unexplored. For this purpose, we propose a novel Hierarchical Multi-granularity Interaction Graph Convolutional Network (HMIGCN) for long document classification, in which three different granularity graphs, i.e., section graph, sentence graph and word graph, are constructed hierarchically. The section graph encapsulates the macrostructure of a long document, while the sentence and word graphs delve into the document's microstructure. Notably, within the sentence graph, we introduce a Global-Local Graph Convolutional (GLGC) block to adaptively capture both global and local dependency structures among sentence nodes. Additionally, to integrate the three graph networks as a whole, two well-designed techniques, namely section-guided pooling block and transfer fusion block, are proposed to train the model jointly by promoting each other. Extensive experiments on five long document datasets show that our model outperforms the existing state-of-the-art LDC models. Tengfei Liu 0005, Yongli Hu, Junbin Gao |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2024 | Hierarchical Multi-Modal Prompting Transformer for Multi-Modal Long Document ClassificationabstractIn the context of long document classification (LDC), effectively utilizing multi-modal information encompassing texts and images within these documents has not received adequate attention. This task showcases several notable characteristics. Firstly, the text possesses an implicit or explicit hierarchical structure consisting of sections, sentences, and words. Secondly, the distribution of images is dispersed, encompassing various types such as highly relevant topic images and loosely related reference images. Lastly, intricate and diverse relationships exist between images and text at different levels. To address these challenges, we propose a novel approach called Hierarchical Multi-modal Prompting Transformer (HMPT). Our proposed method constructs the uni-modal and multi-modal transformers at both the section and sentence levels, facilitating effective interaction between features. Notably, we design an adaptive multi-scale multi-modal transformer tailored to capture the multi-granularity correlations between sentences and images. Additionally, we introduce three different types of shared prompts, i.e., shared section, sentence, and image prompts, as bridges connecting the isolated transformers, enabling seamless information interaction across different levels and modalities. To validate the model performance, we conducted experiments on two newly created and two publicly available multi-modal long document datasets. The obtained results show that our method outperforms state-of-the-art single-modality and multi-modality classification methods. Tengfei Liu 0005, Yongli Hu, Junbin Gao |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | MADE: Multicurvature Adaptive Embedding for Temporal Knowledge Graph CompletionabstractTemporal knowledge graphs (TKGs) are receiving increased attention due to their time-dependent properties and the evolving nature of knowledge over time. TKGs typically contain complex geometric structures, such as hierarchical, ring, and chain structures, which can often be mixed together. However, embedding TKGs into Euclidean space, as is typically done with TKG completion (TKGC) models, presents a challenge when dealing with high-dimensional nonlinear data and complex geometric structures. To address this issue, we propose a novel TKGC model called multicurvature adaptive embedding (MADE). MADE models TKGs in multicurvature spaces, including flat Euclidean space (zero curvature), hyperbolic space (negative curvature), and hyperspherical space (positive curvature), to handle multiple geometric structures. We assign different weights to different curvature spaces in a data-driven manner to strengthen the ideal curvature spaces for modeling and weaken the inappropriate ones. Additionally, we introduce the quadruplet distributor (QD) to assist the information interaction in each geometric space. Ultimately, we develop an innovative temporal regularization to enhance the smoothness of timestamp embeddings by strengthening the correlation of neighboring timestamps. Experimental results show that MADE outperforms the existing state-of-the-art TKGC models. Jiapu Wang, Boyue Wang, Junbin Gao, Shirui Pan, Tengfei Liu 0005, Wen Gao 0001 |
IEEE Trans. Cybern. | 5 |
| 2024 | Cross-modal Multiple Granularity Interactive Fusion Network for Long Document ClassificationabstractLong Document Classification (LDC) has attracted great attention in Natural Language Processing and achieved considerable progress owing to the large-scale pre-trained language models. In spite of this, as a different problem from the traditional text classification, LDC is far from being settled. Long documents, such as news and articles, generally have more than thousands of words with complex structures. Moreover, compared with flat text, long documents usually contain multi-modal content of images, which provide rich information but not yet being utilized for classification. In this article, we propose a novel cross-modal method for long document classification, in which multiple granularity feature shifting networks are proposed to integrate the multi-scale text and visual features of long documents adaptively. Additionally, a multi-modal collaborative pooling block is proposed to eliminate redundant fine-grained text features and simultaneously reduce the computational complexity. To verify the effectiveness of the proposed model, we conduct experiments on the Food101 dataset and two constructed multi-modal long document datasets. The experimental results show that the proposed cross-modal method outperforms the single-modal text methods and defeats the state-of-the-art related multi-modal baselines. Tengfei Liu 0005, Yongli Hu, Junbin Gao |
ACM Trans. Knowl. Discov. Data | 1 |
| 2023 | Zero-Shot Text Classification with Semantically Extended Textual EntailmentabstractZero-shot text classification (0SHOT-TC) aims to detect classes that the model never seen in the training set, and has attracted much attention in the research community of Natural Language Processing (NLP). The emergence of pre-trained language models has fostered the progress of 0SHOT-TC, which turns the task into a textual entailment problem of binary classification. It learns an entailment relatedness (yes/no) between the given sentence (premise) and each category (hypothesis) separately. However, the hypothesis generation paradigms need to be further studied, since the label itself or the label descriptions have limited ability to fully express the category space. Conversely, humans can easily extend a set of words describing the categories to be classified. In this paper, we propose a novel zero-shot text classification method called Semantically Extended Textual Entailment (SETE), which imitates the human's ability in knowledge extension. In the proposed method, three semantic extension methods are used to enrich the categories through a combination of static knowledge (e.g. expert knowledge, knowledge graph) and dynamic knowledge (e.g. language models), and the textual entailment model is finally used for 0SHOT-TC. The experimental results on the benchmarks show that our approach significantly outperforms the current methods in both generalized and non-generalized 0SHOT-TC. Tengfei Liu 0005, Yongli Hu, Puman Chen |
IJCNN | 1 |
| 2023 | Hierarchical Graph Convolutional Networks for Structured Long Document ClassificationabstractLong document classification (LDC) has been a focused interest in natural language processing (NLP) recently with the exponential increase of publications. Based on the pretrained language models, many LDC methods have been proposed and achieved considerable progression. However, most of the existing methods model long documents as sequences of text while omitting the document structure, thus limiting the capability of effectively representing long texts carrying structure information. To mitigate such limitation, we propose a novel hierarchical graph convolutional network (HGCN) for structured LDC in this article, in which a section graph network is proposed to model the macrostructure of a document and a word graph network with a decoupled graph convolutional block is designed to extract the fine-grained features of a document. In addition, an interaction strategy is proposed to integrate these two networks as a whole by propagating features between them. To verify the effectiveness of the proposed model, four structured long document datasets are constructed, and the extensive experiments conducted on these datasets and another unstructured dataset show that the proposed method outperforms the state-of-the-art related classification methods. Tengfei Liu 0005, Yongli Hu, Boyue Wang, Junbin Gao |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Hierarchical Multiple Granularity Attention Network for Long Document ClassificationabstractLong document classification has aroused tremendous attention in the field of Nature Language Processing, due to the exponential increasing of publications. Although the common text classification methods can be extended for long document classification, they are confined to the length of text and do not have enough expressiveness to model the structure of the long document. To solve these problems, we proposed a hierarchical multiple granularity attention network for long document classification, in which the word and section level features are extracted and fused to represent the complex structure of the long document. Furthermore, a feature-based section pooling module is adopted to eliminate redundant text information and accelerate the computing. A series of experiments are conducted to evaluate the proposed method. The experimental results verify that our method is effective, efficient and competitive compared with the related state-of-the-art methods. Yongli Hu, Tengfei Liu 0005, Junbin Gao |
IJCNN | 3 |
| 2021 | Hierarchical Attention Transformer Networks for Long Document ClassificationabstractProfiting from the pre-trained language representation models like BERT, the recently proposed document classification methods have obtained considerable improvement. However, most of these methods usually model the document as a sequence of text and omit the structure information, which appears obviously in long document composed of several sections with assigned relations. For this purpose, we propose a novel Hierarchical Attention Transformer Network (HATN) for long document classification, which extracts the structure of the long document by intra- and inter-section attention transformers, and further strengths the feature interaction by two fusion gates: the Residual Fusion Gate (RFG) and the Feature Fusion Gate (FFG). The proposed method is evaluated on three long document datasets and the experimental results show that our approach outperforms the related state-of-the-art methods. The code will be available at https://github.com/TengfeiLiu966/HATN Yongli Hu, Puman Chen, Tengfei Liu 0005, Junbin Gao |
IJCNN | 3 |
| 2020 | Zero-Shot Text Classification with Semantically Extended Graph Convolutional NetworkabstractAs a challenging task of Natural Language Processing(NLP), zero-shot text classification has attracted more and more attention recently. It aims to detect classes that the model has never seen in the training set. For this purpose, a feasible way is to construct connection between the seen and unseen classes by semantic extension and classify the unseen classes by information propagation over the connection. Although many related zero-shot text classification methods have been exploited, how to realize semantic extension properly and propagate information effectively are far from solved. In this paper, we propose a novel zero-shot text classification method called Semantically Extended Graph Convolutional Network (SEGCN). In the proposed method, the semantic category knowledge from ConceptNet is utilized to semantic extension for linking seen classes to unseen classes and constructing a graph of all categories. Then, we build upon Graph Convolutional Network (GCN) for predicting the textual classifier for each category, which transfers the category knowledge by the convolution operators on the constructed graph and is trained in a semi-supervised manner using the samples of the seen classes. The experimental results on Dbpedia and 20newsgroup datasets show that our method outperforms the state of the art zero-shot text classification methods. Tengfei Liu 0005, Yongli Hu, Junbin Gao |
ICPR | 1 |