VLDB 2026 Research / reviewers in the wild / expert
Yutai Duan
dblp:304/4957
· DBLP profile ↗
9ranked-venue papers
5as first author
9since 2021 · last 2025
0009-0002-6564-3253ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Community-Aware Graph Transformer: Preserving Community Semantics for Effective Global Aggregation
Yutai Duan |
ECML/PKDD (3) | 1 |
| 2025 | KG-prompt: Interpretable knowledge graph prompt for pre-trained language modelsabstractKnowledge graphs (KGs) can provide rich factual knowledge for language models , enhancing reasoning ability and interpretability . However, existing knowledge injection methods usually ignore the structured information in KGs. Using structured knowledge to enhance pre-trained language models (PLMs) still has a set of challenging issues, including resource consumption of knowledge retraining, heterogeneous information, and knowledge noise. To address these issues, we explore how to flexibly inject structured knowledge into frozen PLMs. Inspired by prompt learning, we propose a novel method K nowledge G raph Prompt (KG-Prompt), which for the first time encodes the KG as structured prompts to enhance the knowledge expression ability of PLMs. KG-Prompt consists of a compressed subgraph construction module and a KG prompt generation module. In the compressed subgraph construction module, we construct compressed subgraphs based on a path-weighting strategy to reduce knowledge noise. In the KG prompt generation module, we propose a multi-hop consistency optimization strategy to learn the representation of compressed subgraphs, and then generate KG prompts based on a knowledge mapper to solve the heterogeneous information problem. The KG prompts can be inserted into the input of PLMs expediently, which decouples from PLMs and the downstream model without knowledge retraining and reduces computational resources . Extensive experiments on three knowledge-driven natural language understanding tasks demonstrate that our approach effectively improves the knowledge reasoning ability of PLMs. Furthermore, we provide a detailed analysis of different KG prompts and discuss the interpretability and generalizability of the proposed method. Liyi Chen 0003, Jie Liu 0007, Yutai Duan |
Knowl. Based Syst. | 3 |
| 2025 | Learning global dependencies via parallelized graph transformer with hybrid attention
Yutai Duan, Jie Liu 0007, Xingyang He |
Knowl. Based Syst. | 1 |
| 2025 | HDANet: Enhancing Underwater Salient Object Detection With Physics-Inspired Multimodal Joint LearningabstractUnderwater salient object detection (USOD) poses a significantly greater challenge than traditional terrestrial scenes, due to both the complex image degradation and the absence of multimodal information in underwater environments. Existing image enhancement methods are not specifically optimized for USOD, while current USOD approaches rarely consider effective extraction and utilization of multimodal information, leading to limited performance. This paper proposes HydroDepthAwareNet (HDANet), which addresses these challenges through developing targeted designs to enhance USOD performance. It first integrates a task-driven underwater image enhancement module, named HydroDepthEnhanceModule (HDEM), which is based on physical models to provide enhanced images and multimodal information optimized for USOD tasks. Furthermore, we develop a physics-inspired three-way unsupervised learning strategy, leveraging the complementary effects of re-enhancement and re-degradation to improve HDEM’s generalization across diverse underwater image degradation scenarios. Additionally, we design a robust cross-attention (RCA) module to effectively fuse multimodal features while mitigating noise and blurring by exploiting channel and spatial cross-attention mechanisms. Extensive experiments on various USOD datasets demonstrate that the proposed HDANet significantly outperforms existing state-of-the-art methods. The source code will be made available at https://github.com/mikurules/USOD-HDANet. Jinchao Zhu, Biting Ma, Yutai Duan, Panlong Tan |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | 2-D Transformer: Extending Large Language Models to Long-Context With Few MemoryabstractThe ability of processing long contexts is crucial for large language models (LLMs), but training LLMs with a long-context window requires substantial computational resources. Many sought to mitigate this through the sparse attention mechanism. However, sparse attention faces a noticeable gap compared with full attention in capturing long-distance information, leading to limited long-context processing capabilities. To effectively address this issue, this article proposes a novel sparse transformer architecture called 2-D transformer (2D-former), aimed at extending the context windows of pretrained LLMs while reducing GPU memory requirements. The 2D-former incorporates a 2-D attention mechanism that consists of a long-distance information compressor (LDIC) and a blockwise attention (BA) mechanism. LDIC can self-adaptively extract blockwise representational features by convolution and compress long-distance information into a set of tokens based on the significance of each block. The BA mechanism integrates these features, enabling each token to directly communicate with any of its preceding tokens during the computation of sparse attention. In this way, sparse attention can fully utilize long-distance information to bridge the gap with full attention while greatly reducing computational requirements. The 2D-former only needs to add less than 0.14% of additional trainable parameters to extend the context length of LLaMA2 7B to 32k on 4 A100 GPUs with 40-GB memory. In addition, it is compatible with most current acceleration techniques and parameter-efficient fine-tuning (PEFT) methods. Furthermore, we conduct supervised fine-tuning with 2D-former using our self-collected long-instruction fine-tuning dataset, named LongTuning, which comprises over 11k long-context question-answer (QA) pairs. Experimental results demonstrate that 2D-former achieves efficient long-context extension with minimal GPU memory and computational time consumption, while maintaining superior performance across both downstream long-context and short-context tasks. Xingyang He, Jie Liu 0007, Yutai Duan |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | G-Prompt: Graphon-based Prompt Tuning for graph classification
Yutai Duan, Jie Liu 0007, Shaowei Chen, Liyi Chen 0003 |
Inf. Process. Manag. | 1 |
| 2024 | Contrastive fine-tuning for low-resource graph-level transfer learning
Yutai Duan, Jie Liu 0007, Shaowei Chen |
Inf. Sci. | 1 |
| 2024 | High-frequency and low-frequency dual-channel graph attention network
Yukuan Sun, Yutai Duan |
Pattern Recognit. | 2 |
| 2021 | Learning Key Actors and Their Interactions for Group Activity Recognition
Yutai Duan |
PRCV (4) | 1 |