VLDB 2026 Research / reviewers in the wild / expert
Jiale Han 0001
dblp:266/2826-1
· DBLP profile ↗
21ranked-venue papers
5as first author
15since 2021 · last 2027
0000-0001-6477-0424ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | SimDiff: Depth pruning via similarity and difference
Yuli Chen 0001, Shuhao Zhang 0011, Fanshen Meng, Bo Cheng 0001, Jiale Han 0001, Qiang Tong 0001, Xiulei Liu |
Expert Syst. Appl. | 5 |
| 2026 | Analyze-Compose-Execute: A Dynamic Dialogue Framework for Multi-Agent Debate
Wenyuan Gu, Jiale Han 0001, Xiang Li 0116, Zhixuan Wu, Hongru Xiao, Bo Cheng 0001 |
AAAI | 3 |
| 2026 | Infrared-assisted cross-modality detection for construction site worker safety monitoring
Hongru Xiao, Bin Yang 0029, Jinming Hu, Junze Zhu, Jiale Han 0001 |
Adv. Eng. Informatics | 6 |
| 2026 | Orientation-aware detection system for real-time monitoring of cracks in steel structures
Hongru Xiao, Bin Yang 0029, Jiale Han 0001, Zhen Lian, Songning Lai |
Expert Syst. Appl. | 4 |
| 2025 | Explain-Analyze-Generate: A Sequential Multi-Agent Collaboration Method for Complex ReasoningabstractExploring effective collaboration among multiple large language models (LLMs) represents an active research direction, with multiagent debate (MAD) emerging as a popular approach. MAD involves LLMs independently generating responses and refining their own responses by incorporating feedback from other agents in a debate manner. However,empirical experiments reveal the suboptimal performance of MAD in complex reasoning scenarios. We attribute this to the potential misleading caused by peer agents with limited individual capabilities. To address this, we propose a novel sequential collaboration framework named Explain-Analyze-Generate(EAG). By decomposing complex tasks into essential subtasks and employing a pipeline approach, EAG enable agents provide constructive assistance to peers, ultimately yielding higher performance. We conduct experiments on the comprehensive complex language reasoning benchmark: BIG-Bench-Hard (BBH). Our method achieves the highest performance on 19 out of 23 tasks, with an average improvement of 8% across all tasks, and incurs lower costs compared to MAD, demonstrating its effectiveness and efficiency. Wenyuan Gu, Jiale Han 0001, Xiang Li 0116, Bo Cheng 0001 |
COLING | 2 |
| 2025 | VideoQA-TA: Temporal-Aware Multi-Modal Video Question AnsweringabstractVideo question answering (VideoQA) has recently gained considerable attention in the field of computer vision, aiming to generate answers rely on both linguistic and visual reasoning. However, existing methods often align visual or textual features directly with large language models, which limits the deep semantic association between modalities and hinders a comprehensive understanding of the interactions within spatial and temporal contexts, ultimately leading to sub-optimal reasoning performance. To address this issue, we propose a novel temporal-aware framework for multi-modal video question answering, dubbed VideoQA-TA, which enhances reasoning ability and accuracy of VideoQA by aligning videos and questions at fine-grained levels. Specifically, an effective Spatial-Temporal Attention mechanism (STA) is designed for video aggregation, transforming video features into spatial and temporal representations while attending to information at different levels. Furthermore, a Temporal Object Injection strategy (TOI) is proposed to align object-level and frame-level information within videos, which further improves the accuracy by injecting explicit temporal information. Experimental results on MSVD-QA, MSRVTT-QA, and ActivityNet-QA datasets demonstrate the superior performance of our proposed method compared with the current SOTAs, meanwhile, visualization analysis further verifies the effectiveness of incorporating temporal information to videos. Zhixuan Wu, Bo Cheng 0001, Jiale Han 0001, Jiabao Ma, Shuhao Zhang 0011, Yuli Chen 0001, Changbo Li |
COLING | 3 |
| 2025 | CEFW: A Comprehensive Evaluation Framework for Watermark in Large Language ModelsabstractText watermarking provides an effective solution for identifying synthetic text generated by large language models. However, existing techniques often focus on satisfying specific criteria while ignoring other key aspects, lacking a unified evaluation. To fill this gap, we propose the Comprehensive Evaluation Framework for Watermark (CEFW), a unified framework that comprehensively evaluates watermarking methods across five key dimensions: ease of detection, fidelity of text quality, minimal embedding cost, robustness to adversarial attacks, and imperceptibility to prevent imitation or forgery. By assessing watermarks according to all these key criteria, CEFW offers a thorough evaluation of their practicality and effectiveness. Moreover, we introduce a simple and effective watermarking method called Balanced Watermark (BW), which guarantees robustness and imperceptibility through balancing the way watermark information is added. Extensive experiments show that BW outperforms existing methods in overall performance across all evaluation dimensions. We release our code to the community for future research1. Shuhao Zhang 0011, Bo Cheng 0001, Jiale Han 0001, Yuli Chen 0001, Zhixuan Wu, Changbao Li, Pingli Gu |
ICME | 3 |
| 2025 | DLP: Dynamic Layerwise Pruning in Large Language ModelsabstractPruning has recently been widely adopted to reduce the parameter scale and improve the inference efficiency of Large Language Models (LLMs). Mainstream pruning techniques often rely on uniform layerwise pruning strategies, which can lead to severe performance degradation at high sparsity levels. Recognizing the varying contributions of different layers in LLMs, recent studies have shifted their focus toward non-uniform layerwise pruning. However, these approaches often rely on pre-defined values, which can result in suboptimal performance. To overcome these limitations, we propose a novel method called Dynamic Layerwise Pruning (DLP). This approach adaptively determines the relative importance of each layer by integrating model weights with input activation information, assigning pruning rates accordingly. Experimental results show that DLP effectively preserves model performance at high sparsity levels across multiple LLMs. Specifically, at 70% sparsity, DLP reduces the perplexity of LLaMA2-7B by 7.79 and improves the average accuracy by 2.7% compared to state-of-the-art methods. Moreover, DLP is compatible with various existing LLM compression techniques and can be seamlessly integrated into Parameter-Efficient Fine-Tuning (PEFT). We release the code\footnote{The code is available at: \url{https://github.com/ironartisan/DLP}.} to facilitate future research. Yuli Chen 0001, Bo Cheng 0001, Jiale Han 0001, Yingting Li, Shuhao Zhang 0011 |
ICML | 3 |
| 2025 | Can Audio Language Models Listen Between the Lines? A Study on Metaphorical Reasoning via UnspokenabstractRecent advancements in Audio Language Models (ALMs) have led to significant improvements in speech-related tasks. However, their capacity for profound metaphorical reasoning, especially when derived from audio-specific cues, has yet to be thoroughly investigated. To address this gap, we introduce Unspoken, a bilingual (Chinese-English) question answering benchmark designed to assess ALMs' comprehension of non-literal, metaphor-rich audio. Unlike prior text-centric evaluations, Unspoken emphasizes prosody, phonetic ambiguity, emotional inflection, and other nuanced acoustic features critical to metaphor understanding but often lost in transcription. We construct a high-quality dataset of 2,764 manually curated and validated QA pairs, spanning three reasoning dimensions: semantic, acoustic, and contextual, and covering six common types of metaphors. Evaluation across 23 mainstream ALMs reveals a substantial performance gap: the best model achieves only 69.5% accuracy, significantly below the human average of 81.1%. By analyzing the error patterns, we identify five key failure modes that reveal fundamental limitations in current models' reasoning capabilities. Unspoken not only sets a new standard for evaluating metaphorical reasoning in audio but also pioneers a novel research direction that moves beyond transcription-based assessments. Grounding metaphor understanding in authentic human communication scenarios offers deep insight for developing more cognitively capable ALMs. The data and codes are available at https://github.com/Hongru0306/UNSPOKEN. Hongru Xiao, Xiang Li 0064, Duyi Pan, ZhixueSong ZhixueSong, Jiale Han 0001, Songning Lai, Wenshuo Chen, Benyou Wang |
ACM Multimedia | 6 |
| 2025 | Generative knowledge-guided review system for construction disclosure documents
Hongru Xiao, Jiankun Zhuang, Bin Yang 0029, Jiale Han 0001, Songning Lai |
Adv. Eng. Informatics | 4 |
| 2024 | Making Pre-trained Language Models Better Continual Few-Shot Relation ExtractorsabstractContinual Few-shot Relation Extraction (CFRE) is a practical problem that requires the model to continuously learn novel relations while avoiding forgetting old ones with few labeled training data. The primary challenges are catastrophic forgetting and overfitting. This paper harnesses prompt learning to explore the implicit capabilities of pre-trained language models to address the above two challenges, thereby making language models better continual few-shot relation extractors. Specifically, we propose a Contrastive Prompt Learning framework, which designs prompt representation to acquire more generalized knowledge that can be easily adapted to old and new categories, and margin-based contrastive learning to focus more on hard samples, therefore alleviating catastrophic forgetting and overfitting issues. To further remedy overfitting in low-resource scenarios, we introduce an effective memory augmentation strategy that employs well-crafted prompts to guide ChatGPT in generating diverse samples. Extensive experiments demonstrate that our method outperforms state-of-the-art methods by a large margin and significantly mitigates catastrophic forgetting and overfitting in low-resource scenarios. Shengkun Ma, Jiale Han 0001, Bo Cheng 0001 |
LREC/COLING | 2 |
| 2023 | Towards Hard Few-Shot Relation ClassificationabstractFew-shot relation classification (FSRC) focuses on recognizing novel relations by learning with merely a handful of annotated instances. Meta-learning has been widely adopted for such a task, which trains on randomly generated few-shot tasks to learn generic data representations. Despite impressive results achieved, existing models still perform suboptimally when handling hard FSRC tasks with similar categories that confuse the model to distinguish correctly. We argue this is largely due to two reasons, 1) ignoring pivotal and discriminate information that is crucial to distinguish confusing classes, and 2) training indiscriminately via randomly sampled tasks of varying difficulty. In this article, we introduce a novel prototypical network approach with contrastive learning that learns more informative and discriminative representations by exploiting relation label information. We further design two strategies that increase the difficulty of training tasks and allow the model to adaptively learn to focus on hard tasks. By doing so, our model can better represent subtle inter-relation variance and grow up through task difficulty. Extensive experiments on three standard benchmarks demonstrate the effectiveness of our method. Jiale Han 0001, Bo Cheng 0001, Zhiguo Wan, Wei Lu 0011 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2021 | Learning Discriminative and Unbiased Representations for Few-Shot Relation ExtractionabstractFew-shot relation extraction (FSRE) aims to predict the relation for a pair of entities in a sentence by exploring a few labeled instances for each relation type. Current methods mainly rely on meta-learning to learn generalized representations by optimizing the network parameters based on various collections of tasks sampled from training data. However, these methods may suffer from two main issues. 1) Insufficient supervision of meta-learning to learn discriminative representations on very few training instances, which are sampled from a large amount of base class data. 2) Spurious correlations between entities and relation types due to the biased training procedure that focuses more on entity pair rather than context. To learn more discriminative and unbiased representations for FSRE, this paper proposes a two-stage approach via supervised contrastive learning and sentence- and entity-level prototypical networks. In the first (pre-training) stage, we introduce a supervised contrastive pre-training method, which is able to yield more discriminative representations by learning from the entire training instances, such that the semantically related representations are close to each other, and far away otherwise. In the second (meta-learning) stage, we propose a novel sentence- and entity-level prototypical network equipped with fine-grained feature-wise fusion strategy to learn unbiased representations, where the networks are initialized with the parameters trained in the first stage. Specifically, the proposed network consists of a sentence branch and an entity branch, taking entire sentences and entity mentions as inputs, respectively. The entity branch explicitly captures the correlation between entity pairs and relations, and then dynamically adjusts the sentence branch's prediction distributions. By doing so, the spurious correlations issue caused by biased training samples can be properly mitigated. Extensive experiments on two FSRE benchmarks demonstrate the effectiveness of our approach. Jiale Han 0001, Bo Cheng 0001, Guoshun Nan |
CIKM | 1 |
| 2021 | Exploring Task Difficulty for Few-Shot Relation ExtractionabstractFew-shot relation extraction (FSRE) focuses on recognizing novel relations by learning with merely a handful of annotated instances.Meta-learning has been widely adopted for such a task, which trains on randomly generated few-shot tasks to learn generic data representations.Despite impressive results achieved, existing models still perform suboptimally when handling hard FSRE tasks, where the relations are fine-grained and similar to each other.We argue this is largely because existing models do not distinguish hard tasks from easy ones in the learning process.In this paper, we introduce a novel approach based on contrastive learning that learns better representations by exploiting relation label information.We further design a method that allows the model to adaptively learn how to focus on hard tasks.Experiments on two standard datasets demonstrate the effectiveness of our method. Jiale Han 0001, Bo Cheng 0001, Wei Lu 0011 |
EMNLP (1) | 1 |
| 2021 | Integrating Subgraph-Aware Relation and Direction Reasoning for Question AnsweringabstractQuestion Answering (QA) models over Knowledge Bases (KBs) are capable of providing more precise answers by utilizing relation information among entities. Although effective, most of these models solely rely on fixed relation representations to obtain answers for different question-related KB subgraphs. Hence, the rich structured information of these subgraphs may be overlooked by the relation representation vectors. Meanwhile, the direction information of reasoning, which has been proven effective for the answer prediction on graphs, has not been fully explored in existing work. To address these challenges, we propose a novel neural model, Relation-updated Direction-guided Answer Selector (RDAS), which converts relations in each subgraph to additional nodes to learn structure information. Additionally, we utilize direction information to enhance the reasoning ability. Experimental results show that our model yields substantial improvements on two widely used datasets. Shuai Zhao 0001, Bo Cheng 0001, Jiale Han 0001, Yingting Li, Hao Yang 0006, Ivan Sekulic, Guoshun Nan |
ICASSP | 4 |
| 2020 | Hypergraph Convolutional Network for Multi-Hop Knowledge Base Question Answering (Student Abstract)abstractGraph convolutional networks (GCN) have been applied in knowledge base question answering (KBQA) task. However, the pairwise connection between nodes of GCN limits the representation capability of high-order data correlation. Furthermore, most previous work does not fully utilize the semantic relation information, which is vital to reasoning. In this paper, we propose a novel multi-hop KBQA model based on hypergraph convolutional network. By constructing a hypergraph, the form of pairwise connection between nodes and nodes is converted to the high-level connection between nodes and edges, which effectively encodes complex related data. To better exploit the semantic information of relations, we apply co-attention method to learn similarity between relation and query, and assign weights to different relations. Experimental results demonstrate the effectivity of the model. Jiale Han 0001, Bo Cheng 0001 |
AAAI | 1 |
| 2020 | HGMAN: Multi-Hop and Multi-Answer Question Answering Based on Heterogeneous Knowledge Graph (Student Abstract)abstractMulti-hop question answering models based on knowledge graph have been extensively studied. Most existing models predict a single answer with the highest probability by ranking candidate answers. However, they are stuck in predicting all the right answers caused by the ranking method. In this paper, we propose a novel model that converts the ranking of candidate answers into individual predictions for each candidate, named heterogeneous knowledge graph based multi-hop and multi-answer model (HGMAN). HGMAN is capable of capturing more informative representations for relations assisted by our heterogeneous graph, which consists of multiple entity nodes and relation nodes. We rely on graph convolutional network for multi-hop reasoning and then binary classification for each node to get multiple answers. Experimental results on MetaQA dataset show the performance of our proposed model over all baselines. Shuai Zhao 0001, Bo Cheng 0001, Jiale Han 0001, Yingting Li, Hao Yang 0006, Guoshun Nan |
AAAI | 4 |
| 2020 | Modelling Long-distance Node Relations for KBQA with Global Dynamic GraphabstractThe structural information of Knowledge Bases (KBs) has proven effective to Question Answering (QA).Previous studies rely on deep graph neural networks (GNNs) to capture rich structural information, which may not model node relations in particularly long distance due to oversmoothing issue.To address this challenge, we propose a novel framework GlobalGraph, which models long-distance node relations from two views: 1) Node type similarity: GlobalGraph assigns each node a global type label and models long-distance node relations through the global type label similarity; 2) Correlation between nodes and questions: we learn similarity scores between nodes and the question, and model long-distance node relations through the sum score of two nodes.We conduct extensive experiments on two widely used multi-hop KBQA datasets to prove the effectiveness of our method. Shuai Zhao 0001, Jiale Han 0001, Bo Cheng 0001, Hao Yang 0006, Jianchang Ao, Zhenzi Li |
COLING | 3 |
| 2020 | Two-Phase Hypergraph Based Reasoning with Dynamic Relations for Multi-Hop KBQAabstractMulti-hop knowledge base question answering (KBQA) aims at finding the answers to a factoid question by reasoning across multiple triples. Note that when human performs multi-hop reasoning, one tends to concentrate on specific relation at different hops and pinpoint a group of entities connected by the relation. Hypergraph convolutional networks (HGCN) can simulate this behavior by leveraging hyperedges to connect more than two nodes more than pairwise connection. However, HGCN is for undirected graphs and does not consider the direction of information transmission. We introduce the directed-HGCN (DHGCN) to adapt to the knowledge graph with directionality. Inspired by human's hop-by-hop reasoning, we propose an interpretable KBQA model based on DHGCN, namely two-phase hypergraph based reasoning with dynamic relations, which explicitly updates relation information and dynamically pays attention to different relations at different hops. Moreover, the model predicts relations hop-by-hop to generate an intermediate relation path. We conduct extensive experiments on two widely used multi-hop KBQA datasets to prove the effectiveness of our model. Jiale Han 0001, Bo Cheng 0001 |
IJCAI | 1 |
| 2020 | DVKCM: Knowledge-guided Conversation Generation with Dynamic VocabularyabstractKnowledge-guided conversation models, whose inputs are current input sentence with its background knowledge, make the generation of responses more informative and meaningful. Existing methods assume that words in responses come from the vocabulary of the whole corpus. However, for specific input and knowledge, only a small vocabulary is useful in prediction and other words lead to uncorrelated noise. In this paper, we propose a Dynamic Vocabulary based Knowledge-guided Conversation Model (DVKCM). Inspired by dynamic vocabulary mechanism, DVKCM adopts the vocabulary construction module to allocate the sentence-level vocabulary which relates to the input sentence and background knowledge, and then only uses the small vocabulary to execute the decoding part. Through the sentence-level vocabulary mechanism, we reduce the generation of noise effectively. Experiments on both automatic and human evaluation verify the performance of our model compared with previous models. Moreover, we find that dynamic vocabulary can be applied to other conversation models to improve their performance. Shuai Zhao 0001, Bo Cheng 0001, Jiale Han 0001, Xiangsheng Wei, Hao Yang 0006 |
IJCNN | 4 |
| 2020 | Multiview Clustering by Joint Latent Representation and Similarity LearningabstractSubspace learning-based multiview clustering has achieved impressive experimental results. However, the similarity matrix, which is learned by most existing methods, cannot well characterize both the intrinsic geometric structure of data and the neighbor relationship between data. To consider the fact that original data space does not well characterize the intrinsic geometric structure, we learn the latent representation of data, which is shared by different views, from the latent subspace rather than the original data space by linear transformation. Thus, the learned latent representation has a low-rank structure without solving the nuclear-norm. This reduces the computational complexity. Then, the similarity matrix is adaptively learned from the learned latent representation by manifold learning which well characterizes the local intrinsic geometric structure and neighbor relationship between data. Finally, we integrate clustering, manifold learning, and latent representation into a unified framework and develop a novel subspace learning-based multiview clustering method. Extensive experiments on benchmark datasets demonstrate the superiority of our method. De-Yan Xie, Quanxue Gao, Jiale Han 0001, Xinbo Gao 0001 |
IEEE Trans. Cybern. | 4 |