VLDB 2026 Research / reviewers in the wild / expert
Linhao Zhang
dblp:204/8272
· DBLP profile ↗
16ranked-venue papers
6as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 5 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SADA: Bridging In-Context Learning and Fine-Tuning via State-Aligned Distillation AdaptersabstractPrompt-based in-context learning (ICL) and parameter fine-tuning are two dominant paradigms for incorporating external information into large language models (LLMs), but they incur high inference costs or require expensive retraining.To bridge this gap, context-to-parameter mapping converts prompts into temporary adapter weights.However, we identify a critical failure mode in existing methods: hiddenstate collapse, where the adapter-augmented model's internal states diverge sharply from the full-context oracle in deeper layers.We trace this failure to two coupled gaps: suboptimal Input-Selection and inadequate Supervision-Signal.To address these issues, we propose SADA (State-Aligned Distillation Adapters).We establish the attention-block output as a principled feature interface to improve input selection and introduce statealignment distillation to enforce consistency between the adapter-augmented model and the full-context oracle.Experiments on long-context language modeling (PG19) and downstream NLU and summarization benchmarks show that SADA consistently outperforms strong baselines like StreamAdapter and GenerativeAdapter, achieving performance comparable to ICL while significantly reducing memory footprint and latency.We further analyze when parameterized context compression is effective and when explicit context retention remains preferable. Tianlong Wang, Linhao Zhang, Aiwei Liu, Xiao Zhou 0004 |
ACL (1) | 4 |
| 2026 | Multi-grained dynamic feature resorter with curriculum bootstrapping for multimodal entity and relation extraction
Zhicong Lu, Li Jin 0001, Linhao Zhang, Kaiwen Wei, Qing Liu 0021, Yanfeng Hu |
Neurocomputing | 4 |
| 2025 | U-MERE: Unconstrained Multimodal Entity and Relation Extraction with Collaborative Modeling and Order-Sensitive OptimizationabstractExisting multimodal entity and relation extraction tasks primarily focus on text-to-text or text-to-visual entity relations, overlooking real-world complexities involving visual-to-text and visual-to-visual cases, thus failing to capture the richer semantic structures in complex cross-modal interactions. To address the limitations, we propose a new task, Unconstrained Multimodal Entity and Relation Extraction (U-MERE), which jointly extracts arbitrary visual and textual entities, and their relations from image-text pairs. To accomplish U-MERE, we construct UMERE-Bench, a benchmark with over 9,000 samples that comprehensively covers four cross-modal entity relation directions and three task settings. Given the difficulty of jointly modeling diverse directions of cross-modal entity relations, we introduce Collaborative Modeling and Order-Sensitive (CMOS), which collaboratively guides large vision-language models (LVLMs) to decompose task complexity and mitigates generation order bias from fixed target relation sequences. CMOS employs small models to generate candidate entities, guiding LVLMs to capture key information and jointly optimizes multiple feasible relation orderings to reduce order dependency. Additionally, we design a Multimodal Order-aware Matching (MOM) evaluation method to align predictions with ground truth for precise assessment. Experimental results reveal that current LVLMs show limited performance on U-MERE, underscoring its inherent challenges, while CMOS consistently achieves superior performance across multiple advanced LVLMs, demonstrating its effectiveness and generalization capability. The dataset and code will be available in https://github.com/jiaweidoris/U-MERE. Li Jin 0001, Kaiwen Wei, Yuying Shang, Nayu Liu, Zhicong Lu, Qing Liu 0021, Linhao Zhang, Yanfeng Hu |
ACM Multimedia | 8 |
| 2025 | Multi-SWE-bench: A Multilingual Benchmark for Issue ResolvingabstractThe task of issue resolving aims to modify a codebase to generate a patch that addresses a given issue. However, most existing benchmarks focus almost exclusively on Python, making them insufficient for evaluating Large Language Models (LLMs) across different programming languages. To bridge this gap, we introduce a multilingual issue-resolving benchmark, called Multi-SWE-bench, covering 8 languages of Python, Java, TypeScript, JavaScript, Go, Rust, C, and C++. In particular, this benchmark includes a total of 2,132 high-quality instances, carefully curated by 68 expert annotators, ensuring a reliable and accurate evaluation of LLMs on the issue-resolving task. Based on human-annotated results, the issues are further classified into three difficulty levels. We evaluate a series of state-of-the-art models on Multi-SWE-bench, utilizing both procedural and agent-based frameworks for issue resolving. Our experiments reveal three key findings: (1) Limited generalization across languages: While existing LLMs perform well on Python issues, their ability to generalize across other languages remains limited; (2) Performance aligned with human-annotated difficulty: LLM-based agents' performance closely aligns with human-assigned difficulty, with resolution rates decreasing as issue complexity rises; and (3) Performance drop on cross-file issues: The performance of current methods significantly deteriorates when handling cross-file issues. These findings highlight the limitations of current LLMs and underscore the need for more robust models capable of handling a broader range of programming languages and complex issue scenarios. Daoguang Zan, Zhirong Huang, Hanwu Chen, Shulin Xin, Linhao Zhang, Aoyan Li, Xiaojian Zhong, Yongsheng Xiao, Liangqiang Chen, Yuyu Zhang, Rui Long |
NeurIPS | 6 |
| 2025 | Flexible Optimal Transport With Contrastive Graphical Modeling for Multimodal Hate DetectionabstractMultimodal hate detection plays a crucial role in maintaining harmonious online environments by identifying harmful content, such as hateful memes. Although previous research has made significant progress in detecting explicit hate speech, there remains a critical gap in analyzing implicit hate, which is particularly challenging due to the absence of explicit harmful text claims or demographic visual cues. Despite the promising results based on cross-modal attention, previous methods may suffer from the distributional modality gap caused by the non-literal associations between multimodal elements, which lacks apparent alignment in implicit hateful contents. In this work, we propose a novel framework: Flexible Optimal Transport (FLOT) to capture the non-literal cross-modal alignment for multimodal hate in the context of memes. FLOT formulates the problem of cross-modal alignment as finding optimal transportation plans, which leverages a kernel method to capture complementary information from multiple modalities. The kernel embeddings reproduce a kernel Hilbert space (RKHS) to serve as a non-linear transformation of alignment, which effectively reduces the distributional modality gap with more interpretability. Moreover, we established topological structures with contrastive modeling for the aligned representations, which are optimized to achieve comprehensive alignment between different modalities, and facilitate local reasoning based on multimodal elements. Experimental results have demonstrated that our FLOT achieved state-of-the-art performance on three publicly available benchmark datasets. Furthermore, extensive qualitative analysis confirms the superior ability of FLOT in capturing implicit cross-modal alignment. Linhao Zhang, Li Jin 0001, Xiaoyu Li 0004, Xian Sun 0001, Xin Wang 0117, Zequn Zhang, Jian Liu 0032, Zhicong Lu, Guangluan Xu |
IEEE Trans. Multim. | 1 |
| 2024 | Video Event Extraction with Multi-View Interaction Knowledge DistillationabstractVideo event extraction (VEE) aims to extract key events and generate the event arguments for their semantic roles from the video. Despite promising results have been achieved by existing methods, they still lack an elaborate learning strategy to adequately consider: (1) inter-object interaction, which reflects the relation between objects; (2) inter-modality interaction, which aligns the features from text and video modality. In this paper, we propose a Multi-view Interaction with knowledge Distillation (MID) framework to solve the above problems with the Knowledge Distillation (KD) mechanism. Specifically, we propose the self-Relational KD (self-RKD) to enhance the inter-object interaction, where the relation between objects is measured by distance metric, and the high-level relational knowledge from the deeper layer is taken as the guidance for boosting the shallow layer in the video encoder. Meanwhile, to improve the inter-modality interaction, the Layer-to-layer KD (LKD) is proposed, which integrates additional cross-modal supervisions (i.e., the results of cross-attention) with the textual supervising signal for training each transformer decoder layer. Extensive experiments show that without any additional parameters, MID achieves the state-of-the-art performance compared to other strong methods in VEE. Kaiwen Wei, Runyan Du, Li Jin 0001, Jian Liu 0032, Jianhua Yin 0001, Linhao Zhang, Nayu Liu, Zhi Guo |
AAAI | 6 |
| 2024 | CAMEL: Capturing Metaphorical Alignment with Context Disentangling for Multimodal Emotion RecognitionabstractUnderstanding the emotional polarity of multimodal content with metaphorical characteristics, such as memes, poses a significant challenge in Multimodal Emotion Recognition (MER). Previous MER researches have overlooked the phenomenon of metaphorical alignment in multimedia content, which involves non-literal associations between concepts to convey implicit emotional tones. Metaphor-agnostic MER methods may be misinformed by the isolated unimodal emotions, which are distinct from the real emotions blended in multimodal metaphors. Moreover, contextual semantics can further affect the emotions associated with similar metaphors, leading to the challenge of maintaining contextual compatibility. To address the issue of metaphorical alignment in MER, we propose to leverage a conditional generative approach for capturing metaphorical analogies. Our approach formulates schematic prompts and corresponding references based on theoretical foundations, which allows the model to better grasp metaphorical nuances. In order to maintain contextual sensitivity, we incorporate a disentangled contrastive matching mechanism, which undergoes curricular adjustment to regulate its intensity during the learning process. The automatic and human evaluation experiments on two benchmarks prove that, our model provides considerable and stable improvements in recognizing multimodal emotion with metaphor attributes. Linhao Zhang, Li Jin 0001, Guangluan Xu, Xiaoyu Li 0004, Kaiwen Wei, Nayu Liu |
AAAI | 1 |
| 2024 | Rethinking the Reversal Curse of LLMs: a Prescription from Human Knowledge ReversalabstractZhicong Lu, Li Jin, Peiguang Li, Yu Tian, Linhao Zhang, Sirui Wang, Guangluan Xu, Changyuan Tian, Xunliang Cai. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Zhicong Lu, Li Jin 0001, Peiguang Li, Linhao Zhang, Guangluan Xu, Changyuan Tian 0001 |
EMNLP | 5 |
| 2024 | GOME: Grounding-based Metaphor Binding With Conceptual Elaboration For Figurative Language IllustrationabstractThe illustration or visualization of figurative language, such as linguistic metaphors, is an emerging challenge for existing Large Language Models (LLMs) and multimodal models. Linhao Zhang, Li Jin 0001, Kaiwen Wei, Guangluan Xu |
EMNLP | 1 |
| 2024 | DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads FusionabstractLarge language models (LLMs) with billions of parameters demonstrate impressive performance. However, the widely used Multi-Head Attention (MHA) in LLMs incurs substantial computational and memory costs during inference. While some efforts have optimized attention mechanisms by pruning heads or sharing parameters among heads, these methods often lead to performance degradation or necessitate substantial continued pre-training costs to restore performance. Based on the analysis of attention redundancy, we design a Decoupled-Head Attention (DHA) mechanism. DHA adaptively configures group sharing for key heads and value heads across various layers, achieving a better balance between performance and efficiency. Inspired by the observation of clustering similar heads, we propose to progressively transform the MHA checkpoint into the DHA model through linear fusion of similar head parameters step by step, retaining the parametric knowledge of the MHA checkpoint. We construct DHA models by transforming various scales of MHA checkpoints given target head budgets. Our experiments show that DHA remarkably requires a mere 0.25\% of the original model's pre-training budgets to achieve 96.1\% of performance while saving 75\% of KV cache. Compared to Group-Query Attention (GQA), DHA achieves a 5$\times$ training acceleration, a maximum of 13.93\% performance improvement under 0.01\% pre-training budget, and 5\% relative improvement under 0.05\% pre-training budget. Linhao Zhang, Junyuan Shang, Zhenyu Zhang 0006, Tingwen Liu, Shuohuan Wang |
NeurIPS | 2 |
| 2023 | TOT:Topology-Aware Optimal Transport for Multimodal Hate DetectionabstractMultimodal hate detection, which aims to identify the harmful content online such as memes, is crucial for building a wholesome internet environment. Previous work has made enlightening exploration in detecting explicit hate remarks. However, most of their approaches neglect the analysis of implicit harm, which is particularly challenging as explicit text markers and demographic visual cues are often twisted or missing. The leveraged cross-modal attention mechanisms also suffer from the distributional modality gap and lack logical interpretability. To address these semantic gap issues, we propose TOT: a topology-aware optimal transport framework to decipher the implicit harm in memes scenario, which formulates the cross-modal aligning problem as solutions for optimal transportation plans. Specifically, we leverage an optimal transport kernel method to capture complementary information from multiple modalities. The kernel embedding provides a non-linear transformation ability to reproduce a kernel Hilbert space (RKHS), which reflects significance for eliminating the distributional modality gap. Moreover, we perceive the topology information based on aligned representations to conduct bipartite graph path reasoning. The newly achieved state-of-the-art performance on two publicly available benchmark datasets, together with further visual analysis, demonstrate the superiority of TOT in capturing implicit cross-modal alignment. Linhao Zhang, Li Jin 0001, Xian Sun 0001, Guangluan Xu, Zequn Zhang, Xiaoyu Li 0004, Nayu Liu, Qing Liu 0021, Shiyao Yan |
AAAI | 1 |
| 2023 | Guide the Many-to-One Assignment: Open Information Extraction via IoU-aware Optimal TransportabstractKaiwen Wei, Yiran Yang, Li Jin, Xian Sun, Zequn Zhang, Jingyuan Zhang, Xiao Li, Linhao Zhang, Jintao Liu, Guo Zhi. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Kaiwen Wei, Li Jin 0001, Xian Sun 0001, Zequn Zhang, Linhao Zhang, Zhi Guo |
ACL (1) | 8 |
| 2020 | Graph LSTM with Context-Gated Mechanism for Spoken Language UnderstandingabstractMuch research in recent years has focused on spoken language understanding (SLU), which usually involves two tasks: intent detection and slot filling. Since Yao et al.(2013), almost all SLU systems are RNN-based, which have been shown to suffer various limitations due to their sequential nature. In this paper, we propose to tackle this task with Graph LSTM, which first converts text into a graph and then utilizes the message passing mechanism to learn the node representation. Not only the Graph LSTM addresses the limitations of sequential models, but it can also help to utilize the semantic correlation between slot and intent. We further propose a context-gated mechanism to make better use of context information for slot filling. Our extensive evaluation shows that the proposed model outperforms the state-of-the-art results by a large margin. Linhao Zhang, Dehong Ma, Xiaodong Zhang 0022, Houfeng Wang |
AAAI | 1 |
| 2020 | Syntax-Aware Graph Attention Network for Aspect-Level Sentiment ClassificationabstractAspect-level sentiment classification aims to distinguish the sentiment polarities over aspect terms in a sentence.Existing approaches mostly focus on modeling the relationship between the given aspect words and their contexts with attention, and ignore the use of more elaborate knowledge implicit in the context.In this paper, we exploit syntactic awareness to the model by the graph attention network on the dependency tree structure and external pre-training knowledge by BERT language model, which helps to model the interaction between the context and aspect words better.And the subwords of BERT are integrated into the dependency tree graphs, which can obtain more accurate representations of words by graph attention.Experiments demonstrate the effectiveness of our model. Lianzhe Huang, Xin Sun 0013, Sujian Li, Linhao Zhang, Houfeng Wang |
COLING | 4 |
| 2019 | Using Bidirectional Transformer-CRF for Spoken Language Understanding
Linhao Zhang, Houfeng Wang |
NLPCC (1) | 1 |
| 2018 | The Research of Mobile Location Privacy Protection Access Control Method Based on Game TheoryabstractIn recent years, the Internet of things has developed rapidly. And the location‐based service (LBS) is becoming more and more extensive. Service providers hold a large number of users’ information. In order to improve the quality of service, service providers increasingly use big data technology to provide more accurate services for users. At the same time, it aggravates the information disclosure of users’ privacy. From the perspective of the service provider, a mobile location privacy access control method based on game theory is proposed to solve the access control problem of mobile location privacy information. Firstly, the weight coefficient is set according to the location privacy influence factors, and then the access control threshold is calculated according to the privacy location leakage situation of the mobile location. Different visitor levels are set according to the threshold. In the process of access control, the prejudgement of the access behaviour is performed, and then the privacy information amount of the information requested for access is calculated according to the weights of the different information. Compare the result with the threshold and get the access control strategy. The strategy set is selected based on strategy matrix of game theory and thresholds are adjusted based on the calculation returns of strategy matrix. The effectiveness and practicability of the method are verified through the security analysis. Lijuan Zheng, Linhao Zhang, Ning Cao 0002, Jianrui Ding, Leul Yalemshet, Tsepo Nyakonda, Shepard Musasike |
Wirel. Commun. Mob. Comput. | 2 |