VLDB 2026 Research / reviewers in the wild / expert
Dingkun Long
dblp:190/7094
· DBLP profile ↗
19ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0001-6570-9406ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ERank: Fusing Supervised Fine-Tuning and Reinforcement Learning for Effective and Efficient Text RerankingabstractText reranking models are a crucial component in modern systems like Retrieval-Augmented Generation, tasked with selecting the most relevant documents prior to generation. However, current Large Language Models (LLMs) powered rerankers often face a fundamental trade-off. On one hand, Supervised Fine-Tuning based pointwise methods that frame relevance as a binary classification task lack the necessary scoring discrimination, particularly for those built on reasoning LLMs. On the other hand, approaches designed for complex reasoning often employ powerful yet inefficient listwise formulations, rendering them impractical for low latency applications. To resolve this dilemma, we introduce ERank, a highly Effective and Efficient pointwise reranker built from a reasoning LLM that excels across diverse relevance scenarios. We propose a novel two-stage training pipeline that begins with Supervised Fine-Tuning (SFT). In this stage, we move beyond binary labels and train the model generatively to output fine grained integer scores, which significantly enhances relevance discrimination. The model is then further refined using Reinforcement Learning (RL) with a novel, listwise derived reward. This technique instills global ranking awareness into the efficient pointwise architecture. We evaluate the ERank reranker on the BRIGHT, FollowIR, TREC DL, and BEIR benchmarks, demonstrating superior effectiveness and robustness compared to existing approaches. On the reasoning-intensive BRIGHT benchmark, our ERank-4B achieves an nDCG@10 of 38.7, while a larger 32B variant reaches a state of the art nDCG@10 of 40.2. Yuzheng Cai, Yanzhao Zhang, Dingkun Long, Pengjun Xie, Weiguo Zheng |
AAAI | 3 |
| 2026 | Rethinking Composed Image Retrieval Evaluation: A Fine-Grained Benchmark from Image EditingabstractTingyu Song, Yanzhao Zhang, Mingxin Li, Zhuoning Guo, Dingkun Long, Pengjun Xie, Siyue Zhang, Yilun Zhao, Shu Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Tingyu Song, Yanzhao Zhang, Zhuoning Guo, Dingkun Long, Pengjun Xie, Siyue Zhang, Yilun Zhao 0001 |
ACL (1) | 5 |
| 2026 | Internalizing Explicit Reasoning into Latent Space for Dense RetrievalabstractLarge Language Models (LLMs) have fundamentally transformed dense retrieval, upgrading backbones from discriminative encoders to generative architectures. However, a critical disconnect remains: while LLMs possess strong reasoning capabilities, current retrievers predominantly utilize them as static encoders, leaving their potential for complex reasoning unexplored. To address this, existing approaches typically adopt ''rewrite-then-retrieve'' pipelines to generate explicit Chain-of-Thought (CoT) rationales before retrieval. However, this incurs prohibitive latency. Conversely, implicit reasoning methods utilizing latent tokens offer efficiency but often suffer from semantic degeneration due to the lack of explicit supervision. In this paper, we propose LaSER, a novel self-distillation framework that internalizes explicit reasoning into the latent space of dense retrievers. Operating on a shared LLM backbone, LaSER introduces a dual-view training mechanism: an Explicit view that explicitly encodes ground-truth reasoning paths, and a Latent view that performs implicit latent thinking. To bridge the gap between these views, we design a multi-grained alignment strategy. Beyond standard output alignment, we introduce a trajectory alignment mechanism that synchronizes the intermediate latent states of the latent path with the semantic progression of the explicit reasoning segments. This allows the retriever to ''think'' silently and effectively without autoregressive text generation. Extensive experiments on both in-domain and out-of-domain reasoning-intensive benchmarks demonstrate that LaSER significantly outperforms state-of-the-art baselines. Furthermore, analyses across diverse backbones and model scales validate the robustness of our approach, confirming that our unified learning framework is essential for eliciting effective latent thinking. Our method successfully combines the reasoning depth of explicit CoT pipelines with the inference efficiency of standard dense retrievers. The code, model, and training data are available at https://github.com/RUC-NLPIR/LaSER. Jiajie Jin, Yanzhao Zhang, Dingkun Long, Pengjun Xie, Yutao Zhu 0001, Zhicheng Dou |
SIGIR | 4 |
| 2026 | Progressive Adaptation of Large Language Models for Multilingual Text RankingabstractDespite increasing research attention to text ranking, most studies focus on monolingual scenarios, with a particular emphasis on English-language contexts. This narrow focus limits the applicability of ranking models in cross-lingual contexts, such as ranking Chinese documents based on English queries. Recent advances in large language models (LLMs) have significantly reduced inter-language barriers through pre-training on extensive multilingual corpora, thus facilitating the study of multilingual text ranking (MTR). In this work, we explore the potential of LLMs in MTR tasks. Specifically, we first introduce an MTR benchmark encompassing both monolingual and cross-lingual scenarios. Then, we propose a two-stage training pipeline to alleviate the misalignment between LLMs and text ranking. Lastly, we adapt this training pipeline to multilingual scenarios from the perspective of training data and methods. Our experiments on the MTR benchmark demonstrate that the proposed multilingual two-stage training pipeline significantly improves LLM ranking performance in both monolingual and cross-lingual scenarios, particularly in out-domain settings. We complement these findings with a thorough analysis to deepen the understanding of our approach. Longhui Zhang, Yanzhao Zhang, Dingkun Long, Pengjun Xie, Meishan Zhang, Jing Li 0034, Min Zhang 0005 |
ACM Trans. Inf. Syst. | 3 |
| 2025 | Towards Text-Image Interleaved RetrievalabstractXin Zhang, Ziqi Dai, Yongqi Li, Yanzhao Zhang, Dingkun Long, Pengjun Xie, Meishan Zhang, Jun Yu, Wenjie Li, Min Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xin Zhang 0097, Ziqi Dai, Yongqi Li 0001, Yanzhao Zhang, Dingkun Long, Pengjun Xie, Meishan Zhang, Jun Yu 0002, Wenjie Li 0002, Min Zhang 0005 |
ACL (1) | 5 |
| 2025 | Bridging Modalities: Improving Universal Multimodal Retrieval by Multimodal Large Language ModelsabstractUniversal Multimodal Retrieval (UMR) aims to enable search across various modalities using a unified model, where queries and candidates can consist of pure text, images, or a combination of both. Previous work has attempted to adopt multimodal large language models (MLLMs) to realize UMR using only text data. However, our preliminary experiments demonstrate that more diverse multimodal training data can further unlock the potential of MLLMs. Despite its effectiveness, the existing multimodal training data is highly imbalanced in terms of modality, which motivates us to develop a training data synthesis pipeline and construct a large-scale, high-quality fused-modal training dataset. Based on the synthetic training data, we develop the General Multimodal Embedder (GME), an MLLM-based dense retriever designed for UMR. Furthermore, we construct a comprehensive UMR Benchmark (UMRB) to evaluate the effectiveness of our approach. Experimental results show that our method achieves state-of-the-art performance among existing UMR methods. Last, we provide in-depth analyses of model scaling and training strategies, and perform ablation studies on both the model and synthetic data. Xin Zhang 0097, Yanzhao Zhang, Wen Xie 0006, Ziqi Dai, Dingkun Long, Pengjun Xie, Meishan Zhang, Wenjie Li 0002, Min Zhang 0005 |
CVPR | 6 |
| 2025 | An End-to-End Model for Photo-Sharing Multi-Modal Dialogue GenerationabstractPhoto-Sharing Multi-modal dialogue generation requires a dialogue agent not only to generate text responses but also to share photos at the proper moment. Using image text caption as the bridge, a pipeline model integrates an image caption model, a text generation model, and an image generation model to handle this complex multi-modal task. However, representing the images with text captions may lose important visual details and information and cause error propagation in the complex dialogue system. Besides, the pipeline model isolates the three models separately because discrete image text captions hinder end-to-end gradient propagation. We propose the first end-to-end model for photo-sharing multi-modal dialogue generation, which integrates an image perceptron and an image generator with a large language model. The large language model employs the vision encoder to perceive visual images in the input end. For image generation in the output end, we propose a dynamic vocabulary transformation matrix and use straight-through and gumbel-softmax techniques to align the large language model and stable diffusion model and achieve end-to-end gradient propagation. We perform experiments on PhotoChat and DialogCC datasets to evaluate our end-to-end model. Compared with pipeline models, the end-to-end model gains state-of-the-art performances on various metrics of text and image generation. Peiming Guo, Sinuo Liu, Yanzhao Zhang, Dingkun Long, Pengjun Xie, Meishan Zhang, Min Zhang 0005 |
ICME | 4 |
| 2025 | Breaking the Modality Barrier: Universal Embedding Learning with Multimodal LLMsabstractThe Contrastive Language-Image Pre-training (CLIP) framework has become a widely used approach for multimodal representation learning, particularly in image-text retrieval and clustering. However, its efficacy is constrained by three key limitations: (1) text token truncation, (2) isolated image-text encoding, and (3) deficient compositionality due to bag-of-words behavior. While recent Multimodal Large Language Models (MLLMs) have demonstrated significant advances in generalized vision-language understanding, their potential for learning transferable multimodal representations remains underexplored. In this work, we present UniME (Universal Multimodal Embedding), a novel two-stage framework that leverages MLLMs to learn discriminative representations for diverse downstream tasks. In the first stage, we perform textual discriminative knowledge distillation from a powerful LLM-based teacher model to enhance the embedding capability of the MLLM's language component. In the second stage, we introduce hard negative enhanced instruction tuning to further advance discriminative representation learning. Specifically, we initially mitigate false negative contamination and then sample multiple hard negatives per instance within each batch, forcing the model to focus on challenging samples. This approach not only improves discriminative power but also enhances instruction-following ability in downstream tasks. We conduct extensive experiments on the MMEB benchmark and multiple retrieval tasks, including short & long caption retrieval and compositional retrieval. Results demonstrate that UniME achieves consistent performance improvement across all tasks, exhibiting superior discriminative and compositional capabilities. The code will be released in https://garygutc.github.io/UniME. Tiancheng Gu, Kaicheng Yang 0002, Ziyong Feng, Yanzhao Zhang, Dingkun Long, Yingda Chen, Tom Weidong Cai, Jiankang Deng |
ACM Multimedia | 6 |
| 2025 | SSRB: Direct Natural Language Querying to Massive Heterogeneous Semi-Structured DataabstractSearching over semi-structured data with natural language (NL) queries has attracted sustained attention, enabling broader audiences to access information easily. As more applications, such as LLM agents and RAG systems, emerge to search and interact with semi-structured data, two major challenges have become evident: (1) the increasing diversity of domains and schema variations, making domain-customized solutions prohibitively costly; (2) the growing complexity of NL queries, which combine both exact field matching conditions and fuzzy semantic requirements, often involving multiple fields and implicit reasoning. These challenges make formal language querying or keyword-based search insufficient. In this work, we explore neural retrievers as a unified non-formal querying solution by directly index semi-structured collections and understand NL queries. We employ LLM-based automatic evaluation and build a large-scale semi-structured retrieval benchmark (SSRB) using LLM generation and filtering, containing 14M semi-structured objects from 99 different schemas across 6 domains, along with 8,485 test queries that combine both exact and fuzzy matching conditions. Our systematic evaluation of popular retrievers shows that current state-of-the-art models could achieve acceptable performance, yet they still lack precise understanding of matching constraints. While by in-domain training of dense retrievers, the performance can be significantly improved. We believe that our SSRB could serve as a valuable resource for future research in this area, and we hope to inspire further exploration of semi-structured retrieval with complex queries. Xin Zhang 0097, Yanzhao Zhang, Dingkun Long, Yongqi Li 0001, Pengjun Xie, Meishan Zhang, Wenjie Li 0002, Min Zhang 0005, Philip S. Yu |
NeurIPS | 4 |
| 2024 | Chinese Sequence Labeling with Semi-Supervised Boundary-Aware Language Model Pre-trainingabstractChinese sequence labeling tasks are sensitive to word boundaries. Although pretrained language models (PLM) have achieved considerable success in these tasks, current PLMs rarely consider boundary information explicitly. An exception to this is BABERT, which incorporates unsupervised statistical boundary information into Chinese BERT’s pre-training objectives. Building upon this approach, we input supervised high-quality boundary information to enhance BABERT’s learning, developing a semi-supervised boundary-aware PLM. To assess PLMs’ ability to encode boundaries, we introduce a novel “Boundary Information Metric” that is both simple and effective. This metric allows comparison of different PLMs without task-specific fine-tuning. Experimental results on Chinese sequence labeling datasets demonstrate that the improved BABERT version outperforms the vanilla version, not only in these tasks but also in broader Chinese natural language understanding tasks. Additionally, our proposed metric offers a convenient and accurate means of evaluating PLMs’ boundary awareness. Longhui Zhang, Dingkun Long, Meishan Zhang, Yanzhao Zhang, Pengjun Xie, Min Zhang 0005 |
LREC/COLING | 2 |
| 2023 | Text Representation Distillation via Information Bottleneck PrincipleabstractPre-trained language models (PLMs) have recently shown great success in text representation field.However, the high computational cost and high-dimensional representation of PLMs pose significant challenges for practical applications.To make models more accessible, an effective method is to distill large models into smaller representation models.In order to relieve the issue of performance degradation after distillation, we propose a novel Knowledge Distillation method called IBKD.This approach is motivated by the Information Bottleneck principle and aims to maximize the mutual information between the final representation of the teacher and student model, while simultaneously reducing the mutual information between the student model's representation and the input data.This enables the student model to preserve important learned information while avoiding unnecessary information, thus reducing the risk of over-fitting.Empirical studies on two main downstream applications of text representation (Semantic Textual Similarity and Dense Retrieval tasks) demonstrate the effectiveness of our proposed approach 1 . Yanzhao Zhang, Dingkun Long, Zehan Li, Pengjun Xie |
EMNLP | 2 |
| 2023 | Fine-Grained Domain Adaptation for Chinese Syntactic ProcessingabstractSyntactic processing is fundamental to natural language processing. It provides rich and comprehensive syntax information in sentences that could be potentially beneficial for downstream tasks. Recently, pretrained language models have shown great success in Chinese syntactic processing, which typically involves word segmentation, POS tagging, and dependency parsing. However, the on-going research never ends since performance would be degraded drastically when tested on a highly-discrepant domain. This problem is widely accepted as domain adaptation, where the test domain differs from the training domain in supervised learning. Self-training is one promising solution for it, and straightforward source-to-target adaptation has already shown remarkable effectiveness in previous work. While this strategy ignores the fact that sentences of the target domain sentences may have very different gaps from the source training domain. More specifically, sentences with large gaps might fail by direct self-training adaptation. To this end, we propose fine-grained domain adaptation for Chinese syntactic processing in this work, aiming to model the gaps between the source and the target domains accurately and progressively. The key idea is to divide the target domain into fine-grained subdomains by using a specified domain distance metric, and then perform gradual self-training on the subdomains. We further offer an intuitive theoretical illustration based on the theory of Kumar et al. (2020) approximately. In addition, a novel representation learning framework is proposed to encode fine-grained subdomains effectively, aiming to utilize the above idea fully. Experimental results on benchmark datasets show that our method can achieve significant improvements over a variety of baselines. Meishan Zhang, Peiming Guo, Peijie Jiang, Dingkun Long, Yueheng Sun, Pengjun Xie, Min Zhang 0005 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2022 | Unsupervised Boundary-Aware Language Model Pretraining for Chinese Sequence LabelingabstractBoundary information is critical for various Chinese language processing tasks, such as word segmentation, part-of-speech tagging, and named entity recognition.Previous studies usually resorted to the use of a high-quality external lexicon, where lexicon items can offer explicit boundary information.However, to ensure the quality of the lexicon, great human effort is always necessary, which has been generally ignored.In this work, we suggest unsupervised statistical boundary information instead, and propose an architecture to encode the information directly into pre-trained language models, resulting in Boundary-Aware BERT (BABERT).We apply BABERT for feature induction of Chinese sequence labeling tasks.Experimental results on ten benchmarks of Chinese sequence labeling demonstrate that BABERT can provide consistent improvements on all datasets.In addition, our method can complement previous supervised lexicon exploration, where further improvements can be achieved when integrated with external lexicon information. Peijie Jiang, Dingkun Long, Yanzhao Zhang, Pengjun Xie, Meishan Zhang, Min Zhang 0005 |
EMNLP | 2 |
| 2022 | Multi-CPR: A Multi Domain Chinese Dataset for Passage RetrievalabstractPassage retrieval is a fundamental task in information retrieval (IR) research, which has drawn much attention recently. In the English field, the availability of large-scale annotated dataset (e.g, MS MARCO) and the emergence of deep pre-trained language models (e.g, BERT) has resulted in a substantial improvement of existing passage retrieval systems. However, in the Chinese field, especially for specific domains, passage retrieval systems are still immature due to quality-annotated dataset being limited by scale. Therefore, in this paper, we present a novel multi-domain Chinese dataset for passage retrieval (Multi-CPR). The dataset is collected from three different domains, including E-commerce, Entertainment video and Medical. Each dataset contains millions of passages and a certain amount of human annotated query-passage related pairs. We implement various representative passage retrieval methods as baselines. We find that the performance of retrieval models trained on dataset from general domain will inevitably decrease on specific domain. Nevertheless, a passage retrieval system built on in-domain annotated dataset can achieve significant improvement, which indeed demonstrates the necessity of domain labeled data for further optimization. We hope the release of the Multi-CPR dataset could benchmark Chinese passage retrieval task in specific domain and also make advances for future studies. Dingkun Long, Qiong Gao, Kuan Zou, Pengjun Xie, Ruijie Guo, Guanjun Jiang, Luxi Xing |
SIGIR | 1 |
| 2021 | A Fine-Grained Domain Adaption Model for Joint Word Segmentation and POS TaggingabstractDomain adaption for word segmentation and POS tagging is a challenging problem for Chinese lexical processing.Self-training is one promising solution for it, which struggles to construct a set of high-quality pseudo training instances for the target domain.Previous work usually assumes a universal sourceto-target adaption to collect such pseudo corpus, ignoring the different gaps from the target sentences to the source domain.In this work, we start from joint word segmentation and POS tagging, presenting a fine-grained domain adaption method to model the gaps accurately.We measure the gaps by one simple and intuitive metric, and adopt it to develop a pseudo target domain corpus based on finegrained subdomains incrementally.A novel domain-mixed representation learning model is proposed accordingly to encode the multiple subdomains effectively.The whole process is performed progressively for both corpus construction and model training.Experimental results on a benchmark dataset show that our method can gain significant improvements over a vary of baselines.Extensive analyses are performed to show the advantages of our final domain adaption model as well. Peijie Jiang, Dingkun Long, Yueheng Sun, Meishan Zhang, Pengjun Xie |
EMNLP (1) | 2 |
| 2020 | Coupling Distant Annotation and Adversarial Training for Cross-Domain Chinese Word SegmentationabstractFully supervised neural approaches have achieved significant progress in the task of Chinese word segmentation (CWS).Nevertheless, the performance of supervised models tends to drop dramatically when they are applied to outof-domain data.Performance degradation is caused by the distribution gap across domains and the out of vocabulary (OOV) problem.In order to simultaneously alleviate these two issues, this paper proposes to couple distant annotation and adversarial training for crossdomain CWS.For distant annotation, we rethink the essence of "Chinese words" and design an automatic distant annotation mechanism that does not need any supervision or pre-defined dictionaries from the target domain.The approach could effectively explore domain-specific words and distantly annotate the raw texts for the target domain.For adversarial training, we develop a sentence-level training procedure to perform noise reduction and maximum utilization of the source domain information.Experiments on multiple realworld datasets across various domains show the superiority and robustness of our model, significantly outperforming previous state-ofthe-art cross-domain CWS methods. Ning Ding 0002, Dingkun Long, Muhua Zhu, Pengjun Xie, Xiaobin Wang, Hai-Tao Zheng 0002 |
ACL | 2 |
| 2020 | Hierarchy-Aware Global Model for Hierarchical Text ClassificationabstractHierarchical text classification is an essential yet challenging subtask of multi-label text classification with a taxonomic hierarchy.Existing methods have difficulties in modeling the hierarchical label structure in a global view.Furthermore, they cannot make full use of the mutual interactions between the text feature space and the label space.In this paper, we formulate the hierarchy as a directed graph and introduce hierarchy-aware structure encoders for modeling label dependencies.Based on the hierarchy encoder, we propose a novel end-to-end hierarchy-aware global model (Hi-AGM) with two variants.A multi-label attention variant (HiAGM-LA) learns hierarchyaware label embeddings through the hierarchy encoder and conducts inductive fusion of labelaware text features.A text feature propagation model (HiAGM-TP) is proposed as the deductive variant that directly feeds text features into hierarchy encoders.Compared with previous works, both HiAGM-LA and HiAGM-TP achieve significant and consistent improvements on three benchmark datasets. Jie Zhou 0013, Chunping Ma, Dingkun Long, Ning Ding 0002, Pengjun Xie, Gongshen Liu |
ACL | 3 |
| 2020 | Learning with Noise: Improving Distantly-Supervised Fine-grained Entity Typing via Automatic RelabelingabstractFine-grained entity typing (FET) is a fundamental task for various entity-leveraging applications. Although great success has been made, existing systems still have challenges in handling noisy samples in training data introduced by distant supervision methods. To address these noise, previous studies either focus on processing the clean samples (i,e., have only one label) and noisy samples (i,e., have multiple labels) with different strategies or filtering the noisy labels based on the assumption that the distantly-supervised label set certainly contains the correct type label. In this paper, we propose a probabilistic automatic relabeling method which treats all training samples uniformly. Our method aims to estimate the pseudo-truth label distribution of each sample, and the pseudo-truth distribution will be treated as part of trainable parameters which are jointly updated during the training process. The proposed approach does not rely on any prerequisite or extra supervision, making it effective on real applications. Experiments on several benchmarks show that our method outperforms previous approaches and alleviates the noisy labeling problem. Dingkun Long, Muhua Zhu, Pengjun Xie, Fei Huang 0002, Ji Wang 0001 |
IJCAI | 2 |
| 2018 | Prototypical recurrent unit
Dingkun Long, Richong Zhang, Yongyi Mao |
Neurocomputing | 1 |