VLDB 2026 Research / reviewers in the wild / expert
Xin Zhang 0097
dblp:76/1584-97
· DBLP profile ↗
13ranked-venue papers
7as first author
13since 2021 · last 2025
0000-0002-2550-3056ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards Text-Image Interleaved RetrievalabstractXin Zhang, Ziqi Dai, Yongqi Li, Yanzhao Zhang, Dingkun Long, Pengjun Xie, Meishan Zhang, Jun Yu, Wenjie Li, Min Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xin Zhang 0097, Ziqi Dai, Yongqi Li 0001, Yanzhao Zhang, Dingkun Long, Pengjun Xie, Meishan Zhang, Jun Yu 0002, Wenjie Li 0002, Min Zhang 0005 |
ACL (1) | 1 |
| 2025 | Bridging Modalities: Improving Universal Multimodal Retrieval by Multimodal Large Language ModelsabstractUniversal Multimodal Retrieval (UMR) aims to enable search across various modalities using a unified model, where queries and candidates can consist of pure text, images, or a combination of both. Previous work has attempted to adopt multimodal large language models (MLLMs) to realize UMR using only text data. However, our preliminary experiments demonstrate that more diverse multimodal training data can further unlock the potential of MLLMs. Despite its effectiveness, the existing multimodal training data is highly imbalanced in terms of modality, which motivates us to develop a training data synthesis pipeline and construct a large-scale, high-quality fused-modal training dataset. Based on the synthetic training data, we develop the General Multimodal Embedder (GME), an MLLM-based dense retriever designed for UMR. Furthermore, we construct a comprehensive UMR Benchmark (UMRB) to evaluate the effectiveness of our approach. Experimental results show that our method achieves state-of-the-art performance among existing UMR methods. Last, we provide in-depth analyses of model scaling and training strategies, and perform ablation studies on both the model and synthetic data. Xin Zhang 0097, Yanzhao Zhang, Wen Xie 0006, Ziqi Dai, Dingkun Long, Pengjun Xie, Meishan Zhang, Wenjie Li 0002, Min Zhang 0005 |
CVPR | 1 |
| 2025 | Mieb: Massive Image Embedding BenchmarkabstractImage representations are often evaluated through disjointed, task-specific protocols, leading to a fragmented understanding of model capabilities. For instance, it is unclear whether an image embedding model adept at clustering images is equally good at retrieving relevant images given a piece of text. We introduce the Massive Image Embedding Benchmark (MIEB) to evaluate the performance of image and image-text embedding models across the broadest spectrum to date. MIEB spans 38 languages across 130 individual tasks, which we group into 8 high-level categories. We benchmark 50 models across our benchmark, finding that no single method dominates across all task categories. We reveal hidden capabilities in advanced vision models such as their accurate visual representation of texts, and their yet limited capabilities in interleaved encodings and matching images and texts in the presence of confounders. We also show that the performance of vision encoders on MIEB correlates highly with their performance when used in multimodal large language models. Our code, dataset, and leaderboard are publicly available at https://github.com/embeddings-benchmark/mteb. Chenghao Xiao, Isaac Chung, Imene Kerboua, Jamie Stirling, Xin Zhang 0097, Márton Kardos, Roman Solomatin, Noura Al Moubayed, Kenneth C. Enevoldsen, Niklas Muennighoff |
ICCV | 5 |
| 2025 | R$^2$ec: Towards Large Recommender Models with ReasoningabstractLarge recommender models have extended LLMs as powerful recommenders via encoding or item generation, and recent breakthroughs in LLM reasoning synchronously motivate the exploration of reasoning in recommendation.
In this work, we propose R$^2$ec, a unified large recommender model with intrinsic reasoning capability.
R$^2$ec introduces a dual-head architecture that supports both reasoning chain generation and efficient item prediction in a single model, significantly reducing inference latency. To overcome the lack of annotated reasoning data, we design RecPO, a reinforcement learning framework that optimizes reasoning and recommendation jointly with a novel fused reward mechanism.
Extensive experiments on three datasets demonstrate that R$^2$ec outperforms traditional, LLM-based, and reasoning-augmented recommender baselines, while further analyses validate its competitive efficiency among conventional LLM-based recommender baselines
and strong adaptability to diverse recommendation scenarios. Code and checkpoints available at https://github.com/YRYangang/RRec. Runyang You, Yongqi Li 0001, Xinyu Lin 0001, Xin Zhang 0097, Wenjie Wang 0007, Wenjie Li 0002, Liqiang Nie |
NeurIPS | 4 |
| 2025 | SSRB: Direct Natural Language Querying to Massive Heterogeneous Semi-Structured DataabstractSearching over semi-structured data with natural language (NL) queries has attracted sustained attention, enabling broader audiences to access information easily. As more applications, such as LLM agents and RAG systems, emerge to search and interact with semi-structured data, two major challenges have become evident: (1) the increasing diversity of domains and schema variations, making domain-customized solutions prohibitively costly; (2) the growing complexity of NL queries, which combine both exact field matching conditions and fuzzy semantic requirements, often involving multiple fields and implicit reasoning. These challenges make formal language querying or keyword-based search insufficient. In this work, we explore neural retrievers as a unified non-formal querying solution by directly index semi-structured collections and understand NL queries. We employ LLM-based automatic evaluation and build a large-scale semi-structured retrieval benchmark (SSRB) using LLM generation and filtering, containing 14M semi-structured objects from 99 different schemas across 6 domains, along with 8,485 test queries that combine both exact and fuzzy matching conditions. Our systematic evaluation of popular retrievers shows that current state-of-the-art models could achieve acceptable performance, yet they still lack precise understanding of matching constraints. While by in-domain training of dense retrievers, the performance can be significantly improved. We believe that our SSRB could serve as a valuable resource for future research in this area, and we hope to inspire further exploration of semi-structured retrieval with complex queries. Xin Zhang 0097, Yanzhao Zhang, Dingkun Long, Yongqi Li 0001, Pengjun Xie, Meishan Zhang, Wenjie Li 0002, Min Zhang 0005, Philip S. Yu |
NeurIPS | 1 |
| 2025 | Retrieving and Reading Multimodal Documents for Knowledge-Based VQA
Wen Xie 0006, Xin Zhang 0097, Meishan Zhang, Min Zhang 0005 |
NLPCC (3) | 2 |
| 2025 | Dependency Scoring Learning and Corpus Boosting for Translation-Based Cross-Lingual Dependency ParsingabstractDependency parsing is a fundamental task in natural language processing that involves identifying the grammatical relationships between words in a sentence. One promising approach for performing this task in languages lacking annotated treebanks is treebank translation, which utilizes word alignments to map dependencies from a source treebank to the corresponding target translation. However, due to language differences and the limitations of word alignment tools, this method would inevitably generate noise during mapping. To reduce the effect of noise, we first exploit MetaNet to compute quality scores for each dependency and identify low-score ones as noise. MetaNet is a fake teacher that learns to score homework (dependencies) by comparing answers from the top student (strong parser) and the regular student (weak parser) without knowing the correct answer (gold-standard). With the scoring capability of MetaNet, we design an iterative algorithm to boost the target treebank quality, which trains with high-quality dependencies and relabels the low-quality dependencies. Our method achieves better results than the originally translated treebanks and shows highly competitive performances with prior methods on the Universal Dependency Treebanks v2.2. We also provide detailed analysis and discussions. Huiyao Chen, Xin Zhang 0097, Jing Chen 0062, Meishan Zhang, Min Zhang 0005 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2024 | AutoSurvey: Large Language Models Can Automatically Write SurveysabstractThis paper introduces AutoSurvey, a speedy and well-organized methodology for automating the creation of comprehensive literature surveys in rapidly evolving fields like artificial intelligence. Traditional survey paper creation faces challenges due to the vast volume and complexity of information, prompting the need for efficient survey methods. While large language models (LLMs) offer promise in automating this process, challenges such as context window limitations, parametric knowledge constraints, and the lack of evaluation benchmarks remain. AutoSurvey addresses these challenges through a systematic approach that involves initial retrieval and outline generation, subsection drafting by specialized LLMs, integration and refinement, and rigorous evaluation and iteration. Our contributions include a comprehensive solution to the survey problem, a reliable evaluation method, and experimental validation demonstrating AutoSurvey's effectiveness. Yidong Wang 0003, Wenjin Yao, Xin Zhang 0097, Zhen Wu 0002, Meishan Zhang, Xinyu Dai, Min Zhang 0005, Qingsong Wen, Wei Ye 0004, Shikun Zhang, Yue Zhang 0004 |
NeurIPS | 5 |
| 2023 | Finetuning Language Models for Multimodal Question AnsweringabstractTo achieve multi-modal intelligence, AI must be able to process and respond to inputs from multimodal sources. However, many current question answering models are limited to specific types of answers, such as yes/no and number, and require additional human assessments. Recently, Visual-Text Question Answering (VQTA) dataset has been proposed to fix this gap. In this paper, we conduct an exhaustive analysis and exploration of this task. Specifically, we implement a T5-based multi-modal generative network that overcomes the limitations of traditional labeling space and provides more freedom in responses. Our approach achieve the best performance in both English and Chinese tracks in the VTQA challenge. Xin Zhang 0097, Wen Xie 0006, Ziqi Dai, Jun Rao, Haokun Wen, Meishan Zhang, Min Zhang 0005 |
ACM Multimedia | 1 |
| 2022 | Identifying Chinese Opinion Expressions with Extremely-Noisy Crowdsourcing AnnotationsabstractRecent works of opinion expression identification (OEI) rely heavily on the quality and scale of the manually-constructed training corpus, which could be extremely difficult to satisfy.Crowdsourcing is one practical solution for this problem, aiming to create a large-scale but quality-unguaranteed corpus.In this work, we investigate Chinese OEI with extremelynoisy crowdsourcing annotations, constructing a dataset at a very low cost.Following Zhang et al. (2021), we train the annotator-adapter model by regarding all annotations as goldstandard in terms of crowd annotators, and test the model by using a synthetic expert, which is a mixture of all annotators.As this annotatormixture for testing is never modeled explicitly in the training phase, we propose to generate synthetic training samples by a pertinent mixup strategy to make the training and testing highly consistent.The simulation experiments on our constructed dataset show that crowdsourcing is highly promising for OEI, and our proposed annotator-mixup can further enhance the crowdsourcing modeling. Xin Zhang 0097, Yueheng Sun, Meishan Zhang, Xiaobin Wang, Min Zhang 0005 |
ACL (1) | 1 |
| 2022 | Domain-Specific NER via Retrieving Correlated SamplesabstractSuccessful Machine Learning based Named Entity Recognition models could fail on texts from some special domains, for instance, Chinese addresses and e-commerce titles, where requires adequate background knowledge. Such texts are also difficult for human annotators. In fact, we can obtain some potentially helpful information from correlated texts, which have some common entities, to help the text understanding. Then, one can easily reason out the correct answer by referencing correlated samples. In this paper, we suggest enhancing NER models with correlated samples. We draw correlated samples by the sparse BM25 retriever from large-scale in-domain unlabeled data. To explicitly simulate the human reasoning process, we perform a training-free entity type calibrating by majority voting. To capture correlation features in the training stage, we suggest to model correlated samples by the transformer-based multi-instance cross-encoder. Empirical results on datasets of the above two domains show the efficacy of our methods. Xin Zhang 0097, Yong Jiang 0005, Xiaobin Wang, Xuming Hu, Yueheng Sun, Pengjun Xie, Meishan Zhang |
COLING | 1 |
| 2022 | Extending Phrase Grounding with Pronouns in Visual DialoguesabstractConventional phrase grounding aims to localize noun phrases mentioned in a given caption to their corresponding image regions, which has achieved great success recently.Apparently, sole noun phrase grounding is not enough for cross-modal visual language understanding.Here we extend the task by considering pronouns as well.First, we construct a dataset of phrase grounding with both noun phrases and pronouns to image regions.Based on the dataset, we test the performance of phrase grounding by using a state-of-the-art literature model of this line.Then, we enhance the baseline grounding model with coreference information which should help our task potentially, modeling the coreference structures with graph convolutional networks.Experiments on our dataset, interestingly, show that pronouns are easier to ground than noun phrases, where the possible reason might be that these pronouns are much less ambiguous.Additionally, our final model with coreference information can significantly boost the grounding performance of both noun phrases and pronouns. Panzhong Lu, Xin Zhang 0097, Meishan Zhang, Min Zhang 0005 |
EMNLP | 2 |
| 2021 | Crowdsourcing Learning as Domain Adaptation: A Case Study on Named Entity RecognitionabstractXin Zhang, Guangwei Xu, Yueheng Sun, Meishan Zhang, Pengjun Xie. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Xin Zhang 0097, Yueheng Sun, Meishan Zhang, Pengjun Xie |
ACL/IJCNLP (1) | 1 |