VLDB 2026 Research / reviewers in the wild / expert
Ziyin Zhang
dblp:16/7694
· DBLP profile ↗
10ranked-venue papers
5as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 10 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multilingual Encoder Knows more than You Realize: Shared Weights Pretraining for Extremely Low-Resource LanguagesabstractWhile multilingual language models like XLM-R have advanced multilingualism in NLP, they still perform poorly in extremely low-resource languages.This situation is exacerbated by the fact that modern LLMs such as LLaMA and Qwen support far fewer languages than XLM-R, making text generation models non-existent for many languages in the world.To tackle this challenge, we propose a novel framework for adapting multilingual encoders to text generation in extremely low-resource languages.By reusing the weights between the encoder and the decoder, our framework allows the model to leverage the learned semantic space of the encoder, enabling efficient learning and effective generalization in low-resource languages.Applying this framework to four Chinese minority languages, we present XLM-SWCM, and demonstrate its superior performance on various downstream tasks even when compared with much larger models. Zeli Su, Ziyin Zhang, Yushuang Dong |
ACL (1) | 2 |
| 2025 | GALLa: Graph Aligned Large Language Models for Improved Source Code UnderstandingabstractProgramming languages possess rich semantic information - such as data flow - that is represented by graphs and not available from the surface form of source code. Recent code language models have scaled to billions of parameters, but model source code solely as text tokens while ignoring any other structural information. Conversely, models that do encode structural information of code make modifications to the Transformer architecture, limiting their scale and compatibility with pretrained LLMs. In this work, we take the best of both worlds with GALLa - Graph Aligned Large Language Models. GALLa utilizes graph neural networks and cross-modal alignment technologies to inject the structural information of code into LLMs as an auxiliary task during finetuning. This framework is both model-agnostic and task-agnostic, as it can be applied to any code LLM for any code downstream task, and requires the structural graph data only at training time from a corpus unrelated to the finetuning data, while incurring no cost at inference time over the baseline LLM. Experiments on five code tasks with six different baseline LLMs ranging in size from 350M to 14B validate the effectiveness of GALLa, demonstrating consistent improvement over the baseline, even for powerful models such as LLaMA3 and Qwen2.5-Coder. Ziyin Zhang, Hang Yu 0002, Sage Lee, Peng Di, Rui Wang 0015 |
ACL (1) | 1 |
| 2025 | CMHG: A Dataset and Benchmark for Headline Generation of Minority Languages in ChinaabstractMinority languages in China, such as Tibetan, Uyghur, and Traditional Mongolian, face significant challenges due to their unique writing systems, which differ from international standards.This discrepancy has led to a severe lack of relevant corpora, particularly for supervised tasks like headline generation.To address this gap, we introduce a novel dataset, Chinese Minority Headline Generation (CMHG), which includes 100,000 entries for Tibetan, and 50,000 entries each for Uyghur and Mongolian, specifically curated for headline generation tasks.Additionally, we propose a high-quality test set annotated by native speakers, designed to serve as a benchmark for future research in this domain.We hope this dataset will become a valuable resource for advancing headline generation in Chinese minority languages and contribute to the development of related benchmarks. Zeli Su, Ziyin Zhang, Yushuang Dong |
EMNLP | 3 |
| 2025 | Draft Model Knows When to Stop: Self-Verification Speculative Decoding for Long-Form GenerationabstractConventional speculative decoding (SD) methods utilize a predefined length policy for proposing drafts, which implies the premise that the target model smoothly accepts the proposed draft tokens.However, reality deviates from this assumption: the oracle draft length varies significantly, and the fixed-length policy hardly satisfies such a requirement.Moreover, such discrepancy is further exacerbated in scenarios involving complex reasoning and long-form generation, particularly under testtime scaling for reasoning-specialized models.Through both theoretical and empirical estimation, we establish that the discrepancy between the draft and target models can be approximated by the draft model's prediction entropy: a high entropy indicates a low acceptance rate of draft tokens, and vice versa.Based on this insight, we propose SVIP: Self-Verification Length Policy for Long-Context Speculative Decoding, which is a training-free dynamic length policy for speculative decoding systems that adaptively determines the lengths of draft sequences by referring to the draft entropy.Experimental results on mainstream SD benchmarks as well as reasoning-heavy benchmarks demonstrate the superior performance of SVIP, achieving up to 17% speedup on MT-Bench at 8K context compared with fixed draft lengths, and 22% speedup for QwQ in long-form reasoning. Ziyin Zhang, Zhiwei He 0002, Rui Wang 0015, Zhaopeng Tu |
EMNLP | 1 |
| 2025 | Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering TasksabstractRecent advances in Large Language Models (LLMs) have shown promise in function-level code generation, yet repository-level software engineering tasks remain challenging. Current solutions predominantly rely on proprietary LLM agents, which introduce unpredictability and limit accessibility, raising concerns about data privacy and model customization. This paper investigates whether open-source LLMs can effectively address repository-level tasks without requiring agent-based approaches. We demonstrate this is possible by enabling LLMs to comprehend functions and files within codebases through their semantic information and structural dependencies. To this end, we introduce Code Graph Models (CGMs), which integrate repository code graph structures into the LLM's attention mechanism and map node attributes to the LLM's input space using a specialized adapter. When combined with an agentless graph RAG framework, our approach achieves a 43.00% resolution rate on the SWE-bench Lite benchmark using the open-source Qwen2.5-72B model. This performance ranks first among open weight models, second among methods with open-source systems, and eighth overall, surpassing the previous best open-source model-based method by 12.33%. Hongyuan Tao, Ying Zhang 0090, Zhenhao Tang, Hongen Peng, Xukun Zhu, Bingchang Liu, Yingguang Yang, Ziyin Zhang, Zhaogui Xu, Haipeng Zhang 0004, Linchao Zhu, Rui Wang 0015, Hang Yu 0002, Peng Di |
NeurIPS | 8 |
| 2024 | Self-Distillation Regularized Connectionist Temporal Classification Loss for Text Recognition: A Simple Yet Effective ApproachabstractText recognition methods are gaining rapid development. Some advanced techniques, e.g., powerful modules, language models, and un- and semi-supervised learning schemes, consecutively push the performance on public benchmarks forward. However, the problem of how to better optimize a text recognition model from the perspective of loss functions is largely overlooked. CTC-based methods, widely used in practice due to their good balance between performance and inference speed, still grapple with accuracy degradation. This is because CTC loss emphasizes the optimization of the entire sequence target while neglecting to learn individual characters. We propose a self-distillation scheme for CTC-based model to address this issue. It incorporates a framewise regularization term in CTC loss to emphasize individual supervision, and leverages the maximizing-a-posteriori of latent alignment to solve the inconsistency problem that arises in distillation between CTC-based models. We refer to the regularized CTC loss as Distillation Connectionist Temporal Classification (DCTC) loss. DCTC loss is module-free, requiring no extra parameters, longer inference lag, or additional training data or phases. Extensive experiments on public benchmarks demonstrate that DCTC can boost text recognition model accuracy by up to 2.6%, without any of these drawbacks. Ziyin Zhang, Ning Lu 0003, Minghui Liao, Yongshuai Huang, Cheng Li 0040, Wei Peng 0011 |
AAAI | 1 |
| 2024 | MELA: Multilingual Evaluation of Linguistic AcceptabilityabstractIn this work, we present the largest benchmark to date on linguistic acceptability: Multilingual Evaluation of Linguistic Acceptability-MELA, with 46K samples covering 10 languages from a diverse set of language families.We establish LLM baselines on this benchmark, and investigate cross-lingual transfer in acceptability judgements with XLM-R.In pursuit of multilingual interpretability, we conduct probing experiments with fine-tuned XLM-R to explore the process of syntax capability acquisition.Our results show that GPT-4o exhibits a strong multilingual ability, outperforming fine-tuned XLM-R, while open-source multilingual models lag behind by a noticeable gap.Cross-lingual transfer experiments show that transfer in acceptability judgment is non-trivial: 500 Icelandic fine-tuning examples lead to 23 MCC performance in a completely unrelated language-Chinese.Results of our probing experiments indicate that training on MELA improves the performance of XLM-R on syntaxrelated tasks. Ziyin Zhang, Yikang Liu 0002, Weifang Huang, Junyu Mao, Rui Wang 0015, Hai Hu 0001 |
ACL (1) | 1 |
| 2022 | MUST Augment: Efficient Augmentation with Multi-stage Stochastic Strategy
Qingrui Li, Song Xie, Anil Oymagil, Ziyin Zhang, Mustafa Eseoglu, Choonmeng Lee |
ICANN (1) | 4 |
| 2021 | CATNet: Scene Text Recognition Guided by Concatenating Augmented Text Features
Ziyin Zhang, Lemeng Pan, Lin Du 0010, Qingrui Li |
ICDAR (1) | 1 |
| 2021 | Handwritten Mathematical Expression Recognition with Bidirectionally Trained Transformer
Wenqi Zhao, Liangcai Gao, Zuoyu Yan, Shuai Peng, Ziyin Zhang |
ICDAR (2) | 6 |