VLDB 2026 Research / reviewers in the wild / expert
Jiarui Zhang 0003
dblp:194/0368-3 · also JiaRui Zhang 0003
· DBLP profile ↗
11ranked-venue papers
7as first author
10since 2021 · last 2026
0000-0003-4274-1846ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 5 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Neuromodulated Delta Adapters: Stabilizing Test-Time Adaptation via Gated Error CorrectionabstractStatic neural networks degrade under distribution shift, and existing Test-Time Adaptation (TTA) either backpropagates during inference or relies on unstable Hebbian dynamics.We propose the Neuromodulated Delta Adapter (NDA), a plug-and-play PEFT module that inserts a rank-r fast-weight bottleneck into frozen Transformers.NDA couples a gated Delta rule with a three-factor "surprise" signal, providing adaptive gain control that keeps fast weights Lyapunov-stable.On FLORES-101 (continuous) NDA surpasses TENT by +0.9 spBLEU and remains stable on English-to-Yoruba.On PG-19 it lowers perplexity to 36.9 at 8k tokens and recalls 94% of long-range "needle" facts. Related WorkDynamic Evaluation & TTA: Dynamic evaluation [7] adapts via backpropagation but costs roughly 3× inference time and needs careful tuning.NDA achieves similar adaptation with O(1) forward updates.Fast Weight Programmers: Linear Transformers [8], Fast Weight Programmers [2], Mamba [5], and Gated DeltaNet [6] stabilize long sequences by redesigning backbone layers.Unlike their input-only gates, NDA conditions 327 Jiarui Zhang 0003 |
ESANN | 1 |
| 2026 | NRD: A Hybrid Disentanglement Framework for Mitigating Interference in Multilingual Machine Translation
Jiarui Zhang 0003 |
LREC | 1 |
| 2026 | RouteLlama: Proactive Disentanglement for Robust Multi-domain Text Mining
Jiarui Zhang 0003, Mingzhe Lu |
PAKDD (3) | 1 |
| 2026 | Unlocking the Multilingual Long-Tail Web: A Fused Macro-Micro Framework for Scalable Content AnalysisabstractLarge Language Models (LLMs) enable Web-scale multilingual content analysis but face critical challenges in scaling to long-tail languages and ensuring robustness. Current research is split between two isolated trajectories: a Macro-Paradigm (system-level engineering) and a Micro-Paradigm (internal model intervention). We argue that a true Web-scale solution requires their systematic fusion, balancing large-scale data processing with fine-grained model control. We introduce the Control-Tower Framework (CTF), a novel methodology designed to systematically enhance powerful, pre-trained base models. Inspired by control-theoretic ideas, CTF transforms a base model into a controllable analysis engine via three synergistic stages: (1) Micro-enhanced pre-training that injects linguistic priors (e.g., syntax) to build a robust semantic foundation; (2) a control-inspired fine-tuning stage where a heuristic dynamic feedback loop, driven by micro-level error signals (e.g., knowledge editing loss), actively adjusts the macro-scale learning curriculum; and (3) Macro-optimized inference using Minimum Bayes Risk (MBR) decoding to enhance robustness on noisy user-generated content (UGC). Extensive experiments show that CTF surpasses the leading open-weights model, Tower+ 9B FT, by a substantial margin of +2.18 XCOMET-XXL on low-resource languages (WMT24++). Crucially, CTF unlocks large-scale cross-lingual Web mining by converting unstructured Web text into machine-analyzable assets. We evidence this with substantial gains across both document-level (on MARC) and aspect-based (on SemEval-2016) sentiment analysis tasks. Our work offers a practical pathway toward building more reliable, scalable, and controllable global information ecosystems. Jiarui Zhang 0003, Qihao Wang |
WWW | 1 |
| 2024 | Teaching Large Language Models to Translate on Low-resource Languages with Textbook PromptingabstractLarge Language Models (LLMs) have achieved impressive results in Machine Translation by simply following instructions, even without training on parallel data. However, LLMs still face challenges on low-resource languages due to the lack of pre-training data. In real-world situations, humans can become proficient in their native languages through abundant and meaningful social interactions and can also learn foreign languages effectively using well-organized textbooks. Drawing inspiration from human learning patterns, we introduce the Translate After LEarNing Textbook (TALENT) approach, which aims to enhance LLMs’ ability to translate low-resource languages by learning from a textbook. TALENT follows a step-by-step process: (1) Creating a Textbook for low-resource languages. (2) Guiding LLMs to absorb the Textbook’s content for Syntax Patterns. (3) Enhancing translation by utilizing the Textbook and Syntax Patterns. We thoroughly assess TALENT’s performance using 112 low-resource languages from FLORES-200 with two LLMs: ChatGPT and BLOOMZ. Evaluation across three different metrics reveals that TALENT consistently enhances translation performance by 14.8% compared to zero-shot baselines. Further analysis demonstrates that TALENT not only improves LLMs’ comprehension of low-resource languages but also equips them with the knowledge needed to generate accurate and fluent sentences in these languages. Ping Guo 0002, Yubing Ren, Yue Hu 0002, Yunpeng Li 0006, Jiarui Zhang 0003, Xingsheng Zhang, Heyan Huang |
LREC/COLING | 5 |
| 2024 | Enhancing Zero-Shot Translation in Multilingual Neural Machine Translation: Focusing on Obtaining Location-Agnostic Representations
Jiarui Zhang 0003, Heyan Huang, Yue Hu 0002, Ping Guo 0002 |
ICANN (7) | 1 |
| 2023 | Mitigating Long-Tail Language Representation Collapsing via Cross-Lingual Bootstrapped Unsupervised Fine-TuningabstractLarge Language Models have shown great capability to comprehend natural language and provide reasonable responses. However, previous researches have shown weak performance of these models on low-resource (long-tail) languages. It remains to be a problem to mitigate the performance gap between long-tail languages and rich-resource ones, which is referred to as long-tail language representation collapsing. Though some previous works can generate pseudo-parallel corpora with the auto-regressive generation, this generation progress is time-consuming and remains low quality, particularly for long-tail languages. In this paper, we propose a (X) Cross-lingual Bootstrapped Unsupervised Fine-tuning Framework (X-BUFF) to mitigate long-tail language representation collapsing. X-BUFF iteratively updates cross-lingual PLMs in a curriculum way. In each iteration of X-BUFF, we (1) select sentences with complementary semantics from monolingual corpora in long-tail languages. (2) match these selected sentences with semantic equivalent sentences in many other languages to create parallel sentence pairs, which we then merge with previous sentence pairs to build a larger and more difficult bootstrapped parallel queue. (3) fine-tune the PLMs with the bootstrapped parallel queue. Extensive experiments show that X-BUFF can mitigate the long-tail language representation collapsing problem in cross-lingual PLMs and achieve significant improvements over the previous baselines on several cross-lingual evaluation benchmarks. Ping Guo 0002, Yue Hu 0002, Yubing Ren, Yunpeng Li 0006, Jiarui Zhang 0003, Xingsheng Zhang |
ECAI | 5 |
| 2023 | Importance-Based Neuron Selective Distillation for Interference Mitigation in Multilingual Neural Machine Translation
Jiarui Zhang 0003, Heyan Huang, Yue Hu 0002, Ping Guo 0002, Yuqiang Xie |
KSEM (4) | 1 |
| 2023 | Helping Language Models Learn More: Multi-Dimensional Task Prompt for Few-shot TuningabstractLarge language models (LLMs) can be used as accessible and intelligent chatbots by constructing natural language queries and directly inputting the prompt into the large language model. However, different prompt' constructions often lead to uncertainty in the answers and thus make it hard to utilize the specific knowledge of LLMs (like ChatGPT). To alleviate this, we use an interpretable structure to explain the prompt learning principle in LLMs, which certificates that the effectiveness of language models is determined by position changes of the task's related tokens. Therefore, we propose MTPrompt, a multi-dimensional task prompt learning method consisting based on task-related object, summary, and task description information. By automatically building and searching for appropriate prompts, our proposed MTPrompt achieves the best results on few-shot samples setting and five different datasets. In addition, we demonstrate the effectiveness and stability of our method in different experimental settings and ablation experiments. In interaction with large language models, embedding more task-related information into prompts will make it easier to stimulate knowledge embedded in large language models. Jinta Weng, Jiarui Zhang 0003, Yue Hu 0002, Daidong Fa, Heyan Huang |
SMC | 2 |
| 2021 | A Relation-aware Attention Neural Network for Modeling the Usage of Scientific Online ResourcesabstractMore and more online resources for computer science are introduced, used and released in scientific literature in recent years. Knowledge about the usage of these online resources can help researchers easily find the applicable resources for their works. However, most existing methods ignore the importance of the content of the online resource citations. To this end, we manually create SciR, a dataset that contains 3,012 annotation sentences for this task, and introduce a multi-task learning framework to automatically extract the entities and relations from the context of online resource citations in scientific papers. Furthermore, considering the words in a sentence usually play different roles under different relations. In this paper, we treat different relations as distinctive sub-spaces and model the correlations between words in sentence for each relation type by a supervised biaffine attention network. Based on this relation-aware attention network, our model can not only effectively obtain the word-level correlations under each relation, but also naturally avoid the problem of overlapping relations. To evaluate the effectiveness of our model, we conduct comprehensive experiments on three datasets and the experimental results demonstrate that our model outperforms other state-of-the-art methods on the two tasks of entity recognition and relation extraction. Yongxiu Xu, Heyan Huang, Chong Feng 0001, Chuan Zhou 0001, Jiarui Zhang 0003, Yue Hu 0002 |
IJCNN | 5 |
| 2020 | Dynamic Attention Aggregation with BERT for Neural Machine TranslationabstractThe recently proposed BERT has demonstrated great power in various natural language processing tasks. However, the model does not perform effectively on cross-lingual tasks, especially on machine translation. In this work, we propose three methods to introduce pre-trained BERT into neural machine translation without fine-tuning. Our approach consists of a) a linear-attention aggregation that leverages a parameter matrix to capture the key knowledge of BERT, b) a self-attention aggregation which aims to learn what is vital for input and output, and c) a switch-gate aggregation to dynamically control the balance of the information flowing from the pre-trained BERT or the NMT model. We conduct experiments on several translation benchmarks and substantially improve over 2 BELU points on the IWSLT'14 English - German task with switch-gate aggregation method compared to a strong baseline, while our proposed model also performs remarkably on the other tasks. Jiarui Zhang 0003, Hongzheng Li, Shumin Shi, Heyan Huang, Yue Hu 0002, Xiangpeng Wei |
IJCNN | 1 |