Ran Zhang 0013

dblp:23/4835-13 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Trustworthy machine learning · 44% Language models and text generation · 35% Machine translation · 22%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
interpretability
0.912025
Graph-Guided Textual Explanation Generation Framework · EMNLP 2025
Natural language and speech › Language models and text generation › natural language understanding › question answering
LLM-based question answering
0.912025
LiTransProQA: An LLM-based Literary Translation Evaluation Metric with Professional Question Answering · EMNLP 2025
Natural language and speech › Machine translation
machine translation evaluation
0.912025
LiTransProQA: An LLM-based Literary Translation Evaluation Metric with Professional Question Answering · EMNLP 2025
Machine learning › Trustworthy machine learning › interpretability
natural language explanation
0.912025
Graph-Guided Textual Explanation Generation Framework · EMNLP 2025
Natural language and speech › Language models and text generation
large language model
0.312025
LiTransProQA: An LLM-based Literary Translation Evaluation Metric with Professional Question Answering · EMNLP 2025
Natural language and speech › Language models and text generation
text generation
0.312025
Graph-Guided Textual Explanation Generation Framework · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

reference-free evaluation · 0.9question answering · 0.9human-in-the-loop · 0.9highlight explanation extraction · 0.9graph neural network · 0.9
YearPublicationVenuePosition
2025 Graph-Guided Textual Explanation Generation Framework
abstract
Natural language explanations (NLEs) are commonly used to provide plausible free-text explanations of a model's reasoning about its predictions.However, recent work has questioned their faithfulness, as they may not accurately reflect the model's internal reasoning process regarding its predicted answer.In contrast, highlight explanations-input fragments critical for the model's predicted answers-exhibit measurable faithfulness.Building on this foundation, we propose G-TEx, a Graph-Guided Textual Explanation Generation framework designed to enhance the faithfulness of NLEs.Specifically, highlight explanations are first extracted as faithful cues reflecting the model's reasoning logic toward answer prediction.They are subsequently encoded through a graph neural network layer to guide the NLE generation, which aligns the generated explanations with the model's underlying reasoning toward the predicted answer.Experiments on both encoder-decoder and decoder-only models across three reasoning datasets demonstrate that G-TEx improves NLE faithfulness by up to 12.18% compared to baseline methods.Additionally, G-TEx generates NLEs with greater semantic and lexical similarity to human-written ones.Human evaluations show that G-TEx can decrease redundant content and enhance the overall quality of NLEs.Our work presents a novel method for explicitly guiding NLE generation to enhance faithfulness, serving as a foundation for addressing broader criteria in NLE and generated text.
Shuzhou Yuan, Ran Zhang 0013, Michael Färber 0001, Steffen Eger, Pepa Atanasova, Isabelle Augenstein
EMNLP3
2025 LiTransProQA: An LLM-based Literary Translation Evaluation Metric with Professional Question Answering
abstract
The impact of Large Language Models (LLMs) has extended into literary domains.However, existing evaluation metrics for literature prioritize mechanical accuracy over artistic expression and tend to overrate machine translation as being superior to human translation from experienced professionals.In the long run, this bias could result in an irreversible decline in translation quality and cultural authenticity.In response to the urgent need for a specialized literary evaluation metric, we introduce LITRANSPROQA, a novel, referencefree, LLM-based question-answering framework designed for literary translation evaluation.LITRANSPROQA integrates humans in the loop to incorporate insights from professional literary translators and researchers, focusing on critical elements in literary quality assessment such as literary devices, cultural understanding, and authorial voice.Our extensive evaluation shows that while literaryfinetuned XCOMET-XL yields marginal gains, LITRANSPROQA substantially outperforms current metrics, achieving up to 0.07 gain in correlation and surpassing the best state-of-theart metrics by over 15 points in adequacy assessments.Incorporating professional translator insights as weights further improves performance, highlighting the value of translator inputs.Notably, LITRANSPROQA reaches an adequacy performance comparable to trained linguistic student evaluators, though it still falls behind experienced professional translators.LITRANSPROQA shows broad applicability to open-source models like LLaMA3.3-70b and Qwen2.5-32b,indicating its potential as an accessible and training-free tool for evaluating literary translations that require local processing due to copyright or ethical considerations.
Ran Zhang 0013, Lieve Macken, Steffen Eger
EMNLP1
2025 How Good Are LLMs for Literary Translation, Really? Literary Translation Evaluation with Humans and LLMs
abstract
Ran Zhang, Wei Zhao, Steffen Eger. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Ran Zhang 0013, Steffen Eger
NAACL (Long Papers)1
2024 Cross-lingual Cross-temporal Summarization: Dataset, Models, Evaluation
abstract
Abstract While summarization has been extensively researched in natural language processing (NLP), cross-lingual cross-temporal summarization (CLCTS) is a largely unexplored area that has the potential to improve cross-cultural accessibility and understanding. This article comprehensively addresses the CLCTS task, including dataset creation, modeling, and evaluation. We (1) build the first CLCTS corpus with 328 instances for hDe-En (extended version with 455 instances) and 289 for hEn-De (extended version with 501 instances), leveraging historical fiction texts and Wikipedia summaries in English and German; (2) examine the effectiveness of popular transformer end-to-end models with different intermediate fine-tuning tasks; (3) explore the potential of GPT-3.5 as a summarizer; and (4) report evaluations from humans, GPT-4, and several recent automatic evaluation metrics. Our results indicate that intermediate task fine-tuned end-to-end models generate bad to moderate quality summaries while GPT-3.5, as a zero-shot summarizer, provides moderate to good quality outputs. GPT-3.5 also seems very adept at normalizing historical text. To assess data contamination in GPT-3.5, we design an adversarial attack scheme in which we find that GPT-3.5 performs slightly worse for unseen source documents compared to seen documents. Moreover, it sometimes hallucinates when the source sentences are inverted against its prior knowledge with a summarization accuracy of 0.67 for plot omission, 0.71 for entity swap, and 0.53 for plot negation. Overall, our regression results of model performances suggest that longer, older, and more complex source texts (all of which are more characteristic for historical language variants) are harder to summarize for all models, indicating the difficulty of the CLCTS task. Regarding evaluation, we observe that both the GPT-4 and BERTScore correlate moderately with human evaluations, implicating great potential for future improvement.
Ran Zhang 0013, Jihed Ouni, Steffen Eger
Comput. Linguistics1