Jiazheng Li 0002

dblp:155/6074-2 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
12since 2021 · last 2025
0000-0002-9610-9232ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2025 An Automated Explainable Educational Assessment System Built on LLMs
abstract
In this demo, we present AERA Chat, an automated and explainable educational assessment system designed for interactive and visual evaluations of student responses. This system leverages large language models (LLMs) to generate automated marking and rationale explanations, addressing the challenge of limited explainability in automated educational assessment and the high costs associated with annotation. Our system allows users to input questions and student answers, providing educators and researchers with insights into assessment accuracy and the quality of LLM-assessed rationales. Additionally, it offers advanced visualization and robust evaluation tools, enhancing the usability for educational assessment and facilitating efficient rationale verification.
Jiazheng Li 0002, Artem Bobrov, Cesare Aloisi, Yulan He 0001
AAAI1
2025 ExDDI: Explaining Drug-Drug Interaction Predictions with Natural Language
abstract
Predicting unknown drug-drug interactions (DDIs) is crucial for improving medication safety. Previous efforts in DDI prediction have typically focused on binary classification or predicting DDI categories, with the absence of explanatory insights that could enhance trust in these predictions. In this work, we propose to generate natural language explanations for DDI predictions, enabling the model to reveal the underlying pharmacodynamics and pharmacokinetics mechanisms simultaneously as making the prediction. To do this, we have collected DDI explanations from DDInter and DrugBank and developed various models for extensive experiments and analysis. Our models can provide accurate explanations for unknown DDIs between known drugs. This paper contributes new tools to the field of DDI prediction and lays a solid foundation for further research on generating explanations for DDI predictions.
Zhaoyue Sun, Jiazheng Li 0002, Gabriele Pergola, Yulan He 0001
AAAI2
2025 Drift: Enhancing LLM Faithfulness in Rationale Generation via Dual-Reward Probabilistic Inference
abstract
As Large Language Models (LLMs) are increasingly applied to complex reasoning tasks, achieving both accurate task performance and faithful explanations becomes crucial.However, LLMs often generate unfaithful explanations, partly because they do not consistently adhere closely to the provided context.Existing approaches to this problem either rely on superficial calibration methods, such as decomposed Chain-of-Thought prompting, or require costly retraining to improve model faithfulness.In this work, we propose a probabilistic inference paradigm that leverages taskspecific and lookahead rewards to ensure that LLM-generated rationales are more faithful to model decisions and align better with input context.These rewards are derived from a domainspecific proposal distribution, allowing for optimized sequential Monte Carlo approximations.Our evaluations across three different reasoning tasks show that this method, which allows for controllable generation during inference, improves both accuracy and faithfulness of LLMs.This method offers a promising path towards making LLMs more reliable for reasoning tasks without sacrificing performance.
Jiazheng Li 0002, Hanqi Yan, Yulan He 0001
ACL (1)1
2025 Two Heads Are Better Than One: Dual-Model Verbal Reflection at Inference-Time
abstract
Although preference optimization methods have improved reasoning performance in Large Language Models (LLMs), they often lack transparency regarding why one reasoning outcome is preferred over another.This limitation is especially critical in Automated Student Answer Scoring (ASAS), where explainability is essential to justify assessment outcomes.Verbal reinforcement learning offers the potential to generate explicit reflection, but it tends to produce superficial critiques that can harm assessment performance.Existing LLMs also struggle to reliably detect subtle reasoning errors in ASAS tasks.Moreover, manually identifying intermediate reasoning errors is expensive and difficult to scale.To address these challenges, we introduce a contrastive reflection synthesis pipeline that generates precise verbal feedback by identifying discrepancies in structure reasoning graph paths.Leveraging these synthetic reflection data, we propose DARS, a Dual-model Reflective Scoring framework featuring a dedicated Critic model trained for effective reflection.DARS achieves strong performance and consistently outperforms existing ASAS baselines across all evaluation metrics.Extensive experiments further provide novel insights into the value of reflection data, framework design, and the scaling behavior of DARS. 1 * Now at Google DeepMind. 1 We release the DARS code at https://github.com/ lijiazheng99/DARS.Student Answer: the mRNA, Is met by tRNA with anticodons, that link to the mRNA's codons.This tRNA carries an amino acid that goes with the mRNA's codon.This is done multiple times until there is a stranhd of amino acids.Ribosomes attach these amino acids together.Reasoner: The student's answer scored 3 points according to the marking rubric, which awards 3 points for addressing four key elements.The student successfully described mRNA exiting the nucleus, tRNA carrying complementary anticodons recognizing the mRNA codons, the corresponding amino acids on tRNA being bonded, and the continuation of amino acid linkage until a stop codon is reached, as indicated by '...tRNA with anticodons, that link to the mRNA's codons.' and '...until there is a strand of amino acids.... Critic: It appears there was a misunderstanding in your assessment of the student's answer.You credited the student for mentioning that the'corresponding amino acids on tRNA are bonded to adjacent tRNA's amino acids' and that 'amino acids continue to be linked until a STOP codon is read on the mRNA.'However, upon reviewing the student's response, these elements were not explicitly covered.The student's description primarily focused on the interaction between mRNA and tRNA and the eventual formation of a peptide bond, but did not detail the process of amino acids being bonded to adjacent tRNA's amino acids nor the specific mechanism of elongation ... Reasoner: The student's answer scored 2 points for achieving three key answer elements.The student accurately described the roles of mRNA and tRNA in protein synthesis, mentioning that 'the mRNA, Is met by tRNA with anticodons, that link to the mRNA's codons' which covers the key elements of mRNA bases being read in triplets and tRNA carrying the complementary anticodon recognizing the complementary codon of the mRNA.Additionally, the student implied the bonding of amino acids by stating 'Ribosomes attach these amino acids together.'However, the answer did not explicitly mention the exit of mRNA from the nucleus, or the elongation and termination at a STOP codon which resulted in losing points for these key elements. Critic: Rationale Looks Good! [STOP]
Jiazheng Li 0002, Junru Lu, Gladys Tyen, Lin Gui 0003, Cesare Aloisi, Yulan He 0001
EMNLP1
2024 NarrativePlay: An Automated System for Crafting Visual Worlds in Novels for Role-Playing
abstract
In this demo, we present NarrativePlay -- an innovative system enabling users to role-play a fictional character and interact with dynamically generated narrative environments. Unlike existing predefined sandbox approaches, NarrativePlay centres around the main storyline events extracted from the narrative, allowing users to experience the story from the perspective of a character they chose. To design versatile AI agents for diverse scenarios, we employ a framework built on a Large Language Models (LLMs) to extract detailed character traits from text. We also incorporate automatically generated visual displays of narrative settings, character portraits, and character speech, greatly enhancing the overall user experience.
Runcong Zhao, Jiazheng Li 0002, Lixing Zhu, Yanran Li, Yulan He 0001, Lin Gui 0003
AAAI3
2024 Eliminating Biased Length Reliance of Direct Preference Optimization via Down-Sampled KL Divergence
abstract
Direct Preference Optimization (DPO) has emerged as a prominent algorithm for the direct and robust alignment of Large Language Models (LLMs) with human preferences, offering a more straightforward alternative to the complex Reinforcement Learning from Human Feedback (RLHF).Despite its promising efficacy, DPO faces a notable drawback: "verbosity", a common over-optimization phenomenon also observed in RLHF.While previous studies mainly attributed verbosity to biased labels within the data, we propose that the issue also stems from an inherent algorithmic length reliance in DPO.Specifically, we suggest that the discrepancy between sequencelevel Kullback-Leibler (KL) divergences between chosen and rejected sequences, used in DPO, results in overestimated or underestimated rewards due to varying token lengths.Empirically, we utilize datasets with different label lengths to demonstrate the presence of biased rewards.We then introduce an effective downsampling approach, named SamPO, to eliminate potential length reliance.Our experimental evaluations, conducted across three LLMs of varying scales and a diverse array of conditional and open-ended benchmarks, highlight the efficacy of SamPO in mitigating verbosity, achieving improvements of 5% to 12% over DPO through debaised rewards 1 .
Junru Lu, Jiazheng Li 0002, Siyu An, Yulan He 0001, Xing Sun 0001
EMNLP2
2024 The Mystery of In-Context Learning: A Comprehensive Survey on Interpretation and Analysis
abstract
Understanding in-context learning (ICL) capability that enables large language models (LLMs) to excel in proficiency through demonstration examples is of utmost importance.This importance stems not only from the better utilization of this capability across various tasks, but also from the proactive identification and mitigation of potential risks, including concerns regarding truthfulness, bias, and toxicity, that may arise alongside the capability.In this paper, we present a thorough survey on the interpretation and analysis of in-context learning.First, we provide a concise introduction to the background and definition of in-context learning.Then, we give an overview of advancements from two perspectives: 1) the theoretical perspective, emphasizing studies on mechanistic interpretability and delving into the mathematical foundations behind ICL; and 2) the empirical perspective, concerning studies that empirically analyze factors associated with ICL.We conclude by discussing open questions and the challenges encountered and, by suggesting potential avenues for future research.We believe that our work establishes the basis for further exploration into the interpretation of incontext learning.To aid this effort, we have created a repository 1 containing resources that will be continually updated.1 https://github.com/zyxnlp
Jiazheng Li 0002, Yanzheng Xiang, Hanqi Yan, Lin Gui 0003, Yulan He 0001
EMNLP2
2023 CUE: An Uncertainty Interpretation Framework for Text Classifiers Built on Pre-Trained Language Models
abstract
Text classifiers built on Pre-trained Language Models (PLMs) have achieved remarkable progress in various tasks including sentiment analysis, natural language inference, and question-answering. However, the occurrence of uncertain predictions by these classifiers poses a challenge to their reliability when deployed in practical applications. Much effort has been devoted to designing various probes in order to understand what PLMs capture. But few studies have delved into factors influencing PLM-based classifiers’ predictive uncertainty. In this paper, we propose a novel framework, called CUE, which aims to interpret uncertainties inherent in the predictions of PLM-based models. In particular, we first map PLM-encoded representations to a latent space via a variational auto-encoder. We then generate text representations by perturbing the latent space which causes fluctuation in predictive uncertainty. By comparing the difference in predictive uncertainty between the perturbed and the original text representations, we are able to identify the latent dimensions responsible for uncertainty and subsequently trace back to the input features that contribute to such uncertainty. Our extensive experiments on four benchmark datasets encompassing linguistic acceptability classification, emotion classification, and natural language inference show the feasibility of our proposed framework. Our source code is available at https://github.com/lijiazheng99/CUE.
Jiazheng Li 0002, Zhaoyue Sun, Bin Liang 0004, Lin Gui 0003, Yulan He 0001
UAI1
2022 NumHTML: Numeric-Oriented Hierarchical Transformer Model for Multi-Task Financial Forecasting
abstract
Financial forecasting has been an important and active area of machine learning research because of the challenges it presents and the potential rewards that even minor improvements in prediction accuracy or forecasting may entail. Traditionally, financial forecasting has heavily relied on quantitative indicators and metrics derived from structured financial statements. Earnings conference call data, including text and audio, is an important source of unstructured data that has been used for various prediction tasks using deep earning and related approaches. However, current deep learning-based methods are limited in the way that they deal with numeric data; numbers are typically treated as plain-text tokens without taking advantage of their underlying numeric structure. This paper describes a numeric-oriented hierarchical transformer model (NumHTML) to predict stock returns, and financial risk using multi-modal aligned earnings calls data by taking advantage of the different categories of numbers (monetary, temporal, percentages etc.) and their magnitude. We present the results of a comprehensive evaluation of NumHTML against several state-of-the-art baselines using a real-world publicly available dataset. The results indicate that NumHTML significantly outperforms the current state-of-the-art across a variety of evaluation metrics and that it has the potential to offer significant financial gains in a practical trading context.
Linyi Yang, Jiazheng Li 0002, Ruihai Dong, Yue Zhang 0004, Barry Smyth
AAAI2
2022 PHEE: A Dataset for Pharmacovigilance Event Extraction from Text
abstract
The primary goal of drug safety researchers and regulators is to promptly identify adverse drug reactions.Doing so may in turn prevent or reduce the harm to patients and ultimately improve public health.Evaluating and monitoring drug safety (i.e., pharmacovigilance) involves analyzing an ever growing collection of spontaneous reports from health professionals, physicians, and pharmacists, and information voluntarily submitted by patients.In this scenario, facilitating analysis of such reports via automation has the potential to rapidly identify safety signals.Unfortunately, public resources for developing natural language models for this task are scant.We present PHEE, a novel dataset for pharmacovigilance comprising over 5000 annotated events from medical case reports and biomedical literature, making it the largest such public dataset to date.We describe the hierarchical event schema designed to provide coarse and fine-grained information about patients' demographics, treatments and (side) effects.Along with the discussion of the dataset, we present a thorough experimental evaluation of current state-of-the-art approaches for biomedical event extraction, point out their limitations, and highlight open challenges to foster future research in this area 1 .
Zhaoyue Sun, Jiazheng Li 0002, Gabriele Pergola, Byron C. Wallace, Bino John, Nigel Greene, Joseph Kim, Yulan He 0001
EMNLP2
2022 AQP: an open modular Python platform for objective speech and audio quality metrics
abstract
Audio quality assessment has been widely researched in the signal processing area. Full-reference objective metrics (e.g., POLQA, ViSQOL) have been developed to estimate the audio quality relying only on human rating experiments. To evaluate the audio quality of novel audio processing techniques, researchers constantly need to compare objective quality metrics. Testing different implementations of the same metric and evaluating new datasets are fundamental and ongoing iterative activities. In this paper, we present AQP - an open-source, node-based, light-weight Python pipeline for audio quality assessment. AQP allows researchers to test and compare objective quality metrics helping to improve robustness, reproducibility and development speed. We introduce the platform, explain the motivations, and illustrate with examples how, using AQP, objective quality metrics can be (i) compared and benchmarked; (ii) prototyped and adapted in a modular fashion; (iii) visualised and checked for errors. The code has been shared on GitHub to encourage adoption and contributions from the community.
Jack Geraghty, Jiazheng Li 0002, Alessandro Ragano, Andrew Hines
MMSys2
2021 Exploring the Efficacy of Automatically Generated Counterfactuals for Sentiment Analysis
abstract
Linyi Yang, Jiazheng Li, Padraig Cunningham, Yue Zhang, Barry Smyth, Ruihai Dong. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Linyi Yang, Jiazheng Li 0002, Padraig Cunningham, Yue Zhang 0004, Barry Smyth, Ruihai Dong
ACL/IJCNLP (1)2
2020 MAEC: A Multimodal Aligned Earnings Conference Call Dataset for Financial Risk Prediction
abstract
In the area of natural language processing, various financial datasets have informed recent research and analysis including financial news, financial reports, social media, and audio data from earnings calls. We introduce a new, large-scale multi-modal, text-audio paired, earnings-call dataset named MAEC, based on S&P 1500 companies. We describe the main features of MAEC, how it was collected and assembled, paying particular attention to the text-audio alignment process used. We present the approach used in this work as providing a suitable framework for processing similar forms of data in the future. The resulting dataset is more than six times larger than those currently available to the research community and we discuss its potential in terms of current and future research challenges and opportunities. All resources of this work are available at https://github.com/Earnings-Call-Dataset/
Jiazheng Li 0002, Linyi Yang, Barry Smyth, Ruihai Dong
CIKM1