Yuquan Le

dblp:220/7267 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0001-6283-9037ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 On Imbalance in Case Types: Evaluating and Enhancing PLMs for Criminal Court View Generation
abstract
The criminal court view generation (CCVG) task aims to produce succinct and coherent summaries of fact descriptions, providing interpretable opinions for verdicts. Traditional text generation evaluation metrics, such as ROUGE, BLEU, and BERTSCORE, are extensively employed for this task and measure performance by averaging the assessment scores of all samples within the test set. However, these sample-averaged metrics encounter two primary dilemmas: 1) they fail to fairly assess overall evaluation scores across different case types and 2) they overlook the measurement of the degree of performance imbalance between case types. To fill this research gap, we propose two novel case-type-oriented evaluation metrics: Case-type-oriented Text Generation (CTG) and Case-type-oriented Imbalance Performance (CIP). First, CTG mitigates the unfair assessment among different case types by assigning equal weight to each type. Second, CIP evaluates performance imbalance by measuring the distance between the performance of each case type and the overall performance. We provide three theorems to elucidate the properties of CIP, demonstrating that CIP can effectively identify the extent to which a CCVG model achieves balanced generation performance across different case types. Furthermore, we propose an embarrassingly simple and effective charge-guided encoder-decoder (CGED) framework to enhance performance fairly across different case types in encoder-decoder pretrained language models (PLMs). Code is available at https://yuquanle.github.io/Case-type-oriented-metrics-homepage/.
Yuquan Le, Yan Ding 0004, Chng Eng Siong, Kenli Li 0001
IEEE Trans. Neural Networks Learn. Syst.1
2025 TermDiffuSum: A Term-guided Diffusion Model for Extractive Summarization of Legal Documents
abstract
Extractive summarization for legal documents aims to automatically extract key sentences from legal texts to form concise summaries. Recent studies have explored diffusion models for extractive summarization task, showcasing their remarkable capabilities. Despite these advancements, these models often fall short in effectively capturing and leveraging the specialized legal terminology crucial for accurate legal summarization. To address the limitation, this paper presents a novel term-guided diffusion model for extractive summarization of legal documents, named TermDiffuSum. It incorporates legal terminology into the diffusion model via a well-designed multifactor fusion noise weighting schedule, which allocates higher attention weight to sentences containing a higher concentration of legal terms during the diffusion process. Additionally, TermDiffuSum utilizes a re-ranking loss function to refine the model’s selection of more relevant summaries by leveraging the relationship between the candidate summaries generated by the diffusion process and the reference summaries. Experimental results on a self-constructed legal summarization dataset reveal that TermDiffuSum outperforms existing diffusion-based summarization models, achieving improvements of 3.10 in ROUGE-1, 2.84 in ROUGE-2, and 2.89 in ROUGE-L. To further validate the generalizability of TermDiffuSum, we conduct experiments on three public datasets from news and social media domains, with results affirming the scalability of our approach.
Xiangyun Dong, Yuquan Le, Zhangyue Jiang, Junxi Zhong
COLING3
2025 Neural Causal Graph for Interpretable and Intervenable Classification
abstract
Advancements in neural networks have significantly enhanced the performance of classification models, achieving remarkable accuracy across diverse datasets. However, these models often lack transparency and do not support interactive reasoning with human users, which are essential attributes for applications that require trust and user engagement. To overcome these limitations, we introduce an innovative framework, Neural Causal Graph (NCG), that integrates causal inference with neural networks to enable interpretable and intervenable reasoning. We then propose an intervention training method to model the intervention probability of the prediction, serving as a contextual prompt to facilitate the fine-grained reasoning and human-AI interaction abilities of NCG. Our experiments show that the proposed framework significantly enhances the performance of traditional classification baselines. Furthermore, NCG achieves nearly 95\% top-1 accuracy on the ImageNet dataset by employing a test-time intervention method. This framework not only supports sophisticated post-hoc interpretation but also enables dynamic human-AI interactions, significantly improving the model's transparency and applicability in real-world scenarios.
Jiawei Wang 0025, Shaofei Lu, Da Cao, Yuquan Le, Zhe Quan, Tat-Seng Chua
ICLR5
2025 Spatial-temporal video grounding with cross-modal understanding and enhancement
Shu Luo, Jingyu Pan, Da Cao, Jiawei Wang 0025, Yuquan Le, Meng Liu 0006
Expert Syst. Appl.5
2025 Graph Reasoning With Supervised Contrastive Learning for Legal Judgment Prediction
abstract
Given the fact descriptions of legal cases, the legal judgment prediction (LJP) problem aims to determine three judgment tasks of law articles, charges, and the term of penalty. Most existing studies have considered task dependencies while neglecting the prior dependencies of labels among different tasks. Therefore, how to make better use of the information on the relation dependencies among tasks and labels becomes a crucial issue. To this end, we transform the text classification problem into a node classification framework based on graph reasoning and supervised contrastive learning (SCL) techniques, named GraSCL. Specifically, we first design a graph reasoning network to model the potential dependency structures and facilitate relational learning under various graph topologies. Then, we introduce the SCL method for the LJP task to further leverage the label relation on the graph. To accommodate the node classification settings, we extend the traditional SCL method to novel variants for SCL at the node level, which allows the GraSCL framework to be trained efficiently even with small batches. Furthermore, to recognize the importance of hard negative samples in contrastive learning, we introduce a simple yet effective technique called online hard negative mining (OHNM) to enhance our SCL approach. This technique complements our SCL method and enables us to control the number and complexity of negative samples, leading to further improvements in the model's performance. Finally, extensive experiments are conducted on two well-known benchmarks, demonstrating the effectiveness and rationality of our proposed SCL approach as compared to the state-of-the-art competitors.
Jiawei Wang 0025, Yuquan Le, Da Cao, Shaofei Lu, Zhe Quan, Meng Wang 0001
IEEE Trans. Neural Networks Learn. Syst.2
2024 Topology-aware Multi-task Learning Framework for Civil Case Judgment Prediction
Yuquan Le, Sheng Xiao, Kenli Li 0001
Expert Syst. Appl.1
2024 $\boldsymbol{R}^{2}$: A Novel Recall & Ranking Framework for Legal Judgment Prediction
abstract
The legal judgment prediction (LJP) task is to automatically decide appropriate law articles, charges, and term of penalty for giving the fact description of a law case. It considerably influences many real legal applications and has thus attracted the attention of legal practitioners and AI researchers in recent years. In real scenarios, many confusing charges are encountered, which makes LJP challenging. Intuitively, for a controversial legal case, legal practitioners usually first obtain various possible judgment results as candidates based on the fact description of the case; then these candidates generally need to be carefully considered based on the facts and the rationality of the candidates. Inspired by this observation, this paper presents a novelRecall &Ranking framework, dubbed as$\boldsymbol{R}^{2}$, which attempts to formalize LJP as a two-stage problem. The recall stage is designed to collect high-likelihood judgment results for a given case; these results are regarded as candidates for the ranking stage. The ranking stage introduces a verification technique to learn the relationships between the fact description and the candidates. It treats the partially correct candidates as semi-negative samples, and thus has a certain ability to distinguish confusing candidates. Moreover, we devise a comprehensive judgment strategy to refine the final judgment results by comprehensively considering the rationality of multiple probable candidates. We carry out numerous experiments on two widely used benchmark datasets. The experimental results demonstrate our proposed approach's effectiveness compared to the other competitive baselines.
Yuquan Le, Zhe Quan, Jiawei Wang 0025, Da Cao, Kenli Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 Deconfounded Multimodal Learning for Spatio-temporal Video Grounding
abstract
The task of spatio-temporal video grounding involves identifying the spatial and temporal regions in a video that correspond to the objects or actions described in a given textual description. However, current models used for spatio-temporal video grounding often rely heavily on spatio-temporal priors to make the predictions. As a result, they may suffer from spurious correlations and lack the ability to generalize well to new or diverse scenarios. To overcome this limitation, we introduce a deconfounded multimodal learning framework, which utilizes a structural causal model to treat dataset biases as a confounder and subsequently remove their confounding effect. Through this framework, we can perform causal intervention on the multimodal input and derive an unbiased estimation formula through the do-calculus technique. In order to tackle the challenge of diverse and often unobservable confounders, we further propose a novel retrieval-based approach with a causal mask mechanism. The proposed method leverages analogical reasoning to facilitate deconfounded learning and mitigate dataset biases, enabling unbiased spatio-temporal prediction without explicitly modeling the confounding factors. Extensive experiments on two challenging benchmarks have well verified the effectiveness and rationality of our proposed solution.
Jiawei Wang 0025, Zhanchang Ma, Da Cao, Yuquan Le, Junbin Xiao, Tat-Seng Chua
ACM Multimedia4
2023 Loader: A Log Anomaly Detector Based on Transformer
abstract
Detecting anomalies in logs is crucial for service and system management, since logs are widely used to record the runtime status, and are often the only data available for postmortem analysis. Since anomalies are usually rare in real-world services and systems, a common and feasible practice is to mine or learn normal patterns from logs, and deem those violating the normal patterns as anomalies. As log sequences are a kind of time series data, RNN (Recurrent Neural Network) and its variants have been extensively employed to capture the normal patterns. Nevertheless, the sequential nature of RNN and its variants makes them hard to parallelize and capture long-term dependencies, which may hinder their performance. To address this issue, in this paper we propose Loader, a novel semi-supervisedloganomalydetector based on Transformer, because the Transformer architecture eschews recurrence and is able to draw global dependencies. Loader leverages the Transformer encoder to capture normal patterns from normal log sequences. When detecting, it gives a set of candidate log templates, that may appear after the input log substring under normal conditions. If the template of the actual next log message is not within the candidate set, this implies an anomaly. Previous similar methods select the most possible$k$log templates as candidates in any case, so the performance is sensitive to$k$, and it is nontrivial to pick a proper$k$. To alleviate this, we design a more flexible and robust ‘top-$p$’ algorithm, which determines the candidate set based on the cumulative probability of the most possible log templates. Extensive experiments are conducted based on three public log datasets, the experimental results validate the effectiveness and competitiveness of our approach.
Tong Xiao 0002, Zhe Quan, Zhi-Jie Wang 0009, Yuquan Le, Yunfei Du 0001, Xiangke Liao, Kenli Li 0001, Keqin Li 0001
IEEE Trans. Serv. Comput.4
2022 Legal Charge Prediction via Bilinear Attention Network
abstract
The legal charge prediction task aims to judge appropriate charges according to the given fact description in cases. Most existing methods formulate it as a multi-class text classification problem and have achieved tremendous progress. However, the performance on low-frequency charges is still unsatisfactory. Previous studies indicate leveraging the charge label information can facilitate this task, but the approaches to utilizing the label information are not fully explored. In this paper, inspired by the vision-language information fusion techniques in the multi-modal field, we propose a novel model (denoted as LeapBank) by fusing the representations of text and labels to enhance the legal charge prediction task. Specifically, we devise a representation fusion block based on the bilinear attention network to interact the labels and text tokens seamlessly. Extensive experiments are conducted on three real-world datasets to compare our proposed method with state-of-the-art models. Experimental results show that LeapBank obtains up to 8.5% Macro-F1 improvements on the low-frequency charges, demonstrating our model's superiority and competitiveness.
Yuquan Le, Meng Chen 0006, Zhe Quan, Xiaodong He 0001, Kenli Li 0001
CIKM1
2022 Building Robust Spoken Language Understanding by Cross Attention Between Phoneme Sequence and ASR Hypothesis
abstract
Building Spoken Language Understanding (SLU) robust to Automatic Speech Recognition (ASR) errors is an essential issue for various voice-enabled virtual assistants. Considering that most ASR errors are caused by phonetic confusion between similar-sounding expressions, intuitively, leveraging the phoneme sequence of speech can complement ASR hypothesis and enhance the robustness of SLU. This paper proposes a novel model with Cross Attention for SLU (denoted as CASLU). The cross attention block is devised to catch the fine-grained interactions between phoneme and word embeddings in order to make the joint representations catch the phonetic and semantic features of input simultaneously and for overcoming the ASR errors in downstream natural language understanding (NLU) tasks. Extensive experiments are conducted on three datasets, showing the effectiveness and competitiveness of our approach. Additionally, We also validate the universality of CASLU and prove its complementarity when combining with other robust SLU techniques.
Zexun Wang, Yuquan Le, Mingchao Feng, Meng Chen 0006, Xiaodong He 0001
ICASSP2
2022 IACN: Interactive attention capsule network for similar case matching
abstract
The similar case matching task aims to detect which two cases are more similar for a given triplet. It plays a significant role in the legal industry and thus has gained much attention. Due to the rapid development of natural language processing technology, various deep learning techniques have been applied to similar case matching task and obtained attractive performance. Most existing researches usually focus on encoding legal documents into a continuous vector. However, a unified vector is difficult to model multiple elements of the case. In the real world, cases contain numerous elements, which are the basis for legal practitioners to judge the similarity among cases. Legal experts usually focus on whether the two cases have similar legal elements. It makes this task especially challenging. In this paper, we propose a novel model, namely Interactive Attention Capsule Network (dubbed as IACN). It attempts to simulate the process of judgment by legal experts, which captures fine-grained elements similarity to make an interpretable judgment. In other words, the IACN judges the similarity of the case pairs based on the legal elements. The more similar legal elements of a case pair, the higher the degree of similarity of the case pair. In addition, we devise an interactive dynamic routing mechanism, which can better learn the interactive representation of legal elements among cases than the vanilla dynamic routing. We conduct extensive experiments based on a real-world dataset. The experimental results consistently demonstrate the superiorities and competitiveness of our proposed model.
Yuquan Le, Jiawei He 0003
Intell. Data Anal.3
2022 Enhancing N-Gram Based Metrics with Semantics for Better Evaluation of Abstractive Text Summarization
Jiawei He 0003, Guo-Bang Chen, Yuquan Le, Xiaofei Ding
J. Comput. Sci. Technol.4
2020 Learning to Predict Charges for Legal Judgment via Self-Attentive Capsule Network
abstract
With the rapid development of deep learning technology, more and more traditional industries are changed by Artificial Intelligence. The legal industry is such a popular scenario which attracts lots of researchers' interests. In this work, we focus on automatic charge prediction, which predicts the final charges according to the given fact descriptions in criminal cases. It is crucial for legal assistant systems and can help the judges improve work efficiency greatly. However, extremely imbalanced data distribution and lengthy fact descriptions make this task especially challenging. To tackle these two issues, we propose a novel model, namely Self-Attentive Capsule Network (dubbed as SAttCaps). In particular, we devise a self-attentive dynamic routing, which can not only capture long-range dependency more directly than vanilla dynamic routing, but also learn the high-level generalized features better. The experimental results on three real-world datasets demonstrate that our model significantly outperforms the baselines and creates new state-of-the-art performance. Moreover, our model performs much better than the baselines especially in the low-frequency charges and can bring 5.7% absolute improvement under F1 score.
Yuquan Le, Congqing He, Meng Chen 0006, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001
ECAI1
2019 SECaps: A Sequence Enhanced Capsule Model for Charge Prediction
Congqing He, Yuquan Le, Jiawei He 0003
ICANN (4)3
2019 Dynamically Weighted Multi-View Semi-Supervised Learning for CAPTCHA
Congqing He, Yuquan Le, Jiawei He 0003
PAKDD (2)3
2019 An Efficient Framework for Sentence Similarity Modeling
abstract
Sentence similarity modeling lies at the core of many natural language processing applications, and thus has received much attention. Owing to the success of word embeddings, recently, popular neural network methods achieved sentence embedding. Most of them focused on learning semantic information and modeling it as a continuous vector, yet the syntactic information of sentences has not been fully exploited. On the other hand, prior works have shown the benefits of structured trees that include syntactic information, while few methods in this branch utilized the advantages of word embeddings and another powerful technique-attention weight mechanism. This paper suggests to absorb their advantages by merging these techniques in a unified structure, dubbed as attention constituency vector tree (ACVT). Meanwhile, this paper develops a new tree kernel, known as ACVT kernel, which is tailored for sentence similarity measure based on the proposed structure. The experimental results, based on 19 widely used semantic textual similarity datasets, demonstrate that our model is effective and competitive, when compared against state-of-the-art models. Additionally, the experimental results validate that many attention weight mechanisms and word embedding techniques can be seamlessly integrated into our model, demonstrating the robustness and universality of our model.
Zhe Quan, Zhi-Jie Wang 0009, Yuquan Le, Bin Yao 0002, Kenli Li 0001, Jian Yin 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2018 ACV-tree: A New Method for Sentence Similarity Modeling
abstract
Sentence similarity modeling lies at the core of many natural language processing applications, and thus has received much attention. Owing to the success of word embeddings, recently, popular neural network methods have achieved sentence embedding, obtaining attractive performance. Nevertheless, most of them focused on learning semantic information and modeling it as a continuous vector, while the syntactic information of sentences has not been fully exploited. On the other hand, prior works have shown the benefits of structured trees that include syntactic information, while few methods in this branch utilized the advantages of word embeddings and another powerful technique ? attention weight mechanism. This paper makes the first attempt to absorb their advantages by merging these techniques in a unified structure, dubbed as ACV-tree. Meanwhile, this paper develops a new tree kernel, known as ACVT kernel, that is tailored for sentence similarity measure based on the proposed structure. The experimental results, based on 19 widely-used datasets, demonstrate that our model is effective and competitive, compared against state-of-the-art models.
Yuquan Le, Zhi-Jie Wang 0009, Zhe Quan, Jiawei He 0003, Bin Yao 0002
IJCAI1
2017 An Improved LDA Multi-document Summarization Model Based on TensorFlow
abstract
Latent Dirichlet Allocation (LDA), has been recently used to automatically generate text corpora topics, and applied to sentences extraction based multi-document summarization algorithms. In this paper, we propose a novel approach to automatic generation of aspect-oriented summaries from multiple documents. Our approach is to combine the traditional summary generation algorithm and the the abstract generation algorithm based on deep learning.We employ the improved traditional summary generation algorithm to convert multiple documents into a single document, and then using the resulting single document with the deep learning method to extract the final summary. At first, we apply improved LDA model to cluster sentences in all documents. Second, We employ the extended LexRank algorithm to sort the sentences in each cluster. Third, we use extended Hedge Trimmer algorithm for sentence compression. Fourth, We apply Integer Linear Programming for sentence selection, and in this step ,we get the single document. Finally, We employ the textum on TensorFlow to get the final abstract. The experiments showed that the proposed algorithm achieved better performance compared the other state-of-the-art algorithms on DUC2005 and TAC2010 corpus.
Ying Zhong 0008, Zhuo Tang, Xiaofei Ding, Yuquan Le, Kenli Li 0001, Keqin Li 0001
ICTAI5