Rujing Yao

dblp:189/6903 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
7since 2021 · last 2025
0009-0001-2738-5786ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Deep Interaction Timing: FOL-based Complexity Differentiation for Legal Queries
abstract
The emergence of large language models (LLMs) has made legal consultation resources more accessible. However, due to the inherent ability of LLMs to response to any input, in practical use, if the query itself is incomplete or complex, the model’s response may generate hallucinations, misleading the user. In fact, LLMs are best suited to answer complete and simple legal questions, while complex issues should be handled by legal experts. Therefore, it is essential to route queries as simple or complex before the LLM provides an answer. Yet, current approaches rely solely on internal confidence metrics, overlooking the inherent complexity of legal queries. To address this limitation, we propose a novel method that incorporates first-order logic (FOL) rules as additional evidence to assess the complexity of queries. Our approach consists of two key stages: FOL-based Matching and Inference-driven Routing. In the first stage, LLMs extract key information from user inputs and map it to a pool of FOL rules to match relevant legal evidence. In the second stage, symbolic reasoning is applied to the matched evidence to derive logical inferences and make routing decisions. We fine-tune a 7B model (Qwen2-7B-Instruct) to combine FOL generation with query routing, enhancing the model’s reasoning capabilities and interpretability. This approach achieves a 24.77% improvement.
Tong Zhang 0005, Yiquan Wu 0001, Rujing Yao, Changlong Sun, Xiaozhong Liu 0001
ICAIL3
2025 Elevating Legal LLM Responses: Harnessing Trainable Logical Structures and Semantic Knowledge with Legal Reasoning
abstract
Rujing Yao, Yang Wu, Chenghao Wang, Jingwei Xiong, Fang Wang, Xiaozhong Liu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Rujing Yao, Jingwei Xiong, Fang Wang 0014, Xiaozhong Liu 0001
NAACL (Long Papers)1
2025 Data Optimization in Deep Learning: A Survey
abstract
Large-scale, high-quality data are considered an essential factor for the successful application of many deep learning techniques. Meanwhile, numerous real-world deep learning tasks still have to contend with the lack of sufficient amounts of high-quality data. Additionally, issues such as model robustness, fairness, and trustworthiness are also closely related to training data. Consequently, a huge number of studies in the existing literature have focused on the data aspect in deep learning tasks. Some typical data optimization techniques include data augmentation, logit perturbation, sample weighting, and data condensation. These techniques usually come from different deep learning divisions and their theoretical inspirations or heuristic motivations may seem unrelated to each other. This study aims to organize a wide range of existing data optimization methodologies for deep learning from the previous literature, and makes the effort to construct a comprehensive taxonomy for them. The constructed taxonomy considers the diversity of split dimensions, and deep sub-taxonomies are constructed for each dimension. On the basis of the taxonomy, connections among the extensive data optimization methods for deep learning are built in terms of five aspects. We probe into rendering several promising and interesting future directions. The constructed taxonomy and the revealed connections will enlighten the better understanding of existing methods and the design of novel data optimization techniques. Furthermore, our aspiration for this survey is to promote data optimization as an independent subdivision of deep learning.
Ou Wu 0001, Rujing Yao
IEEE Trans. Knowl. Data Eng.2
2024 A Taxonomy for Learning with Perturbation and Algorithms
abstract
Weighting strategy prevails in machine learning. For example, a common approach in robust machine learning is to exert low weights on samples which are likely to be noisy or quite hard. This study summarizes another less-explored strategy, namely, perturbation. Various incarnations of perturbation have been utilized but it has not been explicitly revealed. Learning with perturbation is called perturbation learning and a systematic taxonomy is constructed for it in this study. In our taxonomy, learning with perturbation is divided on the basis of the perturbation targets, directions, inference manners, and granularity levels. Many existing learning algorithms including some classical ones can be understood with the constructed taxonomy. Alternatively, these algorithms share the same component, namely, perturbation in their procedures. Furthermore, a family of new learning algorithms can be obtained by varying existing learning algorithms with our taxonomy. Specifically, three concrete new learning algorithms are proposed for robust machine learning. Extensive experiments on image classification and text sentiment analysis verify the effectiveness of the three new algorithms. Learning with perturbation can also be used in other various learning scenarios, such as imbalanced learning, clustering, regression, and so on.
Rujing Yao, Ou Wu 0001
ACM Trans. Knowl. Discov. Data1
2023 Exploring developments of the AI field from the perspective of methods, datasets, and metrics
Rujing Yao, Yingchun Ye, Ji Zhang 0001, Shuxiao Li, Ou Wu 0001
Inf. Process. Manag.1
2022 Method and dataset entity mining in scientific literature: A CNN + BiLSTM model with self-attention
Linlin Hou, Ji Zhang 0001, Ou Wu 0001, Ting Yu 0004, Zhen Wang 0037, Zhao Li 0007, Jianliang Gao, Yingchun Ye, Rujing Yao
Knowl. Based Syst.9
2022 Deep human answer understanding for natural reverse QA
Rujing Yao, Linlin Hou, Jie Gui, Ou Wu 0001
Knowl. Based Syst.1
2019 Method and Dataset Mining in Scientific Papers
abstract
Literature analysis facilitates researchers better understanding the development of science and technology. The conventional literature analysis focuses on the topics, authors, abstracts, keywords, references, etc., and rarely pays attention to the content of papers. In the field of machine learning, the involved methods (M) and datasets (D) are key information in papers. The extraction and mining of M and D are useful for discipline analysis and algorithm recommendation. In this paper, we propose a novel entity recognition model, called MDER, and constructe datasets from the papers of the PAKDD conferences (2009-2019). Some preliminary experiments are conducted to assess the extraction performance and the mining results are visualized.
Rujing Yao, Linlin Hou, Yingchun Ye, Ji Zhang 0001, Jian Wu 0006
IEEE BigData1