Jingyi Ren

dblp:261/2607 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Language models and text generation · 51% Trustworthy machine learning · 23% Question answering and dialogue systems · 20%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%
Software engineering, system software, and programming languages
1 paper
Empirical software engineering · 100%

Topics — the 7 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
retrieval-augmented generation
1.012026
UR² : Unify RAG and Reasoning through Reinforcement Learning · ACL (1) 2026
Machine learning › Trustworthy machine learning
uncertainty estimation
1.012026
Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty · ACL (1) 2026
Natural language and speech › Question answering and dialogue systems › interactive question answering
conversational question answering
0.912025
Understanding Large Language Model Performance in Software Engineering: A Large-scale Question Answering Benchmark · SIGIR 2025
Bioinformatics and computational biology › molecular property prediction
peptide toxicity prediction
0.912025
An innovative peptide toxicity prediction model based on multi-scale convolutional neural network and residual connection · Bioinform. 2025
Empirical software engineering › benchmarking
software engineering benchmarks
0.912025
Understanding Large Language Model Performance in Software Engineering: A Large-scale Question Answering Benchmark · SIGIR 2025
Machine learning › Reinforcement learning
reinforcement learning for reasoning
0.312026
UR² : Unify RAG and Reasoning through Reinforcement Learning · ACL (1) 2026
Natural language and speech › Language models and text generation
large language model evaluation
0.312025
Understanding Large Language Model Performance in Software Engineering: A Large-scale Question Answering Benchmark · SIGIR 2025

Methods — techniques the papers use, named apart from their topics

benchmark construction · 1.7uncertainty quantification · 1.0reinforcement learning · 1.0word2vec · 0.9residual connections · 0.9multi-scale convolutional neural network · 0.9bidirectional long short-term memory · 0.9SMOTE · 0.9
YearPublicationVenuePosition
2026 UR² : Unify RAG and Reasoning through Reinforcement Learning
abstract
Weitao Li, Boran Xiang, Xiaolong Wang, Jingyi Ren, Ante Wang, Zhinan Gou, Weizhi Ma, Yang Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Boran Xiang, Jingyi Ren, Ante Wang, Zhinan Gou, Weizhi Ma
ACL (1)4
2026 Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty
abstract
Jingyi Ren, Ante Wang, Yunghwei Lai, Xiaolong Wang, Linlu Gong, Weitao Li, Weizhi Ma, Yang Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jingyi Ren, Ante Wang, Yunghwei Lai, Xiaolong Wang 0014, Linlu Gong, Weizhi Ma, Yang Liu 0005
ACL (1)1
2025 Understanding Large Language Model Performance in Software Engineering: A Large-scale Question Answering Benchmark
abstract
In this work, we introduce CodeRepoQA, a large-scale benchmark specifically designed for evaluating repository-level question-answering capabilities in the field of software engineering. CodeRepoQA encompasses five programming languages and covers a wide range of scenarios, enabling comprehensive evaluation of language models.To construct this dataset, we crawl data from 30 well-known repositories in GitHub, the largest platform for hosting and collaborating on code, and carefully filter the raw data.In total, CodeRepoQA is a multi-turn question-answering benchmark with 585,687 entries. It covers a diverse array of software engineering scenarios, with an average of 6.62 dialogue turns per entry.
Ruida Hu, Chao Peng 0002, Jingyi Ren, Xiangxin Meng, Qinyun Wu, Xinchen Wang 0001, Cuiyun Gao 0001
SIGIR3
2025 An innovative peptide toxicity prediction model based on multi-scale convolutional neural network and residual connection
abstract
MOTIVATION: Peptide toxicity is a critical concern in the development of peptide-based therapeutics, as toxic peptides can lead to severe side effects, including organ damage, immune reactions, and cytotoxicity. Predicting peptide toxicity accurately is essential to ensure the safety and efficacy of these drugs. RESULTS: In this study, we propose a novel model, ToxMSRC, to predict peptide toxicity using a combination of the continuous bag of words (CBOW) method from word2vec, synthetic minority over-sampling technique (SMOTE), multi-scale convolutional neural networks (CNN), and bidirectional long short-term memory (BiLSTM). This approach addresses the challenge of data imbalance by augmenting positive samples and improves feature extraction through multi-scale convolution. Furthermore, the model incorporates a residual connection that helps prevent overfitting and enhances generalization ability, improving classification performance. The model is evaluated on benchmark and independent test sets, achieving BACC scores of 92.17% on independent test1 and 86.89% on independent test2, outperforming existing state-of-the-art models. Additionally, ToxMSRC provides valuable insights into the relationship between peptide toxicity and amino acid sequences, demonstrating its potential and practical value in peptide-based drug development. AVAILABILITY AND IMPLEMENTATION: The complete datasets, source code, and pre-trained models are made available at https://github.com/Renjingyi123/ToxMSRC and https://doi.org/10.5281/zenodo.15668530.
Shengli Zhang 0002, Jingyi Ren, Yunyun Liang
Bioinform.2
2024 Large-scale spatial data visualization method based on augmented reality
abstract
A task assigned to space exploration satellites involves detecting the physical environment within a certain space. However, space detection data are complex and abstract. These data are not conducive for researchers' visual perceptions of the evolution and interaction of events in the space environment. A time-series dynamic data sampling method for large-scale space was proposed for sample detection data in space and time, and the corresponding relationships between data location features and other attribute features were established. A tone-mapping method based on statistical histogram equalization was proposed and applied to the final attribute feature data. The visualization process is optimized for rendering by merging materials, reducing the number of patches, and performing other operations. The results of sampling, feature extraction, and uniform visualization of the detection data of complex types, long duration spans, and uneven spatial distributions were obtained. The real-time visualization of large-scale spatial structures using augmented reality devices, particularly low-performance devices, was also investigated. The proposed visualization system can reconstruct the three-dimensional structure of a large-scale space, express the structure and changes in the spatial environment using augmented reality, and assist in intuitively discovering spatial environmental events and evolutionary rules.
Xiaoning Qiao, Wenming Xie, Xiaodong Peng, Guangyun Li, Dalin Li, Yingyi Guo, Jingyi Ren
Virtual Real. Intell. Hardw.7