Rongjin Li

dblp:243/6781 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Vision and language · 30% Question answering and dialogue systems · 28% Language models and text generation · 26%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 100%

Topics — the 6 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
retrieval-augmented generation
1.322026
NeocorRAG: Less Irrelevant Information, More Explicit Evidence, and More Effective Recall via Evidence Chains · WWW 2026
FinMMDocR: Benchmarking Financial Multimodal Reasoning with Scenario Awareness, Document Understanding, and Multi-Step Computation · AAAI 2026
Computer vision › Vision and language
multimodal benchmark
1.012026
Decoding Scientific Experimental Images: The SPUR Benchmark for Perception, Understanding, and Reasoning · ACL (1) 2026
Computer vision › Vision and language › vision-language model
multimodal large language model
1.012026
FinMMDocR: Benchmarking Financial Multimodal Reasoning with Scenario Awareness, Document Understanding, and Multi-Step Computation · AAAI 2026
Natural language and speech › Question answering and dialogue systems
financial multimodal reasoning
0.912025
$\mathcal{F}_{M}$ FinMMR: Make Financial Numerical Reasoning More Multimodal, Comprehensive, and Challenging · ICCV 2025
Natural language and speech › Language models and text generation › mathematical reasoning › numerical reasoning
financial numerical reasoning
0.912025
FinanceReasoning: Benchmarking Financial Numerical Reasoning More Credible, Comprehensive and Challenging · ACL (1) 2025
Natural language and speech › Language models and text generation › mathematical reasoning
numerical reasoning
0.912025
FinanceReasoning: Benchmarking Financial Numerical Reasoning More Credible, Comprehensive and Challenging · ACL (1) 2025

Methods — techniques the papers use, named apart from their topics

retrieval-augmented generation · 2.0multimodal learning · 1.7large language model · 1.7multimodal large language model · 1.0constrained decoding · 1.0activated search · 1.0benchmark construction · 0.9
YearPublicationVenuePosition
2026 FinMMDocR: Benchmarking Financial Multimodal Reasoning with Scenario Awareness, Document Understanding, and Multi-Step Computation
abstract
We introduce FinMMDocR, a novel bilingual multimodal benchmark for evaluating multimodal large language models (MLLMs) on real-world financial numerical reasoning. Compared to existing benchmarks, our work delivers three major advancements. (1) Scenario Awareness: 57.9% of 1,200 expert-annotated problems incorporate 12 types of implicit financial scenarios (e.g., Portfolio Management), challenging models to perform expert-level reasoning based on assumptions; (2) Document Understanding: 837 Chinese/English documents spanning 9 types (e.g., Company Research) average 50.8 pages with rich visual elements, significantly surpassing existing benchmarks in both breadth and depth of financial documents; (3) Multi-Step Computation: Problems demand 11-step reasoning on average (5.3 extraction + 5.7 calculation steps), with 65.0% requiring cross-page evidence (2.4 pages average). The best-performing MLLM achieves only 58.0% accuracy, and different retrieval-augmented generation (RAG) methods show significant performance variations on this task. We expect FinMMDocR to drive improvements in MLLMs and reasoning-enhanced methods on complex multimodal reasoning tasks in real-world scenarios.
Zichen Tang, Haihong E, Rongjin Li, Linwei Jia, Zhuodi Hao, Zhongjun Yang, Yuanze Li, Haolin Tian, Peizhi Zhao, Xianghe Wang, Xueyuan Lin, Ruofei Bai, Zijian Xie, Ruining Cao, Haocheng Gao
AAAI3
2026 Decoding Scientific Experimental Images: The SPUR Benchmark for Perception, Understanding, and Reasoning
abstract
Junpeng Ding, Zichen Tang, Haihong E, Mengyuan Ji, Yang Liu, Haolin Tian, Haiyang Sun, Pengqi Sun, Yang Xu, Yichen Liu, Haocheng Gao, Zijie Xi, Ruomeng Jiang, Peizhi Zhao, Rongjin Li, Yuanze Li, Jiacheng Liu, Zhongjun Yang, Jintong Chen, Siying Lin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Junpeng Ding, Zichen Tang, Haihong E, Mengyuan Ji, Haolin Tian, Pengqi Sun, Haocheng Gao, Zijie Xi, Ruomeng Jiang, Peizhi Zhao, Rongjin Li, Yuanze Li, Zhongjun Yang, Jintong Chen, Siying Lin
ACL (1)15
2026 NeocorRAG: Less Irrelevant Information, More Explicit Evidence, and More Effective Recall via Evidence Chains
abstract
Although precise recall is a core objective in Retrieval-Augmented Generation (RAG), a critical oversight persists in the field: improvements in retrieval performance do not consistently translate to commensurate gains in downstream reasoning. To diagnose this gap, we propose the Recall Conversion Rate (RCR), a novel evaluation metric to quantify the contribution of retrieval to reasoning accuracy. Our quantitative analysis of mainstream RAG methods reveals that as Recall@5 improves, the RCR exhibits a near-linear decay. We identify the neglect of retrieval quality in these methods as the underlying cause. In contrast, approaches that focus solely on quality optimization often suffer from inferior recall performance. Both categories lack a comprehensive understanding of retrieval quality optimization, resulting in a trade-off dilemma. To address these challenges, we propose comprehensive retrieval quality optimization criteria and introduce the NeocorRAG framework. This framework achieves holistic retrieval quality optimization by systematically mining and utilizing Evidence Chains. Specifically, NeocorRAG first employs an innovative activated search algorithm to obtain a refined candidate space. Then it ensures precise evidence chain generation through constrained decoding. Finally, the retrieved set of evidence chains guides the retrieval optimization process. Evaluated on benchmarks including HotpotQA, 2WikiMultiHopQA, MuSiQue, and NQ, NeocorRAG achieves SOTA performance on both 3B and 70B parameter models, while consuming less than 20% of tokens used by comparable methods. This study presents an efficient, training-free paradigm for RAG enhancement that effectively optimizes retrieval quality while maintaining high recall. Our code is released at https://github.com/BUPT-Reasoning-Lab/NeocorRAG.
Shiyao Peng, Qianhe Zheng, Zhuodi Hao, Zichen Tang, Rongjin Li, Yifan Zhu 0001, Haihong E
WWW5
2025 FinanceReasoning: Benchmarking Financial Numerical Reasoning More Credible, Comprehensive and Challenging
abstract
Zichen Tang, Haihong E, Ziyan Ma, Haoyang He, Jiacheng Liu, Zhongjun Yang, Zihua Rong, Rongjin Li, Kun Ji, Qing Huang, Xinyang Hu, Yang Liu, Qianhe Zheng. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zichen Tang, Haihong E, Ziyan Ma, Haoyang He, Zhongjun Yang, Zihua Rong, Rongjin Li, Kun Ji, Xinyang Hu, Qianhe Zheng
ACL (1)8
2025 SoCov: Semi-Orthogonal Parametric Pooling of Covariance Matrix for Speaker Recognition
abstract
In conventional deep speaker embedding frameworks, the pooling layer aggregates all frame-level features over time and computes their mean and standard deviation statistics as inputs to subsequent segment-level layers. Such statistics pooling strategy produces fixed-length representations from variable-length speech segments. However, this method treats different frame-level features equally and discards covariance information. In this paper, we propose the Semi-orthogonal parameter pooling of Covariance matrix (SoCov) method. The SoCov pooling computes the covariance matrix from the self-attentive frame-level features and compresses it into a vector using the semi-orthogonal parametric vectorization, which is then concatenated with the weighted standard deviation vector to form inputs to the segment-level layers. Deep embedding based on SoCov is called "sc-vector". The proposed sc-vector is compared to several different baselines on the SRE21 development and evaluation sets. The sc-vector system significantly outperforms the conventional x-vector system, with a relative reduction in EER of 15.5% on SRE21Eval. When using self-attentive deep feature, SoCov helps to reduce EER on SRE21Eval by about 30.9% relatively to the conventional "mean + standard deviation" statistics.
Rongjin Li, Dongpeng Chen, Jintao Kang, Xiaofen Xing
ICASSP1
2025 $\mathcal{F}_{M}$ FinMMR: Make Financial Numerical Reasoning More Multimodal, Comprehensive, and Challenging
Zichen Tang, Haihong E, Zhongjun Yang, Rongjin Li, Zihua Rong, Haoyang He, Zhuodi Hao, Xinyang Hu, Kun Ji, Ziyan Ma, Mengyuan Ji, Chenghao Ma, Qianhe Zheng, Zijian Xie, Shiyao Peng
ICCV5
2022 The Coral++ Algorithm for Unsupervised Domain Adaptation of Speaker Recognition
abstract
State-of-the-art speaker recognition systems are trained with a large amount of human-labeled training data set. Such a training set is usually composed of various data sources to enhance the modeling capability of models. However, in practical deployment, unseen condition is almost inevitable. Domain mismatch is a common problem in real-life applications due to the statistical difference between the training and testing data sets. To alleviate the degradation caused by domain mismatch, we propose a new feature-based unsupervised domain adaptation algorithm. The algorithm we propose is a further optimization based on the well-known CORrelation ALignment (CORAL), so we call it CORAL++. On the NIST 2019 Speaker Recognition Evaluation (SRE19), we use SRE18 CTS set as the development set to verify the effectiveness of CORAL++. With the typical x-vector/PLDA setup, the CORAL++ outperforms the CORAL by 9.40% relatively on EER.
Rongjin Li, Dongpeng Chen
ICASSP1
2020 Voiceai Systems to NIST Sre19 Evaluation: Robust Speaker Recognition on Conversational Telephone Speech
abstract
In this study, we present the VoiceAI (VAI) submissions to NIST SRE 2019 challenge on the task of speaker recognition using conversational telephone speech. Domain mismatching remains a challenging problem on SRE19. However, participants are unconstrained to use any public or proprietary data to mitigate this problem. Using larger scale and more diverse training data can effectively improve the performance of the front-end neural networks (NN). In our experiments, we focus on constructing robust systems using x-vectors. Different input acoustic features are also investigated. In addition, we propose hybrid neural architectures to utilizing the strength of different neural networks such as long short-term memory (LSTM), extended time delay neural network (ETDNN) and factorized TDNN (FTDNN). We also explore many back-end strategies to make full use of the development data and to relief the domain mismatching problem. Our best network topology, a FTDNN with two LSTMP layers, significantly outperforms the baseline on NIST SRE18 evaluation set (SRE18Eval). The final system, a fusion of five systems, yields EER 2.59%, minimum cost 0.153, and actual cost 0.155 on the progress data set, ranking second on the leaderboard among all participants.
Rongjin Li, Dongpeng Chen
ICASSP1
2019 Boundary Discriminative Large Margin Cosine Loss for Text-independent Speaker Verification
abstract
Deep neural network based speaker embeddings have attracted much attention in text-independent speaker verification task. In addition to the network architecture, an appropriate design of the loss function is crucial for the deep discriminative embedding extractor. Inspired by the success of Large Margin Cosine Loss (LMCL) in face recognition, we propose an enhanced LMCL named boundary discriminative LMCL (BD-LMCL) to emphasize the discriminative information inherited in the speaker boundaries. Unlike LMCL, where all training samples contribute equally for the objective function, only the samples around the speaker boundaries are considered during the network training with BD-LMCL. Specifically, those samples close to the boundaries are dynamically selected using top-k zero-one loss. Experimental results on a short duration corpus Android Cellphone and NIST SRE 2012 demonstrate better performance compared to LMCL and other popular loss functions.
Rongjin Li, Na Li 0012, Deyi Tuo, Meng Yu 0003, Dan Su 0002, Dong Yu 0001
ICASSP1
2019 Anti-Spoofing Speaker Verification System with Multi-Feature Integration and Multi-Task Learning
Rongjin Li, Miao Zhao, Lin Li 0032, Qingyang Hong
INTERSPEECH1