Lei Li 0009

dblp:13/7007-9 · DBLP profile ↗
← Back
23ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0002-3204-6527ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Computer networks · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Speech recognition and synthesis · 50% Transfer learning and domain adaptation · 25% Trustworthy machine learning · 25%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 77% Energy-efficient computing · 23%

Topics — the 4 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › language model interpretability
brain alignment
1.012026
Temporal Precision Matters: Brain-Tuning Speech Language Models with Millisecond-Resolution Neural Signals · ACL (1) 2026
Natural language and speech › Speech recognition and synthesis
brain-tuning
1.012026
Temporal Precision Matters: Brain-Tuning Speech Language Models with Millisecond-Resolution Neural Signals · ACL (1) 2026
Natural language and speech › Speech recognition and synthesis
speech language model
1.012026
Temporal Precision Matters: Brain-Tuning Speech Language Models with Millisecond-Resolution Neural Signals · ACL (1) 2026
Energy-efficient computing
demand response
0.212023
Toward Sustainable AI: Federated Learning Demand Response in Cloud-Edge Systems via Auctions · INFOCOM 2023

Methods — techniques the papers use, named apart from their topics

skill injection · 1.0fine-tuning · 1.0electrocorticography · 1.0auction · 0.7
YearPublicationVenuePosition
2026 Investigating Cross-Modal Skill Injection: Scenarios, Methods, and Hyperparameters
abstract
Zhiyu Xu, Lean Wang, Yuanxin Liu, Lei Li, Hao Zhou, Fandong Meng, Jie Zhou, Xu Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Lean Wang, Yuanxin Liu, Lei Li 0009, Hao Zhou 0012, Fandong Meng, Jie Zhou 0016, Xu Sun 0001
ACL (1)4
2026 Temporal Precision Matters: Brain-Tuning Speech Language Models with Millisecond-Resolution Neural Signals
abstract
Brain-tuning enhances brain alignment and downstream performance by fine-tuning speech language models with neural recordings.However, previous work relies primarily on fMRI, whose temporal resolution integrates neural activity over seconds, blending distinct processing stages into a single supervision signal and precluding temporally targeted training.We introduce ECoG-tuning, which leverages electrocorticography's millisecond precision to train speech language models.We design temporally targeted windows-a speech window capturing acoustic-phonetic encoding and a language window capturing higher-order linguistic processing-grounded in neuroscientific findings about temporal encoding hierarchies.Evaluating three models on the Podcast ECoG dataset, we find that ECoG-tuning significantly improves brain alignment over pretrained and distillation baselines.Notably, full spatiotemporal dynamics yield 7-17% higher alignment than time-averaged supervision across models, and language-window tuning produces larger gains in higher-order language regions, indicating that temporal precision provides additional training value.Moreover, ECoG-tuned models consistently improve or maintain downstream performance.Overall, our work provides initial evidence that electrophysiology is a viable brain-tuning modality, demonstrating how neuroscientific insights into processing hierarchies can inform principled model training strategies.
Zhejun Zhang, Wenqing Zhou, Haozhe Xu, Lin Zhang 0013, Lei Li 0009
ACL (1)5
2026 MMRet3D: A Multi-Modal Matching Framework for 3D Object Retrieval from Multi-View Images
abstract
Retrieval-based 3D scene reconstruction has become a practical paradigm for building indoor environments from RGB observations by assembling objects from large CAD repositories. However, accurate instance-level CAD retrieval remains a critical bottleneck. Existing RGB-based matching pipelines are vulnerable to appearance-shape ambiguity, where visually similar objects may correspond to substantially different 3D geometries, and the problem is amplified by occlusion and viewpoint bias in real-world scans. We propose MMRet3D, a geometry-consistent framework for instance-level CAD retrieval from multi-view images. MMRet3D learns a retrieval-oriented geometric descriptor from surface-normal cues via Geometric Knowledge Distillation (GKD) and performs viewpoint-aligned multi-modal matching by comparing multi-view observations with pose-aligned renderings of candidate CAD models. In addition, 3D Gaussian Clustering fuses multi-view semantics in 3D to obtain view-consistent object instances as retrieval queries. Experiments on Scan2CAD, ScanNet, and Sintel demonstrate that MMRet3D enhances the accuracy of instance-level retrieval and has a continuous positive impact on the quality of downstream scene reconstruction.
Lei Li 0009
ICMR2
2026 SKG-VLA: Scene Knowledge Graph Priors for Structured Scene Semantics and Multimodal Reasoning for Decision Making
abstract
Decision making in large-scale complaint handling systems increasingly relies on heterogeneous evidence, including complaint narratives, screenshots, order metadata, historical interactions, and platform policies. Existing complaint understanding systems mainly perform shallow classification or template matching over isolated modalities, while underutilizing explicit scene structure, rule knowledge, and cross-evidence dependencies. To address this limitation, we present SKG-VLA for multimodal complaint decision making. The core idea is to model each case as a structured complaint scene and represent its decision-relevant semantics with a Scene Knowledge Graph (SKG), which organizes complaint entities, evidence items, policy clauses, temporal events, transactional states, and action-relevant relations into a unified graph. Based on SKG, we build a data synthesis pipeline that generates complaint scene descriptions, rule-consistent graph generalizations, question–answer supervision, and decision recommendations. We further construct a large-scale complaint scene dataset with both text-only and multimodal in-domain benchmarks. Finally, we adopt a three-stage training strategy—domain-adaptive pre-training, task-oriented instruction fine-tuning, and end-to-end multimodal alignment—to inject structured scene priors into a multimodal decision model. Experiments show that SKG-VLA consistently improves policy-grounded reasoning, complaint decision accuracy, long-tail generalization, and robustness under incomplete evidence.
Lei Li 0009
ICMR2
2026 Scheduling cloud-edge federated learning under demand response with carbon neutrality
Fei Wang 0136, Lei Jiao 0002, Konglin Zhu, Jiayuan Du, Xiaojun Lin 0001, Lei Li 0009
Comput. Networks6
2026 Neural feature alignment between large language models and brain activities: A knowledge-based framework for cross-modal analysis
Zhejun Zhang, Wenqing Zhou, Xinhang Li 0003, Lin Zhang 0013, Lei Li 0009
Neural Networks6
2025 AGAM-EEG: Activation-Gradient Attribution Map for Enhancing Explainability in EEG Decoding
abstract
The explainability of electroencephalogram (EEG) decoding significantly contributes to brain-computer interfaces (BCI) and cognitive neuroscience. Although a few explainable artificial intelligence (XAI) methods have attempted to associate EEG features with the spatial and temporal characteristics of EEG signals, a model-agnostic explainability framework across various deep learning (DL) models is still lacking. In this work, AGAM-EEG (activation-gradient attribution map in EEG) is proposed, which offers adaptive model-agnostic explainability across DL models for EEG decoding. AGAM-EEG initially calculates the gradients of feature maps with respect to the predicted class at the target layer. Subsequently, AGAM-EEG infers the spatial-temporal relevance distribution by integrating activation responses with class-conditional gradients. The integration enables a refined propagation of discriminative cues across the spatial and temporal dimensions of the EEG signals. A novel technique based on the normalized discounted cumulative gain (NDCG) is designed to select the best XAI methods. AGAM-EEG is validated across 6 DL models on 2 public cognitive datasets. Compared with other XAI methods, AGAM-EEG achieves improvements of 4.0% and 3.8% in the spatial and temporal dimensions on NDCG, demonstrating superior explainability.
Hengyi Shao, Lin Zhang 0013, Lei Li 0009
BIBM4
2024 Improving Factual Consistency in Abstractive Summarization with Sentence Structure Pruning
abstract
State-of-the-art abstractive summarization models still suffer from the content contradiction between the summaries and the input text, which is referred to as the factual inconsistency problem. Recently, a large number of works have also been proposed to evaluate factual consistency or improve it by post-editing methods. However, these post-editing methods typically focus on replacing suspicious entities, failing to identify and modify incorrect content hidden in sentence structures. In this paper, we first verify that the correctable errors can be enriched by leveraging sentence structure pruning operation, and then we propose a post-editing method based on that. In the correction process, the pruning operation on possible errors is performed on the syntactic dependency tree with the guidance of multiple factual evaluation metrics. Experimenting on the FRANK dataset shows a great improvement in factual consistency compared with strong baselines and, when combined with them, can achieve even better performance. All the codes and data will be released on paper acceptance.
Dingxin Hu, Xingyue Zhang, Marina Litvak, Natalia Vanetik, Qing Yang 0033, Dongliang Xu, Yanquan Zhou, Lei Li 0009, Yingqi Zhu
LREC/COLING11
2024 Explainable Semantic Communication for Text Tasks
abstract
Task-oriented semantic communication has gained increasing attention due to its ability to reduce the amount of transmitted data without sacrificing task performance. Although some prior efforts have been dedicated to developing semantic communications, the semantics in these works remains to be unexplainable. Challenges related to explainable semantic representation and knowledge-based semantic compression have yet to be explored. In this article, we propose a triplet-based explainable semantic communication (TESC) scheme for representing text semantics efficiently. Specifically, we develop a semantic extraction method to convert text into triplets while using syntactic dependency analysis to enhance semantic completeness. Then, we design a semantic filtering method to further compress the duplicate and task-irrelevant triplets based on prior knowledge. The filtered triplets are encoded and transmitted to the receiver for completing intelligent tasks. Furthermore, we apply the proposed TESC scheme to two emblematic text tasks: 1) sentiment analysis and 2) question answering, in which the semantic codec is meticulously customized for each task. Experimental results demonstrate that 1) the TESC scheme outperforms benchmarks in terms of Top-1 accuracy and transmission efficiency and 2) the TESC scheme enjoys about 150% performance gain compared to the traditional communication method.
Chuanhong Liu, Caili Guo, Yang Yang 0057, Wanli Ni, Yanquan Zhou, Lei Li 0009, Tony Q. S. Quek
IEEE Internet Things J.6
2024 Progression Cognition Reinforcement Learning With Prioritized Experience for Multi-Vehicle Pursuit
abstract
Multi-vehicle pursuit (MVP) such as autonomous police vehicles pursuing suspects is important but very challenging due to its mission and safety-critical nature. While multi-agent reinforcement learning (MARL) algorithms have been proposed for MVP in structured grid-pattern roads, the existing algorithms use random training samples in centralized learning, which leads to homogeneous agents showing low collaboration performance. For the more challenging problem of pursuing multiple evaders, these algorithms typically select a fixed target evader for pursuers without considering dynamic traffic situation, which significantly reduces pursuing success rate. To address the above problems, this paper proposes a Progression Cognition Reinforcement Learning with Prioritized Experience for MVP (PEPCRL-MVP) in urban multi-intersection dynamic traffic scenes. PEPCRL-MVP uses a prioritization network to assess the transitions in the global experience replay buffer according to each MARL agent’s parameters. With the personalized and prioritized experience set selected via the prioritization network, diversity is introduced to the MARL learning process, which can improve collaboration and task-related performance. Furthermore, PEPCRL-MVP employs an attention module to extract critical features from dynamic urban traffic environments. These features are used to develop a progression cognition method to adaptively group pursuing vehicles. Each group efficiently targets one evading vehicle. Extensive experiments conducted with a simulator over unstructured roads of an urban area show that PEPCRL-MVP is superior to other state-of-the-art methods. Specifically, PEPCRL-MVP improves pursuing efficiency by 3.95$\%$over Twin Delayed Deep Deterministic policy gradient-Decentralized Multi-Agent Pursuit and its success rate is 34.78$\%$higher than that of Multi-Agent Deep Deterministic Policy Gradient. Codes are open-sourced.
Xinhang Li 0003, Zheng Yuan 0010, Zhe Wang 0064, Qinwen Wang, Chen Xu 0002, Lei Li 0009, Jianhua He 0001, Lin Zhang 0013
IEEE Trans. Intell. Transp. Syst.7
2023 HoloCV: A Head-Mounted Mixed Reality System for Contactless Vital Signs Monitoring in Medical Emergency Situations
abstract
In medical situations, head-mounted displays have gained significant attention for providing physicians with immediate access to patient information. Traditional vital signs monitoring needs physical contact to patient and separates display to physicians. However, it is difficult to integrate the advanced computer-human interaction with contactless vital signs monitoring device. This paper discusses the design and implementation of HoloCV, a scalable mixed reality-based system that supports contactless vital signs monitoring and presents real-time pre-diagnostic results of patient to the visual field of physicians. HoloCV is supported by a distributed architecture divided into three interrelated applications responsible for advanced computer-human interaction, radar and thermal infrared device integration, computation and communication infrastructure, respectively. To monitor vital signs contactlessly and continuously, the HoloCV utilizes Impulse-Radio Ultra-WideBand(IR-UWB) radar and thermal infrared camera to achieve monitoring respiration rate, heart rate, blood pressure, and temperature. By running the YOLOv7-Tiny model directly on Microsoft Hololens2 head-mounted displays, HoloCV can achieve real-time object detection and present real-time visualizations of vital signs data on the detected targets. The overall weight of the system is less than or equal to 5 kilograms. It has a detection range of 0.5m to 3m, and can continuously operate for more than 1 hour with a standby time of 5 hours. The working environment temperature ranges from 0°C to 46°C. The heart rate measurement range is from 50 to 150 BPM, the respiratory rate measurement range is from 10 to 50 BPM. The temperature measurement range is from 32 degrees Celsius to 42 degrees Celsius, and the blood pressure measurement range is from 50 to 180 mmHg. With its capability for real-time and non-contact monitoring of vital signs, this HoloCV system differentiates itself from traditional contact-based devices, providing physicians with a convenient means to assess patients' conditions and administer necessary treatments without physical contact. The codes of the proposed HoloCV are made publicly available.
Xu Song, Xikang Jiang, Lei Li 0009, Lin Zhang 0013
HealthCom5
2023 Toward Sustainable AI: Federated Learning Demand Response in Cloud-Edge Systems via Auctions
Fei Wang 0136, Lei Jiao 0002, Konglin Zhu, Xiaojun Lin 0001, Lei Li 0009
INFOCOM5
2022 MINIPI: A MultI-scale Neural Network Based Impulse Radio Ultra-Wideband Radar Indoor Personnel Identification Method
Lingyi Meng, Xikang Jiang, Wenyao Mu, Yili Wang 0003, Lei Li 0009, Lin Zhang 0013
PRCV (2)6
2021 DP-YOLOv5: Computer Vision-Based Risk Behavior Detection in Power Grids
Zhe Wang 0064, Yubo Zheng, Xinhang Li 0003, Xikang Jiang, Zheng Yuan 0010, Lei Li 0009, Lin Zhang 0013
PRCV (1)6
2020 Interpretable Machine Learning Based on Integration of NLP and Psychology in Peer-to-Peer Lending Risk Evaluation
Lei Li 0009, Tianyuan Zhao, Yanjie Feng
NLPCC (2)1
2019 In Conclusion Not Repetition: Comprehensive Abstractive Summarization with Diversified Attention Based on Determinantal Point Processes
abstract
Various Seq2Seq learning models designed for machine translation were applied for abstractive summarization task recently.Despite these models provide high ROUGE scores, they are limited to generate comprehensive summaries with a high level of abstraction due to its degenerated attention distribution.We introduce Diverse Convolutional Seq2Seq Model(DivCNN Seq2Seq) using Determinantal Point Processes methods(Micro DPPs and Macro DPPs) to produce attention distribution considering both quality and diversity.Without breaking the end to end architecture, Di-vCNN Seq2Seq achieves a higher level of comprehensiveness compared to vanilla models and strong baselines.All the reproducible codes and datasets are available online 1 .1 available at https://github.com/thinkwee/DPPCNN Sum marization Article: marseille , france the french prosecutor leading an investigation into the crash of germanwings flight 9525 insisted wednesday that he was not aware of any video footage from on board the plane .marseille prosecutor brice robin told cnn that so far no videos were used in the crash investigation ...... of a cell phone video showing the harrowing final seconds from on board germanwings flight 9525 as it crashed into the french alps .......paris match and bild reported that the video was recovered from a phone at the wreckage site ....... cnn 's frederik pleitgen , pamela boykoff , antonia mortensen , sandrine amiel and anna-maja rappard contributed to this report .CNN Seq2Seq: french prosecutor UNK robin says he was not aware of any video
Lei Li 0009, Wei Liu 0161, Marina Litvak, Natalia Vanetik, Zuying Huang
CoNLL1
2019 Quality-Diversity Summarization with Unsupervised Autoencoders
Lei Li 0009, Zuying Huang, Natalia Vanetik, Marina Litvak
ICANN (4)1
2019 Dense People Counting Using IR-UWB Radar With a Hybrid Feature Extraction Method
abstract
People counting is one of the hottest issues in sensing applications. Impulse radio ultrawideband radar has been extensively adopted to count people because it provides a device-free solution without illumination and privacy concerns. However, current solutions have limited performances in congested environments due to signal superpositions and obstructions. In this letter, a hybrid feature extraction method based on the curvelet transform and the distance bin is proposed. First, 2-D radar matrix features are extracted at multiple scales and multiple angles by applying the curvelet transform. Then, the distance bin concept is introduced by dividing each row of the matrix into several bins along the propagating distance to select features. A radar signal data set is constructed for three density scenarios, including people randomly walking in a constrained area at densities of three and four persons per square meter and people in a queue with an average between-person distance of 10 cm. The number of people in the data set scenarios varies from 0 to 20. Four classifiers-a decision tree, an AdaBoost classifier, a random forest, and a neural network-are compared to validate the hybrid features. The random forest achieves the highest accuracy of above 97% in the three density scenarios. To further investigate the reliability of the hybrid features, they are compared with three other features: cluster features, activity features, and features extracted by a convolutional neural network. The comparison results reveal that the proposed hybrid features are stable, and their performance is substantially more effective than that of the others.
Xiuzhu Yang, Wenfeng Yin, Lei Li 0009, Lin Zhang 0013
IEEE Geosci. Remote. Sens. Lett.3
2018 Sociability-based Influence Diffusion Probability Model to evaluate influence of BBS post
Lei Li 0009, MengChu Zhou, LiLi Fu
Neurocomputing1
2017 Summarizing Weibo with Topics Compression
Marina Litvak, Natalia Vanetik, Lei Li 0009
CICLing (2)3
2017 Research on machine learning algorithms and feature extraction for time series
abstract
This paper aims to use various machine learning algorithms and explore the influence between different algorithms and multi-feature in the time series. The real consumption records constitute the time series as the research object. We extract consumption mark, frequency and other features. Moreover, we utilize support vector machine (SVM), long short-term memory (LSTM) and other algorithms to predict the user's consumption behavior. Besides, we have also implemented multi-feature fusion and multi-algorithm fusion with LSTM and SVM. Eventually, the experimental results show that LSTM algorithms is advantageous in prediction when the data is sparse. In the other hand, the SVM is beneficial when the data is more abundant. What's more, LSTM-SVM fusion model has advantages on the extracting features of LSTM and on the classification of SVM. In most cases, LSTM-SVM is most outstanding in prediction.
Lei Li 0009, Yabin Wu, Yihang Ou, Qi Li 0057, Yanquan Zhou, Daoxin Chen
PIMRC1
2014 UGC collection strategy oriented user influence classification and feature analysis
abstract
With the development of Social Networking Services, users can publish and receive information expediently. With the scale of information to be processed is becoming bigger, the information acquisition rules adopted by traditional search engine met with many difficulties. To solve the problem, we must develop new UGC collection strategy. We think there are three main factors: user, UGC and the interactive relationship between user and UGC. In this paper, we propose an algorithm named UIC to calculate user influence value. Based on the results, we classify the user into three types: star user, usual user, zombie user. In order to make further analysis of features of all kinds of users, based on preliminary classification, we make sampling and content analysis of the three types of user in order to optimize the user classification, and also make better support the UGC acquisition strategy.
Yue Zhai, Lei Li 0009
ASONAM2
2013 Predicting Stay Time of Mobile Users With Contextual Information
abstract
Mobile service providers and manufacturers continue to provide services and devices that take advantage of the location information associated with devices to provide a more personalized experience for users. For many such services, the user experience can be dramatically improved if a mobile device can predict how long a mobile user will stay at the current location. In this paper, we propose to take advantage of contextual information for predicting the stay time of mobile users. Specially, we investigate two strategies for modeling the relevance between it and contextual information, i.e., Stay Status Prediction (SSP) and Stay Time Prediction (STP). SSP is to predict whether a mobile user will stay at the current location at time point ti+naccording to the contextual information at ti, while STP is to directly predict how long a mobile user will stay at the current location. Moreover, we study several typical machine learning models which can be extended for implementing SSP and STP and evaluate their performance with respect to prediction accuracy. We also conduct extensive experiments on real data sets to evaluate several implementations of the proposed strategies in terms of both effectiveness and efficiency for STP.
Huanhuan Cao, Lei Li 0009, MengChu Zhou
IEEE Trans Autom. Sci. Eng.3