Lingfeng Shen

dblp:240/5490 · DBLP profile ↗
← Back
31ranked-venue papers
12as first author
27since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 7 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 HNI-SLAM: Neural implicit SLAM for high quality reconstruction
Lina Duan, Yejun Shou, Lingfeng Shen, Yanlong Cao
Comput. Graph.6
2026 Towards efficient privacy-preserving keyword search for outsourced data in intelligent transportation systems
Guanghui Wang 0003, Lingfeng Shen, Shuang Ding, Xin He 0021, Zhonghao Zhai, Zongqi Shi
Future Gener. Comput. Syst.3
2026 Physical Layer Security of Coupled Phase Shifts STAR-RIS-Aided NOMA System Under Hybrid Far- and Near-Field Scenarios
abstract
Near-field (NF) communications have attracted considerable interest, particularly with the implementation of extremely large-scale antenna arrays (ELAA). Additionally, the increase in communication frequencies and the expansion of reconfigurable intelligent surface (RIS) apertures contribute to this growing field. This paper investigates the synergy of simultaneously transmitting and reflecting (STAR)-RIS and non-orthogonal multiple access (NOMA) for secure transmission under hybrid far-field (FF) and NF scenarios. The secrecy sum rate (SSR) maximization problem is formulated by joint optimization of the power allocation, the beamforming at access point (AP), and the transmission/reflection coefficients (TRCs). Specifically, we consider the transmit power budget, unit-norm conditions, coupled phase shifts (CPS), quality of service requirements, and decoding order. To tackle this extremely challenging problem, we combine the successive convex approximation (SCA), Riemannian exact penalty method via smoothing, and penalty dual decomposition (PDD) and successfully develop an efficient iterative algorithm. Simulation results reveal that the proposed design exhibits superior effectiveness when compared to other traditional benchmarks.
Lei Shi 0001, Zhiqing Tang, Lingfeng Shen, Wanming Hao, Jie Li 0002
IEEE Internet Things J.4
2026 RA-GAN: Region-adaptive GAN for binarization on degraded document images
Menghui Liu, Lang Yu, Guanghui Wang 0003, Lingfeng Shen
Knowl. Based Syst.5
2026 RGD-SLAM: Robust Gaussian splatting SLAM for dynamic environments
Yejun Shou, Lingfeng Shen, Yanlong Cao
Pattern Recognit.3
2026 Cost-Efficient Open Vocabulary 3D Scene Understanding Based on Semantic Probability
abstract
Traditional 3D scene understanding methods heavily depend on 3D annotation and training, which allow for the identification of seen classes but struggle to recognize unseen classes. In this paper, we leverage the open vocabulary inference capabilities of pre-trained models, enabling the encoding of open vocabulary concepts. However, unlike existing open vocabulary 3D scene understanding methods, we propose a framework based on semantic probability. This innovation significantly reduces computational cost and is compatible with state-of-the-art two-stage 2D pre-trained models. Specifically, we align the text features from the CLIP model with the pixel features from the 2D pre-trained models, inferring semantic probability of image pixels based on similarity and projecting it onto 3D points. Subsequently, we introduce a point cloud pairs semantic fusion method to merge the point clouds, reducing the semantic probability of erroneous 3D points. Based on probability scores, we achieve 3D semantic segmentation on open vocabularies without any supervision or training. In addition, the semantic probability of 3D points can serve as pseudo-labels for 3D distillation, and the geometric features of the 3D scene can be exploited to improve the segmentation performance. Experimental results demonstrate that the proposed method exhibits competitive performance on publicly available benchmark datasets, including ScanNet, Matterport3D, and nuScenes.
Lingfeng Shen, Xiaoyao Wei, Gang Pan 0001, Yanlong Cao
IEEE Trans. Image Process.1
2026 Fast Photometric Stereo by Time- and Spectral-Multiplexing With Crosstalk Handling
abstract
Photometric stereo is widely used to recover detailed surface normals. However, previous methods fail to balance the accuracy and efficiency. Conventional photometric stereo achieves high accuracy but suffers from low efficiency due to spectral-multiplexing and inefficient algorithms. In contrast, multispectral photometric stereo captures images efficiently with spectral-multiplexing, but its accuracy is harmed by crosstalk. In this paper, we aim to resolve the crosstalk issue to achieve fast photometric stereo (FPS) at low cost. First, we analyze the formulation and impact of crosstalk, showing that it significantly affects normal estimation, with external factors being primary contributors to crosstalk and internal factors being the secondary. Subsequently, we propose the FPS framework with a fast data capture scheme that combines time- and spectral-multiplexing to introduce constraints on crosstalk regarding both internal and external factors, along with a lightweight network, FPS-Net, to remove crosstalk caused by those factors based on constraints under such scheme. Finally, we build a real-world crosstalk-affected FPS dataset to evaluate the performance in handling crosstalk for normal estimation. Experimental results show the superior accuracy and efficiency of our method. The code and dataset are available at https://github.com/wxy-zju/FPS-Net.
Xiaoyao Wei, Lingfeng Shen, Zhijie Xu, Yanlong Cao
IEEE Trans. Image Process.2
2025 Unsupervised Rgb-D Point Cloud Registration for Scenes With Low Overlap and Photometric Inconsistency
Yejun Shou, Lingfeng Shen, Gang Pan 0001, Yanlong Cao
ICCV3
2025 E-FCOS: Enhanced Historical Text Detection with Fast Fourier Transform Denoising and Adaptive Multi-scale Fusion
Menghui Liu, Lang Yu, Yilan Yang, Lingfeng Shen
ICDAR (4)5
2025 MIR-Bench: Can Your LLM Recognize Complicated Patterns via Many-Shot In-Context Reasoning?
abstract
The ability to recognize patterns from examples and apply them to new ones is a primal ability for general intelligence, and is widely studied by psychology and AI researchers. Many benchmarks have been proposed to measure such ability for Large Language Models (LLMs); however, they focus on few-shot (usually <10) setting and lack evaluation for aggregating many pieces of information from long contexts. On the other hand, the ever-growing context length of LLMs have brought forth the novel paradigm of many-shot In-Context Learning (ICL), which addresses new tasks with hundreds to thousands of examples without expensive and inefficient fine-tuning. However, many-shot evaluations often focus on classification, and popular long-context LLM tasks such as Needle-In-A-Haystack (NIAH) seldom require complicated intelligence for integrating many pieces of information. To fix the issues from both worlds, we propose MIR-Bench, the first many-shot in-context reasoning benchmark for pattern recognition that asks LLM to predict output via input-output examples from underlying functions with diverse data format. Based on MIR-Bench, we study many novel problems for many-shot in-context reasoning, and acquired many insightful findings including scaling effect, robustness, inductive vs. transductive reasoning, retrieval Augmented Generation (RAG), coding for inductive reasoning, cross-domain generalizability, etc. Our dataset is available at https://huggingface.co/datasets/kaiyan289/MIR-Bench.
Zhan Ling, Ting-Han Fan, Lingfeng Shen, Zhengyin Du, Jiecao Chen
NeurIPS6
2025 Towards Resilient Federated Learning Against Collusion Attacks
abstract
Federated learning has significant application value in the field of data privacy protection for Internet of Things (IoT) applications. However, collusion attacks, in which multiple clients collaboratively engage in adversarial behavior, may arise among clients in federated learning, posing serious threats to model training. To address this issue, this paper proposes a Resilient Federated Learning (RFL) algorithm by introducing a multidimensional risk-aware mechanism. RFL assesses the risks of collusion using the reputation, distance, and vulnerability level. Theoretical analysis shows that RFL exhibits convergence and satisfies ε-security under collusion attacks. Experimental results in the MNIST and CIFAR-10 datasets demonstrate that RFL can effectively defend against collusion attacks and consistently outperforms baseline schemes under various attack intensities.
Youjuan Zhu, Jingyao Xu 0005, Lingfeng Shen, Guanghui Wang 0003
VTC2025-Fall4
2025 Enhanced multi-scale feature adaptive fusion sparse convolutional network for large-scale scenes semantic segmentation
Lingfeng Shen, Yanlong Cao, Yejun Shou, Zhijie Xu
Comput. Graph.1
2025 Achieving efficient and accurate privacy-preserving localization for internet of things: A quantization-based approach
Guanghui Wang 0003, Xueyuan Zhang, Lingfeng Shen, Shengbo Chen, Fei Tong 0001, Xin He 0021
Future Gener. Comput. Syst.3
2025 Weakly supervised point cloud semantic segmentation using pseudo-label reliability and consistency regularization
Lingfeng Shen, Yanlong Cao, Xiaoyao Wei
Neurocomputing1
2025 Unleashing powerful generalization for point cloud registration
Yejun Shou, Lingfeng Shen, Yanlong Cao
Knowl. Based Syst.3
2024 iS-MAP: Neural Implicit Mapping and Positioning for Structural Environments
Yanlong Cao, Yejun Shou, Lingfeng Shen, Xiaoyao Wei, Zhijie Xu
ACCV (9)4
2024 AnaloBench: Benchmarking the Identification of Abstract and Long-context Analogies
abstract
Xiao Ye, Andrew Wang, Jacob Choi, Yining Lu, Shreya Sharma, Lingfeng Shen, Vijay Murari Tiyyala, Nicholas Andrews, Daniel Khashabi. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Jacob Choi, Yining Lu, Lingfeng Shen, Vijay Murari Tiyyala, Nicholas Andrews, Daniel Khashabi
EMNLP6
2024 The Trickle-down Impact of Reward Inconsistency on RLHF
abstract
Standard practice within Reinforcement Learning from Human Feedback (RLHF) involves optimizing against a Reward Model (RM), which itself is trained to reflect human preferences for desirable generations. A notable subject that is understudied is the (in-)consistency of RMs --- whether they can recognize the semantic changes to different prompts and appropriately adapt their reward assignments --- and their impact on the downstream RLHF model. In this paper, we visit a series of research questions relevant to RM inconsistency: (1) How can we measure the consistency of reward models? (2) How consistent are the existing RMs and how can we improve them? (3) In what ways does reward inconsistency influence the chatbots resulting from the RLHF model training? We propose **Contrast Instruction** -- a benchmarking strategy for the consistency of RM. Each example in **Contrast Instruction** features a pair of lexically similar instructions with different ground truth responses. A consistent RM is expected to rank the corresponding instruction and response higher than other combinations. We observe that current RMs trained with the standard ranking objective fail miserably on \contrast{} compared to average humans. To show that RM consistency can be improved efficiently without using extra training budget, we propose two techniques **ConvexDA** and **RewardFusion**, which enhance reward consistency through extrapolation during the RM training and inference stage, respectively. We show that RLHF models trained with a more consistent RM yield more useful responses, suggesting that reward inconsistency exhibits a trickle-down effect on the downstream RLHF process.
Lingfeng Shen, Linfeng Song, Lifeng Jin, Baolin Peng, Haitao Mi, Daniel Khashabi, Dong Yu 0001
ICLR1
2024 Position: Do pretrained Transformers Learn In-Context by Gradient Descent?
abstract
The emergence of In-Context Learning (ICL) in LLMs remains a remarkable phenomenon that is partially understood. To explain ICL, recent studies have created theoretical connections to Gradient Descent (GD). We ask, do such connections hold up in actual pre-trained language models? We highlight the limiting assumptions in prior works that make their setup considerably different from the practical setup in which language models are trained. For example, their experimental verification uses ICL objective (training models explicitly for ICL), which differs from the emergent ICL in the wild. Furthermore, the theoretical hand-constructed weights used in these studies have properties that don’t match those of real LLMs. We also look for evidence in real models. We observe that ICL and GD have different sensitivity to the order in which they observe demonstrations. Finally, we probe and compare the ICL vs. GD hypothesis in a natural setting. We conduct comprehensive empirical analyses on language models pre-trained on natural data (LLaMa-7B). Our comparisons of three performance metrics highlight the inconsistent behavior of ICL and GD as a function of various factors such as datasets, models, and the number of demonstrations. We observe that ICL and GD modify the output distribution of language models differently. These results indicate that the equivalence between ICL and GD remains an open hypothesis and calls for further studies.
Lingfeng Shen, Aayush Mishra, Daniel Khashabi
ICML1
2024 Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation
abstract
Moderate-sized large language models (LLMs) -- those with 7B or 13B parameters -- exhibit promising machine translation (MT) performance. However, they do not match the performance of state-of-the-art conventional encoder-decoder translation models or larger-scale LLMs such as GPT-4. In this study, we bridge this performance gap. We first assess the shortcomings of supervised fine-tuning for LLMs in the MT task, emphasizing the quality issues present in the reference data, despite being human-generated. Then, in contrast to supervised fine-tuning which mimics reference translations, we introduce Contrastive Preference Optimization (CPO), a novel approach that trains models to avoid generating adequate but not perfect translations. Applying CPO to ALMA models with only 22K parallel sentences and 0.1% parameters yields significant improvements. The resulting model, called ALMA-R, can match or exceed the performance of the WMT competition winners and GPT-4 on WMT'21, WMT'22 and WMT'23 test datasets.
Amr Sharaf, Yunmo Chen, Weiting Tan, Lingfeng Shen, Benjamin Van Durme, Kenton Murray, Young Jin Kim 0006
ICML5
2024 SemStamp: A Semantic Watermark with Paraphrastic Robustness for Text Generation
abstract
Abe Hou, Jingyu Zhang, Tianxing He, Yichen Wang, Yung-Sung Chuang, Hongwei Wang, Lingfeng Shen, Benjamin Van Durme, Daniel Khashabi, Yulia Tsvetkov. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Abe Bohan Hou, Tianxing He, Yichen Wang 0002, Yung-Sung Chuang, Lingfeng Shen, Benjamin Van Durme, Daniel Khashabi, Yulia Tsvetkov
NAACL-HLT7
2024 DiffNorm: Self-Supervised Normalization for Non-autoregressive Speech-to-speech Translation
abstract
Non-autoregressive Transformers (NATs) are recently applied in direct speech-to-speech translation systems, which convert speech across different languages without intermediate text data. Although NATs generate high-quality outputs and offer faster inference than autoregressive models, they tend to produce incoherent and repetitive results due to complex data distribution (e.g., acoustic and linguistic variations in speech). In this work, we introduce DiffNorm, a diffusion-based normalization strategy that simplifies data distributions for training NAT models. After training with a self-supervised noise estimation objective, DiffNorm constructs normalized target data by denoising synthetically corrupted speech features. Additionally, we propose to regularize NATs with classifier-free guidance, improving model robustness and translation quality by randomly dropping out source information during training. Our strategies result in a notable improvement of about $+7$ ASR-BLEU for English-Spanish (En-Es) translation and $+2$ ASR-BLEU for English-French (En-Fr) on the CVSS benchmark, while attaining over $14\times$ speedup for En-Es and $5 \times$ speedup for En-Fr translations compared to autoregressive baselines.
Weiting Tan, Lingfeng Shen, Daniel Khashabi, Philipp Koehn
NeurIPS3
2024 Structerf-SLAM: Neural implicit representation SLAM for structural environments
Yanlong Cao, Xiaoyao Wei, Yejun Shou, Lingfeng Shen, Zhijie Xu
Comput. Graph.5
2023 TextShield: Beyond Successfully Detecting Adversarial Sentences in text classification
Lingfeng Shen, Haiyun Jiang
ICLR1
2022 KATG: Keyword-Bias-Aware Adversarial Text Generation for Text Classification
abstract
Recent work has shown that current text classification models are vulnerable to small adversarial perturbation to inputs, and adversarial training that re-trains the models with the support of adversarial examples is the most popular way to alleviate the impact of the perturbation. However, current adversarial training methods have two principal problems: worse model generalization and ineffective defending against other text attacks. In this paper, we propose a Keyword-bias-aware Adversarial Text Generation model (KATG) that implicitly generates adversarial sentences using a generator-discriminator structure. Instead of using a benign sentence to generate an adversarial sentence, the KATG model utilizes extra multiple benign sentences (namely prior sentences) to guide adversarial sentence generation. Furthermore, to cover more perturbation used in existing attacks, a keyword-bias-aware sampling is proposed to select sentences containing biased words as prior sentences. Besides, to effectively utilize prior sentences, a generative flow mechanism is proposed to construct latent semantic space and learn a latent representation for the prior sentences. Experiments demonstrate that adversarial sentences generated by our KATG model can strengthen the victim model's robustness and generalization.
Lingfeng Shen, Shoushan Li, Ying Chen 0012
AAAI1
2022 DRCNet: Dynamic Image Restoration Contrastive Network
Fei Li 0022, Lingfeng Shen, Yang Mi
ECCV (19)2
2022 On the Evaluation Metrics for Paraphrase Generation
abstract
In this paper we revisit automatic metrics for paraphrase evaluation and obtain two findings that disobey conventional wisdom:(1) Reference-free metrics achieve better performance than their reference-based counterparts.(2) Most commonly used metrics do not align well with human annotation.Underlying reasons behind the above findings are explored through additional experiments and in-depth analyses.Based on the experiments and analyses, we propose ParaScore, a new evaluation metric for paraphrase generation.It possesses the merits of referencebased and reference-free metrics and explicitly models lexical divergence.Based on our analysis and improvements, our proposed reference-based outperforms than referencefree metrics.Experimental results demonstrate that ParaScore significantly outperforms existing metrics.Our codes and toolkit are released in https://github.com/ shadowkiller33/ParaScore.
Lingfeng Shen, Lemao Liu, Haiyun Jiang, Shuming Shi 0001
EMNLP1
2020 Forecasting People's Action via Social Media Data
abstract
In this paper, we explored predicting human activity based on users' content from social media. The activities we do are related to our interests, personalities, political preferences and our future decisions. The data set we collect contains examples of social media users related to many daily business activities. Then, we use the latest tailored sentence embedding framework to recognize the semantics of human activities and automatically cluster these activities. We proposed a neural network architecture to predict all activities performed by a specific user based on previous posts and self-describing text. Besides, we explore how adding inferred user characteristics into our model contributes to this forecasting task.
Lingfeng Shen, Zhuoming Liu 0003, Xiongtao Zhou
IEEE BigData1
2020 UAV-enabled Data Collection for mMTC Networks: AEM Modeling and Energy-Efficient Trajectory Design
abstract
Massive machine-type communications (mMTC) is a new key feature of 5G cellular and is expected to be further improved in future evolutions of the cellular standards. Data collection from machine-type communication devices (MTCDs), which can be achieved by various approaches, is important to operation of mMTC networks. This work studies data collection for mMTC networks enabled by unmanned aerial vehicle (UAV) stations moving in the air. Consider the limitation in battery lifetime at both the MTCDs and the UAV station, the UAV trajectory design problem is investigated from an energy efficiency perspective. In a generalized model where the target MTCDs are grouped into multiple clusters, the UAV station travels across the clusters and collect data from each cluster while hovering above the cluster. The corresponding MTCD clustering strategy, UAV hovering strategy and UAV flying strategy all have impacts on the energy consumption of the system, which results in a strongly coupled energy minimization problem that is difficult to solve. The sub-problems obtained through decomposition are decoupled in the proposed solution approach. Clustering of the MTCDs is done by a greedy learning clustering (GLC) algorithm. A novel modeling technique based on the idea of artificial energy map (AEM) is proposed to find the optimal hovering position within a cluster. The flying strategy that minimizes the energy consumption is equivalently transformed into a classic travelling salesman problem that is readily solved by the genetic algorithm (GA). Through alternating iterative optimization of the clustering and hovering strategies, the communication energy consumption and the UAV hovering energy consumption are monotonically decreasing until convergence.
Lingfeng Shen, Ning Wang 0004, Zhengyu Zhu 0001, Yajun Fan, Xiaomin Mu
ICC1
2020 Energy-Awareness Dynamic Trajectory Planning for UAV-Enabled Data Collection in mMTC Networks
abstract
Massive machine-type communications (mMTC) is a key enabling technology for Internet of Things (IoT) services in 5G and beyond. Efficient data collection from massive machine-type communication devices (MTCDs) performing sensing tasks is an important part of the service. In this paper, we consider an unmanned aerial vehicle (UAV) being deployed to facilitate data collection from MTCDs. Taking into account the limited energy of battery-powered MTCDs, the UAV trajectory is optimized to improve the energy efficiency of data collection. By fixing the starting and ending points of the UAV trajectory, a globally optimal (GO) trajectory can be obtained, based on the assumption that the UAV's serving radius and access capacity (number of served MTCDs) are unlimited. Interestingly, it is shown that the optimal trajectory always exists as long as the UAV flying height is greater than its service radius multiplied by a constant. However, the increase of the UAV flying height deteriorates the channel, leading to reduced efficiency of energy consumption. Alternatively, a greedy dynamic (GD) trajectory optimization scheme with limited UAV service radius and access capacity is then investigated, resulting in the optimal service location of the UAV being at a lower flying height, and the energy consumption for accomplishing the data collection task being reduced. Specifically, the UAV sorts the MTCDs within its serving radius based on the distance and selects its closest serving MTCD set. In a serving MTCD set, there is an optimal hovering location that maximizes the data collection efficiency. The UAV dynamically adjusts its service set and the optimal data collection location when MTCDs finish their data transmission and exit the service set. The process continues until all the MTCDs are served and the UAV arrives at the ending point of the trajectory. Simulation results show that both the GO and GD algorithm can improve the efficiency of overall energy consumption. In particular, the online dynamic trajectory optimization scheme is less restrictive and achieves higher efficiency.
Lingfeng Shen, Ning Wang 0004, Jun Chen 0005, Xiaomin Mu, Kon Max Wong
VTC Fall1
2019 Trajectory Optimization for Physical Layer Secure Buffer-Aided UAV Mobile Relaying
abstract
In this work, we study the buffer-aided relaying mechanism in a UAV-enabled mobile relaying system assisting the terrestrial communications. Optimal UAV trajectory design against a randomly located eavesdropper is investigated from the physical layer (PHY) security perspective considering the wireless channel dynamics as the UAV relay moves in the air. Specifically, we maximize the sum secrecy rate by optimizing the discrete trajectory anchor points based on the information causality and UAV mobility constraints. To make the non- convex problem tractable, the increments of the trajectory anchor points are optimized instead through an iterative updating procedure, and successive convex approximation technique is applied for progressive optimization. The convergence of the proposed iterative optimization technique is proved by introducing additional rate bound constraints and employing the squeeze principle. Simulation results show that the proposed optimal trajectory finding algorithm is effective and fast converging. Simulation results also reveal that the distribution of the eavesdropper location has a significant impact on the PHY security performance.
Lingfeng Shen, Zhengyu Zhu 0001, Ning Wang 0004, Xiaomin Mu, Lin Cai 0001
VTC Fall1