Lanshun Nie

dblp:30/2639 · DBLP profile ↗
← Back
25ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 1 since 2021Computer networks · 4 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CRED: Calibrated Relational Enhanced Distillation for LLM-Based Pointwise Reranking
abstract
In document reranking, rerankers based on Large Language Models (LLMs) demonstrate superior performance but are constrained by high memory consumption and latency. To develop lightweight yet high-performance LLM-based pointwise rerankers through knowledge distillation, we identify two critical limitations: teachers often yield over-smoothed and inaccurate supervision on hard negative samples, thereby hindering the student's optimization; furthermore, traditional methods underutilize the relevance score differences between candidates, which are crucial for ranking tasks. To address these challenges, we propose CRED (Calibrated Relational Enhanced Distillation), which integrates Adaptive Teacher Calibration (ATC) to calibrate teacher predictions and amplify score margins, while employing Preference Relation Alignment (PRA) to align the distributional patterns of relevance score differences, enabling the student to capture precise ranking structures. To support this approach, we also construct FineDistill, a dataset of 1M samples providing fine-grained score supervision. We distill an 8B teacher into a 0.6B pointwise student. Extensive experiments on TREC and BEIR benchmarks show that our model outperforms leading baselines in both performance and generalization.
Qingran Yang 0001, Xue Li 0011, Lanshun Nie
SIGIR5
2026 Uncovering Underexplored Runtime Behaviors in ROS2-Based Autonomous Systems
abstract
Autonomous systems built on ROS2 are increasingly deployed in safety- and performance-critical domains such as autonomous driving and mobile robotics. While existing research has proposed various timing analyses and scheduling strategies for ROS2, many rely on simplified assumptions that do not hold in real-world applications. In this article, we present a detailed empirical study of ROS2-based autonomous applications, uncovering underexplored runtime behaviors that significantly impact both real-time and functional performance. These include the importance of partial cause-effect chains, dynamic execution paths and timing variability, non-FIFO data access patterns, and computation threads uncontrolled by ROS2 executors. We extend an existing tracing tool to support Transform Library and ROS2’s Action entity, enabling reconstruction and analysis of realistic cause-effect chains. Our findings are validated through experiments in simulated autonomous robot scenarios and a case study using the Autoware autonomous driving framework. Together, our results highlight the need for rethinking ROS2 modeling, scheduling and analysis to better reflect the realities of autonomous systems.
Chenghao Fan, Lanshun Nie, Jing Li 0025
ACM Trans. Internet Things2
2025 LLaPa: A Vision-Language Model Framework for Counterfactual-Aware Procedural Planning
abstract
While large language models (LLMs) have advanced procedural planning for embodied AI systems through strong reasoning abilities, the integration of multimodal inputs and counterfactual reasoning remains underexplored. To tackle these challenges, we introduce LLaPa, a vision-language model framework designed for multimodal procedural planning. LLaPa generates executable action sequences from textual task descriptions and visual environmental images using vision-language models (VLMs). Furthermore, we enhance LLaPa with two auxiliary modules to improve procedural planning. The first module, the Task-Environment Reranker (TER), leverages task-oriented segmentation to create a task-sensitive feature space, aligning textual descriptions with visual environments and emphasizing critical regions for procedural execution. The second module, the Counterfactual Activities Retriever (CAR), identifies and emphasizes potential counterfactual conditions, enhancing the model's reasoning capability in counterfactual scenarios. Extensive experiments on ActPlan-1K and ALFRED benchmarks demonstrate that LLaPa generates higher-quality plans with superior LCS and correctness, outperforming advanced models. The code and models are available https://github.com/sunshibo1234/LLaPa.
Shibo Sun, Xue Li 0011, Donglin Di, Lanshun Nie, Weinan Zhang 0003, Dechen Zhan, Yang Song 0001, Lei Fan 0007
ACM Multimedia5
2025 SAGE: A Visual Language Model for Anomaly Detection via Fact Enhancement and Entropy-aware Alignment
abstract
While Vision-Language Models (VLMs) have shown promising progress in general multimodal tasks, they often struggle with industrial anomaly detection and reasoning, particularly in delivering interpretable explanations and generalizing to unseen categories. This limitation stems from the inherently domain-specific nature of anomaly detection, which hinders the applicability of existing VLMs in industrial scenarios that require precise, structured, and context-aware analysis. To address these challenges, we propose SAGE, a VLM-based framework that enhances anomaly reasoning through Self-Guided Fact Enhancement (SFE) and Entropy-aware Direct Preference Optimization (E-DPO). SFE integrates domain-specific knowledge into visual reasoning via fact extraction and fusion, while E-DPO aligns model outputs with expert preferences using entropy-aware optimization. Additionally, we introduce AD-PL, a preference-optimized dataset tailored for industrial anomaly reasoning, consisting of 28,415 question-answering instances with expert-ranked responses. To evaluate anomaly reasoning models, we develop Multiscale Logical Evaluation (MLE), a quantitative framework analyzing model logic and consistency. SAGE demonstrates superior performance on industrial anomaly datasets under zero-shot and one-shot settings. The code, model, and dataset are available at https://github.com/amoreZgx1n/SAGE.
Guoxin Zang, Xue Li 0011, Donglin Di, Lanshun Nie, Dechen Zhan, Yang Song 0001, Lei Fan 0007
ACM Multimedia4
2025 S2C-HAR: A Semi-Supervised Human Activity Recognition Framework Based on Contrastive Learning
abstract
ABSTRACT Human activity recognition (HAR) has emerged as a critical element in various domains, such as smart healthcare, smart homes, and intelligent transportation, owing to the rapid advancements in wearable sensing technology and mobile computing. Nevertheless, existing HAR methods predominantly rely on deep supervised learning algorithms, necessitating a substantial supply of high‐quality labeled data, which significantly impacts their accuracy and reliability. Considering the diversity of mobile devices and usage environments, the quest for optimizing recognition performance in deep models while minimizing labeled data usage has become a prominent research area. In this paper, we propose a novel semi‐supervised HAR framework based on contrastive learning named S2C‐HAR, which is capable of generating accurate pseudo‐labels for unlabeled data, thus achieving comparable performance with supervised learning with only a few labels applied. First, a contrastive learning model for HAR (CLHAR) is designed for more general feature representations, which contains a contrastive augmentation transformer pre‐trained exclusively on unlabeled data and fine‐tuned in conjunction with a model‐agnostic classification network. Furthermore, based on the FixMatch technique, unlabeled data with two different perturbations imposed are fed into the CLHAR to produce pseudo‐labels and prediction results, which effectively provides a robust self‐training strategy and improves the quality of pseudo‐labels. To validate the efficacy of our proposed model, we conducted extensive experiments, yielding compelling results. Remarkably, even with only 1% labeled data, our model achieves satisfactory recognition performance, outperforming state‐of‐the‐art methods by approximately 5%.
Xue Li 0011, Mingxing Liu, Lanshun Nie, Wenxiao Cheng, Xiaohe Wu, Dechen Zhan
Concurr. Comput. Pract. Exp.3
2023 Generalizable Reinforcement Learning-Based Coarsening Model for Resource Allocation over Large and Diverse Stream Processing Graphs
abstract
Resource allocation for stream processing graphs on computing devices is critical to the performance of stream processing. Efficient allocations need to balance workload distribution and minimize communication simultaneously and globally. Since this problem is known to be NP-complete, recent machine learning solutions were proposed based on an encoder-decoder framework, which predicts the device assignment of computing nodes sequentially as an approximation. However, for large graphs, these solutions suffer from the deficiency in handling long-distance dependency and global information, resulting in suboptimal predictions. This work proposes a new paradigm to deal with this challenge, which first coarsens the graph and conducts assignments on the smaller graph with existing graph partitioning methods. Unlike existing graph coarsening works, we leverage the theoretical insights in this resource allocation problem, formulate the coarsening of stream graphs as edge-collapsing predictions, and propose an edge-aware coarsening model. Extensive experiments on various datasets show that our framework significantly improves over existing learning-based and heuristic-based baselines with up to 56% relative improvement on large graphs.
Lanshun Nie, Yuqi Qiu, Mo Yu, Jing Li 0025
IPDPS1
2022 Few shot learning-based fast adaptation for human activity recognition
Lanshun Nie, Xue Li 0011, Tianying Gong, Dechen Zhan
Pattern Recognit. Lett.1
2022 Holistic Resource Allocation Under Federated Scheduling for Parallel Real-time Tasks
abstract
With the technology trend of hardware and workload consolidation for embedded systems and the rapid development of edge computing, there has been increasing interest in supporting parallel real-time tasks to better utilize the multi-core platforms while meeting the stringent real-time constraints. For parallel real-time tasks, the federated scheduling paradigm, which assigns each parallel task a set of dedicated cores, achieves good theoretical bounds by ensuring exclusive use of processing resources to reduce interferences. However, because cores share the last-level cache and memory bandwidth resources, in practice tasks may still interfere with each other despite executing on dedicated cores. Such resource interferences due to concurrent accesses can be even more severe for embedded platforms or edge servers, where the computing power and cache/memory space are limited. To tackle this issue, in this work, we present a holistic resource allocation framework for parallel real-time tasks under federated scheduling. Under our proposed framework, in addition to dedicated cores, each parallel task is also assigned with dedicated cache and memory bandwidth resources. Further, we propose a holistic resource allocation algorithm that well balances the allocation between different resources to achieve good schedulability. Additionally, we provide a full implementation of our framework by extending the federated scheduling system with Intel’s Cache Allocation Technology and MemGuard. Finally, we demonstrate the practicality of our proposed framework via extensive numerical evaluations and empirical experiments using real benchmark programs.
Lanshun Nie, Chenghao Fan, Shuang Lin, Jing Li 0025
ACM Trans. Embed. Comput. Syst.1
2021 Enhancing Representation of Deep Features for Sensor-Based Activity Recognition
Xue Li 0011, Lanshun Nie, Xiandong Si, Renjie Ding, Dechen Zhan
Mob. Networks Appl.2
2020 An Improved Weighted-Removal Sentence Embedding Based Approach for Service Recommendation
abstract
Currently, there is a large amount of information about user requirements and service in natural language. How to measure the semantic similarity between user requirements and service description is a critical issue in service recommendation and service solution construction. In this paper, we propose a service recommendation method based on the improved Weighted-Removal(WR) sentence embedding to solve the shortcomings of traditional information retrieval methods. After data preprocessing, we use the GloVe method to obtain the word vectors and use the improved WR sentence embedding method to obtain the sentence vectors. The similarity between the vectors can be better measured. The experimental results show that the proposed improved WR method is significantly better than the traditional methods in terms of recommendation accuracy, richness, and ranking.
Hanchuan Xu, Xiao Wang 0097, Lanshun Nie, Xiaofei Xu 0001
ICSS4
2020 A Class Incremental Temporal-Spatial Model Based on Wireless Sensor Networks for Activity Recognition
Xue Li 0011, Lanshun Nie, Xiandong Si, Dechen Zhan
WASA (1)2
2019 Balanced Sparsity for Efficient DNN Inference on GPU
abstract
In trained deep neural networks, unstructured pruning can reduce redundant weights to lower storage cost. However, it requires the customization of hardwares to speed up practical inference. Another trend accelerates sparse model inference on general-purpose hardwares by adopting coarse-grained sparsity to prune or regularize consecutive weights for efficient computation. But this method often sacrifices model accuracy. In this paper, we propose a novel fine-grained sparsity approach, Balanced Sparsity, to achieve high model accuracy with commercial hardwares efficiently. Our approach adapts to high parallelism property of GPU, showing incredible potential for sparsity in the widely deployment of deep learning services. Experiment results show that Balanced Sparsity achieves up to 3.1x practical speedup for model inference on GPU, while retains the same high model accuracy as finegrained sparsity.
Zhuliang Yao, Shijie Cao, Wencong Xiao, Chen Zhang 0001, Lanshun Nie
AAAI5
2019 SeerNet: Predicting Convolutional Neural Network Feature-Map Sparsity Through Low-Bit Quantization
abstract
In this paper we present a novel and general method to accelerate convolutional neural network (CNN) inference by taking advantage of feature map sparsity. We experimentally demonstrate that a highly quantized version of the original network is sufficient in predicting the output sparsity accurately, and verify that leveraging such sparsity in inference incurs negligible accuracy drop compared with the original network. To accelerate inference, for each convolution layer our approach first obtains a binary sparsity mask of the output feature maps by running inference on a quantized version of the original network layer, and then conducts a full-precision sparse convolution to find out the precise values of the non-zero outputs. Compared with existing work, our approach avoids the overhead of training additional auxiliary networks, while is still applicable to general CNN networks without being limited to certain application domains.
Shijie Cao, Lingxiao Ma, Wencong Xiao, Chen Zhang 0001, Yunxin Liu 0001, Lanshun Nie, Zhi Yang 0001
CVPR7
2019 Poster: IEEE 802.11ax User Scheduling Algorithm for Low Latency
Yingjie Zeng, Lanshun Nie
EWSN2
2019 Poster: Cross-Technology Communication via Phase Shift Emulation
Yingjie Zeng, Lanshun Nie
EWSN2
2019 Poster: Channel Prediction Based on BP Neural Network for Backscatter Communication Networks
Yingjie Zeng, Lanshun Nie
EWSN2
2019 Poster: Fast and Reliable Burst Data Transmission for Backscatter Communications
Yingjie Zeng, Lanshun Nie
EWSN2
2019 Poster: Intelligent Management of Chemical Warehouse with RFID Systems
Yingjie Zeng, Lanshun Nie
EWSN2
2019 Poster: Deep Gait Recognition via Millimeter Wave
Yingjie Zeng, Lanshun Nie
EWSN2
2019 Efficient and Effective Sparse LSTM on FPGA with Bank-Balanced Sparsity
abstract
Neural networks based on Long Short-Term Memory (LSTM) are widely deployed in latency-sensitive language and speech applications. To speed up LSTM inference, previous research proposes weight pruning techniques to reduce computational cost. Unfortunately, irregular computation and memory accesses in unrestricted sparse LSTM limit the realizable parallelism, especially when implemented on FPGA. To address this issue, some researchers propose block-based sparsity patterns to increase the regularity of sparse weight matrices, but these approaches suffer from deteriorated prediction accuracy. This work presents Bank-Balanced Sparsity (BBS), a novel sparsity pattern that can maintain model accuracy at a high sparsity level while still enable an efficient FPGA implementation. BBS partitions each weight matrix row into banks for parallel computing, while adopts fine-grained pruning inside each bank to maintain model accuracy. We develop a 3-step software-hardware co-optimization approach to apply BBS in real FPGA hardware. First, we propose a bank-balanced pruning method to induce the BBS pattern on weight matrices. Then we introduce a decoding-free sparse matrix format, Compressed Sparse Banks (CSB), that transparently exposes inter-bank parallelism in BBS to hardware. Finally, we design an FPGA accelerator that takes advantage of BBS to eliminate irregular computation and memory accesses. Implemented on Intel Arria-10 FPGA, the BBS accelerator can achieve 750.9 GOPs on sparse LSTM networks with a batch size of 1. Compared to state-of-the-art FPGA accelerators for LSTM with different compression techniques, the BBS accelerator achieves 2.3 ~ 3.7x improvement on energy efficiency and 7.0 ~ 34.4x reduction on latency with negligible loss of model accuracy.
Shijie Cao, Chen Zhang 0001, Zhuliang Yao, Wencong Xiao, Lanshun Nie, Dechen Zhan, Yunxin Liu 0001, Ming Wu 0007
FPGA5
2019 FlexSaaS: A Reconfigurable Accelerator for Web Search Selection
abstract
Web search engines deploy large-scale selection services on CPUs to identify a set of web pages that match user queries. An FPGA-based accelerator can exploit various levels of parallelism and provide a lower latency, higher throughput, more energy-efficient solution than commodity CPUs. However, maintaining such a customized accelerator in a commercial search engine is challenging because selection services are changed often. This article presents our design for FlexSaaS (Flexible Selection as a Service), an FPGA-based accelerator for web search selection. To address efficiency and flexibility challenges, FlexSaaS abstracts computing models and separates memory access from computation. Specifically, FlexSaaS (i) contains a reconfigurable number of matching processors that can handle various possible query plans, (ii) decouples index stream reading from matching computation to fetch and decode index files, and (iii) includes a universal memory accessor that hides the complex memory hierarchy and reduces host data access latency. Evaluated on FPGAs in the selection service of a commercial web search--the Bing web search engine—FlexSaaS can be evolved quickly to adapt to new updates. Compared to the software baseline, FlexSaaS on Arria 10 reduces average latency by 30% and increases throughput by 1.5×.
Shijie Cao, Lanshun Nie, Dechen Zhan, Ningyi Xu, Ramashis Das, Ming Wu 0007, Derek Chiou
ACM Trans. Reconfigurable Technol. Syst.2
2018 Collaborative Fall Detection Using Smart Phone and Kinect
Xue Li 0001, Lanshun Nie, Hanchuan Xu, Xianzhi Wang 0001
Mob. Networks Appl.2
2016 Real-Time Wireless Sensor-Actuator Networks for Industrial Cyber-Physical Systems
abstract
With recent adoption of wireless sensor-actuator networks (WSANs) in industrial automation, industrial wireless control systems have emerged as a frontier of cyber-physical systems. Despite their success in industrial monitoring applications, existing WSAN technologies face significant challenges in supporting control systems due to their lack of real-time performance and dynamic wireless conditions in industrial plants. This article reviews a series of recent advances in real-time WSANs for industrial control systems: 1) real-time scheduling algorithms and analyses for WSANs; 2) implementation and experimentation of industrial WSAN protocols; 3) cyber-physical codesign of wireless control systems that integrate wireless and control designs; and 4) a wireless cyber-physical simulator for codesign and evaluation of wireless control systems. This article concludes by highlighting research directions in industrial cyber-physical systems.
Chenyang Lu 0001, Abusayeed Saifullah, Bo Li 0020, Mo Sha 0001, Humberto González, Dolvara Gunatilaka, Chengjie Wu, Lanshun Nie, Yixin Chen 0001
Proc. IEEE8
2010 Comprehensive Practice on Service Engineering: An Experimental Solution
abstract
To develop a well-designed curriculum of service engineering is a prerequisite for training students get new skills and ability to design and develop globally effective solutions in a service business environment. As service engineering is an application-oriented discipline, sufficient practice is extremely important. In this paper, we show an experimental solution of a course "comprehensive practice on service engineering", which has been explored for three years in Harbin Institute of Technology. Its objective, process, execution program, and detailed design of each phase, are elaborately presented. Lessons learned from the explorations and future improvement on this solution, are briefly discussed.
Zhongjie Wang 0003, Xiaofei Xu 0001, Lanshun Nie, Tianyi Zang
ICSS4
2006 Model for Negotiating Prices and Due Dates with Suppliers in Make-to-Order Supply Chains
Lanshun Nie, Xiaofei Xu 0001, Dechen Zhan
PRIMA1