VLDB 2026 Research / reviewers in the wild / expert
Jiangwei Li
dblp:38/6086
· DBLP profile ↗
21ranked-venue papers
7as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-authorArtificial intelligence and machine learning · 5 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MCD-YOLO: An improved method for detecting mixed-class defects on steel plate surfaces
Jiangwei Li, Maoxiang Chu, Peinan Zong, Simin Ma |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | Optimizing Containerized Edge Service Migration Through File-Level Storage Sharing
Jiangwei Li, Zhangbing Zhou, Sami Yangui, Deng Zhao, Ruixi Pan, Walid Gaaloul |
ICSOC (1) | 1 |
| 2025 | SCSTL: Spatial Composite Signal Temporal Logic for IoT Service Monitoring
Ruixi Pan, Zhangbing Zhou, Deng Zhao, Sami Yangui, Jiangwei Li |
ICSOC (1) | 5 |
| 2025 | BSTL: Bayesian Signal Temporal Logic for Predictive Edge Service MonitoringabstractEdge service monitoring is crucial for ensuring the robustness and reliability of service executions. Predictive monitoring, in particular, enables proactive detection of potential service violations. Existing predictive monitoring approaches, often leveraging Signal Temporal Logic (STL) for requirement specification, primarily focus on deterministic signals, and thus, lack probabilistic guarantees for uncertainty interpretation. To address these challenges, this paper introduces Bayesian STL (BSTL), an extension of STL that enables probabilistic reasoning over stochastic signals. Specifically, Bayesian Neural Networks (BNNs) are utilized to transform deterministic sequential predictions into sequences of posterior probability distributions. Uncertainty interpretation over these distribution predictions is achieved by a novel expected robustness metric that jointly quantifies both the degree and probability of service satisfaction. Thereafter, a BSTL-based predictive monitoring framework is developed, wherein service constraints are formally specified by BSTL formulae and interpreted with both qualitative and quantitative semantics. Besides, confidence levels and constraint thresholds ensuring robust satisfaction of BSTL formulae are rigorously estimated. Extensive experiments on publicly available datasets demonstrate that BSTL outperforms baseline techniques in expressiveness, robustness, and applicability. Deng Zhao, Zhangbing Zhou, Shuiguang Deng, Xiao Xue 0001, Ruixi Pan, Jiangwei Li, Sami Yangui |
ICWS | 6 |
| 2024 | Energy-Aware Service Migration in End-Edge-Cloud Collaborative NetworksabstractEmpowered by edge computing, resources and computation capabilities provided by edge devices can be encapsulated as containerized services. When burst requests are coming, there may have edge devices which are overloaded, since most requests are spatially and temporally constrained, and edge devices are resource-scarceness and capacity-limited. Overloaded devices should be relieved through optimally migrating one or more activated services to contiguous edge devices. Besides, sensory data gathered by original edge devices should be periodically transmitted to migrated devices for data analysis purpose. To mitigate this issue, this paper proposes an Energy-efficient Online Service Migration (EOSM) mechanism to conduct the migration of multiple services simultaneously. Extensive experimental results show that our EOSM mechanism outperforms the state of arts techniques in mitigating overloaded services in terms of access latency, energy consumption, and request success rate. Jiangwei Li, Zhangbing Zhou, Deng Zhao, Zhensheng Shi, Lin Meng 0001, Walid Gaaloul |
ICWS | 1 |
| 2024 | Lightweight Secure Communication Scheme Based on PUF for In-vehicle Controller Area NetworksabstractThe Controller Area Network (CAN) is the most widely used protocol for data transmission in in-vehicle networks. However, the lack of a robust security scheme makes it vulnerable to various cyber-attacks, posing significant threats to automotive functional safety. The limited computing resources of Electronic Control Units (ECUs), restricted network bandwidth, and high real-time requirements make it challenging to apply conventional encryption and authentication mechanisms. This paper proposes a lightweight secure communication scheme using encryption and authentication based on Physically Unclonable Function (PUF) technology. We first designed a ring oscillator PUF with a multi-loop comparison circuit (MLC-ROPUF) suitable for in-vehicle networks with constrained resources. The MLC-ROPUF is used to generate key parameter for the lightweight encryption algorithm Salsa20, which is applied to distribute shared key parameters to different ECUs. After receiving the CAN message, ECU utilizes the output response of MLC-ROPUF to generate the key stream of Salsa20 and decrypt the data segment of the CAN message, obtain the shared key parameters. The SM3 algorithm is then utilized to generate shared keys for confidential communication based on the obtained key parameters. Simultaneously, ECU identity authentication is completed to reduce communication costs. To validate the proposed scheme, we conducted experiments on a designed prototype system, and the results show that the proposed method performs well in terms of resource consumption and real-time capabilities. Yehua Wei, Wenjia Li, Jiangwei Li |
IPCCC | 4 |
| 2024 | Fast detection method for pedestrian video abnormal behavior based on keyframe extraction and multi-task mixed modelabstractIn recent years, many video anomaly detection methods have mainly used reconstruction and prediction based methods. However, due to the powerful encoding and decoding capabilities of autoencoders in reconstruction methods, the misjudgment rate of abnormal samples is high, and prediction methods are easily affected by environmental changes and data noise. Real time and accurate detection of pedestrian abnormal events still faces huge challenges. This article proposes a fast method VAD-KEMM for detecting abnormal behavior, which uses keyframe extraction and a multi task hybrid model. Firstly, key frames are extracted through segmented clustering and inter frame differences to improve detection efficiency; Then, a dual branch hybrid model is constructed using human skeletal information to improve reconstruction accuracy, and multi task learning is used to enhance prediction ability. The experimental results show that the AUC values of this method on the HR Shanghai Tech and HR Avenue datasets are 76.8% and 87.1%, respectively, indicating high detection efficiency and accuracy. Huiyu Mu, Jiangwei Li, Jiashuai Su, Luhui Wang, Lanxue Dang |
ISPA | 2 |
| 2024 | STA-VAD: Video Anomaly Detection Utilizing Spatio-Temporal Memory and Adaptive DeformationabstractClassical unsupervised video anomaly detection methods learn normal patterns from normal video frames and assume significant reconstruction deviations for unforeseen anomalous video frames. These methods often introduce additional constraints to distinguish anomalies, potentially leading to higher reconstruction errors for normal instances. While specific methods have alleviated this issue by estimating multiscale deformation fields, these rely on additional background information that is susceptible to changes in the real world. In this paper, we propose a spatio-temporal relationship-based video anomaly detection STA-VAD method. By emphasizing spatio-temporal relationships in videos, our model better captures contextual information, reducing the impact of irrelevant background information. Additionally, we enhance the multiscale deformation field by adaptively adjusting it to quantify the severity of anomalies in video frames, thereby improving discrimination between normal and anomalous samples. In the widely recognized video anomaly detection datasets UCSD Ped2 and CUHK Avenue, our method demonstrates significant enhancements in accuracy and reliability, achieving state-of-the-art performance overall. Huiyu Mu, Jiashuai Su, Luhui Wang, Jiangwei Li, Xianyu Zuo |
ISPA | 4 |
| 2024 | Energy-Efficient Online Service Migration in Edge NetworksabstractEmpowered by edge computing, resources and computation capabilities provided by edge devices can be encapsulated as containerized services, and domain applications can be achieved through service compositions. When burst requests are coming to be satisfied, there may exist edge devices which are overloaded, since requests are mostly spatially and temporally constrained, and edge devices are resource-scarceness and capacity-limited. In this setting, overloaded devices should be relieved through optimally migrating one or more activated services to contiguous edge devices. Besides, sensory data gathered by original edge devices should be periodically transmitted to migrated devices for data analysis purpose. To mitigate this issue, this paper proposes an Energy-efficient Online Service Migration (EOSM) mechanism to conduct the migration of multiple services simultaneously. Specifically, a light service sharing strategy is developed to only transmit the top container layer, and a modified NSGA-II algorithm is adopted to generate one or multiple paths for the container layer and time-series sensory data migration of each migrated service. Extensive experimental results show that our EOSM strategy outperforms the state of arts techniques in mitigating overloading devices in terms of access latency, energy consumption, and request success rate. Jiangwei Li, Deng Zhao, Zhensheng Shi, Lin Meng 0001, Walid Gaaloul, Zhangbing Zhou |
IEEE Internet Things J. | 1 |
| 2023 | Visual Localization Based on Multiple MapsabstractThis paper proposes a multi-map based visual localization method for image sequences. Given multiple single-map based localization results, we combine them with SLAM to estimate robust and accurate camera poses under challenging conditions. Our method comprises three modules connected in a sequence. First, we reconstruct multiple reference maps using the Structure-from-Motion technique, one map for each reference sequence. A single-image-based localization pipeline is performed to estimate 6-DoF camera poses for each query image, one for each map. Second, a consensus set maximization module is proposed to select the best camera poses from multi-map poses, estimating one 6-DoF camera pose for each query image. Finally, a robust pose refinement module is proposed to optimize 6-DoF camera poses of query images, combining map-based localization and local SLAM information. Experiments show that the proposed pipeline achieves state-of-the-art performance on challenging map-based localization benchmarks. Demonstrating the broad applicability of our method, we obtained first place in the challenge of Map-Based Localization for Autonomous Driving at ECCV2022. Yukai Lin, Liu Liu 0009, Jiangwei Li |
IROS | 4 |
| 2022 | Building an Efficient Retrieval-based Dialogue System with Contrastive LearningabstractWe focus on retrieval-based dialogue systems. Such a system aims to select an appropriate response from a candidate pool for a given context. Recent methods commonly utilize powerful interaction-based pre-trained language models like BERT to achieve the goal. However, their time cost is usually not satisfying since the procedure of computing relevance scores is not efficient, especially in scenarios that require online response selection. We propose an efficient dialogue system that utilizes a representation-based BERT to address this issue, which can produce an independent representation for every response candidate and context. The relevance score can be simply calculated by the dot product. We further enhance the representation ability of this model by applying domain adaptive post-training and supervised contrastive learning fine-tuning. Experimental results on two benchmark datasets show that our method achieves competitive performance with other interaction-based models while retaining the advantage of time efficiency. We also provide an empirical and theoretical analysis of time efficiency between representation-based models and interaction-based models. The main contribution of this paper is to propose a novel methodology to build a simple but efficient dialogue system. Jiangwei Li, Jinghui Zhong |
IJCNN | 1 |
| 2019 | MSCap: Multi-Style Image Captioning With Unpaired Stylized TextabstractIn this paper, we propose an adversarial learning network for the task of multi-style image captioning (MSCap) with a standard factual image caption dataset and a multi-stylized language corpus without paired images. How to learn a single model for multi-stylized image captioning with unpaired data is a challenging and necessary task, whereas rarely studied in previous works. The proposed framework mainly includes four contributive modules following a typical image encoder. First, a style dependent caption generator to output a sentence conditioned on an encoded image and a specified style. Second, a caption discriminator is presented to distinguish the input sentence to be real or not. The discriminator and the generator are trained in an adversarial manner to enable more natural and human-like captions. Third, a style classifier is employed to discriminate the specific style of the input sentence. Besides, a back-translation module is designed to enforce the generated stylized captions are visually grounded, with the intuition of the cycle consistency for factual caption and stylized caption. We enable an end-to-end optimization of the whole model with differentiable softmax approximation. At last, we conduct comprehensive experiments using a combined dataset containing four caption styles to demonstrate the outstanding performance of our proposed method. Longteng Guo, Jing Liu 0001, Jiangwei Li, Hanqing Lu |
CVPR | 4 |
| 2019 | Aligning Linguistic Words and Visual Semantic Units for Image CaptioningabstractImage captioning attempts to generate a sentence composed of several linguistic words, which are used to describe objects, attributes, and interactions in an image, denoted as visual semantic units in this paper. Based on this view, we propose to explicitly model the object interactions in semantics and geometry based on Graph Convolutional Networks (GCNs), and fully exploit the alignment between linguistic words and visual semantic units for image captioning. Particularly, we construct a semantic graph and a geometry graph, where each node corresponds to a visual semantic unit, i.e., an object, an attribute, or a semantic (geometrical) interaction between two objects. Accordingly, the semantic (geometrical) context-aware embeddings for each unit are obtained through the corresponding GCN learning processers. At each time step, a context gated attention module takes as inputs the embeddings of the visual semantic units and hierarchically align the current word with these units by first deciding which type of visual semantic unit (object, attribute, or interaction) the current word is about, and then finding the most correlated visual semantic units under this type. Extensive experiments are conducted on the challenging MS-COCO image captioning dataset, and superior results are reported when comparing to state-of-the-art approaches. Longteng Guo, Jing Liu 0001, Jinhui Tang 0001, Jiangwei Li, Hanqing Lu |
ACM Multimedia | 4 |
| 2013 | One-against-all frame differences based hand detection for human and mobile interaction
Kongqiao Wang, Lei Xu 0009, Yikai Fang, Jiangwei Li |
Neurocomputing | 4 |
| 2012 | An MDE-based approach to the verification of SysML state machine diagramabstractState Machine Diagram (SMD) is one of the SysML behavior diagrams, but it is a kind of semi-formal model language. As a consequence, models can not be verified conveniently and efficiently, especially in real-time embedded system (RTES) field as there are no descriptions of time and probability in SMD. To address these problems, we extend SMD with time and probability elements extracted from MARTE and propose a transformation algorithm based on MDE. With the algorithm, we transform the extended SMD to timed automata (TA) and then analyze and verify the transformation result using existing tools. So at the very beginning of system design, errors and deficiencies can be found. At last, we construct an instance to illustrate the validity of our approach. Xiaopu Huang, Qingqing Sun 0003, Jiangwei Li, Minxue Pan, Tian Zhang 0001 |
Internetware | 3 |
| 2012 | Plug&touch: a mobile interaction solution for large display via vision-based hand gesture detectionabstractSmart phones and external large displays are a superior combination for content sharing and multiple person interaction. However, the potential of such a combination is not fully explored due to the lack of effective and portable interaction solutions. To solve this problem, we present a novel mobile solution called Plug&Touch, which is very easy to be set up and enables mobile users to directly and naturally operate on the large display surface. Plug&Touch utilizes vision-based methods to detect the user's finger tip and then control user interface. The Local Binary Pattern asymmetric texture difference is proposed to extract the foreground region, and the coarse-to-fine intersection process is adopted to precisely localize the finger tip. Plug&Touch has been implemented on the Symbian3 platform with two defined gestures - click and scroll. Experimental results and user evaluation has shown the practicability of Plug&Touch. Lei Xu 0009, Yikai Fang, Kongqiao Wang, Jiangwei Li |
ACM Multimedia | 4 |
| 2011 | Real-time and multi-view face tracking on mobile platformabstractThe development of mobile platform has raised an emergent requirement for face-related multimedia applications. However, as the basis of such applications, face detection and tracking still suffers from large facial pose variation and limited computation resource. In this paper, a multi-view and real-time face tracking system is presented for mobile platform. First, three basic detectors are trained by local binary pattern (LBP) and Boosting algorithm. These detectors are then flexibly expanded for multi-view face detection through the rotation facility of LBP. To further accelerate the face tracking process, a robust facial pose estimation algorithm and the face matrix partition scheme is proposed. The experimental results on handset show the superior performance of the proposed method. Lei Xu 0009, Jiangwei Li, Kongqiao Wang |
ICASSP | 2 |
| 2011 | Local primitive code mining for fast and accurate face recognitionabstractThis paper proposes a new feature descriptor named as Local Primitive Code (LPC), which exhibits impressively discriminative capability on various face datasets. Essentially, LPC descriptors are somewhat like filter banks of Gaussian derivatives to capture multi-scale and multi-orientation image local textures but meanwhile enables faster feature extraction. It employs a framework composed of two stages: Firstly facial images are preprocessed by the pool of differential and quotient filters to generate numerous filtered images with various textures, and then directional binary encoding (DBE) operates on filtered images for primitive code mining. With such stages, feature maps with complementary discriminative information will be generated, and the fusion of them can greatly improve face recognition performance. Experimental results verify its performance even on some challenging databases. The algorithm is transplanted into mobile platform and achieves real-time performance on Nokia N82. Jiangwei Li, Lei Xu 0009, Kongqiao Wang |
ICIP | 1 |
| 2010 | Local binary pattern probability model based facial feature localizationabstractIn this paper, an active shape model (ASM) based facial feature localization strategy is proposed, which employs a local binary pattern (LBP) probability model. Due to the computation simplicity and illumination insensitivity of LBP texture descriptor and the learning ability of the probability model, the algorithm is robust and fast. In addition, component-based ASM is used to impose reasonable constraints on the shape. Multi-state shape and texture models with state classifier are trained to handle highly flexible components, i.e. eyes and mouth. Our database consisting of tens of persons with various expressions and illuminations is used to train and verify the proposed algorithm. The experiments demonstrate its accuracy, efficiency and robustness. Lei Xu 0009, Kongqiao Wang, Jiangwei Li |
ICIP | 4 |
| 2008 | Face indexing and searching from videosabstractFace indexing and searching from videos can be widely used in multimedia applications, information mining and biometric recognition. This paper presents a novel method for effective indexing and searching in videos. Differing from traditional methods, we sequentially integrate visual appearance with temporal continuity and position continuity for face comparison. Subgroups are so formed. Then, exemplars attached corresponding weights are selected from subgroups to form signatures. Lastly, subgroups are incorporated into groups with Earth Mover's Distance for hierarchical indexing and searching purposes. Experiments on some videos have demonstrated that the algorithm can achieve satisfying performance by full exploiting visual and timing information in videos. Jiangwei Li, Yunhong Wang 0001 |
ICIP | 1 |
| 2005 | A Secure Routing Protocol SDSR for Mobile Ad Hoc Networks
Chuanhe Huang, Jiangwei Li, Xiaohua Jia |
MSN | 2 |