VLDB 2026 Research / reviewers in the wild / expert
Hailong Shi
dblp:125/2282
· DBLP profile ↗
21ranked-venue papers
5as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-authorComputer networks · 4 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorSecurity and privacy · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | STKPS-Net: Spatio-Temporal Key Patch Selection Network for Few Shot Anomalous Action RecognitionabstractFor providing timely warnings and preventing potential damages, it is crucial to detect anomalous actions that threaten public safety through surveillance cameras. Compared to normal actions, anomalous actions often occupy only a small portion of surveillance videos and exhibit more complex manifestations in terms of time and space. Considering that normal action recognition methods fail to highlight crucial information from small-sized patches, we propose the Spatio-temporal Key Patch Selection Network (STKPS-Net). It includes a spatially adaptive key patch selection module to select small but informative patches, and a long-short feature map spatio-temporal relation module to capture dynamic changes in anomalous actions. Additionally, a spatio-temporal refined loss is introduced to enhance fine-grained feature learning. Experimental results on the HMDB51, Kinetics, and UCF-Crime v2 datasets show that our STKPS-Net achieves state-of-the-art performance in few-shot anomalous action recognition, outperforming the most competitive methods by 1.2% on the anomalous action dataset UCF-Crime v2. Jinsheng Xiao, Ruidi Chen, Xingyu Gao 0001, Hailong Shi, Zhongyuan Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2026 | Multi-Perspective Visual Contrastive Decoding for Reliable AssistanceabstractMultimodal Large Language Models (MLLMs) offer promising capabilities for assisting individuals with blindness and low vision (BLV), but their effectiveness is compromised when processing BLV-captured images, which typically suffer from three fundamental challenges: quality degradation, object incompleteness, and spatial misalignment. This article presents MPVCD (Multi-Perspective Visual Contrastive Decoding), a novel framework that addresses these challenges through visual contrastive decoding techniques. MPVCD implements three specialized perspectives: Noise Contrastive Decoding addresses quality issues by comparing predictions between original and noise-injected images; Retrieval Contrastive Decoding tackles object incompleteness by retrieving semantically similar images from a memory bank; and Focus Contrastive Decoding resolves spatial misalignment by focusing on detected object regions. These perspectives are dynamically balanced through an Adaptive Perspective Integration that optimizes token selection based on prediction confidence. Our comprehensive experiments across diverse datasets demonstrate MPVCD’s effectiveness in reducing hallucinations under varied scenarios. By generating more accurate and reliable visual descriptions, MPVCD represents a significant advancement toward assistive technologies that BLV users can confidently rely on for environmental understanding and decision-making. Bocheng Pan, Hailong Shi, Xingyu Gao 0001 |
ACM Trans. Internet Things | 2 |
| 2026 | mmWave Radar-based Personalized Multi-object Vital Signs MonitoringabstractFrequency Modulated Continuous Wave (FMCW)-based mmWave radar has attracted widespread attention because of its non-contact and high spatial resolution for vital signs monitoring. Meanwhile, current studies focus mainly on how to improve the detection performance of steady multiple objects or unsteady single objects. In this work, we propose an innovative method for identity-based multi-object vital signs monitoring under unsteady scenarios. The method automatically distinguishes between steady and motion states, and conducts a best-effort vital signs monitoring during unsteady scenarios. To this end, we design a weight vector enhancement method combined with object spatial positioning for differentiating multiple objects, and identify each object according to the gait-based EfficientNet model. We also design a steady-state detector based on the MobileNet-V2 network to find the slots of object keeping steady for vital signs monitoring and then apply the variational mode decomposition (VMD) algorithm to extract the respiratory and heart rates of a single object. The experimental results showed that the mean absolute error of respiratory rate and heart rate decreased to 1.37 bpm and 2.56 bpm respectively in the case of multiple objects. In addition, the steady-state detector achieves close to 98.1% accuracy in recognizing motion types, and the average recognition rate of identity recognition based on gait features reaches about 93.26%. Jiefan Qiu, Xingyu Gao 0001, Dongfu Zhu, Mengqi Jiang, Jiahan Song, Hailong Shi |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2025 | Efficient Parallel Training Methods for Spiking Neural Networks with Constant Time ComplexityabstractSpiking Neural Networks (SNNs) often suffer from high time complexity $O(T)$ due to the sequential processing of $T$ spikes, making training computationally expensive.
In this paper, we propose a novel Fixed-point Parallel Training (FPT) method to accelerate SNN training without modifying the network architecture or introducing additional assumptions.
FPT reduces the time complexity to $O(K)$, where $K$ is a small constant (usually $K=3$), by using a fixed-point iteration form of Leaky Integrate-and-Fire (LIF) neurons for all $T$ timesteps.
We provide a theoretical convergence analysis of FPT and demonstrate that existing parallel spiking neurons can be viewed as special cases of our approach.
Experimental results show that FPT effectively simulates the dynamics of original LIF neurons, significantly reducing computational time without sacrificing accuracy.
This makes FPT a scalable and efficient solution for real-world applications, particularly for long-duration simulations. Wanjin Feng, Xingyu Gao 0001, Wenqian Du 0005, Hailong Shi, Peilin Zhao, Chunyan Miao |
ICML | 4 |
| 2025 | DR-VQA: Decompose-then-Reconstruct for Visual Question Answering in BLV AssistanceabstractVisual impairment affects over 200 million individuals globally, creating significant challenges in daily visual tasks. While vision-language models offer transformative assistive potential, existing systems based on Multimodal Large Language Models (MLLMs) face a serious cross-contamination problem when processing real-world images captured by blind and low-vision (BLV) users: when jointly processing imperfect images and specific questions, current models are often misled by question assumptions rather than adhering to visual facts, generating hallucinations about objects not present in the image. We introduce DR-VQA (Decompose-then-Reconstruct Visual Question Answering), a novel framework that balances user intent with visual facts. Our approach prevents cross-contamination through structured reasoning. Our approach deliberately separates image processing from question analysis, ensuring model-generated descriptions are strictly based on image facts without being influenced by questions. Subsequently, through a structured decomposition mechanism, the system generates targeted sub-questions relevant to user intent, gradually aligning visual descriptions with user needs while minimizing question bias. During final synthesis, a memory-reset LLM reconstructs the reasoning chain with detailed information to generate responses that either provide evidence-supported conclusions or transparently acknowledge information limitations. Experimental evaluations demonstrate our framework's effectiveness in reducing hallucination risks while improving answer accuracy. By systematically balancing user intent with factual visual evidence, this work advances BLV-assistive technologies from probabilistic outputs to reliable visual assistance services. Bocheng Pan, Hailong Shi, Xingyu Gao 0001 |
ACM Multimedia | 2 |
| 2025 | Scribble-Supervised Video Object Segmentation via Scribble EnhancementabstractCurrent video object segmentation methods heavily rely on pixel-level mask annotations when training, which are expensive and time-consuming to acquire. To address this problem, some approaches try to train with sparse scribble annotations and take sparse target scribble as initial information for inference. However, due to the sparsity of scribble annotations, the performance is often limited, and the corresponding loss function needs to be designed. Inspired by the powerful ability of Segment Anything Model (SAM) to leverage prompt for segmentation, we argue that this problem can be alleviated by improving the quality of scribble. Therefore, we propose SEVOS, a framework for scribble-supervised video object segmentation, which contains a scribble enhancement algorithm and an semi-supervised video object segmentation network. Specifically, the scribble enhancement algorithm first samples corresponding positive sample points and negative sample points from target scribbles, and then feeds them into the SAM in turn, achieving high-quality scribble enhancement without human intervention. This algorithm augments the scribble-annotated video dataset, which is used for additional training of the model. Furthermore, we design a post-processing enhancement algorithm to further improve the prediction results. The obtained model outperforms state-of-the-art methods with a considerable performance gap, indicating the generalization and effectiveness of the proposed model. Xingyu Gao 0001, Zuolei Li, Hailong Shi, Zhenyu Chen 0003, Peilin Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Brain-Inspired Fast- and Slow-Update Prompt Tuning for Few-Shot Class-Incremental LearningabstractFew-shot class-incremental learning (FSCIL) aims to learn new classes incrementally with a limited number of samples per class. Foundation models combined with prompt tuning showcase robust generalization and zero-shot learning (ZSL) capabilities, endowing them with potential advantages in transfer capabilities for FSCIL. However, existing prompt tuning methods excel in optimizing for stationary datasets, diverging from the inherent sequential nature in the FSCIL paradigm. To address this issue, taking inspiration from the "fast and slow mechanism" of the complementary learning systems (CLSs) in the brain, we present fast- and slow-update prompt tuning FSCIL (FSPT-FSCIL), a brain-inspired prompt tuning method for transferring foundation models to the FSCIL task. We categorize the prompts into two groups: fast-update prompts and slow-update prompts, which are interactively trained through meta-learning. Fast-update prompts aim to learn new knowledge within a limited number of iterations, while slow-update prompts serve as meta-knowledge and aim to strike a balance between rapid learning and avoiding catastrophic forgetting. Through experiments on multiple benchmark tests, we demonstrate the effectiveness and superiority of FSPT-FSCIL. The code is available at https://github.com/qihangran/FSPT-FSCIL. Hang Ran, Xingyu Gao 0001, Lusi Li, Weijun Li 0002, Songsong Tian, Gang Wang 0023, Hailong Shi, Xin Ning 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2025 | Maximizing Long-Term Task Completion Ratio of UAV-Enabled Wirelessly Powered MEC SystemsabstractUnmanned Aerial Vehicle (UAV)-enabled wirelessly powered Mobile Edge Computing (MEC) is emerging as a powerful technology for boosting computational capability and energy supplementation in Internet of Things (IoT). This work addresses the long-term task completion ratio maximization problem in UAV-enabled wirelessly powered MEC systems. Besides the large number of optimization parameters, the environment can only be partially observed as the UAVs cannot cover the whole network area. Then, it is very challenging to obtain good solutions due to the lack of global information. We introduce a novel distributed Multi-Agent Deep Reinforcement Learning (MADRL) framework for optimizing UAVs’ actions and resource allocation, considering the constraints of tasks that vary in size, arrival times, and required computation completion time. To decouple the complicated parameters, we divide the problem into two manageable subproblems—UAVs’ action decision and resource allocation under a given UAV’s action. We employ a distributed Deep Reinforcement Learning (DRL) scheme for the former subproblem to cope with the partially observable nature. By revealing some important properties of the later subproblem, we design an efficient two-stage optimal algorithm to minimize the total consumed energy of nodes while maximizing the task-completing number. Extensive simulations validate the effectiveness of the proposed framework, achieving over a 50% improvement in task completion ratio compared to baseline schemes in some scenarios. Shaojun Zhu, Bingcheng Zhu, Kaikai Chi, Jiefan Qiu, Hailong Shi, Xingyu Gao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | REDIR: Refocus-Free Event-Based De-occlusion Image Reconstruction
Hailong Shi, Jinsheng Xiao, Xingyu Gao 0001 |
ECCV (80) | 2 |
| 2024 | An Adaptive Framework of Geographical Group-Specific Network on O2O Recommendation
Luo Ji, Jiayu Mao, Hailong Shi, Yunfei Chu, Hongxia Yang |
ECIR (3) | 3 |
| 2024 | A Coarse-to-Fine Fusion Network for Event-Based Image Deblurring
Hailong Shi, Xingyu Gao 0001 |
IJCAI | 2 |
| 2023 | Mixtron: Bandit Online Multiclass Prediction with Implicit FeedbackabstractThe exploitation and exploration dilemma is a crucial issue in bandit online multiclass prediction. Conventional algorithms typically resort to either random sample or estimate uncertainty for exploration. In contrast, we propose a novel scheme that focuses solely on exploitation with implicit feedback. To ensure efficient information feedback even when predictions are incorrect, we introduce mixed losses into our proposed scheme. We derive two mixed versions of the multiclass hinge loss and the logistic loss, along with their corresponding algorithms. Context-free and context-aware experiments are conducted to evaluate the performance of our proposed algorithms. Experiment results demonstrate that random sampling are unnecessary if a reasonable loss function is employed. By compared with several state-of-the-art baselines on both synthetic and real-world datasets, our proposed algorithms show their superior performance. Remarkably, they even outperform the Perceptron algorithm with full-information feedback in some cases. Wanjin Feng, Hailong Shi, Peilin Zhao, Xingyu Gao 0001 |
ICDM | 2 |
| 2016 | EasiTMC: Transportation Mode Classification with a High Accuracy Trajectory Detection MethodabstractRecently, smartphones are increasingly used in situation-aware applications, e.g. determining the way of transportation when an individual is going outside. However the existing methods of transportation mode classification often suffer from high computing complexity and unsatisfactory accuracy. In this paper we present EasiTMC, a two-stage hierarchical classification framework which consists of coarse- grained and fine-grained classifiers to automatically infer transportation modes from a GPS receiver and an accelerometer. We aim to classify different forms of transportation, including non-motorized modes such as walking, running, and biking, as well as motorized modes such as bus and car rides. The primary contributions of our work include a novel segment-based partition algorithm for achieving high accuracy with low computational complexity, a feature selection method based on overlap entropy for improving the accuracy of classification and a merge-amendment algorithm for decreasing the classification error. EasiTMC is evaluated with over 200 hours of transportation traces from seven individuals. Compared with two other segmentation methods, namely the uniform distance based and the uniform time interval based methods, the proposed partition method achieved higher degree of partition accuracy with lower computing complexity. The overall accuracy in identifying ways of transportation is about 93.3% for different users. Changtian Yuan, Hailong Shi |
GLOBECOM | 4 |
| 2015 | A volume correlation subspace detector for signals buried in unknown clutterabstractDetecting the presence of target subspace signals with unknown clutters is a well-known hard problem encountered in various signal processing applications. Traditional methods fails to solve this problem because prior knowledge of clutter subspace is required, which can not be obtained when target and clutter are intimately mixed. In this paper, we propose a novel subspace detector that can detect target signal buried in clutter without knowledge of clutter subspace. This detector makes use of the geometrical relation between target and clutter subspaces and is derived based upon the calculation of volume of high dimensional geometrical objects. Moreover, the proposed detector can accomplish the detection simultaneously with the learning processes of clutter, a property called “detecting while learning”. The performance of detector was showed by theoretical analysis and numerical simulation. Hailong Shi, Hao Zhang 0005, Xiqin Wang |
ISIT | 1 |
| 2015 | Stable Embedding of Grassmann Manifold via Gaussian Random MatricesabstractCompressive sensing (CS) provides a new perspective for data reduction without compromising performance when the signal of interest is sparse or has intrinsically low-dimensional structure. The theoretical foundation for most of the existing studies on CS is based on the stable embedding (i.e., a distance-preserving property) of vectors that are sparse or in a union of subspaces via random measurement matrices. To the best of our knowledge, few existing literatures of CS have clearly discussed the stable embedding of linear subspaces via compressive measurement systems. In this paper, we explore a volume-based stable embedding of multidimensional signals based on Grassmann manifold, via Gaussian random measurement matrices. The Grassmann manifold is a topological space, in which each point is a linear vector subspace, and is widely regarded as an ideal model for multidimensional signals generated from linear subspaces. In this paper, we formulate the linear subspace spanned by multidimensional signal vectors as points on the Grassmann manifold, and use the volume and the product of sines of principal angles (also known as the product of principal sines) as the generalized norm and distance measure for the space of Grassmann manifold. We prove a volume-preserving embedding property for points on the Grassmann manifold via Gaussian random measurement matrices, i.e., the volumes of all parallelotopes from a finite set in Grassmann manifold are preserved upon compression. This volume-preserving embedding property is a multidimensional generalization of the conventional stable embedding properties, which only concern the approximate preservation of lengths of vectors in certain unions of subspaces. In addition, we use the volume-preserving embedding property to explore the stable embedding effect on a generalized distance measure of Grassmann manifold induced from volume. It is proved that the generalized distance measure, i.e., the product of principal sines between different points on the Grassmann manifold, is well preserved in the compressed domain via Gaussian random measurement matrices. Numerical simulations are also provided for validation. Hailong Shi, Hao Zhang 0005, Gang Li 0008, Xiqin Wang |
IEEE Trans. Inf. Theory | 1 |
| 2014 | EasiCAE: A runtime framework for efficient sensor sharing among concurrent IoT applicationsabstractTraditional wireless sensor networks (WSNs) can be integrated into Internet and be regarded as its sensing infrastructure, which supports development and running of multiple third-party applications simultaneously. Therefore, due to constrained resource of sensor nodes, it is necessary to establish a runtime framework to improve sensor sharing efficiency for concurrent third-party applications. This paper presents EasiCAE, a concurrent applications runtime framework, to enhance sensor sharing efficiency greatly by incorporating task allocation with redundancy elimination. In brief, EasiCAE decompose the applications into tasks and distributes tasks to the sensors which will bring the least energy to run them. EasiCAE has three salient features. Firstly, we define task-sensor correlation to indicate how many samplings of a sensor can be shared with the new task. Secondly, EasiCAE reduces energy consumption by assigning tasks to a sensor with higher task-sensor correlation. Finally, a light-weight merging algorithm is proposed to eliminate redundant samplings for the assigned sensors. Experimental results show that EasiCAE reduces energy consumption by 31% to 79% compared with existing methods, while introducing tolerable overheads. We also evaluate EasiCAE with various influencing parameters, showing that the performance of EasiCAE increases stably as the network scale and the number of concurrent applications increases. Hailong Shi, Dong Li 0008, Haiming Chen 0002, Jiefan Qiu |
ICPADS | 1 |
| 2014 | Stable grassmann manifold embedding via Gaussian random matricesabstractCompressive Sensing (CS) provides a new perspective for dimensionnality reduction without compromising performance. The theoretical foundation for most of existing studies of CS is a stable embedding (i.e., a distance-preserving property) of certain low-dimensional signal models such as sparse signals or signals in a union of linear subspaces. However, few existing literatures clearly discussed the embedding effect of points on the Grassmann manifold in under-sampled linear measurement systems. In this paper, we explore the stable embedding property of multi-dimensional signals based on Grassmann manifold, which is a topological space with each point being a linear subspace of ℝN(or ℂN), via the Gaussian random matrices. It should be noted that the stability mentioned here is about the volume-preserving instead of distance-preserving, because volume is the key characteristic for linear subspace spanned by multiple vectors. The theorem of the volume-preserving stable embedding property is proposed, and sketched proofs as well as discussions about our theorem is also given. Hailong Shi, Hao Zhang 0005, Gang Li 0008, Xiqin Wang |
ISIT | 1 |
| 2014 | SeaHttp: A Resource-Oriented Protocol to Extend REST Style for Web of Things
Chen-Da Hou, Dong Li 0008, Jiefan Qiu, Hailong Shi |
J. Comput. Sci. Technol. | 4 |
| 2014 | EasiSMP: A Resource-Oriented Programming Framework Supporting Runtime Propagation of RESTful Resources
Jiefan Qiu, Dong Li 0008, Hailong Shi, Chen-Da Hou |
J. Comput. Sci. Technol. | 3 |
| 2014 | A Task Execution Framework for Cloud-Assisted Sensor Networks
Hailong Shi, Dong Li 0008, Jiefan Qiu, Chen-Da Hou |
J. Comput. Sci. Technol. | 1 |
| 2014 | Fusing multi-cues description for partial-duplicate image retrieval
Chenggang Yan 0001, Liang Li 0003, Jian Yin 0003, Hailong Shi, Shuqiang Jiang, Qingming Huang |
J. Vis. Commun. Image Represent. | 5 |