Hai-Sheng Li 0002

dblp:145/2836-2 · also Haisheng Li 0002 · DBLP profile ↗
← Back
53ranked-venue papers
1as first author
36since 2021 · last 2027
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 26 · 1 first-author · 19 since 2021Artificial intelligence and machine learning · 21 · 14 since 2021Computer networks · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
YearPublicationVenuePosition
2027 A decision-oriented approach for employee turnover prediction with counterfactual explanations
Yi Chen 0007, Hai-Sheng Li 0002, Yu Dong 0001
Expert Syst. Appl.3
2026 Not All Distortions Are Created Equal: Distortion-Selective Domain Adaptation for Point Cloud Quality Assessment
abstract
Point cloud quality assessment (PCQA) has advanced significantly with synthetic datasets offering diverse distortion coverage for model training. However, when applied to new application scenarios, models often suffer from performance drops due to mismatched distortion characteristics between source and target domains. Most current methods use all available synthetic distortions, which may introduce irrelevant features and hinder generalization. To address this, we propose DST-PCQA, a distortion-selective training framework for PCQA. Unlike previous approaches that treat all distortions equally, DST-PCQA identifies and selects distortion types most relevant to a target domain by analyzing inter-domain distortion similarity. This selective strategy reduces negative transfer and enables efficient domain-specific training. To fully leverage the selected distortions for both classification and quality prediction, we adopt a dual-branch architecture that fuses 2D visual cues and 3D geometric structure via cross-modal attention. This design supports multi-level feature alignment across modalities and enables fine-grained distortion understanding. Extensive evaluations across three target domains have verified the effectiveness of DST-PCQA over full-set training baselines. Moreover, its distortion-selective strategy is orthogonal to existing model-based PCQA methods, enabling improved cross-domain performance and reduced training costs across a wide range of architectures.
Yangwei Li, Xin Shang, Hai-Sheng Li 0002
AAAI4
2026 HFR-MKGC: Hierarchical Fusion Reasoning with MLLMs for Multi-modal Knowledge Graph Completion
abstract
Multi-modal knowledge graph completion (MMKGC) aims to infer missing entities of triples by leveraging heterogeneous information in knowledge graph (KG). However, existing approaches often struggle with inconsistent modality alignment, limited reasoning depth, and insufficient negative sample quality. In this work, we propose HFR-MKGC, a novel framework that integrates hierarchical modal fusion and Multimodal Large Language Model (MLLM) reasoning for robust and expressive MMKGC. Specifically, we introduce a relation-guided hierarchical modal fusion module, which conducts fine-grained intra-visual fusion and relation-guided cross-modal integration to yield rich entity representations. HFR-MKGC employs a fine-tuned MLLM to perform instruction-based triple reasoning, producing candidate entities for completion. Then, it constructs hard negative samples through textual perturbation by MLLM and visual feature augmentation with rotation and noise. HFR-MKGC optimizes the model via adversarial training. Extensive experiments on three MMKGC benchmarks demonstrate that our method outperforms state-of-the-art methods, validating its effectiveness in MMKGC.
Junping Du 0001, Zhe Xue, Meiyu Liang, Guanhua Ye, Yingxia Shao, Hai-Sheng Li 0002
AAAI7
2026 3D Gaussian splatting for reconstruction: Methods, datasets, and comparisons
Jiulin Liang, Hai-Sheng Li 0002, Wenshan Shen, Qianwen Yao
Expert Syst. Appl.3
2026 Music-driven dance generation: A comprehensive review
Li Sun 0004, Hai-Sheng Li 0002
Neurocomputing2
2026 Deep Learning for 3-D Lane Detection in Autonomous Driving: A Survey
abstract
3D lane detection has become a critical component in the perception task of autonomous vehicles. Unlike 2D lane detection, which operates in the image plane, 3D lane detection estimates the spatial layout of lanes in real-world coordinates, enabling fine-grained localization, map construction, and planning. However, the task remains challenging due to depth ambiguity, sensor limitations, and diverse road conditions. Existing surveys mostly focus on 2D or organize 3D lane detection by sensor modality, lacking a systematic treatment of algorithmic designs. In this paper, we present a comprehensive survey of deep learning-based 3D lane detection methods. We introduce a dual-axis taxonomy that jointly considers modeling paradigms and representation spaces. Based on this framework, we categorize existing methods into four primary paradigms: geometry-based, end-to-end, query-based, and implicit field-based. We analyze how each paradigm interacts with spatial representations such as image, BEV, 3D, and topological spaces. For each category, we review representative frameworks, architectural principles, and performance trade-offs. We also provide an extensive summary of public datasets, evaluation metrics, and state-of-the-art results across multiple benchmarks. Finally, we identify current limitations and outline future research directions toward robust, scalable, and interpretable 3D lane detection.
Xiaoqiang Teng, Zuo Chen, Shunpeng Chen, Shibiao Xu, Zhihao Hao, Deke Guo, Hai-Sheng Li 0002
IEEE Internet Things J.8
2026 RGDF: a structure-aware lightweight fusion for efficient and accurate knowledge graph completion
Qiang Cai 0001, Yuxiao Wu, Yanzhao Ren, Hai-Sheng Li 0002
J. Intell. Inf. Syst.4
2026 DPR-Net: dual-branch probabilistic regression for no-reference point cloud quality assessment
Yangwei Li, Xin Shang, Hai-Sheng Li 0002
Multim. Syst.5
2025 Leveraging the Dual Capabilities of LLM: LLM-Enhanced Text Mapping Model for Personality Detection
abstract
Personality detection aims to deduce a user’s personality from their published posts. The goal of this task is to map posts to specific personality types. Existing methods encode post information to obtain user vectors, which are then mapped to personality labels. However, existing methods face two main issues: first, only using small models makes it hard to accurately extract semantic features from multiple long documents. Second, the relationship between user vectors and personality labels is not fully considered. To address the issue of poor user representation, we utilize the text embedding capabilities of LLM. To solve the problem of insufficient consideration of the relationship between user vectors and personality labels, we leverage the text generation capabilities of LLM. Therefore, we propose the LLM-Enhanced Text Mapping Model (ETM) for Personality Detection. The model applies LLM’s text embedding capability to enhance user vector representations. Additionally, it uses LLM’s text generation capability to create multi-perspective interpretations of the labels, which are then used within a contrastive learning framework to strengthen the mapping of these vectors to personality labels. Experimental results show that our model achieves state-of-the-art performance on benchmark datasets.
Weihong Bi, Feifei Kou, Lei Shi 0030, Yawen Li 0001, Hai-Sheng Li 0002, Jinpeng Chen 0001, Mingying Xu
AAAI5
2025 DCHM: Dynamic Collaboration of Heterogeneous Models Through Isomerism Learning in a Blockchain-Powered Federated Learning Framework
abstract
Solutions to time-varying problems are crucial for research areas such as predicting changes in human body shape over time. While recurrent neural networks have made significant advancements in this field, their reliance on centralized processing has led to challenges such as model silos and data isolation. In response, distributed AI systems like federated learning have emerged to facilitate dynamic collaboration among models; however, they still depend on central coordinators, which pose risks to system security and efficiency. Moreover, traditional federated learning primarily supports homogeneous models and lacks effective strategies for the interaction of heterogeneous models. To address these limitations, we propose a novel method called Dynamic Collaboration of Heterogeneous Models (DCHM), based on Isomerism Learning, which leverages a consortium blockchain network to enhance model credibility and facilitate coordination among heterogeneous models. Additionally, we introduce a Distributed Hierarchical Aggregation (DHA) algorithm that enables permissioned nodes within each group to aggregate local model results and share them for standardized processing. After several iterative cycles, these nodes perform secondary integration of local results to produce global outcomes. Experimental results demonstrate that DCHM effectively analyzes the temporal variability of body shape changes with high efficiency.
Zhihao Hao, Bob Zhang 0001, Hai-Sheng Li 0002
AAAI3
2025 Transforming Classification with Federated Learning on Blockchain: A Unique Model Integration Approach
abstract
The need for robust machine learning models is particularly evident in the realm of biological pattern recognition. Traditional centralized methods often struggle, as they frequently depend on large datasets that are challenging to gather due to stringent data privacy regulations. To address these limitations while maintaining classification accuracy, we propose an innovative approach that unifies models with diverse architectures within a federated learning framework built on a blockchain network. This decentralized and trustworthy system fosters effective collaboration among various models. Furthermore, we have implemented a weight distribution mechanism designed to maximize the individual strengths of each model. By leveraging blockchains inherent transparency and auditability, this approach also ensures secure and traceable data exchanges among participants. Additionally, the adaptability of the framework allows it to be extended to other domains where privacy-preserving data sharing is critical. Our experimental results showcase that the proposed methodology significantly enhances performance in classification tasks compared to existing alternatives.
Zhihao Hao, Bob Zhang 0001, Hai-Sheng Li 0002
ICASSP3
2025 Star Operation in Self-Attention for 3D Human Pose Estimation
abstract
Recent transformer-based methods have achieved notable success in 3D human pose estimation. However, the most utilized self-attention mechanisms compute the attention matrix by performing a dot product on inter-vector features, which may overlook finer element-wise interactions. In this paper, we introduce star-attention, which integrates the star operation into the self-attention module. This approach retains the non-linearity and high dimensionality characteristics of matrix multiplication in self-attention while enhancing the feature representation by capturing interactions at a finer granularity. Specifically, we validated the promotion of the proposed star-attention module on MixSTE as an example. Experimental results demonstrate that our approach achieves competitive performance on benchmarks compared to state-of-the-art methods, producing smoother 3D poses across successive frames.
Xiaoming Chen 0006, Hai-Sheng Li 0002
ICASSP5
2025 DiffusionIMU: Diffusion-Based Inertial Navigation with Iterative Motion Refinement
abstract
Inertial navigation enables self-contained localization using only Inertial Measurement Units (IMUs), making it widely applicable in various domains such as navigation, augmented reality, and robotics. However, existing methods suffer from drift accumulation due to the sensor noise and difficulty capturing long-range temporal dependencies, limiting their robustness and accuracy. To address these challenges, we propose DiffusionIMU, a novel diffusion-based framework for inertial navigation. DiffusionIMU enhances direct velocity regression from IMU data through an iterative generative denoising process, progressively refining motion state estimation. It integrates the noise-adaptive feature modulation for sensor variability handling, the feature alignment mechanism for representation consistency, and the diffusion-based temporal modeling to decrease accumulated drift. Experiments show that DiffusionIMU consistently outperforms existing methods, demonstrating superior generalization to unseen users while alleviating the impact of the sensor noise.
Xiaoqiang Teng, Shibiao Xu, Zhihao Hao, Deke Guo, Hai-Sheng Li 0002, Weiliang Meng, Xiaopeng Zhang 0001
IJCAI7
2025 Towards a Global Spatial-Temporal Food Memory: A Vision for Privacy-Preserving Collaborative Multimedia Analysis
abstract
The dynamic variations of food quality across spatial and temporal scales pose significant challenges for global food safety and nutrition research, requiring comprehensive analysis of diverse, multi-modal, and distributed data while preserving privacy. Existing centralized approaches suffer from data silos and limited collaboration, and although federated learning and blockchain technologies have shown promise independently, their combined potential for incentivized, privacy-preserving, and heterogeneous model collaboration remains underexplored. In this paper, we propose the concept of a Global Spatial-Temporal Food Memory-a novel research paradigm that envisions secure, decentralized, and incentivized collaboration among multiple stakeholders worldwide, leveraging blockchain-enabled token-based rewards integrated with federated learning of heterogeneous models. We discuss the scientific challenges and opportunities inherent in this vision, including multi-modal data fusion, trustworthy incentive mechanisms, and scalable long-term temporal analysis. This work aims to open new avenues in multimedia research by bridging decentralized AI, blockchain, and spatiotemporal food quality monitoring, providing a foundation for future explorations in privacy-preserving, collaborative, and large-scale multimedia data analysis.
Zhihao Hao, Bob Zhang 0001, Hai-Sheng Li 0002
ACM Multimedia3
2025 OSTAR: Optimized Statistical Text-classifier with Adversarial Resistance
abstract
The advancements in generative models and the real-world attack of machine-generated text(MGT) create a demand for more robust detection methods. The existing MGT detection methods for adversarial environments primarily consist of manually designed statistical-based methods and fine-tuned classifier-based approaches. Statistical-based methods extract intrinsic features but suffer from rigid decision boundaries vulnerable to adaptive attacks, while fine-tuned classifiers achieve outstanding performance at the cost of overfitting to superficial textual feature. We argue that the key to detection in current adversarial environments lies in how to extract intrinsic invariant features and ensure that the classifier possesses dynamic adaptability. In that case, we propose OSTAR, a novel MGT detection framework designed for adversarial environments which composed of a statistical enhanced classifier and a Multi-Faceted Contrastive Learning(MFCL). In the classifier aspect, our Multi-Dimensional Statistical Profiling (MDSP) module extracts intrinsic difference between human and machine texts, complementing classifiers with useful stable features. In the model optimization aspect, the MFCL strategy enhances robustness by contrasting feature variations before and after text attacks, jointly optimizing statistical feature mapping and baseline pre-trained models. Experimental results on three public datasets under various adversarial scenarios demonstrate that our framework outperforms existing MGT detection methods, achieving state-of-the-art performance and robust against attacks.The code is available at https://github.com/BUPT-SN/OSTAR.
Yuhan Yao 0001, Feifei Kou, Lei Shi 0030, Zhongbao Zhang, Suguo Zhu, Jiwei Zhang 0007, Lirong Qiu, Hai-Sheng Li 0002
NeurIPS9
2025 MeshKINN: A self-supervised mesh generation model based on Kolmogorov-Arnold-Informed neural network
Haoxuan Zhang, Hai-Sheng Li 0002, Nan Li 0031
Expert Syst. Appl.3
2025 VANE-IN: Velocity Auto-Encoder for Inertial Navigation
abstract
Data-driven inertial navigation is crucial for mobile computing applications, such as navigation, augmented reality, and robotics. It typically depends on a trained velocity regression network (VRN) to estimate velocities from inertial measurement unit (IMU) data, enabling position determination through integration. However, using prior velocity information for feature representation in inertial navigation remains underexplored. This work introduces a framework called velocity auto-encoder for inertial navigation (VANE-IN), which employs a Teacher-Student scheme to enhance VRN performance by encoding the velocity. Specifically, a velocity auto-encoder (VANE) is proposed as a student model to distill prior velocity insights from the training dataset, which is then guided by the VRN acting as the teacher model. Additionally, an attention mechanism is introduced to fuse these insights into the features of the VRN. To this end, the VANE-IN achieves a state-of-the-art position accuracy on the RoNIN benchmarks. Our experimental results demonstrate that the VANE-IN achieves approximately 5% performance improvements over existing methods regarding position accuracy.
Xiaoqiang Teng, Shibiao Xu, Deke Guo, Hai-Sheng Li 0002
IEEE Internet Things J.5
2025 AIGR-LSTM: Land Surface Temperature Interpolation Using Adaptive Inductive Graph Representation Learning
abstract
Land surface temperature (LST) plays an important role in many fields including urban heat island analysis, climate change studies, and agricultural disaster monitoring. However, achieving high-resolution and continuous LST data remains challenging due to the limitations of remote sensing observations and the sparse distribution of ground sensors. Recently, spatio-temporal kriging has been widely used to address such challenges, but often relies on fixed neighborhood locations and typically ignores temporal dependencies in LST data. To address these challenges, we propose AIGR-LSTM, an Adaptive Inductive Graph Representation Learning model based on long short-term memory (LSTM) for LST interpolation. To better determine the spatial proximity relationship between nodes, we propose a threshold-based adaptive neighbor selection module, which effectively captures complex spatial dependencies by adaptively selecting relevant neighbors. Besides, we design a geography-based spatio-temporal modeling module that incorporates geographic spatial relationships, combines temporal encoding information with spatial features, and leverages LSTM to achieve accurate spatio-temporal interpolation. The experiment was conducted on LST and three other datasets, and the results showed that AIGR-LSTM outperformed the six state-of-the-art methods in all evaluation metrics. On the LST dataset, the RMSE was 0.185°C and the MAE was 0.133°C, which improved by 6.09% and 8.28% respectively compared to the best baseline model.
Huanpu Yin, Di Zhang 0024, Hai-Sheng Li 0002
IEEE Geosci. Remote. Sens. Lett.4
2025 Unsupervised arbitrary-scale point cloud upsampling by learning neural gradient function
Jiangshan Feng, Xiaoqun Wu, Hai-Sheng Li 0002
Multim. Syst.4
2025 DualFocus GAN for Robust Watermarking in Transportation Cyber-Physical Systems
abstract
With the advancement of Transportation Cyber-Physical Systems (TCPS), information security has become increasingly critical. Invisible watermarking, which ensures reliable information traceability without compromising carrier quality, holds significant potential for TCPS. However, achieving high robustness in the real world while maintaining imperceptibility is a challenge. To address this, we propose DGWW (Dual-discriminator GAN-based WaveFusion Watermarking), a novel invisible watermarking method that balances robustness and imperceptibility. The GAN-based approach is well suited for TCPS, as it enables adaptive watermark embedding aligned with the dynamic and heterogeneous nature of transportation data, effectively handling diverse noise conditions and data types. DGWW integrates a WaveFusion Encoding Module, a Dual-Focus Discriminator, and a contrastive learning-based optimization strategy to enhance watermark embedding without degrading robustness. These components leverage multi-frequency information, assess local and global impacts on image quality, and guide model optimization. Experimental results show that DGWW outperforms state-of-the-art methods in visual quality and robustness under various noise conditions, offering a robust and scalable solution for image watermarking in TCPS environments. By maintaining data usability and strong resistance to noise attacks, DGWW advances digital watermarking in intelligent transportation systems.
Feifei Kou, Yuhan Yao 0001, Jideng Han, Hai-Sheng Li 0002, Jiwei Zhang 0007
IEEE Trans. Intell. Transp. Syst.5
2025 A Novel Public Sentiment Analysis Method Based on an Isomerism Learning Model via Multiphase Processing
abstract
The dissemination of public opinion in the social media network is driven by public sentiment, which can be used to promote the effective resolution of social incidents. However, public sentiments for incidents are often affected by environmental factors such as geography, politics, and ideology, which increases the complexity of the sentiment acquisition task. Therefore, a hierarchical mechanism is designed to reduce complexity and utilize processing at multiple phases to improve practicality. Through serial processing between different phases, the task of public sentiment acquisition can be decomposed into two subtasks, which are the classification of report text to locate incidents and sentiment analysis of individuals' reviews. Performance has been improved through improvements to the model structure, such as embedding tables and gating mechanisms. That being said, the traditional centralized structure model is not only easy to form model silos in the process of performing tasks but also faces security risks. In this article, a novel distributed deep learning model called isomerism learning based on blockchain is proposed to address these challenges, the trusted collaboration between models can be realized through parallel training. In addition, for the problem of text heterogeneity, we also designed a method to measure the objectivity of events to dynamically assign the weights of models to improve aggregation efficiency. Extensive experiments demonstrate that the proposed method can effectively improve performance and outperform the state-of-the-art methods significantly.
Zhihao Hao, Guan-Cheng Wang 0002, Bob Zhang 0001, Zhuowen Feng, Hai-Sheng Li 0002, Fahui Chong, Wei Li 0016
IEEE Trans. Neural Networks Learn. Syst.5
2025 Potential Features Fusion Network for Multimodal Fake News Detection
abstract
With the popularization of social networks, fake news is also widely and rapidly spreading, which poses a great threat to the Internet. Therefore, how to detect fake news automatically and efficiently has become an urgent problem to be solved. However, the existing approaches mostly focus on the explicit features (images and text) and deep fusions, without considering potential features such as text emotion and image category. To find a solution to this issue, we propose a Potential Features Fusion Network (PFFN), which models the explicit and potential features at the same time. To exploit the potential image features, we introduce a mixture of experts structure to process the news image separately, which can best use the relationships between the news image category and fake news detection. Besides, we also extract emotion features as potential text features and fuse them with explicit text features. Finally, we establish an attention-based feature fusion network to fuse the potential features with the explicit features, which can obtain a multimodal fusion feature of a piece of news and thus further improve the performance. We make experiments on four public datasets (Weibo16, Weibo19, Twitter, and PolitiFact); the results compared with the baseline approaches demonstrate that our PFFN has a better performance. Our code is available at https://github.com/Wang-bupt/PFFN
Feifei Kou, Bingwei Wang, Hai-Sheng Li 0002, Chuangying Zhu, Lei Shi 0030, Jiwei Zhang 0007, Limei Qi
ACM Trans. Multim. Comput. Commun. Appl.3
2024 No-Reference Point Cloud Quality Assessment with Adaptive Keyframe Selection
abstract
In recent years, projection-based no-reference point cloud quality metrics have shown remarkable performance on public datasets. These metrics involve projecting the point cloud onto a 2D space from multiple viewpoints and applying traditional image quality metrics to the resulting projected images. However, these existing methods suffer from a measurable projection strategy, which would induce information loss or redundancy leading to training bias and reduced generalizability. To overcome these limitations, we propose AKSNet, a novel no-reference point cloud quality metric that utilizes an adaptive keyframe selection strategy. In particular, we treat the process of point cloud projection as video sequences captured from successive viewpoints and design a keyframe selection module to pick the most representative frames. Based on the selected frames, we can subsequently predict the quality score of the point cloud. Experimental results demonstrate that our method achieves competitive performance in comparison with state-of-the-art methods. While the proposed adaptive keyframe selecting strategy can effectively promote the generalization performance.
Xianpeng Yuan, Xianming Chen, Hai-Sheng Li 0002
VCIP5
2024 Multivariable High-Dimension Time-Series Prediction in SIoT via Adaptive Dual-Graph-Attention Encoder-Decoder With Global Bayesian Optimization
abstract
In the current intelligent era, high-dimensional multivariate time-series (HMTSs) data are continuously monitored by heterogeneous devices from multiple observers in the Social Internet of Things (SIoT), which forms complicated time-series forecasting issues that must be addressed. It is an emerging paradigm that emphasizes the importance of time-series prediction methods, which play a crucial role in accurately predicting future changes, further facilitating intelligent decision-making. However, it is still challenging to design accurate time-series predictors for HMTS data to handle high-dimensional variable interactions and potential spatio-temporal correlations recorded by different observers embedded in the complex data. To solve the above dilemma, we designed the Dual-graphic Representation Mechanism (DgRM) based on the observatory features and observed variables to simultaneously mine their correlations hidden in abundant HMTS data. Subsequently, a dual-attention mechanism is introduced into the adaptive encoder-decoder module (AEdM) to construct a novel time-series predictor named after DAG-Net. With the assistance of the global Bayesian optimization (GBO) strategy, DAG-Net obtains a preferable balance between performance and robustness. Extensive experiments on three SIoT data sets demonstrated the outstanding performance of DAG-Net, surpassing contrastive prediction methods. DAG-Net achieved preferable results in air quality, transportation, and intelligent agriculture prediction tasks, with a 10.67%, 59.67%, and 19.29% performance improvement in terms of the root mean squared error index. More performance analysis further verified the application prospects of the proposed predictor in practical SIoT and intelligent systems.
Zimeng Dong, Jian-Lei Kong, Xiaoyi Wang 0001, Hai-Sheng Li 0002
IEEE Internet Things J.5
2024 CANet: cross attention network for food image segmentation
Xiaoxiao Dong, Hai-Sheng Li 0002, Junping Du 0001
Multim. Tools Appl.2
2024 DEGAN: Detail-Enhanced Generative Adversarial Network for Monocular Depth-Based 3D Reconstruction
abstract
Although deep networks-based 3D reconstruction methods can recover the 3D geometry given few inputs, they may produce unfaithful reconstruction when predicting occluded parts of 3D objects. To address the issue, we propose Detail-Enhanced Generative Adversarial Network (DEGAN) which consists of Encoder–Decoder-Based Generator (EDGen) and Voxel-Point Embedding Network-Based Discriminator (VPDis) for 3D reconstruction from a monocular depth image of an object. Firstly, EDGen decodes the features from the 2.5D voxel grid representation of an input depth image and generates the 3D occupancy grid under GAN losses and a sampling point loss. The sampling loss can improve the accuracy of predicted points with high uncertainty. VPDis helps reconstruct the details under voxel and point adversarial losses, respectively. Experimental results show that DEGAN not only outperforms several state-of-the-art methods on both public ModelNet and ShapeNet datasets but also predicts more reliable occluded/missing parts of 3D objects.
Yali Chen 0005, Minhong Zhu, Chenhui Hao, Hai-Sheng Li 0002
ACM Trans. Multim. Comput. Commun. Appl.5
2024 Video2Haptics: Converting Video Motion to Dynamic Haptic Feedback with Bio-Inspired Event Processing
abstract
In cinematic VR applications, haptic feedback can significantly enhance the sense of reality and immersion for users. The increasing availability of emerging haptic devices opens up possibilities for future cinematic VR applications that allow users to receive haptic feedback while they are watching videos. However, automatically rendering haptic cues from real-time video content, particularly from video motion, is a technically challenging task. In this article, we propose a novel framework called "Video2Haptics" that leverages the emerging bio-inspired event camera to capture event signals as a lightweight representation of video motion. We then propose efficient event-based visual processing methods to estimate force or intensity from video motion in the event domain, rather than the pixel domain. To demonstrate the application of Video2Haptics, we convert the estimated force or intensity to dynamic vibrotactile feedback on emerging haptic gloves, synchronized with the corresponding video motion. As a result, Video2Haptics allows users not only to view the video but also to perceive the video motion concurrently. Our experimental results show that the proposed event-based processing methods for force and intensity estimation are one to two orders of magnitude faster than conventional methods. Our user study results confirm that the proposed Video2Haptics framework can considerably enhance the users' video experience.
Xiaoming Chen 0006, Zexi Hu, Guangxin Zhao, Hai-Sheng Li 0002, Vera Chung, Aaron J. Quigley
IEEE Trans. Vis. Comput. Graph.4
2023 A novel method using LSTM-RNN to generate smart contracts code templates for improved usability
Zhihao Hao, Bob Zhang 0001, Dianhui Mao, Jerome Yen, Zhihua Zhao 0003, Hai-Sheng Li 0002, Cheng-Zhong Xu 0001
Multim. Tools Appl.7
2023 A multimodal dialogue system for improving user satisfaction via knowledge-enriched response and image recommendation
Jiangnan Wang, Hai-Sheng Li 0002, Leiquan Wang, Chunlei Wu
Neural Comput. Appl.2
2022 Attentive frequency learning network for super-resolution
Fenghai Li, Qiang Cai 0001, Hai-Sheng Li 0002, Jian Cao 0003, Shanshan Li 0005
Appl. Intell.3
2022 Patch-based mesh inpainting via low rank recovery
Xiaoqun Wu, Xiaoyun Lin, Nan Li 0031, Hai-Sheng Li 0002
Graph. Model.4
2022 Generating diverse chinese poetry from images via unsupervised method
Jiangnan Wang, Hai-Sheng Li 0002, Chunlei Wu, Faming Gong, Leiquan Wang
Neurocomputing2
2022 Colorful 3D reconstruction at high resolution using multi-view representation
Yanping Zheng, Hai-Sheng Li 0002, Qiang Cai 0001, Junping Du 0001
J. Vis. Commun. Image Represent.3
2022 SATMAC: Self-Adaptive TDMA-Based MAC Protocol for VANETs
abstract
Rapid development and deployment of vehicular ad-hoc networks (VANETs) require an efficient and scalable media access control (MAC) protocol to support high-priority safety applications and infotainment requirements. This paper proposes SATMAC, a self-adaptive time division multiple access (TDMA)-based MAC protocol for VANETs. In order to improve the stability of the time slot scheduling in VANETs, a slot status updating strategy is carefully designed, which utilizes accurate information of the two-hop neighbors and the rough information of the three-hop neighbors to detect the potential packet collisions and avoids potential collisions by adjusting the occupied time slot. Besides, an adaptive frame length (the number of time slots contained in a frame) approach is proposed on the basis of the slot adjustment to support various densities of vehicles, where the frame length between neighbors can be inconsistent. We conduct theoretical analysis and extensive simulations in a realistic VANET environment to evaluate SATMAC. Simulation results show that compared with IEEE 802.11p and LTE-V2X PC5 Mode 4, SATMAC significantly improves PDR of beacons over 60%. Moreover, our SATMAC design is further implemented and validated on our FPGA-based testbed.
Jingbang Wu, Huimei Lu, Yong Xiang 0002, Feng Wang 0001, Hai-Sheng Li 0002
IEEE Trans. Intell. Transp. Syst.5
2021 Discriminative semantic region selection for fine-grained recognition
Chunjie Zhang 0001, Dahan Wang, Hai-Sheng Li 0002
J. Vis. Commun. Image Represent.3
2021 Label Rectification Learning through Kernel Extreme Learning Machine
abstract
Along with the strong representation of the convolutional neural network (CNN), image classification tasks have achieved considerable progress. However, majority of works focus on designing complicated and redundant architectures for extracting informative features to improve classification performance. In this study, we concentrate on rectifying the incomplete outputs of CNN. To be concrete, we propose an innovative image classification method based on Label Rectification Learning (LRL) through kernel extreme learning machine (KELM). It mainly consists of two steps: (1) preclassification, extracting incomplete labels through a pretrained CNN, and (2) label rectification, rectifying the generated incomplete labels by the KELM to obtain the rectified labels. Experiments conducted on publicly available datasets demonstrate the effectiveness of our method. Notably, our method is extensible which can be easily integrated with off‐the‐shelf networks for improving performance.
Qiang Cai 0001, Fenghai Li, Hai-Sheng Li 0002, Jian Cao 0003, Shanshan Li 0005
Wirel. Commun. Mob. Comput.4
2020 Attention-aware invertible hashing network with skip connections
Shanshan Li 0005, Qiang Cai 0001, Zhuangzi Li, Hai-Sheng Li 0002, Naiguang Zhang, Xiaoyu Zhang 0002
Pattern Recognit. Lett.4
2020 Cross-Media Semantic Correlation Learning Based on Deep Hash Network and Semantic Expansion for Social Network Cross-Media Search
abstract
Cross-media search from large-scale social network big data has become increasingly valuable in our daily life because it can support querying different data modalities. Deep hash networks have shown high potential in achieving efficient and effective cross-media search performance. However, due to the fact that social network data often exhibit text sparsity, diversity, and noise characteristics, the search performance of existing methods often degrades when dealing with this data. In order to address this problem, this article proposes a novel end-to-end cross-media semantic correlation learning model based on a deep hash network and semantic expansion for social network cross-media search (DHNS). The approach combines deep network feature learning and hash-code quantization learning for multimodal data into a unified optimization architecture, which successfully preserves both intramedia similarity and intermedia correlation, by minimizing both cross-media correlation loss and binary hash quantization loss. In addition, our approach realizes semantic relationship expansion by constructing the image-word relation graph and mining the potential semantic relationship between images and words, and obtaining the semantic embedding based on both internal graph deep walk and an external knowledge base. Experimental results demonstrate that DHNS yields better cross-media search performance on standard benchmarks.
Meiyu Liang, Junping Du 0001, Cong-Xian Yang, Zhe Xue, Hai-Sheng Li 0002, Feifei Kou, Yue Geng
IEEE Trans. Neural Networks Learn. Syst.5
2020 A relationship extraction method for domain knowledge graph construction
Haoze Yu, Hai-Sheng Li 0002, Dianhui Mao, Qiang Cai 0001
World Wide Web2
2019 Video-level Multi-model Fusion for Action Recognition
abstract
The approaches based on spatio-temporal features for video action recognition have emerged such as two-stream based methods and 3D convolution based methods. However, current methods suffer from the problems caused by partial observation, or restricted to single information modeling, and so on. Segment-level recognition results obtained from dense sampling can not represent the entire video and, therefore lead to partial observation. And a single model is hard to capture the complementary information on spacial, temporal and spatio-temporal information from video at the same time. Therefore, the challenge is to build the video-level representation and capture multiple information. In this paper, a video-level multi-model fusion action recognition method is proposed to solve these problems. Firstly, an efficient video-level 3D convolution model is proposed to get the global information in the video which assembling segment-level 3D convolution models. Secondly, a multi-model fusion architecture is proposed for video action recognition to capture multiple information. The spatial, temporal and spatio-temporal information are aggregate with SVM classifier. Experimental results show that this method achieves the state-of-the-art performance on the datasets of UCF-101(97.6%) without pre-training on Kinetics.
Junsan Zhang, Leiquan Wang, Philip S. Yu, Hai-Sheng Li 0002
CIKM6
2019 Attention-Aware Invertible Hashing Network
Shanshan Li 0005, Qiang Cai 0001, Zhuangzi Li, Hai-Sheng Li 0002, Naiguang Zhang, Jian Cao 0003
ICIG (3)4
2019 A multi-feature probabilistic graphical model for social network semantic search
Feifei Kou, Junping Du 0001, Cong-Xian Yang, Yan-Song Shi, Meiyu Liang, Zhe Xue, Hai-Sheng Li 0002
Neurocomputing7
2019 Visual analytics of cellular signaling data
Hai-Sheng Li 0002, Yuanjie Huang, Qiang Cai 0001, Junping Du 0001
Multim. Tools Appl.1
2018 Retrieving indoor objects: 2D-3D alignment using single image and interactive ROI-based refinement
Fuchang Liu, Shuangjian Wang, Dandan Ding, Qingshu Yuan, Zhengwei Yao, Hai-Sheng Li 0002
Comput. Graph.7
2018 Generative Adversarial Image Super-Resolution Through Deep Dense Skip Connections
abstract
Abstract Recently, image super‐resolution works based on Convolutional Neural Networks (CNNs) and Generative Adversarial Nets (GANs) have shown promising performance. However, these methods tend to generate blurry and over‐smoothed super‐resolved (SR) images, due to the incomplete loss function and powerless architectures of networks. In this paper, a novel generative adversarial image super‐resolution through deep dense skip connections (GSR‐DDNet), is proposed to solve the above‐mentioned problems. It aims to take advantage of GAN's ability of modeling data distributions, so that GSR‐DDNet can select informative feature representation and model the mapping across the low‐quality and high‐quality images in an adversarial way. The pipeline of the proposed method consists of three main components: 1) The generator of a novel dense skip connection network with the deep structure for learning robust mapping function is proposed to generate SR images from low‐resolution images; 2) The feature extraction network based on VGG‐19 is adopted to capture high frequency feature maps for content loss; and 3) The discriminator with Wasserstein distance is adopted to identify the overall style of SR and ground‐truth images. Experiments conducted on four publicly available datasets demonstrate the superiority against the state‐of‐the‐art methods.
Xiaobin Zhu 0001, Zhuangzi Li, Xiaoyu Zhang 0002, Hai-Sheng Li 0002, Lei Wang 0101
Comput. Graph. Forum4
2018 Common Green Plants Recognition Based on Wavelet Transformation and Varied Local Edge Patterns
abstract
Green plant species identification plays an important role in so many aspects, such as ecological environment protection, Chinese medicine preparation, agricultural and horticultural application, etc. A method on green plants recognition based on wavelet transform and variable local edge patterns (VLEP) is proposed in this paper. Firstly, the original image is decomposed by wavelet transformation. Then texture features are extracted using VLEPs. At the same time, block-based and multi-resolution ideas are considered together to extract features after images are transformed by wavelet. Finally, the fused texture features are classified by the nearest neighbor method. The experimental results show that the proposed method is a promising method for recognizing the common green plants with the natural and complex background compared with the other state-of-the-art methods, and combination of block-based and multi-resolution ideas can further improve the accuracy rate effectively.
Yu Wang 0087, Qiang Cai 0001, Hai-Sheng Li 0002
Int. J. Pattern Recognit. Artif. Intell.4
2018 Geometry of Motion for Video Shakiness Detection
Xiaoqun Wu, Hai-Sheng Li 0002, Jian Cao 0003, Qiang Cai 0001
J. Comput. Sci. Technol.2
2017 Variational reconstruction using subdivision surfaces with continuous sharpness control
abstract
We present a variational method for subdivision surface reconstruction from a noisy dense mesh. A new set of subdivision rules with continuous sharpness control is introduced into Loop subdivision for better modeling subdivision surface features such as semi-sharp creases, creases, and corners. The key idea is to assign a sharpness value to each edge of the control mesh to continuously control the surface features. Based on the new subdivision rules, a variational model with L1 norm is formulated to find the control mesh and the corresponding sharpness values of the subdivision surface that best fits the input mesh. An iterative solver based on the augmented Lagrangian method and particle swarm optimization is used to solve the resulting non-linear, non-differentiable optimization problem. Our experimental results show that our method can handle meshes well with sharp/semi-sharp features and noise.
Xiaoqun Wu, Jianmin Zheng, Yiyu Cai, Hai-Sheng Li 0002
Comput. Vis. Media4
2016 Shape Retrieval of Non-rigid 3D Human Models
abstract
3D models of humans are commonly used within computer graphics and vision, and so the ability to distinguish between body shapes is an important shape retrieval problem. We extend our recent paper which provided a benchmark for testing non-rigid 3D shape retrieval algorithms on 3D human models. This benchmark provided a far stricter challenge than previous shape benchmarks. We have added 145 new models for use as a separate training set, in order to standardise the training data used and provide a fairer comparison. We have also included experiments with the FAUST dataset of human scans. All participants of the previous benchmark study have taken part in the new tests reported here, many providing updated results using the new data. In addition, further participants have also taken part, and we provide extra analysis of the retrieval results. A total of 25 different shape retrieval methods are compared.
David Pickup, Xianfang Sun, Paul L. Rosin, Ralph R. Martin, Zhouhui Lian, Masaki Aono, A. Ben Hamza, Alexander M. Bronstein, Michael M. Bronstein, S. Bu, Umberto Castellani, S. Cheng, Valeria Garro, Andrea Giachetti 0001, Afzal Godil, Luca Isaia, Henry Johan, Long Lai, Bo Li 0013, Chenfeng Li, Hai-Sheng Li 0002, Roee Litman, Yijuan Lu, Li Sun 0004, Gary K. L. Tam, Atsushi Tatsuma, Jianbo Ye
Int. J. Comput. Vis.23
2016 A varied local edge pattern descriptor and its application to texture classification
Yu Wang 0087, Qiang Cai 0001, Hai-Sheng Li 0002, Huaixin Yan
J. Vis. Commun. Image Represent.4
2015 Saliency detection using two-stage scoring
abstract
In this paper, we propose a novel saliency detection approach, which is robust to images with complex background. In our algorithm, an intuitive and straightforward pre-treatment method is formulated for conducting over-segmentation adaptively. To detect saliency effectively, a two-stage scoring method is adopted, in which both background prior and foreground cues are considered. In the first stage, we conduct random walk on absorbing Markov chain with background prior. And in the second stage, we use the saliency scores computed by the first stage scoring as foreground cues for manifold ranking. Experimental results on publicly available datasets demonstrate that our method outperforms the state-of-the-art methods in detecting salient objects.
Qiang Cai 0001, Xiaobin Zhu 0001, Jian Cao 0003, Hai-Sheng Li 0002
ICIP5
2015 A comparison of 3D shape retrieval methods based on a large-scale benchmark supporting multimodal queries
Bo Li 0013, Yijuan Lu, Chunyuan Li, Afzal Godil, Tobias Schreck, Masaki Aono, Martin Burtscher, Nihad Karim Chowdhury, Hongbo Fu 0001, Takahiko Furuya, Hai-Sheng Li 0002, Jianzhuang Liu, Henry Johan, Ryuichi Kosaka, Hitoshi Koyanagi, Ryutarou Ohbuchi, Atsushi Tatsuma, Yajuan Wan, Changqing Zou
Comput. Vis. Image Underst.13
2013 A review of object representation based on local features
abstract
Object representation based on local features is a topical subject in the domain of image understanding and computer vision. We discuss the defects of global features in present methods and the advantages of local features in object recognition, and briefly explore state-of-the-art recognition methods using local features, especially the main approaches of local feature extraction and object representation. To clearly explain these methods, the problem of local feature extraction is divided into feature region detection, feature region description, and feature space optimization. The main components and merits of these steps are presented. Technologies for object presentation are classified into three types: vector space, sliding window, and structure relationship models. Future development trends are discussed briefly.
Jian Cao 0003, Dianhui Mao, Qiang Cai 0001, Hai-Sheng Li 0002, Junping Du 0001
J. Zhejiang Univ. Sci. C4