Shuai Liu 0002

dblp:76/5789-2 · DBLP profile ↗
← Back
73ranked-venue papers
23as first author
45since 2021 · last 2026
0000-0001-9909-0664ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 7 first-author · 20 since 2021Computer networks · 24 · 7 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 3 since 2021Systems, architecture and hardware · 5 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Semantic-Aware Multimodal Collaborative Learning for Unsupervised Visible-Infrared Person Re-Identification
abstract
Unsupervised visible-infrared person re-identification (VI-ReID) is challenging due to the significant modality gap between visible and infrared images. Most existing methods rely on one-hot clustering pseudo-labels as supervision signals, which often fail to capture the full semantic relationships among samples and are highly susceptible to noise. To address these limitations, we propose a Semantic-aware Multimodal Collaborative Learning (SAMCL) framework for unsupervised VI-ReID. Specifically, a Modality-aware Semantic Fusion (MSF) module is designed to bridge the inter-modality gap by integrating complementary semantic details from both visible and infrared modalities, generating enriched cross-modal supervision signals, for cross-modal collaborative learning. Meanwhile, we present a Dynamic Contrastive Learning (DCL) module to refine intra-modality feature learning by dynamically aligning samples with their neighboring centroids in the feature space, improving clustering reliability and intra-modality feature discrimination. By combining the two modules, SAMCL harnesses multimodal collaboration, minimizes dependence on noisy pseudo-labels, and provides a robust approach to unsupervised VI-ReID. Extensive experiments demonstrate the superiority of our proposed method. For instance, on the SYSU-MM01 dataset, our model achieves a Rank-1 accuracy of 68.68% in the All Search setting, surpassing the state-of-the-art (SOTA) by 3.48%. On the RegDB dataset, it achieves a Rank-1 accuracy of 94.47% in the Visible-to-Infrared setting, outperforming the SOTA by 3.57%. On the LLCM dataset, it achieves a Rank-1 accuracy of 50.6% in the Visible-to-Infrared setting, outperforming the SOTA by 3.7%. The code is available at https://github.com/luoshixi123/SAMCL.
Shixi Luo, Min Liu 0008, Gautam Srivastava 0001, Shuai Liu 0002, Yaonan Wang 0001
IEEE Trans. Image Process.5
2025 A Comparative Analysis of AI-Enabled Science Education Research in China and Abroad
Gonghao Sun, Shuai Liu 0002, Gautam Srivastava 0001
IEEE Big Data2
2025 Disco4D: Disentangled 4D Human Generation and Animation from a Single Image
abstract
We present Disco4D, a novel Gaussian Splatting framework for 4D human generation and animation from a single image. Different from existing methods, Disco4D distinctively disentangles clothings (with Gaussian models) from the human body (with SMPL-X model), significantly enhancing the generation details and flexibility. Specifically, 1) Disco4D learns to efficiently fit the clothing Gaussians over the SMPL-X Gaussians. 2) Next, Disco4D adopts diffusion models to enhance the 3D generation process, e.g., modeling occluded parts not visible in the input image. 3) Finally, Disco4D learns an identity encoding for each clothing Gaussian to facilitate the separation and extraction of clothing assets. Furthermore, Disco4D naturally supports 4D human animation with vivid dynamics. Extensive experiments demonstrate the superiority of Disco4D on 4D human generation and animation tasks. Our code is available at https://github.com/disco-4d/Disco4D
Hui En Pang, Shuai Liu 0002, Zhongang Cai, Lei Yang 0045, Tianwei Zhang 0004, Ziwei Liu 0002
CVPR2
2025 EgoLife: Towards Egocentric Life Assistant
abstract
We introduce EgoLife, a project to develop an egocentric life assistant that accompanies and enhances personal efficiency through AI-powered wearable glasses. To lay the foundation for this assistant, we conducted a comprehensive data collection study where six participants lived together for one week, continuously recording their daily activities—including discussions, shopping, cooking, social-izing, and entertainment—using AI glasses for multimodal person-view video references. This effort resulted in EgoLife Dataset, a comprehensive 300-hour egocentric, terpersonal, multiview, and multimodal daily life with intensive annotation. Leveraging this dataset, we troduce EgoLifeQA, a suite of long-context, life-oriented question-answering tasks designed to provide meaningful sistance in daily life by addressing practical questions as recalling past relevant events, monitoring health and offering personalized recommendations.To address the key technical challenges of 1) developing robust visual-audio models for egocentric data, 2) enabling identity recognition, and 3) facilitating long-context question answering over extensive temporal information, we introduce EgoBulter, an integrated system comprising EgoGPT and EgoRAG. EgoGPT is an omni-modal model trained on egocentric datasets, achieving state-of-the-art performance on egocentric video understanding. EgoRAG is a retrieval-based component that supports answering ultra-long-context questions. Our experimental studies verify their working mechanisms and reveal critical factors and bottlenecks, guiding future improvements. By releasing our datasets, models, and benchmarks, we aim to stimulate further research in egocentric AI assistants.
Shuai Liu 0002, Hongming Guo, Yuhao Dong, Xiamengwei Zhang, Pengyun Wang, Zitang Zhou, Binzhu Xie, Bei Ouyang, Zhengyu Lin, Marco Cominelli, Zhongang Cai, Bo Li 0080, Yuanhan Zhang, Peiyuan Zhang, Fangzhou Hong, Jörg Widmer, Francesco Gringoli, Lei Yang 0059, Ziwei Liu 0002
CVPR2
2025 Service Enhancement and System Maintenance for MEIS
Caiyi Li, Changling Peng, Shuai Liu 0002
Mob. Networks Appl.3
2025 PCDPose: enhancing the lightweight 2D human pose estimation model with pose-enhancing attention and context broadcasting
Zhenyuan Tian, Weina Fu, Marcin Wozniak, Shuai Liu 0002
Pattern Anal. Appl.4
2025 Fcdnet: Fuzzy Cognition-Based Dynamic Fusion Network for Multimodal Sentiment Analysis
abstract
Multimodal sentiment analysis (MSA) provides a novel way to understand human sentiments. However, the differences between distribution patterns across modalities bring challenges in this domain. The inconsistency of recognitions with different modalities leads to incorrect final results. Moreover, the gaps between sentiments with different degrees are small in one modality, but the gaps between sentiments with same degree are large across different modalities. The imbalance leads to incorrect recognition for different sentiment degrees. Since the fuzzy network shows excellent performance in integrating data from multiple modalities, this study constructs a fuzzy cognition-based dynamic fusion network (Fcdnet) for MSA. The Fcdnet dynamically integrates sentiment scores across different modalities using a fuzzy cognition fusion mechanism (FCM), significantly enhancing the accuracy of identifying divergent sentiments across modalities. Additionally, a disparity balancing module (DBM) is proposed to normalize the representations between different modality features by penalizing the similarity of sentiments with different degrees and rewarding the separability of sentiments with same degree. Experimental results demonstrate that Fcdnet outperforms state-of-the-art methods on public datasets, validating the superiority and effectiveness.
Shuai Liu 0002, Weina Fu
IEEE Trans. Fuzzy Syst.1
2025 Fuzzy-Assisted Contrastive Decoding Improving Code Generation of Large Language Models
abstract
Large Language Models (LLMs) play a crucial role in intelligent code generation tasks. Most existing work focuses on pre-training or fine-tuning specialized code LLMs, e.g., CodeLlama. However, pre-training or fine-tuning a code LLM requires a vast corpus of data, significant computational resources, and considerable human effort. Compared to pre-training or fine-tuning LLMs, a simple and flexible method of contrastive decoding has garnered widespread attention to improve the text generation quality of LLMs. While contrastive decoding can indeed improve the text generation quality of LLMs, our research has found that directly using contrastive decoding: 1) introduces erroneous information into the logit distribution generated from normal prompts (i.e., user's input), particularly in the code generation of LLMs; 2) significantly impedes the inference and decoding time of LLMs. In this work, the limitations of using contrastive decoding directly are systematically highlighted, and a novel real-time fuzzy-assisted contrastive decoding (FCD) mechanism is proposed to improve the code generation quality of LLMs. The proposed FCD mechanism initially categorises prompts into high-quality and low-quality groups based on the results of the evaluator (i.e., unit test) before integrating the LLM. Next, feature values (e.g., standard deviation, peak value, etc.) related to the logit distribution of predicted tokens during the LLM's inference process for both high-quality and low-quality prompts are extracted. Finally, the extracted feature values are used to train the fuzzy neural network (i.e, fuzzy min-max neural network) offline, allowing for the prejudgement of the reliability of the logit distribution for normal prompt outputs. This prevents the direct use of erroneous information from contrastive decoding and improves the code generation quality of LLMs. Through extensive experiments, it has been demonstrated that the proposed FCD mechanism can significantly improve the code generation quality of LLMs through fuzzy-assisted contrastive decoding. Moreover, the FCD mechanism can also reduce the time required for inference and contrastive decoding. The code and data are publicly available on GitHub11https://github.com/LLMcodegen/Fuzzy_contrastive_decoding.and HuggingFace22https://huggingface.co/wangle123/Fuzzy_contrastive_decoding..
Shuai Wang 0011, Liang Ding 0006, Yibing Zhan, Yong Luo 0002, Shuai Liu 0002, Weiping Ding 0001
IEEE Trans. Fuzzy Syst.5
2024 Octopus: Embodied Vision-Language Programmer from Environmental Feedback
Yuhao Dong, Shuai Liu 0002, Bo Li 0080, Haoran Tan, Chencheng Jiang, Jiamu Kang, Yuanhan Zhang, Kaiyang Zhou, Ziwei Liu 0002
ECCV (1)3
2024 TSR-Jack: An In-Browser Crypto-Jacking Detection Method Based on Time Series Representation Learning
Bo Cui 0005, Shuai Liu 0002
ICICS (2)2
2024 Cefdet: Cognitive Effectiveness Network Based on Fuzzy Inference for Action Detection
abstract
Action detection and understanding provide the foundation for the generation and interaction of multimedia content. However, existing methods mainly focus on constructing complex relational inference networks, overlooking the judgment of detection effectiveness. Moreover, these methods frequently generate detection results with cognitive abnormalities. To solve the above problems, this study proposes a cognitive effectiveness network based on fuzzy inference (Cefdet), which introduces the concept of 'cognition--based detection' to simulate human cognition. First, a fuzzy-driven cognitive effectiveness evaluation module (FCM) is established to introduce fuzzy inference into action detection. FCM is combined with human action features to simulate the cognition-based detection process, which clearly locates the position of frames with cognitive abnormalities. Then, a fuzzy cognitive update strategy (FCS) is proposed based on the FCM, which utilizes fuzzy logic to re-detect the cognition-based detection results and effectively update the results with cognitive abnormalities. Experimental results demonstrate that Cefdet exhibits superior performance against several mainstream algorithms on the public datasets, validating its effectiveness and superiority.
Weina Fu, Shuai Liu 0002, Saeed Anwar, Sambit Bakshi, Khan Muhammad 0001
ACM Multimedia3
2024 Solution of wide and micro background bias in contrastive action representation learning
Shuai Liu 0002, Yunhe Wang 0009, Weina Fu, Weiping Ding 0001
Eng. Appl. Artif. Intell.1
2024 Coverage Path Planning for IoUAVs With Tiny Machine Learning in Complex Areas Based on Convex Decomposition
abstract
For Unmanned Aerial Vehicles (UAVs) with Tiny Machine Learning (TML), there is mutual exclusivity between the energy consumption for flight and the energy consumption to support their computation and processing. IoUAVs integrated with TML systems often consume substantial amounts of energy during flights, particularly when engaged in extended coverage and surveillance missions. The energy consumption of a UAV with TML performing long, wide-area coverage patrols and monitoring missions in complex areas is significant for the flight itself, and the energy required for the TML to perform calculations and processing is not guaranteed. Therefore, to better support TML computations, this study optimizes flight paths to reduce the energy consumption of UAVs while ensuring coverage. Specifically, in this study, the use of concave point elimination algorithms, enhanced convex decomposition algorithms, and determination of flight direction significantly reduced the frequency of UAV turns. The computational cost of obtaining a complete path is reduced by merging the subconvex regions and the weighted minimum traversal of the graph. This novel bidirectional forwarding path coverage path-planning (BFP-CPP) algorithm maximizes the reduction in the number of turns, reduces energy consumption, and achieves global coverage. The simulation experimental results show that compared with the existing methods without concave point elimination, the BFP-CPP algorithm can effectively reduce the number of subregions, minimize the number of drone turns, and lower energy consumption.
Bing Jia, Jianqiang Jing, Baoqi Huang, Shuai Liu 0002, Khan Muhammad 0001, Joel J. P. C. Rodrigues
IEEE Internet Things J.5
2024 DGNet: A Handwritten Mathematical Formula Recognition Network Based on Deformable Convolution and Global Context Attention
Cuihong Wen, Lemin Yin, Shuai Liu 0002
Mob. Networks Appl.3
2024 An algorithm for overlapping chromosome segmentation based on region selection
Xiangbin Liu, Jerry Chun-Wei Lin, Shuai Liu 0002
Neural Comput. Appl.4
2023 Node-Disjoint Paths in Balanced Hypercubes with Application to Fault-Tolerant Routing
Shuai Liu 0002, Yan Wang 0078, Jianxi Fan, Baolei Cheng
ICA3PP (3)1
2023 SynBody: Synthetic Dataset with Layered Human Models for 3D Human Perception and Modeling
abstract
Synthetic data has emerged as a promising source for 3D human research as it offers low-cost access to large-scale human datasets. To advance the diversity and annotation quality of human models, we introduce a new synthetic dataset, SynBody, with three appealing features: 1) a clothed parametric human model that can generate a diverse range of subjects; 2) the layered human representation that naturally offers high-quality 3D annotations to support multiple tasks; 3) a scalable system for producing realistic data to facilitate real-world tasks. The dataset comprises 1.2M images with corresponding accurate 3D annotations, covering 10,000 human body models, 1,187 actions, and various viewpoints. The dataset includes two subsets for human pose and shape estimation as well as human neural rendering. Extensive experiments on SynBody indicate that it substantially enhances both SMPL and SMPL-X estimation. Furthermore, the incorporation of layered annotations offers a valuable training resource for investigating the Human Neural Radiance Fields(NeRF).
Zhitao Yang, Zhongang Cai, Haiyi Mei, Shuai Liu 0002, Zhaoxi Chen 0009, Weiye Xiao, Yukun Wei, Zhongfei Qing, Bo Dai 0002, Wayne Wu, Chen Qian 0006, Dahua Lin, Ziwei Liu 0002, Lei Yang 0059
ICCV4
2023 4D Panoptic Scene Graph Generation
abstract
We are living in a three-dimensional space while moving forward through a fourth dimension: time. To allow artificial intelligence to develop a comprehensive understanding of such a 4D environment, we introduce **4D Panoptic Scene Graph (PSG-4D)**, a new representation that bridges the raw visual data perceived in a dynamic 4D world and high-level visual understanding. Specifically, PSG-4D abstracts rich 4D sensory data into nodes, which represent entities with precise location and status information, and edges, which capture the temporal relations. To facilitate research in this new area, we build a richly annotated PSG-4D dataset consisting of 3K RGB-D videos with a total of 1M frames, each of which is labeled with 4D panoptic segmentation masks as well as fine-grained, dynamic scene graphs. To solve PSG-4D, we propose PSG4DFormer, a Transformer-based model that can predict panoptic segmentation masks, track masks along the time axis, and generate the corresponding scene graphs via a relation component. Extensive experiments on the new dataset show that our method can serve as a strong baseline for future research on PSG-4D. In the end, we provide a real-world application example to demonstrate how we can achieve dynamic scene understanding by integrating a large language model into our PSG-4D system.
Jun Cen, Wenxuan Peng, Shuai Liu 0002, Fangzhou Hong, Xiangtai Li, Kaiyang Zhou, Qifeng Chen 0001, Ziwei Liu 0002
NeurIPS4
2023 Human Inertial Thinking Strategy: A Novel Fuzzy Reasoning Mechanism for IoT-Assisted Visual Monitoring
abstract
Computer vision has always been a hot field of research by contemporary scholars due to its wide range of applications. As an important branch of this field, the visual monitoring technology has shown superior vitality in the actual monitoring environment of the Internet of Things (IoT). However, when the monitoring environment is complex, once the target monitoring fails, the important information related to the target also disappears. At this time, if the existing monitoring method is used, the target cannot be monitored again. Moreover, the current filtering monitoring algorithm also has the problem of poor interpretability. Therefore, this article combines the relevant characteristics of human inertial thinking when dealing with such problems. First, our method screens the movement information of the target and introduces a fuzzy reasoning mechanism to infer the location area of the target through fuzzy thinking. Then, an alternative selection strategy based on the thinking set is applied, which alternates between the location of thinking reasoning and the location of memory to further obtain the effective visual monitoring of the target. The filtering and monitoring algorithm fused with the new mechanism in the OTB-2015 data set, the UVA123 data set, and the TC128 data set all show that the proposed fuzzy inference mechanism has good robustness and universality. Furthermore, our results confirm that it can not only ensure the monitoring speed and overall accuracy but also improve the stability of monitoring in the IoT-assisted monitoring environment, showing its effectiveness compared to state-of-the-art methods. In addition, our results confirm that the integration of the proposed edge learning method with the IoT can be well applied to the construction of smart cities and future generation systems.
Shuai Liu 0002, Shuai Wang 0011, Xinyu Liu 0012, Jianhua Dai 0003, Khan Muhammad 0001, Amir Hossein Gandomi, Weiping Ding 0001, Mohammad Hijji, Victor Hugo C. de Albuquerque
IEEE Internet Things J.1
2023 Multi-modal fusion network with complementarity and importance for emotion recognition
Shuai Liu 0002, Weina Fu, Weiping Ding 0001
Inf. Sci.1
2023 Student behavior recognition for interaction detection in the classroom environment
Abdul Khader Jilani Saudagar, Abdul Malik Badshah, Khan Muhammad 0001, Shuai Liu 0002
Image Vis. Comput.6
2023 Manta ray foraging optimizer-based image segmentation with a two-strategy enhancement
Benedict Jun Ma, João Luiz Junho Pereira, Diego Oliva 0001, Shuai Liu 0002, Yong-Hong Kuo
Knowl. Based Syst.4
2023 Empirical Research of Classroom Behavior Based on Online Education: A Systematic Review
Yishu Huang, Changling Peng, Shuai Liu 0002
Mob. Networks Appl.3
2023 An Effective Learning Evaluation Method Based on Text Data with Real-time Attribution - A Case Study for Mathematical Class with Students of Junior Middle School in China
abstract
In today's intelligent age, the vigorous development of education-based information analysis technology has had a profound impact on the education and teaching process. The use of computational linguistics technology to extract teaching data for learning evaluation is an important hot domain in this research field. Therefore, the study of student learning assessment methods based on text data has become a key issue. The text data extracted from the education process has attributes related to time and operational attributes, which are important indicators to measure the effect of student learning effect. However, these attributes are not focused by the traditional educational effect evaluation method, which make the learning effect of students difficult to measure comprehensively and effectively. In response to this problem, this article first uses perception technology to extract learning text data based on time and operational attributes. Secondly, according to the real-time attributes of text data, such as time and operation attributes, a learning evaluation method based on real-time text data is proposed. Finally, this article compares the traditional evaluation method with the proposed method. The results show that using real-time attribute text data is more effective in students’ learning measure.
Shuai Liu 0002, Tenghui He, Akshi Kumar 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2023 Efficient Visual Tracking Based on Fuzzy Inference for Intelligent Transportation Systems
abstract
Remote monitoring is an important application of intelligent transportation systems (ITSs). The combination of monitoring equipment and tracking algorithms can be used to automatically track moving targets. The tracking algorithm based on the Siamese network is both accurate and efficient, and its development potential is better than that of other algorithms. Its output is a detection map that reflects the probability that any position in the search area is the center of the target’s bounding box, and the maximum value of the detection map is the center of the target’s bounding box predicted by the algorithm. Owing to partial occlusion, target deformation, out-of-view, and background clutter, local maxima in the detection map may also be the center of the target’s bounding box. A tracker’s ability to make accurate judgments is currently limited. Furthermore, previous trackers extracted only the target features in the initial frame as the matching template. Although this matching template is highly reliable, it cannot effectively combine the target features available in the subsequent frames. Therefore, in this study, fuzzy inference is introduced into the tracking process to analyze the reliability of the detection map. When this map is reliable, the target feature of the search area is transformed into a substitute template; otherwise, multiple substitute templates are selected from the template pool for parallel matching as per the set rules. The optimal result is selected from multiple detection results, based on the priority of the detection results when the initial frame is used as the matching template. Experimental results on multiple datasets show that the proposed algorithm is superior to other similar algorithms in terms of multiple assessment metrics and can improve the robustness of remote monitoring tasks in ITSs.
Shuai Liu 0002, Shichen Huang, Xiyu Xu, Jaime Lloret Mauri, Khan Muhammad 0001
IEEE Trans. Intell. Transp. Syst.1
2023 A Reliable Sample Selection Strategy for Weakly Supervised Visual Tracking
abstract
Reliability is an important property in the applied engineering systems, especially in visual tracking. The supervised visual tracking method uses reliable ground truth that is manually annotated, which is hard to get in many applications. However, weakly supervised visual trackings are limited by the low-quality labels. Therefore, a reliable sample selection strategy is the most important issue for the weakly supervised visual trackings. In this article, we propose an optimal sample selection strategy and apply it to the visual tracking system. The strategy first assesses the reliability of the samples according to the score map, where the score map is the pseudolabel generated by the upstream task to meet the needs of the downstream task. Then, the unreliable pseudolabels are replaced by reliable ground truth or discarded to overcome the degraded modeling problem by filtering low-quality samples. Finally, through comparison with multiple selection strategies, it is verified that the model trained using this strategy has the best performance. The proposed visual tracking model achieves the best performance among multiple assessment metrics in multiple datasets. Experiments verify that the scientific sample quality assessment method is very important. It can guide the improvement of model performance, which is of great help to the weakly supervised learning systems based on data.
Shuai Liu 0002, Xiyu Xu, Khan Muhammad 0001, Weina Fu
IEEE Trans. Reliab.1
2023 Key problem on mobile intelligent multimedia system
Weina Fu, Zeshi Chen, Shuai Liu 0002
Wirel. Networks3
2022 A high-precision correction method in non-rigid 3D motion poses reconstruction
abstract
Occlusion, rotation and other factors affect human motion structure because of the incomplete acquired image sequence, resulting in poor performance of non-rigid three-dimensional (3D) motion pose reconstruction. A non-rigid 3D reconstruction and high-precision correction method for motion pose are studied in this paper. A non-rigid imaging model is designed to obtain 3D moving images. According to the frame difference and morphological processing, the background of image is separated and denoised. Combined with motion analysis, 3D motion pose features are extracted as identification of non-rigid 3D motion error actions in a hybrid Convolution Neural Network-Hidden Markov Model to train the correction coefficients, which are used to adjust the pose in 3D motion reconstruction and realise correction. Experimental results show that this method has high precision reconstruction and correction of non-rigid 3D motion pose.
Cuihong Fan, Weina Fu, Shuai Liu 0002
Connect. Sci.3
2022 A fingerprint-based localization algorithm based on LSTM and data expansion method for sparse samples
Bing Jia, Wenling Qiao, Zhaopeng Zong, Shuai Liu 0002, Mohammad Hijji, Javier Del Ser, Khan Muhammad 0001
Future Gener. Comput. Syst.4
2022 An image enhancement algorithm of video surveillance scene based on deep learning
abstract
Abstract Target enhancement is the most important task in a video surveillance system. In order to improve the accuracy and efficiency of target enhancement, and better deal with the subsequent recognition, tracking, behaviour understanding and other processing of targets, a deep learning‐based image enhancement algorithm for video surveillance scenes is proposed. First, the super‐resolution reconstruction of the image is carried out through the image super‐resolution reconstruction method based on the hybrid deep convolutional network to improve the sharpness of the image. Then, for the reconstructed video surveillance scene image, the watershed image enhancement algorithm based on morphology and region merging is used to realize the enhancement of the video surveillance scene image. Deep learning algorithms can improve the accuracy of image enhancement through iterative calculations. Experimental results show that after image enhancement in daytime, night and noisy video surveillance scenes, the maximum enhancement difference rate is less than 0.5%, the cross‐linking degree is close to 1, and the average image enhancement time is less than 1.3 s. It can realize image enhancement of video surveillance scenes and improve the image clarity of the video surveillance scene.
Wei-wei Shen, Shuai Liu 0002, Yudong Zhang 0001
IET Image Process.3
2022 Human-centered attention-aware networks for action recognition
abstract
Action recognition in video is a research hot spot in the field of computer vision. Learning important clues in video context has significant effect to promote the interaction prediction and gesture recognition. Most existing methods infer the interactions between actor and context through relational reasoning methods. While these relational features contribute to improve the salience of action performance, the error will occur when the salient region is irrelevant to the recognized action. Therefore, this paper establishes a human-centered attention mechanism that dynamically highlights regions associated with action recognition according to target appearance to selectively recognize the human-object interaction action. The effectiveness of the proposed mechanism is verified on the AVA2.2 data set, and the visualized attention map further shows that the proposed attention mechanism can effectively recognize human-centered strongly correlated action.
Shuai Liu 0002, Weina Fu
Int. J. Intell. Syst.1
2022 Human Short Long-Term Cognitive Memory Mechanism for Visual Monitoring in IoT-Assisted Smart Cities
abstract
In the industry 4.0 era, the visualization and real-time automatic monitoring of smart cities supported by the Internet of Things is becoming increasingly important. The use of filtering algorithms in smart city monitoring is a feasible method for this purpose. However, maintaining fast and accurate monitoring in complex surveillance environments with restricted resources remains a major challenge. Since the cognitive theory in visual monitoring is difficult to realize in practice, efficient monitoring of complex environments is accordingly hard to be achieved. Moreover, current monitoring methods do not consider the particularities of the human cognitive system, so the remonitoring ability of the process/target is weak in case of monitoring failure by the monitoring system. To overcome these issues, this article proposes a novel human short-long cognitive memory mechanism for video surveillance in smart cities. In this mechanism, a memory with a high reliability target is used as a “long-term memory,” whereas a memory with a low reliability target is used as a “short-term memory.” During the monitoring process, the “short-term memory” and “long-term memory” alternation strategy is combined with the stored target appearance characteristics, ensuring that the original model in the memory will not be contaminated or mislaid by changes in the external environment (occlusion, fast motion, motion blur, and background clutter). Extensive simulations showcase that the algorithm proposed in this article not only improves the monitoring speed without hindering its real-time operation but also monitors and traces the monitored target accurately, ultimately improving the robustness of the detection in complex scenery, and enabling its application to IoT-assisted smart cities.
Shuai Wang 0011, Xinyu Liu 0012, Shuai Liu 0002, Khan Muhammad 0001, Ali Asghar Heidari, Javier Del Ser, Victor Hugo C. de Albuquerque
IEEE Internet Things J.3
2022 Multi-strategy ensemble binary hunger games search for feature selection
Benedict Jun Ma, Shuai Liu 0002, Ali Asghar Heidari
Knowl. Based Syst.2
2022 Advanced Machine Learning Based Mobile Multimedia Application
Weina Fu, Shuai Liu 0002
Mob. Networks Appl.3
2022 An Introduction to Artificial Intelligence and Machine Learning for Online Education
Changling Peng, Xuanyu Zhou, Shuai Liu 0002
Mob. Networks Appl.3
2022 Profile of Intelligent Hybrid Information System in Mobile World
Weina Fu, Shuai Liu 0002
Mob. Networks Appl.3
2022 A closed-loop healthcare processing approach based on deep reinforcement learning
Yinglong Dai, Guojun Wang 0001, Khan Muhammad 0001, Shuai Liu 0002
Multim. Tools Appl.4
2021 Effective template update mechanism in visual tracking with background clutter
Shuai Liu 0002, Dongye Liu, Khan Muhammad 0001, Weiping Ding 0001
Neurocomputing1
2021 An Introduction to Key Technology in Artificial Intelligence and big Data Driven e-Learning and e-Education
Shuai Liu 0002
Mob. Networks Appl.3
2021 A Survey of CRF Algorithm Based Knowledge Extraction of Elementary Mathematics in Chinese
Shuai Liu 0002, Tenghui He, Jianhua Dai 0003
Mob. Networks Appl.1
2021 An Introduction to Multimedia Technology and Enhanced Learning
Liyun Xia, Shuai Liu 0002
Mob. Networks Appl.2
2021 Fuzzy-aided solution for out-of-view challenge in visual tracking under IoT-assisted complex environment
Shuai Liu 0002, Xinyu Liu 0012, Shuai Wang 0011, Khan Muhammad 0001
Neural Comput. Appl.1
2021 Fuzzy Detection Aided Real-Time and Robust Visual Tracking Under Complex Environments
abstract
Today, a new generation of artificial intelligence has brought several new research domains such as computer vision (CV). Thus, target tracking, the base of CV, has been a hotspot research domain. Correlation filter (CF)-based algorithm has been the basis of real-time tracking algorithms because of the high tracking efficiency. However, CF-based algorithms usually failed to track objects in complex environments. Therefore, this article proposes a fuzzy detection strategy to prejudge the tracking result. If the prejudge process determines that the tracking result is not good enough in the current frame, the stored target template is used for following tracking to avoid the template pollution. During testing on the OTB100 dataset, the experimental results show that the proposed auxiliary detection strategy improves the tracking robustness under complex environment by ensuring the tracking speed.
Shuai Liu 0002, Shuai Wang 0011, Xinyu Liu 0012, Chin-Teng Lin, Zhihan Lyu
IEEE Trans. Fuzzy Syst.1
2021 Human Memory Update Strategy: A Multi-Layer Template Update Mechanism for Remote Visual Monitoring
abstract
In the era of rapid development of artificial intelligence, the integration of multimedia and human-artificial intelligence has become an important research hotspot. Especially in the multimedia environment, effective remote visual monitoring has become the exploration direction of many scholars. The use of traditional correlation filtering (CF) algorithm for real-time monitoring in the context of multimedia is a practical strategy. However, most existing filtering-based visual monitoring algorithms still have the problem of insufficient robustness and effectiveness. Therefore, by considering the strategy of updating human memory, this paper proposes a multi-layer template update mechanism to achieve effective monitoring in a multimedia environment. In this strategy, the weighted template of the high-confidence matching memory is used as the confidence memory, and the unweighted template of the low-confidence matching memory is used as the cognitive memory. Through the alternate use of confidence memory, matching memory, and cognitive memory, it is ensured that the target will not be lost during the monitoring process. Experimental results show that this strategy does not affect the speed (still real-time) and improves the robustness in the multimedia background.
Shuai Liu 0002, Shuai Wang 0011, Xinyu Liu 0012, Amir Hossein Gandomi, Mahmoud Daneshmand, Khan Muhammad 0001, Victor Hugo C. de Albuquerque
IEEE Trans. Multim.1
2021 Medical Image Classification based on an Adaptive Size Deep Learning Model
abstract
With the rapid development of Artificial Intelligence (AI), deep learning has increasingly become a research hotspot in various fields, such as medical image classification. Traditional deep learning models use Bilinear Interpolation when processing classification tasks of multi-size medical image dataset, which will cause the loss of information of the image, and then affect the classification effect. In response to this problem, this work proposes a solution for an adaptive size deep learning model. First, according to the characteristics of the multi-size medical image dataset, the optimal size set module is proposed in combination with the unpooling process. Next, an adaptive deep learning model module is proposed based on the existing deep learning model. Then, the model is fused with the size fine-tuning module used to process multi-size medical images to obtain a solution of the adaptive size deep learning model. Finally, the proposed solution model is applied to the pneumonia CT medical image dataset. Through experiments, it can be seen that the model has strong robustness, and the classification effect is improved by about 4% compared with traditional algorithms.
Xiangbin Liu, Jiesheng He, Liping Song, Shuai Liu 0002, Gautam Srivastava 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2020 Parallel generated method of transcriptional regulatory networks
abstract
Summary Generated method of transcriptional regulatory networks remains an important research in biology. Many approaches have been proposed to construct transcriptional regulatory networks. However, with the increase of ChIP‐seq and RNA‐seq data, the speed of constructing transcriptional regulation networks is still a challenge. Moreover, parallel computing lacks application in constructing gene regulatory networks through the analysis of the relationships between transcription factors (TFs) and target genes (TGs). Therefore, in this paper, a parallel generated method of transcriptional regulatory networks was proposed. First, two datasets, Michigan Cancer Foundation – 7 (MCF‐7) and Cardiomyocytes (CM) were applied. Then, a parallel method was used to generate transcriptional regulatory network with their transcription factors (TFs) and target genes (TGs). Finally, experimental results showed that 61% regulatory relations in MCF‐7 were validated in the Gene Expression Omnibus (GEO), while 29% results needed further experimental verification. Besides, 56% regulatory relations in CM were consistent with GEO, while 33% results were not yet verified. Furthermore, speed of parallel algorithm was faster than traditional serial algorithm in generating transcriptional regulatory networks.
Shuai Liu 0002, Na Ta 0005, Mengye Lu, Gaocheng Liu, Weiling Bai, Wenhui Li 0002
Concurr. Comput. Pract. Exp.1
2020 Property of Self-Similarity Between Baseband and Modulated Signals
Shuai Liu 0002, Weiling Bai, Gautam Srivastava 0001, José António Tenreiro Machado
Mob. Networks Appl.1
2020 Recent Advancement in Hybrid Big Data Processing
Shuai Liu 0002, Huiyu Zhou 0001, Xiaochun Cheng
Mob. Networks Appl.1
2020 Recent Progress on the Intelligent Computing for Multimodal Information
Tiejun Zhu, Shuai Liu 0002
Mob. Networks Appl.2
2020 Practical augmented reality (AR) technology and its applications
Jin-Ho Choi, Ruidan Su, Shuai Liu 0002, Hyun-Jong Cha
Multim. Tools Appl.3
2020 An optimized cluster storage method for real-time big data in Internet of Things
Li Tu, Shuai Liu 0002, Yan Wang 0078
J. Supercomput.2
2019 Joint Dictionary Learning for Unsupervised Feature Selection
Jianhua Dai 0003, Qilai Zhang, Shuai Liu 0002
ICANN (2)4
2019 A Robust Parallel Object Tracking Method for Illumination Variations
Shuai Liu 0002, Gaocheng Liu, Huiyu Zhou 0001
Mob. Networks Appl.1
2019 Introduction of Key Problems in Long-Distance Learning and Training
Shuai Liu 0002, Zhaojun Li 0001, Yudong Zhang 0001, Xiaochun Cheng
Mob. Networks Appl.1
2019 A real-time distributed cluster storage optimization for massive data in internet of multimedia things
Shuai Liu 0002
Multim. Tools Appl.2
2019 Utilization of DenseNet201 for diagnosis of breast abnormality
Nianyin Zeng, Shuai Liu 0002, Yudong Zhang 0001
Mach. Vis. Appl.3
2019 Nucleosome positioning based on generalized relative entropy
Mengye Lu, Shuai Liu 0002
Soft Comput.2
2019 Corrections to "Fractal Intelligent Privacy Protection in Online Social Network Using Attribute-Based Encryption Schemes"
abstract
In[1], the financial support information in the first footnote should have read as follows.
Wei Wei 0006, Shuai Liu 0002, Wenjia Li, Ding-Zhu Du
IEEE Trans. Comput. Soc. Syst.2
2018 Visual attention feature (VAF) : A novel strategy for visual tracking based on cloud platform in intelligent surveillance systems
Shuai Liu 0002, Arun Kumar Sangaiah, Khan Muhammad 0001
J. Parallel Distributed Comput.2
2018 Introduction of Recent Advanced Hybrid Information Processing
Shuai Liu 0002, Zhaojun Li 0001, Xiaochun Cheng, Yun Lin 0005
Mob. Networks Appl.1
2018 Fractal Research on the Edge Blur Threshold Recognition in Big Data Classification
Jia Wang 0008, Shuai Liu 0002, Houbing Song
Mob. Networks Appl.2
2018 Fractal Intelligent Privacy Protection in Online Social Network Using Attribute-Based Encryption Schemes
abstract
While the online social network (OSN) has brought much convenience to users, there are still some serious problems, such as personal privacy leaks. Today, OSN security and privacy protection are one of the most important focuses of the research. In this paper, we present an intelligent privacy protection approach to solve problems of security and privacy protection in OSNs. First, the proposed algorithm combines a neural network with a hybrid hierarchy genetic algorithm and radial basis function, which is used to construct a prediction model of OSN security. Then, a support vector machine is applied to preprocess information of the OSN, and the attribute-based encryption scheme is adopted to encrypt the OSN information. Finally, a particle swarm optimization algorithm is used to improve OSN security and privacy protection. The experimental results demonstrate the effectiveness of the proposed method.
Wei Wei 0006, Shuai Liu 0002, Wenjia Li, Ding-Zhu Du
IEEE Trans. Comput. Soc. Syst.2
2017 Digital image watermarking method based on DCT and fractal encoding
abstract
With the rapid development of computer science, problems with digital products piracy and copyright dispute become more serious; therefore, it is an urgent task to find solutions for these problems. In this study, the authors’ develop a digital watermarking algorithm based on a fractal encoding method and the discrete cosine transform (DCT). The proposed method combines fractal encoding method and DCT method for double encryptions to improve traditional DCT method. The image is encoded by fractal encoding as the first encryption, and then encoded parameters are used in DCT method as the second encryption. First, the fractal encoding method is adopted to encode a private image with private scales. Encoding parameters are applied as digital watermarking. Then, digital watermarking is added to the original image to reversibly using DCT, which means the authors can extract the private image from the carrier image with private encoding scales. Finally, attacking experiments are carried out on the carrier image by using several attacking methods. Experimental results show that the presented method has higher performance characteristics such as robustness and peak signal to noise ratio than classical methods.
Shuai Liu 0002, Houbing Song
IET Image Process.1
2017 Multiple Object Detection and Tracking in Complex Background
abstract
Multiple object tracking is a fundamental step for many computer vision applications. However, detecting and tracking objects in complex background is still a challenging task. This paper proposes an approach, which combines an improved Gaussian mixture modeling (GMM) with multiple particle filters (MPFs) for automatic multiple targets detecting and tracking. For GMM, we make improvement on GMM in the phase of model updating by using the expectation maximization algorithm and [Formula: see text] recent frames with weight parameters of Gaussian distributions. In the tracking stage, we integrate multiple features of targets, including color, edge and depth, into MPFs to improve the performance of object tracking. By comparing with various particle filter approaches, the experimental results show that our approach can track multiple targets in complex backgrounds automatically and accurately.
Shuai Liu 0002
Int. J. Pattern Recognit. Artif. Intell.3
2017 Editorial: Multimedia in Technology Enhanced Learning
Zhigao Zheng 0001, Jinming Wen, Shuai Liu 0002
Mob. Networks Appl.3
2017 Erratum to: Editorial: Multimedia in Technology Enhanced Learning
Zhigao Zheng 0001, Jinming Wen, Shuai Liu 0002
Mob. Networks Appl.3
2017 Distribution of primary additional errors in fractal encoding method
Shuai Liu 0002, Weina Fu, Liqiang He, Jiantao Zhou 0002, Ming Ma 0006
Multim. Tools Appl.1
2017 A review of visual moving target tracking
Shuai Liu 0002, Weina Fu
Multim. Tools Appl.2
2017 Development healthcare PC and multimedia software for improvement of health status and exercise habits
Sekyoung Youm, Shuai Liu 0002
Multim. Tools Appl.2
2017 The Fusion Model of Multidomain Context Information for the Internet of Things
abstract
The Internet of Things aims to provide the user with deep adaptive intelligence services according to the user’s personalized characteristics. Most of the characteristics are presented in the form of high-level context. But it often lacks methods to obtain high-level context information directly in the Internet of Things. In this paper, so as to achieve the corresponding high-level context information using the specific low-level multidomain context directly obtained by different sensors in the Internet of Things, we present a machine learning method to construct a context fusion model based on the feature selection algorithm and the multiclassification algorithm. First, we propose a wrapper feature selection method based on the genetic algorithm to obtain a simpler and more important subset of the context features from the low-level multidomain context, by defining a suitable fitness function and a convergence condition. Then, we use the decision tree algorithm which is a multiclassification algorithm, based on the rules obtained by training the subset of context features, to determine which high-level context the record set of the low-level context information belongs to. Experiments confirm that the model can be used to achieve higher classification accuracy without more significant time consumption.
Bing Jia, Shuai Liu 0002, Yushuai Guan, Wuyungerile Li, Weiwu Ren
Wirel. Commun. Mob. Comput.2
2016 Differential trajectory tracking with automatic learning of background reconstruction
Weina Fu, Jiantao Zhou 0002, Shuai Liu 0002, Ming Ma 0006
Multim. Tools Appl.3
2016 A fractal image encoding method based on statistical loss used in agricultural image compression
Shuai Liu 0002, Lingyun Qi, Ming Ma 0006
Multim. Tools Appl.1
2013 PreDNA: accurate prediction of DNA-binding sites in proteins by integrating sequence and geometric structure information
abstract
MOTIVATION: Protein-DNA interactions often take part in various crucial processes, which are essential for cellular function. The identification of DNA-binding sites in proteins is important for understanding the molecular mechanisms of protein-DNA interaction. Thus, we have developed an improved method to predict DNA-binding sites by integrating structural alignment algorithm and support vector machine-based methods. RESULTS: Evaluated on a new non-redundant protein set with 224 chains, the method has 80.7% sensitivity and 82.9% specificity in the 5-fold cross-validation test. In addition, it predicts DNA-binding sites with 85.1% sensitivity and 85.3% specificity when tested on a dataset with 62 protein-DNA complexes. Compared with a recently published method, BindN+, our method predicts DNA-binding sites with a 7% better area under the receiver operating characteristic curve value when tested on the same dataset. Many important problems in cell biology require the dense non-linear interactions between functional modules be considered. Thus, our prediction method will be useful in detecting such complex interactions.
Qianzhong Li, Shuai Liu 0002, Yongchun Zuo
Bioinform.3