Yujie Li 0001

dblp:28/7846-1 · DBLP profile ↗
← Back
77ranked-venue papers
17as first author
38since 2021 · last 2026
0000-0002-0275-2797ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 18 · 5 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 5 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 14 · 3 first-author · 7 since 2021Software engineering, systems software and programming languages · 8 · 1 first-author · 2 since 2021Systems, architecture and hardware · 7 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2026 MSENet: High efficiency video compression via Multivariate Spatiotemporal Entropy Network
Huimin Lu 0001, Liangfan Shi, Yuchao Zheng 0001, Yujie Li 0001
Image Vis. Comput.4
2025 Polygon Mesh Recovery via Segmentation Priors and Neural Radiance Fields
abstract
Neural Radiance Fields (NeRF) has made significant advancements in view synthesis and 3D reconstruction. However, the reconstructed outputs often lack semantic information, limiting its applicability for scene understanding tasks. Additionally, NeRF-based implicit representations focus on novel view synthesis, leading to coarse, non-editable surface meshes. This paper addresses the asset generation task within NeRF, incorporating prior semantic information to guide the reconstruction process. We propose a texture mesh recovery method that combines learned segmentation priors with NeRF. By leveraging the Segmentation Anything Model (SAM) for extracting object masks from RGB images, we ensure semantic consistency. Camera poses, derived from a multi-view reconstruction algorithm, are integrated into a mesh-based NeRF framework to learn explicit scene representations. Using an optimization approach, we disentangle texture and mesh, thereby enabling the creation of editable 3D assets. Experimental results demonstrate the effectiveness of the proposed method across multiple reconstruction datasets.
Jintong Cai, Huimin Lu 0001, Yujie Li 0001
IWCMC3
2025 A Multi-Degree-of-Freedom Wave Energy Harvester for Self-Powered IoT System
abstract
The Internet of Things (IoT) systems are essential for the development of smart cities, enabling efficient management and real-time monitoring of urban water resources. However, IoT systems still face challenges related to power supply in remote or off-grid environments. This study presents the design and development of a hybrid wave energy harvester (H-WEH), which integrates electromagnetic generator (EMG) and triboelectric nanogenerator (TENG) to capture ocean wave energy and provide a reliable power source for IoT systems. The experimental results show that the H-WEH effectively captures wave energy, with stable power output and energy storage capabilities, making it suitable for IoT applications. This study highlights the potential of wave energy as a sustainable solution for powering IoT systems, contributing to the advancement of smart city infrastructure, and providing a new approach to urban water resource monitoring.
Bozhi Ding, Huimin Lu 0001, Yujie Li 0001
IWCMC3
2024 Embracing Unimodal Aleatoric Uncertainty for Robust Multimodal Fusion
abstract
As a fundamental problem in multimodal learning, multimodal fusion aims to compensate for the inherent limitations of a single modality. One challenge of multimodal fusion is that the unimodal data in their unique embedding space mostly contains potential noise, which leads to corrupted cross-modal interactions. However, in this paper, we show that the potential noise in unimodal data could be well quantified and further employed to enhance more stable unimodal embeddings via contrastive learning. Specifically, we propose a novel generic and robust multimodal fusion strategy, termed Embracing Aleatoric Uncertainty (EAU), which is simple and can be applied to kinds of modalities. It consists of two key steps: (1) the Stable Unimodal Feature Augmentation (SUFA) that learns a stable unimodal representation by incorporating the aleatoric uncertainty into self-supervised contrastive learning. (2) Robust Multimodal Feature Integration (RMFI) leveraging an information-theoretic strategy to learn a robust compact joint representation. We evaluate our proposed EAU method on five multimodal datasets, where the video, RGB image, text, audio, and depth image are involved. Extensive experiments demonstrate the EAU method is more noise-resistant than existing multimodal fusion strategies and establishes new state-of-the-art on several benchmarks.
Zixian Gao, Xun Jiang 0001, Xing Xu 0001, Fumin Shen, Yujie Li 0001, Heng Tao Shen
CVPR5
2024 Few-Shot Object Detection Algorithm Based on Geometric Prior and Attention RPN
abstract
Intelligent factories driven by deep vision technology use robotic arms to perform tasks such as picking and assembling in a production environment. In practical applications, with the continuous changes of products on the industrial pipeline, the detection model needs to continuously train new weights to adapt to new application scenarios. It is time-consuming and labor-intensive to manually collect training data when deploying the production line, and it cannot be quickly adapted in industrial scenarios. Therefore, we propose an attention RPN (Region Proposal Network) few-shot object detection algorithm based on geometric prior. The algorithm uses the attention RPN module to strengthen the feature extraction ability of the detection model and uses the virtual simulation software to generate synthetic data similar to the real object geometry as the base class data to train the feature extraction network so that the network obtains the ability to extract geometric features on the base class object. By comparing the learning strategies, only a small number of real data samples are used to train the detection model twice. The experimental results show that the algorithm can detect more objects than the existing few-shot object detection algorithm in the industrial scene with only a small amount of real sample data, and the detection accuracy can reach 97%.
Xiu Chen, Yujie Li 0001, Huimin Lu 0001
IWCMC2
2024 Efficient 3D Object Recognition for Unadjusted Bin Picking Automation
abstract
In light of burgeoning technological progress and burgeoning labor deficits, the adoption of industrial robots has markedly intensified. These sophisticated automatons are pivotal in addressing the growing trend of high-mix, low-volume production, catering to the heterogeneous requisites of end-users. In this domain, it is imperative for industrial robots to facilitate automated bin picking, ensuring versatility and continuity in production workflows. Despite this, extant bin picking modalities fall short in discerning and orienting designated parts accurately. Our study introduces an avant-garde 3D object recognition framework, underpinned by deep learning algorithms, to streamline the bin picking process, obviating the necessity for human intervention. Moreover, while annotated data remains the cornerstone of deep learning paradigms, its procurement through conventional annotation is fraught with challenges. Addressing this bottleneck, we put forth a strategy that exploits training data autonomously generated within a simulation milieu, laying the groundwork for an object recognition model that eschews manual calibration. This model adeptly harnesses both bi-dimensional imagery and tri-dimensional point clouds to refine its recognition capabilities. Our empirical investigation, straddling simulated and authentic settings, substantiates the precision of our proposed methodology.
Yuchao Zheng 0001, Xiu Chen, Yujie Li 0001
IWCMC3
2024 Underwater Visibility Enhancement IoT System in Extreme Environment
abstract
Imagery captured in extreme underwater environments often presents unique challenges, including blurred details, color distortion, and reduced contrast. These discrepancies largely emanate from the intricate interplay of light absorption and scattering within the aquatic medium. Predominant restoration techniques, rather simplistically, apply a static attenuation coefficient, neglecting the dynamic nuances of underwater conditions, leading to an inconsistent restoration outcome. To counter these impediments, we introduce an avant-garde Underwater Internet of Things (Underwater IoT) system, underpinned by a scene-depth fusion paradigm. Our methodology astutely accounts for the spectral decay of light underwater to infer a more refined attenuation coefficient tailored to the specific scene. This system, employing a quadtree decomposition for precise localization coupled with depth mapping, facilitates an astute estimation of prevailing luminescence. This depth map, once synthesized and refined, aids in gauging the precise attenuation dynamics of the aqueous milieu, culminating in a more precise transmission map derivation. Segueing from this, we employ an inverse model to refurbish the original image. Experimental results highlight our system’s prowess in counteracting issues like muddied details and chromatic anomalies while concurrently amplifying contrast. In juxtaposition with a spectrum of existing methodologies, our innovation outshines in terms of finesse and accuracy, underscoring its unparalleled efficacy in the challenging underwater conditions.
Yujie Li 0001, Yuchao Zheng 0001, Huimin Lu 0001, Jianru Li, Zhengxiang Shen
IEEE Internet Things J.1
2024 Underwater image restoration based on light attenuation prior and color-contrast adaptive correction
Jianru Li, Yuchao Zheng 0001, Huimin Lu 0001, Yujie Li 0001
Image Vis. Comput.5
2024 Fuzzy Multimodal Graph Reasoning for Human-Centric Instructional Video Grounding
abstract
Human-centric instructional videos provide opportunities for users to learn real-world multistep tasks, such as cooking, makeup, and using professional tools. However, these lengthy videos always lead to a tedious learning experience, making it challenging for learners to catch specific guidance efficiently. In this article, we present a novel approach, namedfuzzy multimodal graph reasoning (FMGR), to extract target events in long untrimmed human-centric instructional videos using natural language. Specifically, we devise a fuzzy multimodal graph learning layers in our method, which encompass first contextual graph reasoning that transforms the individual features into contextualized features, second cross-modal relation fuzzifier that models the fine-grained matching relationships between two modalities, and third fuzzy graph reasoning that conducts massage passing among cross-modal matching node pairs. Particularly, we integrate fuzzy theory into the cross-modal relation fuzzifier to amplify potential matching pairs, while simultaneously mitigating the interference from ambiguous matches. To validate our method, we conducted evaluations on two human-centric instructional video datasets, i.e., MedVidQA and YouMakeUp. Moreover, we also take further analysis on the impacts of interrogative and declarative queries. Extensive experimental results and further analysis reveal the effectiveness of our proposed FMGR method.
Yujie Li 0001, Xun Jiang 0001, Xing Xu 0001, Huimin Lu 0001, Heng Tao Shen
IEEE Trans. Fuzzy Syst.1
2024 Brain-Inspired Perception Feature and Cognition Model Applied to Safety Patrol Robot
abstract
To satisfy the development trend of a few workers or unmanned production in industries, especially in dangerous mining industry, the safety patrol robot is required to replace the safety inspectors. The challenge is how to model the cognition mechanism of inspectors for the safety patrol. The specific problems involve the cognition modeling scheme, the aliasing of the commonly used empirical mode decomposition (EMD) of the electroencephalograph (EEG) filtering, the multisource EEG feature vector construction, and the brain-inspired modeling method. To this end, this article focuses on the perception feature and cognition model applied to safety patrol robot in mining industry. First, the inspector's cognition modeling scheme is designed by using brain-computer interface. Second, a filtering algorithm is developed by embedding the sample entropy and independent component into the EMD. Third, a multisource EEG feature vector is fused by using the power spectral density and the EEG map and the functional brain connectivity. Fourth, the cognition model is built by a convolutional neural network embedded the inception module. The experiments indicate that the modeling scheme is effective. The developed filtering algorithm increases the signal-to-noise ratio by 4.16%. The integrated model reaches the average accuracy of 88.17%.
Yujie Li 0001, Mei Wang 0002, Xiaoyan Xie, Wenbin Chai, Xiu Chen
IEEE Trans. Ind. Informatics1
2023 DCEL: Deep Cross-modal Evidential Learning for Text-Based Person Retrieval
abstract
Text-based person retrieval aims at searching for a pedestrian image from multiple candidates with textual descriptions. It is challenging due to uncertain cross-modal alignments caused by the large intra-class variations. To address the challenge, most existing approaches rely on various attention mechanisms and auxiliary information, yet still struggle with the uncertain cross-modal alignments arising from significant intra-class variation, leading to coarse retrieval results. To this end, we propose a novel framework termed Deep Cross-modal Evidential Learning (DCEL), which deploys evidential deep learning to consider the cross-modal alignment uncertainty. Our DCEL model comprises three components: (1) Bidirectional Evidential Learning, which models alignment uncertainty to measure and mitigate the influence of large intra-class variation; (2) Multi-level Semantic Alignment, which leverages a proposed Semantic Filtration module and image-text similarity distribution to facilitate cross-modal alignments; (3) Cross-modal Relation Learning, which reasons about latent correspondences between multi-level tokens of image and text. Finally, we integrate the advantages of the three proposed components to enhance the model to achieve reliable cross-modal alignments. Our DCEL method consistently outperforms more than ten state-of-the-art methods in supervised, weakly supervised, and domain generalization settings on three benchmarks: CUHK-PEDES, ICFG-PEDES, and RSTPReid.
Shenshen Li, Xing Xu 0001, Yang Yang 0002, Fumin Shen, Yijun Mo, Yujie Li 0001, Heng Tao Shen
ACM Multimedia6
2023 Potentially Unwanted App Detection for Blockchain-Based Android App Marketplace
abstract
Android is a mobile operating system with a high degree of openness, which attracts an increasing number of developers. Android application (or simply, app) marketplace provides a trusted source of apps for users and a more equitable competition environment for individual developers and commercial teams. Blockchain’s advantages of decentralization and data immutability are suitable for the Android app marketplace, which is mainly characterized by openness, equality, and security. However, this may also facilitate malicious developers to publish low-quality apps to display ads or steal users’ privacy for revenue. Therefore, blockchain-based app marketplaces have a strong need to identify those potentially unwanted apps (PUAs). In this article, we first introduce our blockchain-based app marketplace model. Then, we propose a new PUA detection method, mainly based on metadata and user ratings, and they are easily accessible from blockchain-based app marketplaces. Moreover, we introduce dynamic analysis to check whether the URLs visited by the app are in malicious URL blacklists since apps with massive access to these URLs tend to affect user experience. After that, we preprocess those complex and redundant features and represent each app as an embedding. Finally, to validate the effectiveness of our method, we utilize several clustering algorithms to represent these apps as clusters and search for suspicious PUA clusters. Our study reveals several characteristics of PUA and suggests that PUAs are still present and need to be urgently removed.
Yuning Cui 0002, Yi Sun 0006, Zhaowen Lin, Baoquan Ma, Yujie Li 0001
IEEE Internet Things J.5
2023 Joint Semantic-Instance Segmentation Method for Intelligent Transportation System
abstract
Getting the point cloud data from sensors and correctly understanding the scene is the core of the intelligent transportation system. Point cloud segmentation can help intelligent transportation systems distinguish different objects in the scene. Some methods process the point cloud through a feature extraction network and complete the segmentation task. However, these methods have high requirements on the feature extraction network, and the fineness of the features will directly affect the final segmentation result. In this paper, we propose a new feature extraction network for segmentation by adding an encoder-decoder structure, which can extract the multiscale local feature information from the feature map. In our opinion, the merged multiscale features obtain a better feature matrix, which improves the performance of the segmentation tasks. We report results on the S3DIS dataset, new feature extraction network greatly improves both semantic segmentation and instance segmentation tasks.
Yujie Li 0001, Jintong Cai, Quan Zhou 0004, Huimin Lu 0001
IEEE Trans. Intell. Transp. Syst.1
2023 Pose Estimation of Point Sets Using Residual MLP in Intelligent Transportation Infrastructure
abstract
6D pose estimation of arbitrary objects is a crucial topic for intelligent transportation infrastructure measurement. However, some external environmental factors and the characteristics of the object itself impact the accuracy of the object’s pose estimation in practical applications. In this paper, we propose a new multi-class dataset ICD-4 (Industrial car Components Dataset) for 6D object pose estimation, which mainly includes four component categories, and every category takes 20,000 different scenarios. ICD-4 dataset delivers quite a few research challenges involving the range of object pose transformations and has significant research value for small-scale pose estimation tasks. We also propose an innovative method PoseMLP, a pose estimation network that uses residual MLP (multilayer perceptron) modules to predict the 6D pose estimation directly. Simultaneously, the experimental results demonstrate the effectiveness and reliability of the proposed method.
Yujie Li 0001, Zhiyun Yin, Yuchao Zheng 0001, Huimin Lu 0001, Tohru Kamiya, Yoshihisa Nakatoh, Seiichi Serikawa
IEEE Trans. Intell. Transp. Syst.1
2023 Learning Latent Dynamics for Autonomous Shape Control of Deformable Object
abstract
In recent years, the methods of loading and transporting rigid objects have become more and more perfect. However, in the process of transportation, the shape control of deformable objects has attracted extensive attention because deformable objects have been widely used in intelligent tasks such as packing and sorting cables before transportation. Restricted by the super-degrees of freedom and nonlinear dynamic models of deformable objects, planning the action trajectories to control the shape of deformable objects is a challenging task. In this work, we use contrastive learning to solve the shape control problem of deformable objects. The method jointly optimizes the visual representation model and dynamic model of deformable objects, maps the target nonlinear state to linear latent space which avoids model inference for deformable objects in infinite-dimensional configuration spaces. Furthermore, to extract effective information in the latent space, we construct an encoder with a multi-branch topology to improve the representation ability of the model. Experimentally, we collect dynamic trajectory data for random shape control task involving cloth or rope in a simulated environment. Then we apply it to train the proposed offline method to obtain latent dynamic models for shape control of deformable objects. In comparison with other baseline methods, our proposed method achieves substantial performance improvements.
Huimin Lu 0001, Yadong Teng, Yujie Li 0001
IEEE Trans. Intell. Transp. Syst.3
2023 Multidimensional Deformable Object Manipulation Based on DN-Transporter Networks
abstract
In the process of transportation, the handling and loading methods of rigid objects are becoming more and more perfect. However, whether in today’s transportation system or in daily life, such as packing objects or sorting cables before transportation, the manipulation of deformable objects has been always inevitable and has attracted more and more attention. Due to the super degrees of freedom and the unpredictable physical state of deformed objects. It is difficult for robots to complete tasks under the environment of the deformable object. Therefore, we present a method based on imitation learning. In the generated expert demonstration, the agent is offered to learn the state sequence, and then imitate the expert’s trajectory sequence which avoid the above-mentioned difficulties. In addition, compared with the baseline method, our proposed DN-Transporter Networks are more competitive in a simulation environment involving cloth, ropes or bags.
Yadong Teng, Huimin Lu 0001, Yujie Li 0001, Tohru Kamiya, Yoshihisa Nakatoh, Seiichi Serikawa, Pengxiang Gao
IEEE Trans. Intell. Transp. Syst.3
2022 Grasp Position Estimation from Depth Image Using Stacked Hourglass Network Structure
abstract
In recent years, robots have been used not only in factories. However, most robots currently used in such places can only perform the actions programmed to perform in a predefined space. For robots to become widespread in the future, not only in factories, distribution warehouses, and other places but also in homes and other environments where robots receive complex commands and their surroundings are constantly being updated, it is necessary to make robots intelligent. Therefore, this study proposed a deep learning grasp position estimation model using depth images to achieve intelligence in pick-and-place. This study used only depth images as the training data to build the deep learning model. Some previous studies have used RGB images and depth images. However, in this study, we used only depth images as training data because we expect the inference to be based on the object's shape, independent of the color information of the object. By performing inference based on the target object's shape, the deep learning model is expected to minimize the need for re-training when the target object package changes in the production line since it is not dependent on the RGB image. In this study, we propose a deep learning model that focuses on the stacked encoder-decoder structure of the Stacked Hourglass Network. We compared the proposed method with the baseline method in the same evaluation metrics and a real robot, which shows higher accuracy than other methods in previous studies.
Keisuke Hamamoto, Huimin Lu 0001, Yujie Li 0001, Tohru Kamiya, Yoshihisa Nakatoh, Seiichi Serikawa
COMPSAC3
2022 Robotic Grasp Detection for Parallel Grippers: A Review
abstract
With the continuous progress of robot grasping technology, the application of robots in industrial applications is promoted. However, reliable grasping of any object is still a difficult problem for robot grasping tasks. In this paper, the parallel grabber is studied as the grabber used in robot grabber detection. The grab detection includes the two-dimensional plane grab method and six-degree-of-freedom grab method, in which the former is constrained to grab from one direction. This paper summarizes the development trend of the two methods and analyzes their advantages and disadvantages.
Zhiyun Yin, Yujie Li 0001, Jintong Cai, Huimin Lu 0001
COMPSAC2
2022 Meta-seg: A survey of meta-learning for image segmentation
Yujie Li 0001, Pengxiang Gao, Yichuan Wang 0001, Seiichi Serikawa
Pattern Recognit.2
2022 Weakly-Supervised Semantic Segmentation Network With Iterative dCRF
abstract
This Autonomous driving methods driven by big data are becoming more and more perfect, but the cost of existing data labeling is too high, so how to reduce or even not label data has attracted more and more attention. Semantic segmentation networks supervised by image-level annotations are all trained using pseudo-labels. Most methods use image classification networks to generate class activation maps (CAMs) and start with CAMs to diffuse features to other parts of the target to obtain pseudo-labels. However, due to its weak supervision information, it is difficult for the existing methods to obtain better results. Therefore, we propose a weakly-supervised semantic segmentation network with iterative dCRF based on graph convolution. Specifically, we use ResNet to generate CAMs and node features and then use graph convolution for feature propagation and merge the low-level and high-level semantic information of the image. Then execute dCRF in an iterative manner, and finally obtain refined pseudo-labels. On the PASCAL VOC 2012 data set, our model achieves an mIoU of 63.5%, which is 0.3% higher than the graph convolutional network method.
Yujie Li 0001, Yun Li 0010
IEEE Trans. Intell. Transp. Syst.1
2022 Improved Point-Voxel Region Convolutional Neural Network: 3D Object Detectors for Autonomous Driving
abstract
Recently, 3D object detection based on deep learning has achieved impressive performance in complex indoor and outdoor scenes. Among the methods, the two-stage detection method performs the best; however, this method still needs improved accuracy and efficiency, especially for small size objects or autonomous driving scenes. In this paper, we propose an improved 3D object detection method based on a two-stage detector called the Improved Point-Voxel Region Convolutional Neural Network (IPV-RCNN). Our proposed method contains online training for data augmentation, upsampling convolution and k-means clustering for the bounding box to achieve 3D detection tasks from raw point clouds. The evaluation results on the KITTI 3D dataset show that the IPV-RCNN achieved a 96% mAP, which is 3% more accurate than the state-of-the-art detectors.
Yujie Li 0001, Shuo Yang 0013, Yuchao Zheng 0001, Huimin Lu 0007
IEEE Trans. Intell. Transp. Syst.1
2022 Semantic-Aligned Attention With Refining Feature Embedding for Few-Shot Image Classification
abstract
Autonomous driving relies on trusty visual recognition of surrounding objects. Few-shot image classification is used in autonomous driving to help recognize objects that are rarely seen. Successful embedding and metric-learning approaches to this task normally learn a feature comparison framework between an unseen image and the labeled images. However, these approaches usually have problems with ambiguous feature embedding because they tend to ignore important local visual and semantic information when extracting intra-class common features from the images. In this paper, we introduce a Semantic-Aligned Attention (SAA) mechanism to refine feature embedding and it can be applied to most of the existing embedding and metric-learning approaches. The mechanism highlights pivotal local visual information with attention mechanism and aligns the attentive map with semantic information to refine the extracted features. Incorporating the proposed mechanism into the prototypical network, evaluation results reveal competitive improvements in both few-shot and zero-shot classification tasks on various benchmark datasets.
Xianda Xu, Xing Xu 0001, Fumin Shen, Yujie Li 0001
IEEE Trans. Intell. Transp. Syst.4
2022 VLD-45: A Big Dataset for Vehicle Logo Recognition and Detection
abstract
Vehicle logo detection (VLD) is a special and significant topic in object detection for vehicle identification system applications. Nevertheless, the range of the research and analysis for VLD are seriously narrow in the real complex scenes, although it’s a critical role in the object detection of small sizes. In this paper, we make further analysis work toward vehicle logo recognition and detection in real-world situations. To begin with, we propose a new multi-class VLD dataset, called VLD-45 (Vehicle Logo Dataset), which contains 45000 images and 50359 objects from 45 categories respectively. Our new dataset provides several research challenges involve in small sizes object, shape deformation, low contrast and so on. Meanwhile, we use 6 existing classifiers and 6 detectors to evaluate our dataset and show the baseline performance. According to the result, our dataset has very significant research value for the task of small-scale object detection. The dataset source:https://github.com/YangShuoys/VLD-45-B-DATASET-Detection
Shuo Yang 0013, Chunjuan Bo, Junxing Zhang, Pengxiang Gao, Yujie Li 0001, Seiichi Serikawa
IEEE Trans. Intell. Transp. Syst.5
2022 Global-PBNet: A Novel Point Cloud Registration for Autonomous Driving
abstract
Registration performs an individual and deciding role in multiple intelligent transport systems. The advancement of deep-learning-based methods enhances the robustness and effectiveness of the preliminary registration stage, although the algorithm will effortlessly fall into local optima when improving the ultimate exactitude. Similarly, traditional method based on optimization has a more reliable performance in terms of precision. However, its performance still counts on the quality of initialization. In order to solve the above problems, we propose a PBNet that combines a point cloud network with a global optimization method. This framework uses the feature information of objects to perform high-precision rough registration and then searches the entire 3D motion space to implement branch-and-bound and iterative nearest point methods. The evaluation results show that PBNet significantly reduce the influence of initial values on registration and has good robustness against noise and outliers.
Yuchao Zheng 0001, Yujie Li 0001, Shuo Yang 0013, Huimin Lu 0001
IEEE Trans. Intell. Transp. Syst.2
2022 Virtual Reality Aided High-Quality 3D Reconstruction by Remote Drones
abstract
Artificial intelligence including deep learning and 3D reconstruction methods is changing the daily life of people. Now, an unmanned aerial vehicle that can move freely in the air and avoid harsh ground conditions has been commonly adopted as a suitable tool for 3D reconstruction. The traditional 3D reconstruction mission based on drones usually consists of two steps: image collection and offline post-processing. But there are two problems: one is the uncertainty of whether all parts of the target object are covered, and another is the tedious post-processing time. Inspired by modern deep learning methods, we build a telexistence drone system with an onboard deep learning computation module and a wireless data transmission module that perform incremental real-time dense reconstruction of urban cities by itself. Two technical contributions are proposed to solve the preceding issues. First, based on the popular depth fusion surface reconstruction framework, we combine it with a visual-inertial odometry estimator that integrates the inertial measurement unit and allows for robust camera tracking as well as high-accuracy online 3D scan. Second, the capability of real-time 3D reconstruction enables a new rendering technique that can visualize the reconstructed geometry of the target as navigation guidance in the HMD. Therefore, it turns the traditional path-planning-based modeling process into an interactive one, leading to a higher level of scan completeness. The experiments in the simulation system and our real prototype demonstrate an improved quality of the 3D model using our artificial intelligence leveraged drone system.
Feng Xu 0005, Chi-Man Pun, Yang Yang 0002, Rushi Lan, Yujie Li 0001, Hao Gao 0005
ACM Trans. Internet Techn.7
2022 Cognitive ocean of things: a comprehensive review and future trends
Yujie Li 0001, Shinya Takahashi, Seiichi Serikawa
Wirel. Networks1
2022 Multi-feature fusion point cloud completion network
Xiu Chen, Yujie Li 0001, Yun Li 0010
World Wide Web2
2021 Underwater image enhancement using improved generative adversarial network
abstract
Summary The generative adversarial network is widely used in image generation, and the generation of images with different styles is applied to underwater image enhancement. The existing underwater image generative adversarial network does not realize color correction when processing underwater images Therefore, we propose an improved generative adversarial network for image color restoration. Firstly, the loss function in the network is improved to train the dataset. Then the improved network is used to detect the underwater image. After network testing, the underwater image is more satisfactory than the traditional image. Numerical results show that this method has a good color restoration and sharpening effects.
Yujie Li 0001, Shinya Takahashi
Concurr. Comput. Pract. Exp.2
2021 RFID Reader Anticollision Based on Distributed Parallel Particle Swarm Optimization
abstract
The deployment of a very large number of readers in a limited space may increase the probability of collision among radio-frequency identification (RFID) readers and reduce the dependability and controllability of Internet-of-Things (IoT) systems. Intelligent computing technologies can be used to realize intelligent management by scheduling resources to circumvent collision issues. In this article, an improved RFID reader anticollision model is constructed by modifying the measure index, introducing a constraint function, and simultaneously considering collisions among readers and between readers and tags. The dense deployment of large numbers of readers increases the number of variables to be encoded, resulting in a high-dimensional problem that cannot be effectively and efficiently solved by traditional algorithms. Accordingly, distributed parallel cooperative co-evolution particle swarm optimization (DPCCPSO) is proposed. The inertia weight and learning factors are adjusted during evolution, and an improved grouping strategy is presented. Moreover, various combinations of random number generation functions are tested. For improved efficiency, DPCCPSO is implemented with distributed parallelism. Experimental verification shows that the proposed novel algorithm exhibits superior performance to existing state-of-the-art algorithms, particularly when numerous RFID readers are deployed.
Bin Cao 0005, Yu Gu 0018, Zhihan Lyu, Jianwei Zhao 0001, Yujie Li 0001
IEEE Internet Things J.6
2021 Confused-Modulo-Projection-Based Somewhat Homomorphic Encryption - Cryptosystem, Library, and Applications on Secure Smart Cities
abstract
With the development of cloud computing, the storage and processing of massive visual media data has gradually transferred to the cloud server. For example, if the intelligent video monitoring system cannot process a large amount of data locally, the data will be uploaded to the cloud. Therefore, how to process data in the cloud without exposing the original data has become an important research topic. We propose a single-server version of somewhat homomorphic encryption cryptosystem based on confused modulo projection theorem named CMP-SWHE, which allows the server to complete blind data processing withoutseeingthe effective information of user data. On the client side, the original data is encrypted by amplification, randomization, and setting confusing redundancy. Operating on the encrypted data on the server side is equivalent to operating on the original data. As an extension, we designed and implemented a blind computing scheme of accelerated version based on batch processing technology to improve efficiency. To make this algorithm easy to use, we also designed and implemented an efficient general blind computing library based on CMP-SWHE. We have applied this library to foreground extraction, optical flow tracking, and object detection with satisfactory results, which are helpful for building smart cities. We also discuss how to extend the algorithm to deep learning applications. Compared with other homomorphic encryption cryptosystems and libraries, the results show that our method has obvious advantages in computing efficiency. Although our algorithm has some tiny errors ($10^{-6}$) when the data is too large, it is very efficient and practical, especially suitable for blind image and video processing.
Xin Jin 0015, Xiaodong Li 0013, Beisheng Liu, Shujiang Xie, Amit Kumar Singh 0001, Yujie Li 0001
IEEE Internet Things J.8
2021 Adaptive Square Attack: Fooling Autonomous Cars With Adversarial Traffic Signs
abstract
To better understand the road condition and make correct driving decisions, traffic sign recognition becomes a crucial component commonly equipped in the vision system of modern autonomous cars. The state-of-the-art traffic sign recognition models are designed with the backbones of deep neural networks (DNNs) since DNNs are powerful to extract more effective visual features that benefit recognition performance. As the recent studies on adversarial attacks have shown that DNNs are easy to be fooled by perturbed images and lead to misclassification, in this article, we explore the vulnerability of the DNN-based traffic sign recognition model. Most existing adversarial attack methods limitedly focus on the white-box attack on the recognition models whose underlying configurations (e.g., network architectures and parameters) are accessible. Differently, we propose a novel attacking method dubbed adaptive square attack (ASA) that can accomplish the black-box attack, i.e., bypassing the access the configurations of the recognition models. Specifically, the proposed ASA method employs an efficient sampling strategy that can generate perturbations for traffic sign images with fewer query times. Extensive experiments on the benchmark data set German traffic sign recognition benchmark with large-scale traffic sign images for autonomous cars show that our proposed ASA method is advanced to perform the black-box attack with high efficiency. Although the generated adversarial traffic sign images by the proposed ASA method are visually similar to the raw images with almost imperceptible differences, they can successfully lead to the misclassification of the state-of-the-art recognition model.
Yujie Li 0001, Xing Xu 0001, Jinhui Xiao, Heng Tao Shen
IEEE Internet Things J.1
2021 Few-shot prototype alignment regularization network for document image layout segementation
Yujie Li 0001, Xing Xu 0001, Yi Lai, Fumin Shen, Lijiang Chen, Pengxiang Gao
Pattern Recognit.1
2021 Deep Fuzzy Hashing Network for Efficient Image Retrieval
abstract
Hashing methods for efficient image retrieval aim at learning hash functions that map similar images to semantically correlated binary codes in the Hamming space with similarity well preserved. The traditional hashing methods usually represent image content by hand-crafted features. Deep hashing methods based on deep neural network (DNN) architectures can generate more effective image features and obtain better retrieval performance. However, the underlying data structure is hardly captured by existing DNN models. Moreover, the similarity (either visually or semantically) between pairwise images is ambiguous, even uncertain, to be measured in the existing deep hashing methods. In this article, we propose a novel hashing method termed deep fuzzy hashing network (DFHN) to overcome the shortcomings of existing deep hashing approaches. Our DFHN method combines the fuzzy logic technique and the DNN to learn more effective binary codes, which can leverage fuzzy rules to model the uncertainties underlying the data. Derived from fuzzy logic theory, the generalized hamming distance is devised in the convolutional layers and fully connected layers in our DFHN to model their outputs, which come from an efficientxoroperation on given inputs and weights. Extensive experiments show that our DFHN method obtains competitive retrieval accuracy with highly efficient training speed compared with several state-of-the-art deep hashing approaches on two large-scale image datasets: CIFAR-10 and NUS-WIDE.
Huimin Lu 0001, Xing Xu 0001, Yujie Li 0001, Heng Tao Shen
IEEE Trans. Fuzzy Syst.4
2021 Adversarial Attack Against Urban Scene Segmentation for Autonomous Vehicles
abstract
Understanding the surrounding environment is crucial for autonomous vehicles to make correct driving decisions. In particular, urban scene segmentation is a significant integral module commonly equipped in the perception system of autonomous vehicles to understand the real scene like a human. Any missegmentation of the driving scenario can potentially result in uncontrollable consequences such as serious accidents or the exception of the perception system. In this article, we investigate the vulnerability of the popular scene segmentation models designed with the backbones of deep neural networks (DNNs), which have been shown to be sensitive to adversarial attacks. Specifically, we propose an iterative projected gradient-based attack method that can effectively fool several DNN-based segmentation models with a remarkably higher attacking successful rate, and much smaller adversarial perturbations. Moreover, we also develop an adversarial training algorithm with min-max optimization style to enrich the robustness of the scene segmentation models. Extensive experiments on the Cityscape benchmark dataset consisting of large-scale urban scene images for autonomous vehicles demonstrate the effectiveness of our proposed attack method, as well as the benefit of the adversarial training scheme for the scene segmentation models.
Xing Xu 0001, Jingran Zhang, Yujie Li 0001, Yichuan Wang 0001, Yang Yang 0002, Heng Tao Shen
IEEE Trans. Ind. Informatics3
2021 User-Oriented Virtual Mobile Network Resource Management for Vehicle Communications
abstract
Currently, advanced communications and networks greatly enhance user experiences and have a major impact on all aspects of people's lifestyles in terms of work, society, and the economy. However improving competitiveness and sustainable vehicle network services, such as higher user experience, considerable resource utilization and effective personalized services, is a great challenge. Addressing these issues, this paper proposes a virtual network resource management based on user behavior to further optimize the existing vehicle communications. In particular, ensemble learning is implemented in the proposed scheme to predict the user's voice call duration and traffic usage for supporting user-centric mobile services optimization. Sufficient experiments show that the proposed scheme can significantly improve the quality of services and experiences and that it provides a novel idea for optimizing vehicle networks.
Huimin Lu 0001, Yin Zhang 0002, Yujie Li 0001, Haider Abbas
IEEE Trans. Intell. Transp. Syst.3
2021 Multi-Aspect Aware Session-Based Recommendation for Intelligent Transportation Services
abstract
In the intelligent transportation system, the session data usually represents the users' demand. However, the traditional approaches only focus on the sequence information or the last item clicked by the user, which cannot fully represent user preferences. To address this issue, this paper proposes an Multi-aspect Aware Session-based Recommendation (MASR) model for intelligent transportation services, which comprehensively considers the user's personalized behavior from multiple aspects. In addition, it developed a concise and efficient transformer-style self-attention to analyze the sequence information of the current session, for accurately grasping the user's intention. Finally, the experimental results show that MASR is available to improve user satisfaction with more accurate and rapid recommendations, and reduce the number of user operations to decrease the safety risk during the transportation service.
Yin Zhang 0002, Yujie Li 0001, Ranran Wang 0001, M. Shamim Hossain, Huimin Lu 0001
IEEE Trans. Intell. Transp. Syst.2
2021 Fast Search of Lightweight Block Cipher Primitives via Swarm-like Metaheuristics for Cyber Security
abstract
With the construction and improvement of 5G infrastructure, more devices choose to access the Internet to achieve some functions. People are paying more attention to information security in the use of network devices. This makes lightweight block ciphers become a hotspot. A lightweight block cipher with superior performance can ensure the security of information while reducing the consumption of device resources. Traditional optimization tools, such as brute force or random search, are often used to solve the design of Symmetric-Key primitives. The metaheuristic algorithm was first used to solve the design of Symmetric-Key primitives of SKINNY. The genetic algorithm and the simulated annealing algorithm are used to increase the number of active S-boxes in SKINNY, thus improving the security of SKINNY. Based on this, to improve search efficiency and optimize search results, we design a novel metaheuristic algorithm, named particle swarm-like normal optimization algorithm (PSNO) to design the Symmetric-Key primitives of SKINNY. With our algorithm, one or better algorithm components can be obtained more quickly. The results in the experiments show that our search results are better than those of the genetic algorithm and the simulated annealing algorithm. The search efficiency is significantly improved. The algorithm we proposed can be generalized to the design of Symmetric-Key primitives of other lightweight block ciphers with clear evaluation indicators, where the corresponding indicators can be used as the objective functions.
Xin Jin 0015, Yuwei Duan, Mengdong Li, Ming Mao, Amit Kumar Singh 0001, Yujie Li 0001
ACM Trans. Internet Techn.8
2021 Merging Grid Maps in Diverse Resolutions by the Context-based Descriptor
abstract
Building an accurate map is essential for autonomous robot navigation in the environment without GPS. Compared with single-robot, the multiple-robot system has much better performance in terms of accuracy, efficiency and robustness for the simultaneous localization and mapping (SLAM). As a critical component of multiple-robot SLAM, the problem of map merging still remains a challenge. To this end, this article casts it into point set registration problem and proposes an effective map merging method based on the context-based descriptors and correspondence expansion. It first extracts interest points from grid maps by the Harris corner detector. By exploiting neighborhood information of interest points, it automatically calculates the maximum response radius as scale information to compute the context-based descriptor, which includes eigenvalues and normals computed from local structures of each interest point. Then, it effectively establishes origin matches with low precision by applying the nearest neighbor search on the context-based descriptor. Further, it designs a scale-based corresponding expansion strategy to expand each origin match into a set of feature matches, where one similarity transformation between two grid maps can be estimated by the Random Sample Consensus algorithm. Subsequently, a measure function formulated from the trimmed mean square error is utilized to confirm the best similarity transformation and accomplish the coarse map merging. Finally, it utilizes the scaling trimmed iterative closest point algorithm to refine initial similarity transformation so as to achieve accurate merging. As the proposed method considers scale information in the context-based descriptor, it is able to merge grid maps in diverse resolutions. Experimental results on real robot datasets demonstrate its superior performance over other related methods on accuracy and robustness.
Zhiyang Lin, Jihua Zhu, Zutao Jiang, Yujie Li 0001, Yaochen Li, Zhongyu Li 0002
ACM Trans. Internet Techn.4
2020 Deep Learning for Visual Segmentation: A Review
abstract
Big data-driven deep learning methods have been widely used in image or video segmentation. The main challenge is that a large amount of labeled data is required in training deep learning models, which is important in real-world applications. To the best of our knowledge, there exist few researches in the deep learning-based visual segmentation. To this end, this paper summarizes the algorithms and current situation of image or video segmentation technologies based on deep learning and point out the future trends. The characteristics of segmentation that based on semi-supervised or unsupervised learning, all of the recent novel methods are summarized in this paper. The principle, advantages and disadvantages of each algorithms are also compared and analyzed.
Yujie Li 0001, Huimin Lu 0001, Tohru Kamiya, Seiichi Serikawa
COMPSAC2
2020 Multi-task reading for intelligent legal services
Yujie Li 0001, Jinyang Du, Haider Abbas, Yin Zhang 0002
Future Gener. Comput. Syst.1
2020 Editorial: Cognitive Science and Artificial Intelligence for Human Cognition and Communication
Huimin Lu 0001, Yujie Li 0001
Mob. Networks Appl.2
2020 Cognitive Computing for Intelligence Systems
Huimin Lu 0001, Yujie Li 0001
Mob. Networks Appl.2
2020 Cognitive computing for intelligent application and service
Yin Zhang 0002, Haider Abbas, Yujie Li 0001
Neural Comput. Appl.3
2019 Touch switch sensor for cognitive body sensor networks
Yujie Li 0001, Huimin Lu 0001, Hyoungseop Kim, Seiichi Serikawa
Comput. Commun.1
2019 Secure face retrieval for group mobile users
Xin Jin 0015, Yujie Li 0001, Shiming Ge, Chenggen Song, Xinghui Zhou
Soft Comput.2
2019 Dilated-aware discriminative correlation filter for visual tracking
Guoxia Xu, Hu Zhu, Lizhen Deng, Lixin Han, Yujie Li 0001, Huimin Lu 0001
World Wide Web5
2018 BrainNets: Human Emotion Recognition Using an Internet of Brian Things Platform
abstract
Human wearable helmet is a useful tool for monitoring the status of miners in the mining industry. However, there is little research regarding human emotion recognition in an extreme environment. In this paper, an emotional state evoked paradigm is designed to identify the brain area where the emotion feature is most evident. Next, the correct electrode position is determined for the collection of the negative emotion by the electroencephalograph (EEG) based on the international 10-20 system of electrode placement. And then, a fusion algorithm of the anxiety level is proposed to evaluate the person's mental state using the θ, α, and β rhythms of an EEG. Experiments demonstrate that the position Fp2 is the best electrode position for obtaining the anxiety level parameter. The most visible EEG changes appear within the first two seconds following stimulation. The amplitudes of the θ rhythm increase most significantly in the negative emotional state.
Huimin Lu 0001, Hyoungseop Kim, Yujie Li 0001, Yin Zhang 0002
IWCMC3
2018 Automatic road detection system for an air-land amphibious car drone
Yujie Li 0001, Huimin Lu 0001, Yoshiki Nakayama, Hyoungseop Kim, Seiichi Serikawa
Future Gener. Comput. Syst.1
2018 Low illumination underwater light field images reconstruction using deep convolutional neural networks
Huimin Lu 0001, Yujie Li 0001, Tomoki Uemura, Hyoungseop Kim, Seiichi Serikawa
Future Gener. Comput. Syst.2
2018 Motor Anomaly Detection for Unmanned Aerial Vehicles Using Reinforcement Learning
abstract
Unmanned aerial vehicles (UAVs) are used in many fields including weather observation, farming, infrastructure inspection, and monitoring of disaster areas. However, the currently available UAVs are prone to crashing. The goal of this paper is the development of an anomaly detection system to prevent the motor of the drone from operating at abnormal temperatures. In this anomaly detection system, the temperature of the motor is recorded using DS18B20 sensors. Then, using reinforcement learning, the motor is judged to be operating abnormally by a Raspberry Pi processing unit. A specially built user interface allows the activity of the Raspberry Pi to be tracked on a Tablet for observation purposes. The proposed system provides the ability to land a drone when the motor temperature exceeds an automatically generated threshold. The experimental results confirm that the proposed system can safely control the drone using information obtained from temperature sensors attached to the motor.
Huimin Lu 0001, Yujie Li 0001, Shenglin Mu, Dong Wang 0004, Hyoungseop Kim, Seiichi Serikawa
IEEE Internet Things J.2
2018 Non-uniform de-Scattering and de-Blurring of Underwater Images
Yujie Li 0001, Huimin Lu 0001, Kuanching Li, Hyoungseop Kim, Seiichi Serikawa
Mob. Networks Appl.1
2018 Extraction of GGO Candidate Regions on Thoracic CT Images using SuperVoxel-Based Graph Cuts for Healthcare Systems
Huimin Lu 0001, Masashi Kondo, Yujie Li 0001, Joo Kooi Tan, Hyoungseop Kim, Seiichi Murakami, Takotoshi Aoki, Shoji Kido
Mob. Networks Appl.3
2018 Brain Intelligence: Go beyond Artificial Intelligence
Huimin Lu 0001, Yujie Li 0001, Min Chen 0003, Hyoungseop Kim, Seiichi Serikawa
Mob. Networks Appl.2
2018 Evaluation of Local Features for Structure from Motion
Mingwei Cao, Wei Jia 0001, Yujie Li 0001, Zhihan Lyu, Liping Zheng, Xiaoping Liu 0003
Multim. Tools Appl.4
2018 Guided local laplacian filter-based image enhancement for deep-sea sensor networks
Jianru Li, Yujie Li 0001
Multim. Tools Appl.2
2018 Active contour model-based segmentation algorithm for medical robots recognition
Yujie Li 0001, Yun Li 0010, Hyoungseop Kim, Seiichi Serikawa
Multim. Tools Appl.1
2018 FDCNet: filtering deep convolutional network for marine organism classification
Huimin Lu 0001, Yujie Li 0001, Tomoki Uemura, ZongYuan Ge, Xing Xu 0001, Li He 0001, Seiichi Serikawa, Hyoungseop Kim
Multim. Tools Appl.2
2018 Single slice based detection for Alzheimer's disease via wavelet entropy and multilayer perceptron trained by biogeography-based optimization
Shuihua Wang, Yin Zhang 0002, Yujie Li 0001, Wen-Juan Jia 0001, Fang-Yuan Liu, Yudong Zhang 0001
Multim. Tools Appl.3
2018 Smart pathological brain detection system by predator-prey particle swarm optimization and single-hidden layer neural-network
Hainan Wang, Yi-Ding Lv, Yujie Li 0001, Yin Zhang 0002, Zhihai Lu
Multim. Tools Appl.4
2018 Erratum to: Smart pathological brain detection system by predator-prey particle swarm optimization and single-hidden layer neural-network
Hainan Wang, Yi-Ding Lv, Yujie Li 0001, Yin Zhang 0002, Zhihai Lu
Multim. Tools Appl.4
2018 A Mobile Computing Method Using CNN and SR for Signature Authentication with Contour Damage and Light Distortion
abstract
A signature is a useful human feature in our society, and determining the genuineness of a signature is very important. A signature image is typically analyzed for its genuineness classification; however, increasing classification accuracy while decreasing computation time is difficult. Many factors affect image quality and the genuineness classification, such as contour damage and light distortion or the classification algorithm. To this end, we propose a mobile computing method of signature image authentication (SIA) with improved recognition accuracy and reduced computation time. We demonstrate theoretically and experimentally that the proposed golden global‐local (G‐L) algorithm has the best filtering result compared with the methods of mean filtering, medium filtering, and Gaussian filtering. The developed minimum probability threshold (MPT) algorithm produces the best segmentation result with minimum error compared with methods of maximum entropy and iterative segmentation. In addition, the designed convolutional neural network (CNN) solves the light distortion problem for detailed frame feature extraction of a signature image. Finally, the proposed SIA algorithm achieves the best signature authentication accuracy compared with CNN and sparse representation, and computation times are competitive. Thus, the proposed SIA algorithm can be easily implemented in a mobile phone.
Mei Wang 0002, Ke Zhai 0004, Chi Harold Liu, Yujie Li 0001
Wirel. Commun. Mob. Comput.4
2017 Wound intensity correction and segmentation with convolutional neural networks
abstract
Summary Wound area changes over multiple weeks are highly predictive of the wound healing process. A big data eHealth system would be very helpful in evaluating these changes. We usually analyze images of the wound bed for diagnosing injury. Unfortunately, accurate measurements of wound region changes from images are difficult. Many factors affect the quality of images, such as intensity inhomogeneity and color distortion. To this end, we propose a fast level set model‐based method for intensity inhomogeneity correction and a spectral properties‐based color correction method to overcome these obstacles. State‐of‐the‐art level set methods can segment objects well. However, such methods are time‐consuming and inefficient. In contrast to conventional approaches, the proposed model integrates a new signed energy force function that can detect contours at weak or blurred edges efficiently. It ensures the smoothness of the level set function and reduces the computational complexity of re‐initialization. To increase the speed of the algorithm further, we also include an additive operator‐splitting algorithm in our fast level set model. In addition, we consider using a camera, lighting, and spectral properties to recover the actual color. Numerical synthetic and real‐world images demonstrate the advantages of the proposed method over state‐of‐the‐art methods. Experimental results also show that the proposed model is at least twice as fast as methods used widely. Copyright © 2016 John Wiley & Sons, Ltd.
Huimin Lu 0001, Bin Li 0006, Junwu Zhu, Yujie Li 0001, Yun Li 0010, Xing Xu 0001, Li He 0001, Xin Li 0034, Jianru Li, Seiichi Serikawa
Concurr. Comput. Pract. Exp.4
2017 Underwater Optical Image Processing: a Comprehensive Review
Huimin Lu 0001, Yujie Li 0001, Yudong Zhang 0001, Min Chen 0003, Seiichi Serikawa, Hyoungseop Kim
Mob. Networks Appl.2
2016 Underwater image descattering and quality assessment
abstract
Vision-based underwater navigation and object detection requires robust computer vision algorithms to operate in turbid water. Many conventional methods aimed at improving visibility in low turbid water. In this paper, we propose a novel contrast enhancement to enhance high turbid underwater images using descattering and color correction. The proposed enhancement method removes the scatter and preserves colors. In addition, as a rule to compare the performance of different image enhancement algorithms, a more comprehensive image quality assessment index Qu is proposed. The index combines the benefits of SSIM index and color distance index. Experimental results show that the proposed approach statistically outperforms state-of-the-art general purpose underwater image contrast enhancement algorithms. The experiment also demonstrated that the proposed method performs well for image classification.
Huimin Lu 0001, Yujie Li 0001, Xing Xu 0001, Li He 0001, Yun Li 0010, Donald G. Dansereau, Seiichi Serikawa
ICIP2
2016 Super Resolving of the Depth Map for 3D Reconstruction of Underwater Terrain Using Kinect
abstract
In recent years, sonar has been widely used for restoring the underwater terrain. Sonar imaging has the benefits such as long-range photographing, robust for turbidity water. However, it is not suitable for short-range imaging. Meanwhile, it also cannot meet the need of mining machine. Therefore, it is important to develop a 3D reconstruction method for short-range imaging. In this paper, we propose a Kinect-based underwater 3D image reconstruction method. To overcome the drawbacks of low accuracy of depth maps, we propose a novel super-resolution (SR) method, which uses the underwater dark channel prior dehazing, weight guided image SR, and inpainting. The proposed method considered the influence of mud sediments in water, it performs better than the traditional methods. The experimental results demonstrated that, after inpainting, dehazing and the super-resolution, it can obtain high accuracy depth maps.
Yu Nakagawa, Keita Kihara, Ryunosuke Tadoh, Seiichi Serikawa, Huimin Lu 0001, Yudong Zhang 0001, Yujie Li 0001
ICPADS7
2016 Underwater image enhancement method using weighted guided trigonometric filtering and artificial light correction
Huimin Lu 0001, Yujie Li 0001, Xing Xu 0001, Jian-Ru Lin, Zhifei Liu, Xin Li 0034, Jianmin Yang, Seiichi Serikawa
J. Vis. Commun. Image Represent.2
2016 Single image dehazing through improved atmospheric light estimation
Huimin Lu 0001, Yujie Li 0001, Shota Nakashima, Seiichi Serikawa
Multim. Tools Appl.2
2015 Single underwater image descattering and color correction
abstract
Absorption, scattering, and color distortion are three major issues in underwater optical imaging. Light rays traveling through water are scattered and absorbed according to their wavelength. Scattering is caused by large suspended particles that degrade optical images captured underwater. Color distortion occurs because different wavelengths are attenuated to different degrees in water; consequently, images of ambient underwater environments are dominated by a bluish tone. In the present paper, we propose a novel underwater imaging model that compensates for the attenuation discrepancy along the propagation path. In addition, we develop a fast weighted guided normalized convolution domain filtering algorithm for enhancing underwater optical images in shallow oceans. The enhanced images are characterized by a reduced noised level, better exposure in dark regions, and improved global contrast, by which the finest details and edges are enhanced significantly.
Huimin Lu 0001, Yujie Li 0001, Seiichi Serikawa
ICASSP2
2015 Underwater Image Devignetting and Colour Correction
Yujie Li 0001, Huimin Lu 0001, Seiichi Serikawa
ICIG (3)1
2015 Real-Time Underwater Image Contrast Enhancement Through Guided Filtering
Huimin Lu 0001, Yujie Li 0001, Xuelong Hu, Seiichi Serikawa
ICIG (3)2
2013 Underwater image enhancement using guided trigonometric bilateral filter and fast automatic color correction
abstract
This paper describes a novel method to enhance underwater optical images by guided trigonometric bilateral filters and color correction. Scattering and color distortion are two major problems of distortion for underwater optical imaging. Scattering is caused by large suspended particles, like fog or turbid water which contains abundant particles. Color distortion corresponds to the varying degrees of attenuation encountered by light traveling in the water with different wavelengths, rendering ambient underwater environments dominated by a bluish tone. Our key contributions are proposed a new underwater model to compensate the attenuation discrepancy along the propagation path, and to propose a fast guided trigonometric bilateral filtering enhancing algorithm and a novel fast automatic color enhancement algorithm. The enhanced images are characterized by reduced noised level, better exposedness of the dark regions, improved global contrast while the finest details and edges are enhance significantly. In addition, our enhancement method is comparable to higher quality than the state-of-the-art methods by assuming in the latest image evaluation systems.
Huimin Lu 0001, Yujie Li 0001, Seiichi Serikawa
ICIP2
2013 Underwater optical image dehazing using guided trigonometric bilateral filtering
abstract
This paper describes a novel method to enhance underwater optical images by dehazing. Scattering and color change are two major problems of distortion for underwater imaging. Scattering is caused by large suspended particles, like fog or turbid water which contains abundant particles, plankton etc. Color change corresponds to the varying degrees of attenuation encountered by light traveling in the water with different wavelengths, rendering ambient underwater environments dominated by a bluish tone. Our key contribution is to propose a fast image and video dehazing algorithm, to compensate the attenuation discrepancy along the propagation path, and to take the influence of the possible presence of an artificial lighting source into consideration. The enhanced images are characterized by reduced noised level, better exposedness of the dark regions, improved global contrast while the finest details and edges are enhance significantly. In addition, our enhancement method is comparable to higher quality than the state-of-the-art methods.
Huimin Lu 0001, Yujie Li 0001, Akira Yamawaki 0002, Seiichi Serikawa
ISCAS2
2013 Cross Depth Image Filter-Based Natural Image Matting
abstract
In this paper we propose a novel explicit image filter called guided depth image filter for natural image matting. Different from the traditional matting model, the guided image filter computes the filtering output by considering the content of a depth image. The guided depth image filter can be used as an edge-preserving smoothing operator like bilateral filter, but has better behaviors near edges. The proposed filter by using nonlocal neighborhoods, and contribute a simple and fast algorithm giving competitive results. Experimental results indicate that our matting results are comparable to the state of the art methods.
Yujie Li 0001, Huimin Lu 0001, Seiichi Serikawa
SNPD1
2013 Multiframe Medical Images Enhancement on Dual Tree Complex Wavelet Transform Domain
abstract
As a novel of multi-resolution analysis tool, dual tree complex wavelet transform (DTCWT) provides flexible multiresolution and directional expansion for medical image fusion. In this paper, a novel fusion method for multiframe medical images based on DTCWT is proposed. Contrary to present fusion methods, the proposed algorithm extends the 2 inputs model to multi inputs model. We take 8 input unclearly medical images as input. Then, we estimate the decomposed coefficients through the weighted soft threshold in each image, and choose the weighted coefficients. After DTCWT reconstruction, the clear image is gotten. During abundant experiments, we evaluate the proposed method both human visual and quantitative analysis. Compare with the-state-of-the-art methods, the new strategy for attaining image fusion with satisfactory performance.
Huimin Lu 0001, Yujie Li 0001, Shota Nakashima
SNPD2
2013 Distance Measurement with a General 3D Camera by Using a Modified Phase Only Correlation Method
abstract
This paper proposed a new approach of 3D measurement using a home use 3D camera. Stereo image measurement is different from the other active measurement method like using a laser range finder or an ultra-sonic sensor. It is a passive method, which acquire the object as two images, and then calculate the distance information form the two images according to the principle of triangulation. Ordinarily, such of two images is taken by a special use stereo camera which were settled with a precisely accuracy so that to keep a parallel optical axes, depth of focus and so on. But the accuracy settings of a home use 3D camera cannot satisfy such a requirement. In this paper, phase-only-correlation method which can yield sub-pixel accuracy is used, also with some modification and new approach. The simulation shows a good result.
Yujie Li 0001, Huimin Lu 0001, Seiichi Serikawa
SNPD2
2012 Multimodal Medical Image Fusion in Modified Sharp Frequency Localized Contourlet Domain
abstract
As a novel of multi-resolution analysis tool, the modified sharp frequency localized contour let transforms (MSFLCT) provides flexible multiresolution, anisotropy, and directional expansion for medical images. In this paper, we proposed a new fusion rule for multimodal medical images based on MSFLCT. The multimodal medical images are decomposed by MSFLCT. For the high-pass sub band, the weighted sum modified laplacian (WSML) method is used for choose the high frequency coefficients. For the low pass sub band, the maximum local energy (MLE) method is combined with "region" idea for low frequency coefficient selection. The final fusion image is obtained by applying inverse MSFLCT to fused low pass and high pass sub bands. Abundant experiments have been made on groups of multimodality datasets, both human visual and quantitative analysis show that the new strategy for attaining image fusion with satisfactory performance.
Seiichi Serikawa, Huimin Lu 0001, Yujie Li 0001
SNPD3
2012 Maximum Local Energy Based Multifocus Image Fusion in Mirror Extended Curvelet Transform Domain
abstract
In this paper, we firstly propose the maximum local energy (MLE) method to calculate the low frequency coefficients of images and compare the results with those of mirror extended curve let transform, which enhance the edge features and details of images. An image fusion step was performed as follows: First, we obtained the coefficients of two different types of images through mirror extended curve let transform. Second, we selected the low frequency coefficients by maximum local energy and obtaining the high-frequency coefficients using the absolute maximum value (AMV) method. Finally, the fused image was obtained by performing an inverse mirror extended curve let transform. In addition to human vision analysis, the images were also compared through quantitative analysis. multifocus images were used in the experiments to compare the results among the beyond wavelets. The numerical experiments reveal that maximum local energy is a new strategy for attaining image fusion with satisfactory performance.
Huimin Lu 0001, Yujie Li 0001, Seiichi Serikawa
SNPD3