Huimin Lu 0001

dblp:64/2633-1 · DBLP profile ↗
← Back
167ranked-venue papers
35as first author
93since 2021 · last 2026
0000-0001-9794-3221ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 52 · 9 first-author · 28 since 2021Computer networks · 38 · 14 first-author · 15 since 2021Artificial intelligence and machine learning · 34 · 3 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 27 · 5 first-author · 22 since 2021Systems, architecture and hardware · 12 · 4 first-author · 5 since 2021Software engineering, systems software and programming languages · 9 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021
YearPublicationVenuePosition
2026 MSTrack: Mamba-based Spatio-Temporal Framework with Linear Complexity for Efficient Single Object Tracking
Huimin Lu 0001
IWCMC2
2026 Recent advances of local mechanisms in vision foundation models: A survey and outlook
Qiangchang Wang, Jing Li 0175, Yilong Yin, Huimin Lu 0001
Comput. Vis. Image Underst.4
2026 MSENet: High efficiency video compression via Multivariate Spatiotemporal Entropy Network
Huimin Lu 0001, Liangfan Shi, Yuchao Zheng 0001, Yujie Li 0001
Image Vis. Comput.1
2026 HyperPoint: Multimodal 3D foundation model in hyperbolic space
Haozhe Cheng, Chaoyi Lu, Zhengqiao Li, Minghong Wu, Huimin Lu 0001, Jihua Zhu
Pattern Recognit.6
2026 Curve3D: Curvature-aware masked autoencoder for self-supervised point cloud understanding
Chaoyi Lu, Haozhe Cheng, Huimin Lu 0001, Jihua Zhu
Pattern Recognit.5
2026 Feature-constrained consistency learning for multi-view online action detection
Huimin Lu 0001
Pattern Recognit.5
2026 IGASA: Integrated Geometry-Aware and Skip-Attention Modules for Enhanced Point Cloud Registration
Jihua Zhu, Wenbiao Yan, Peilin Fan, Huimin Lu 0001
IEEE Trans. Circuits Syst. Video Technol.7
2025 Multiple Rotation Averaging with Constrained Reweighting Deep Matrix Factorization
abstract
Multiple rotation averaging plays a crucial role in computer vision and robotics domains. The conventional optimization-based methods optimize a nonlinear cost function based on certain noise assumptions, while most previous learning-based methods require ground truth labels in the supervised training process. Recognizing the handcrafted noise assumption may not be reasonable in all real-world scenarios, this paper proposes an effective rotation averaging method for mining data patterns in a learning manner while avoiding the requirement of labels. Specifically, we apply deep matrix factorization to directly solve the multiple rotation averaging problem in free linear space. For deep matrix factorization, we design a neural network model, which is explicitly low-rank and symmetric to better suit the background of multiple rotation averaging. Meanwhile, we utilize a spanning tree-based edge filtering to suppress the influence of rotation outliers. What's more, we also adopt a reweighting scheme and dynamic depth selection strategy to further improve the robustness. Our method synthesizes the merit of both optimization-based and learning-based methods. Experimental results on various datasets validate the effectiveness of our proposed method.
Jihua Zhu, Naiwen Hu, Mingchen Zhu, Zhongyu Li 0002, Di Wang 0006, Huimin Lu 0001
ICRA8
2025 Generalizable Zero-Shot Object Pose Estimation for Bin-Picking
abstract
Unordered grasping in industrial robotic manipulation requires precise six-degree-of-freedom (6D) pose estimation. However, existing methods often struggle with unknown objects and require retraining, limiting their practicality. Traditional 3D point-pair feature methods, while training-free, perform poorly with textured symmetric objects. We propose a generalizable approach for zero-shot 6 D pose estimation without retraining. Our method consists of two steps: generating CAD-based templates through real-time rendering for coarse pose estimation, and refining poses using semantic point-pair features aligned with the camera viewpoint. We conducted experiments on seven core datasets from the Benchmark for 6D Object Pose Estimation (BOP) challenge, and the results are publicly available on the BOP website. Integration into a robotic grasping system further highlights its high precision and fast execution, making it ideal for applications such as bin-picking. (GZS6D-BP) https://bop.felk.cvut.cz/leaderboards/.
Zijiang Zhang, Huimin Lu 0001, Jintong Cai, Tohru Kamiya, Seiichi Serikawa
ICRA2
2025 Polygon Mesh Recovery via Segmentation Priors and Neural Radiance Fields
abstract
Neural Radiance Fields (NeRF) has made significant advancements in view synthesis and 3D reconstruction. However, the reconstructed outputs often lack semantic information, limiting its applicability for scene understanding tasks. Additionally, NeRF-based implicit representations focus on novel view synthesis, leading to coarse, non-editable surface meshes. This paper addresses the asset generation task within NeRF, incorporating prior semantic information to guide the reconstruction process. We propose a texture mesh recovery method that combines learned segmentation priors with NeRF. By leveraging the Segmentation Anything Model (SAM) for extracting object masks from RGB images, we ensure semantic consistency. Camera poses, derived from a multi-view reconstruction algorithm, are integrated into a mesh-based NeRF framework to learn explicit scene representations. Using an optimization approach, we disentangle texture and mesh, thereby enabling the creation of editable 3D assets. Experimental results demonstrate the effectiveness of the proposed method across multiple reconstruction datasets.
Jintong Cai, Huimin Lu 0001, Yujie Li 0001
IWCMC2
2025 A Multi-Degree-of-Freedom Wave Energy Harvester for Self-Powered IoT System
abstract
The Internet of Things (IoT) systems are essential for the development of smart cities, enabling efficient management and real-time monitoring of urban water resources. However, IoT systems still face challenges related to power supply in remote or off-grid environments. This study presents the design and development of a hybrid wave energy harvester (H-WEH), which integrates electromagnetic generator (EMG) and triboelectric nanogenerator (TENG) to capture ocean wave energy and provide a reliable power source for IoT systems. The experimental results show that the H-WEH effectively captures wave energy, with stable power output and energy storage capabilities, making it suitable for IoT applications. This study highlights the potential of wave energy as a sustainable solution for powering IoT systems, contributing to the advancement of smart city infrastructure, and providing a new approach to urban water resource monitoring.
Bozhi Ding, Huimin Lu 0001, Yujie Li 0001
IWCMC2
2025 LWD-IUM: A Lightweight Detector for Advancing Robotic Grasp in VR-Based Industrial and Underwater Metaverse
abstract
In the burgeoning field of virtual reality (VR) metaverse, the sophistication of interactions between robotic agents and their environment has become a critical concern. In this work, we present LWD-IUM, a novel light-weight detector designed to enhance robotic grasp capabilities in the VR metaverse. LWD-IUM applies deep learning techniques to discern and navigate the complex VR metaverse environment, aiding robotic agents in the identification and grasping of objects with high precision and efficiency. The algorithm is constructed with an advanced lightweight neural network structure based on self-attention mechanism that ensures optimal balance between computational cost and performance, making it highly suitable for real-time applications in VR. Evaluation on the KITTI 3D dataset demonstrated real-time detection capabilities (24-30 fps) of LWD-IUM, with its mean average precision (mAP) remaining 80% above standard 3D detectors, even with a 50% parameter reduction. In addition, we show that LWD-IUM outperforms existing models for object detection and grasping tasks through the real environment testing on a Baxter dual-arm collaborative robot. By pioneering advancements in robotic grasp in the VR metaverse, LWD-IUM promotes more immersive and realistic interactions, pushing the boundaries of what’s possible in virtual experiences.
Liangfan Shi, Yufeng Gu, Yuchao Zheng 0001, Shintaro Kameda, Huimin Lu 0001
IWCMC5
2025 Delving Into Out-of-Distribution Detection with Medical Vision-Language Models
Lie Ju, Sijin Zhou, Huimin Lu 0001, Zhuoting Zhu, Pearse A. Keane, ZongYuan Ge
MICCAI (5)4
2025 SPADesc: Semantic and parallel attention with feature description
Haijun Meng, Huimin Lu 0001, Bozhi Ding, Qiangchang Wang
Neurocomputing2
2025 Turbid Underwater Image Enhancement With Illumination-Constrained and Structure-Preserved Retinex Model
abstract
Turbid underwater images often suffer from color distortion, contrast degradation, and detail loss. To improve the visual quality of these images, this paper proposes an illumination-constrained, structure-preserved retinex variational model. The proposed approach consists of three main components: a nonlinear model based on the classical retinex theory to represent the multiple adverse deformations of turbid underwater images; an adaptive channel compensation method to correct the color cast; and an illumination-constrained structure-preserved variational retinex model that simultaneously estimates a smooth illumination component and a detail display reflection component and uniformly predicts the noise pattern of preprocessed underwater images. Specifically, an adaptive weight matrix is proposed to reveal the structural details in reflectance. The overall smoothness of illumination is constrain by exponential guided filtering and l1/2 norm. The total intensity of the noise pattern is constrained by l2 norm. To solve the resulting optimization problem, we employ alternating direction minimization of logless transformations of Lagrange multipliers. Extensive experiments demonstrate the effectiveness of the proposed method in improving the quality of turbid underwater images. Beyond subjective visual observations, the method also exhibits competitive performance in objective image quality evaluations.
Shuai Liu 0009, Yuchao Zheng 0001, Jianru Li, Huimin Lu 0001, Zhengxiang Shen, Zhanshan Wang 0002
IEEE Trans. Circuits Syst. Video Technol.4
2025 Underwater Image Enhancement via Wavelet Decomposition Fusion of Advantage Contrast
abstract
Underwater images encounter a range of quality degradation issues caused by the differential scattering and absorption of light in water. To address these challenges, we introduce a WFAC method, a wavelet decomposition fusion method that combines global and local contrast for underwater image enhancement. Specifically, we begin with a color transfer compensation strategy to correct the colors in a degraded underwater image. Subsequently, we utilize the pixel gradient distribution to create a matrix weight map that dynamically adjusts the weight distribution in overly bright or dark areas of the color-corrected image, enhancing its global contrast. Simultaneously, we apply a rapid integration statistical strategy to adaptively fine-tune the local contrast of color-corrected images using the local mean and variance statistics. To combine the strengths of various enhanced images, we implement a wavelet decomposition fusion strategy to break down different scale components of globally and locally contrast-enhanced images and merge the benefits of varying scale images to obtain a high-quality underwater image. Comprehensive experimental assessments across three underwater image datasets demonstrate that our WFAC method efficiently recovers colors and boosts contrast in degraded underwater images. The code is publicly available at:https://www.researchgate.net/publication/386508762_2024WFAC.
Weidong Zhang 0007, Qingmin Liu, Huimin Lu 0001, Jianping Wang 0004, Jing J. Liang
IEEE Trans. Circuits Syst. Video Technol.3
2025 High-Turbidity Underwater Image Enhancement via Turbidity Suppression Fusion
abstract
Underwater operations frequently encounter turbid environments, where light absorption and scattering by suspended particles degrade image quality by causing color distortion, uneven brightness, and blurred details. Clear imaging in such conditions is essential for enhancing the efficiency and effectiveness of underwater tasks, including exploration, marine ecological monitoring, and the preservation of underwater cultural heritage. However, existing underwater image enhancement methods struggle to perform well in turbid waters, especially in highly turbid conditions. In this study, we present an advanced method designed to significantly improve the clarity of images captured in turbid water. We begin by introducing an adaptive color correction algorithm that uses the dominant color channel’s pixel values to adjust and restore the colors of other channels, mitigating color distortion in turbid conditions. Subsequently, we apply adaptive threshold segmentation and turbidity assessment to automatically calibrate histogram equalization, which enhances local contrast and suppresses noise. Finally, we develop a dark channel prior based on turbidity background light estimation, which further improves color restoration and detail recovery. Our proposed method outperforms existing state-of-the-art techniques in color restoration, turbidity removal, and detail enhancement. Experimental results demonstrate that our approach effectively enhances imaging performance in turbid waters, thereby significantly improving the operational efficiency of various underwater applications.
Yuchao Zheng 0001, Huimin Lu 0001, Weidong Zhang 0007, Mohsen Guizani
IEEE Trans. Circuits Syst. Video Technol.2
2025 Joint Objective and Subjective Fuzziness Denoising for Multimodal Sentiment Analysis
abstract
Multimodal Sentiment Analysis (MSA) aims at teaching computers or robotics to understand human sentiment with diverse multimodal signals, including audio, vision, and text. Current MSA approaches primarily concentrate on devising fusion strategies for multimodal signals and trying to learn better multimodal joint representations. However, employing multimodal signals directly is not appropriate since the human psychological states are fuzzy and can not be categorized easily, which undermines the effectiveness of existing methods. In this paper, we regard the natural fuzziness of human sentiments can be observed as two types: objective fuzziness introduced by human expression and subjective fuzziness caused by the complexity of human affection. Based on the assumption, we proposed a novel method termedJoint Objective and Subjective Fuzziness Denoising (JOSFD), which introduced fuzzy logic into the multimodal fusion process and sentiment decision process to overcome the objective and subjective fuzziness. Specifically, our JOSFD method contains two key modules: (1) Modality-Specific Fuzzification Module leveraging uncertainty estimation and fuzzy logic to overcome the influence of objective fuzziness in different modalities in multimodal fusion. (2) Attitude-Intensity Representation Disentangling that learns joint representations for human attitude and sentiment strength separately and further employs fuzzy logic to decide the sentiment analysis results. We evaluate our proposed JOSFD method on three widely used MSA benchmark datasets, CMU-MOSI, CMU-MOSEI, and CH-SIMS. Extensive experiments demonstrate our proposed JOSFD method outperforms recent state-of-the-art methods.
Xun Jiang 0001, Xing Xu 0001, Huimin Lu 0001, Lianghua He, Heng Tao Shen
IEEE Trans. Fuzzy Syst.3
2025 Probabilistic Temporal Masked Attention for Cross-View Online Action Detection
abstract
As a critical task in video sequence classification within computer vision, Online Action Detection (OAD) has garnered significant attention. The sensitivity of mainstream OAD models to varying video viewpoints often hampers their generalization when confronted with unseen sources. To address this limitation, we propose a novel Probabilistic Temporal Masked Attention (PTMA) model, which leverages probabilistic modeling to derive latent compressed representations of video frames in a cross-view setting. The PTMA model incorporates a GRU-based temporal masked attention (TMA) cell, which leverages these representations to effectively query the input video sequence, thereby enhancing information interaction and facilitating autoregressive frame-level video analysis. Additionally, multi-view information can be integrated into the probabilistic modeling to facilitate the extraction of view-invariant features. Experiments conducted under three evaluation protocols—cross-subject (cs), cross-view (cv), and cross-subject-view (csv)—demonstrate that the PTMA achieves state-of-the-art performance on the DAHLIA, IKEA ASM, and Breakfast datasets.
Shicheng Jing, Huimin Lu 0001, Kan-Jian Zhang
IEEE Trans. Multim.4
2025 Attention-Guided Multiscale Interaction Network for Face Super-Resolution
abstract
Recently, CNN and Transformer hybrid networks demonstrated excellent performance in face super-resolution (FSR) tasks. Because of numerous features at different scales in hybrid networks, how to fuse these multiscale features and promote their complementarity is crucial for enhancing FSR. However, existing hybrid network-based FSR methods ignore this, only simply combining the Transformer and CNN. To address this issue, we propose an attention-guided multiscale interaction network (AMINet), which incorporates local and global feature interactions, as well as encoder–decoder phase feature interactions. Specifically, we propose a local and global feature interaction (LGFI) module to promote the fusion of global features and the local features extracted from different receptive fields by our residual depth feature extraction (RDFE) module. Additionally, we propose a selective kernel attention fusion (SKAF) module to adaptively select fusions of different features within the LGFI and encoder–decoder phases. Our above design allows the free flow of multiscale features from within modules and between the encoder and decoder, which can promote the complementarity of different scale features to enhance FSR. Comprehensive experiments confirm that our method consistently performs well with less computational consumption and faster inference.
Xujie Wan, Guangwei Gao, Huimin Lu 0001, Jian Yang 0003, Chia-Wen Lin
IEEE Trans. Syst. Man Cybern. Syst.4
2024 NeRF-based Multi-View Synthesis Techniques: A Survey
abstract
In recent years, with the continuous advancement of deep learning technology, the field of computer graphics has begun to use Neural Radiance Fields (NeRF) technology to solve traditional rendering problems. NeRF technology can output high-quality rendering results and generate three-dimensional scenes that are realistic, making it well-suited for 3D reconstruction, virtual reality, and industrial production. This article summarizes the key algorithms and important work in the NeRF field, classifies and discusses key improved models, briefly introduces the datasets and evaluation indicators in this field, and looks forward to the future development of this field, starting from the principles of NeRF technology.
Jintong Cai, Huimin Lu 0001
IWCMC2
2024 Few-Shot Object Detection Algorithm Based on Geometric Prior and Attention RPN
abstract
Intelligent factories driven by deep vision technology use robotic arms to perform tasks such as picking and assembling in a production environment. In practical applications, with the continuous changes of products on the industrial pipeline, the detection model needs to continuously train new weights to adapt to new application scenarios. It is time-consuming and labor-intensive to manually collect training data when deploying the production line, and it cannot be quickly adapted in industrial scenarios. Therefore, we propose an attention RPN (Region Proposal Network) few-shot object detection algorithm based on geometric prior. The algorithm uses the attention RPN module to strengthen the feature extraction ability of the detection model and uses the virtual simulation software to generate synthetic data similar to the real object geometry as the base class data to train the feature extraction network so that the network obtains the ability to extract geometric features on the base class object. By comparing the learning strategies, only a small number of real data samples are used to train the detection model twice. The experimental results show that the algorithm can detect more objects than the existing few-shot object detection algorithm in the industrial scene with only a small amount of real sample data, and the detection accuracy can reach 97%.
Xiu Chen, Yujie Li 0001, Huimin Lu 0001
IWCMC3
2024 A Review of Multimodal Sentiment Analysis: Modal Fusion and Representation
abstract
With the continuous advancement of multimedia technology, various modalities of data such as video, text, and audio are becoming increasingly abundant. Analyzing the unified emotional expression behind this diverse set of modalities has become a crucial research area. The rapid development of deep learning technology has been applied extensively across various domains. This paper delves into various deep learning-based multimodal sentiment analysis methods, focusing on two aspects: the fusion and representation of different modalities of data. Additionally, it outlines commonly used datasets, identifies potential challenges in the field, and provides a comprehensive review. This review is significant for deepening the understanding of multimodal sentiment analysis, summarizing previous work, and inspiring future research initiatives.
Hanxue Fu, Huimin Lu 0001
IWCMC2
2024 A Survey of Deep Learning Technology in Visual SLAM
abstract
Simultaneous localization and mapping (SLAM) is an indispensable component in robot navigation systems. The precision of SLAM in localization and mapping directly impacts the successful execution of subsequent tasks. Conventional SLAM frameworks are built upon manually designed algorithms that depend on explicit physical models. Nevertheless, in complex and dynamic environments with significant lighting changes and severe object occlusion, these traditional methods often fail to deliver the robustness and precision required for specialized robotic tasks. With the advent of deep learning and hardware computing power, an increasing number of researchers are integrating deep learning with SLAM and employing data-driven approaches to compensate for the limitations of handcrafted algorithms, particularly in establishing accurate models. This paper delves into the various aspects of SLAM where deep learning technology has been applied in recent years, elucidating the core technologies of each method. Additionally, it identifies the current challenges that necessitate resolution and suggests potential future development trends and directions.
Haijun Meng, Huimin Lu 0001
IWCMC2
2024 Comprehensive Review of End-to-End Video Compression
abstract
In recent years, end-to-end video compression has emerged as a new and promising solution for video compression. This paper provides a comprehensive review of the development and current state of end-to-end video encoding and decoding technologies, detailing the fundamental principles of both traditional hybrid encoders and end-to-end encoders. It extensively discusses the evolution from traditional video compression algorithms to innovative methods based on deep learning, including the development of various optimization strategies and key technologies using Deep Neural Networks (DNN). The paper especially focuses on the advancement of end-to-end video compression frameworks such as DVC, DCVC, and Transformer-based models, highlighting their impact on enhancing the efficiency and quality of video compression. Additionally, it summarizes the latest research findings in this field and offers a brief outlook on future directions.
Liangfan Shi, Huimin Lu 0001
IWCMC2
2024 Enhanced Experts with Uncertainty-Aware Routing for Multimodal Sentiment Analysis
abstract
Multimodal sentiment analysis, which has garnered widespread attention in recent years, aims to predict human emotional states using multimodal data. Previous studies have primarily focused on enhancing multimodal fusion and integrating information across different modalities, while overlooking the impact of noisy data on the internal features of each single modality. In this paper, we propose the Enhanced experts with Uncertainty-Aware Routing (EUAR) method to address the influence of noisy data on multimodal sentiment analysis by capturing uncertainty and dynamically altering the network. Specifically, we introduce the Mixture of Experts approach into multimodal sentiment analysis for the first time, leveraging its properties under conditional computation to dynamically alter the network in response to different types of noisy data. Particularly, we refine the experts within the MoE framework to capture uncertainty in the data and extract clearer features. Additionally, a novel routing mechanism is introduced. Through our proposed U-loss, which utilizes the quantified uncertainty by experts, the network learns to route different samples to experts with lower uncertainty for processing, thus obtaining clearer, noise-free features. Experimental results demonstrate that our method achieves state-of-the-art performance on three widely used multimodal sentiment analysis datasets. Moreover, experiments on noisy datasets show that our approach outperforms existing methods in handling noisy data.
Zixian Gao, Disen Hu, Xun Jiang 0001, Huimin Lu 0001, Heng Tao Shen, Xing Xu 0001
ACM Multimedia4
2024 A multi-attention and depthwise separable convolution network for medical image segmentation
abstract
Automatic medical image segmentation method is highly needed to help experts in lesion segmentation. The deep learning technology emerging has profoundly driven the development of medical image segmentation. While U-Net and attention mechanisms are widely utilized in this field, the application of attention, albeit successful in natural scene image segmentation, tends to inflate the number of model parameters and neglects the potential for feature fusion between different convolutional layers. In response to these challenges, we present the Multi-Attention and Depthwise Separable Convolution U-Net (MDSU-Net), designed to enhance feature extraction. The multi-attention aspect of our framework integrates dual attention and attention gates, adeptly capturing rich contextual details and seamlessly fusing features across diverse convolutional layers. Additionally, our encoder integrates a depthwise separable convolution layer, streamlining the model’s complexity without sacrificing its efficacy, ensuring versatility across various segmentation tasks. The results demonstrate that our method outperforms state-of-the-art across three diverse medical image datasets.
Fuji Ren, Huimin Lu 0001, Satoshi Nakagawa, Xiao Shan
Neurocomputing4
2024 Underwater Visibility Enhancement IoT System in Extreme Environment
abstract
Imagery captured in extreme underwater environments often presents unique challenges, including blurred details, color distortion, and reduced contrast. These discrepancies largely emanate from the intricate interplay of light absorption and scattering within the aquatic medium. Predominant restoration techniques, rather simplistically, apply a static attenuation coefficient, neglecting the dynamic nuances of underwater conditions, leading to an inconsistent restoration outcome. To counter these impediments, we introduce an avant-garde Underwater Internet of Things (Underwater IoT) system, underpinned by a scene-depth fusion paradigm. Our methodology astutely accounts for the spectral decay of light underwater to infer a more refined attenuation coefficient tailored to the specific scene. This system, employing a quadtree decomposition for precise localization coupled with depth mapping, facilitates an astute estimation of prevailing luminescence. This depth map, once synthesized and refined, aids in gauging the precise attenuation dynamics of the aqueous milieu, culminating in a more precise transmission map derivation. Segueing from this, we employ an inverse model to refurbish the original image. Experimental results highlight our system’s prowess in counteracting issues like muddied details and chromatic anomalies while concurrently amplifying contrast. In juxtaposition with a spectrum of existing methodologies, our innovation outshines in terms of finesse and accuracy, underscoring its unparalleled efficacy in the challenging underwater conditions.
Yujie Li 0001, Yuchao Zheng 0001, Huimin Lu 0001, Jianru Li, Zhengxiang Shen
IEEE Internet Things J.4
2024 Underwater image restoration based on light attenuation prior and color-contrast adaptive correction
Jianru Li, Yuchao Zheng 0001, Huimin Lu 0001, Yujie Li 0001
Image Vis. Comput.4
2024 A Parkinson's Auxiliary Diagnosis Algorithm Based on a Hyperparameter Optimization Method of Deep Learning
abstract
Parkinson's disease is a common mental disease in the world, especially in the middle-aged and elderly groups. Today, clinical diagnosis is the main diagnostic method of Parkinson's disease, but the diagnosis results are not ideal, especially in the early stage of the disease. In this paper, a Parkinson's auxiliary diagnosis algorithm based on a hyperparameter optimization method of deep learning is proposed for the Parkinson's diagnosis. The diagnosis system uses ResNet50 to achieve feature extraction and Parkinson's classification, mainly including speech signal processing part, algorithm improvement part based on Artificial Bee Colony algorithm (ABC) and optimizing the hyperparameters of ResNet50 part. The improved algorithm is called Gbest Dimension Artificial Bee Colony algorithm (GDABC), proposing "Range pruning strategy" which aims at narrowing the scope of search and "Dimension adjustment strategy" which is to adjust gbest dimension by dimension. The accuracy of the diagnosis system in the verification set of Mobile Device Voice Recordings at King's College London (MDVR-CKL) dataset can reach more than 96%. Compared with current Parkinson's sound diagnosis methods and other optimization algorithms, our auxiliary diagnosis system shows better classification performance on the dataset within limited time and resources.
Shujuan Li, Chi-Man Pun, Yijing Guo, Feng Xu 0005, Hao Gao 0005, Huimin Lu 0001
IEEE Trans. Comput. Biol. Bioinform.7
2024 Fuzzy Multimodal Graph Reasoning for Human-Centric Instructional Video Grounding
abstract
Human-centric instructional videos provide opportunities for users to learn real-world multistep tasks, such as cooking, makeup, and using professional tools. However, these lengthy videos always lead to a tedious learning experience, making it challenging for learners to catch specific guidance efficiently. In this article, we present a novel approach, namedfuzzy multimodal graph reasoning (FMGR), to extract target events in long untrimmed human-centric instructional videos using natural language. Specifically, we devise a fuzzy multimodal graph learning layers in our method, which encompass first contextual graph reasoning that transforms the individual features into contextualized features, second cross-modal relation fuzzifier that models the fine-grained matching relationships between two modalities, and third fuzzy graph reasoning that conducts massage passing among cross-modal matching node pairs. Particularly, we integrate fuzzy theory into the cross-modal relation fuzzifier to amplify potential matching pairs, while simultaneously mitigating the interference from ambiguous matches. To validate our method, we conducted evaluations on two human-centric instructional video datasets, i.e., MedVidQA and YouMakeUp. Moreover, we also take further analysis on the impacts of interrogative and declarative queries. Extensive experimental results and further analysis reveal the effectiveness of our proposed FMGR method.
Yujie Li 0001, Xun Jiang 0001, Xing Xu 0001, Huimin Lu 0001, Heng Tao Shen
IEEE Trans. Fuzzy Syst.4
2024 Boundary-Guided Lightweight Semantic Segmentation With Multi-Scale Semantic Context
abstract
Lightweight semantic segmentation plays an essential role in image signal processing that is beneficial to many multimedia applications, such as self-driving, robotic vision, and virtual reality. Due to the powerful capability to encode image details and semantics, many lightweight dual-resolution networks have been proposed in recent years for semantic segmentation. In spite of achieving remarkable progresses, they often ignore semantic context ranged from different scales. Furthermore, most of them always neglect the object boundaries, serving as a significant assistance for lightweight semantic segmentation. To alleviate these problems, this paper develops a Boundary-guide dual-resolution lightweight network with multi-scale Semantic Context, called BSCNet, for semantic segmentation. Specifically, to enhance the capability of feature representation, an Extremely Lightweight Pyramid Pooling Module (ELPPM) is designed to capture multi-scale semantic context at the top of low-resolution branch of BSCNet. In addition, to increase feature similarity of the same object while keeping feature discrimination of different objects, pixel information is propagated throughout the entire object area using a simple Boundary Auxiliary Fusion Module (BAFM), where the predicted object boundaries are served as high-level guidance to refine low-level convolutional features. The comprehensive experimental results have demonstrated that our BSCNet is simple and effective, achieving state-of-the-art trade-off in terms of segmentation accuracy and running efficiency on CityScapes, CamVid, and KITTI datasets.
Quan Zhou 0004, Guangwei Gao, Bin Kang, Weihua Ou, Huimin Lu 0001
IEEE Trans. Multim.6
2024 Deep-sea visual dataset of the South China sea
Jianru Li, Huimin Lu 0001
Wirel. Networks3
2023 Multiscale Shared Learning for Fault Diagnosis of Rotating Machinery in Transportation Infrastructures
abstract
Rotating machinery is ubiquitous, and its failures constitute a major cause of the failures of transportation infrastructures. Most fault-diagnosis methods for rotating machinery are based on vibration-signal analysis because vibrations directly reflect the transient regime of machinery elements. This article proposes a novel multiscale shared-learning network (MSSLN) architecture to extract and classify the fault features inherent to multiscale factors of vibration signals. The architecture fuses layer-wise activations with multiscale flows, to enable the network to fully learn the shared representation with consistency across multiscale factors. This characteristic helps MSSLN provide more faithful diagnoses than existing single- and multiscale methods. Experiments on bearing and gearbox datasets are used to evaluate the fault-diagnosis performance of transportation infrastructures. Extensive experimental results and comprehensive analyses demonstrate the superiority of the proposed MSSLN in fault diagnosis for bearings and gearboxes, the two foundational elements in transportation infrastructures.
Zhe Chen 0004, Shiqing Tian, Xiaotao Shi, Huimin Lu 0001
IEEE Trans. Ind. Informatics4
2023 Edge Computing With Complementary Capsule Networks for Mental State Detection in Underground Mining Industry
abstract
Most safety accidents are caused by human factor in underground resource mining industry. This is because the nonuniform lighted and noisy and dangerous environment easily evokes the negative mental state and causes the nonstandard production operation. Aiming at the difficult problem to be solved urgently, this article proposes an edge computing mental state framework of the Internet of Things in the underground mining industry. Moreover, a filtering algorithm using a defined threshold function is developed. Furthermore, an complemented capsule network model is constructed by using two residual modules. Specially, a two-stage mental state fusion algorithm is proposed with electrocardiogram signals and facial expression. Finally, the mental state variation characteristics are explored with the underground illuminating and coloring. Experiments show that the mental state detection accuracy is increased by 2.6%. The higher mental arousal is at the illumination between 320Lxand 330Lx.
Mei Wang 0002, Yuancheng Li 0006, Huimin Lu 0001
IEEE Trans. Ind. Informatics4
2023 Joint Semantic-Instance Segmentation Method for Intelligent Transportation System
abstract
Getting the point cloud data from sensors and correctly understanding the scene is the core of the intelligent transportation system. Point cloud segmentation can help intelligent transportation systems distinguish different objects in the scene. Some methods process the point cloud through a feature extraction network and complete the segmentation task. However, these methods have high requirements on the feature extraction network, and the fineness of the features will directly affect the final segmentation result. In this paper, we propose a new feature extraction network for segmentation by adding an encoder-decoder structure, which can extract the multiscale local feature information from the feature map. In our opinion, the merged multiscale features obtain a better feature matrix, which improves the performance of the segmentation tasks. We report results on the S3DIS dataset, new feature extraction network greatly improves both semantic segmentation and instance segmentation tasks.
Yujie Li 0001, Jintong Cai, Quan Zhou 0004, Huimin Lu 0001
IEEE Trans. Intell. Transp. Syst.4
2023 Pose Estimation of Point Sets Using Residual MLP in Intelligent Transportation Infrastructure
abstract
6D pose estimation of arbitrary objects is a crucial topic for intelligent transportation infrastructure measurement. However, some external environmental factors and the characteristics of the object itself impact the accuracy of the object’s pose estimation in practical applications. In this paper, we propose a new multi-class dataset ICD-4 (Industrial car Components Dataset) for 6D object pose estimation, which mainly includes four component categories, and every category takes 20,000 different scenarios. ICD-4 dataset delivers quite a few research challenges involving the range of object pose transformations and has significant research value for small-scale pose estimation tasks. We also propose an innovative method PoseMLP, a pose estimation network that uses residual MLP (multilayer perceptron) modules to predict the 6D pose estimation directly. Simultaneously, the experimental results demonstrate the effectiveness and reliability of the proposed method.
Yujie Li 0001, Zhiyun Yin, Yuchao Zheng 0001, Huimin Lu 0001, Tohru Kamiya, Yoshihisa Nakatoh, Seiichi Serikawa
IEEE Trans. Intell. Transp. Syst.4
2023 Learning Latent Dynamics for Autonomous Shape Control of Deformable Object
abstract
In recent years, the methods of loading and transporting rigid objects have become more and more perfect. However, in the process of transportation, the shape control of deformable objects has attracted extensive attention because deformable objects have been widely used in intelligent tasks such as packing and sorting cables before transportation. Restricted by the super-degrees of freedom and nonlinear dynamic models of deformable objects, planning the action trajectories to control the shape of deformable objects is a challenging task. In this work, we use contrastive learning to solve the shape control problem of deformable objects. The method jointly optimizes the visual representation model and dynamic model of deformable objects, maps the target nonlinear state to linear latent space which avoids model inference for deformable objects in infinite-dimensional configuration spaces. Furthermore, to extract effective information in the latent space, we construct an encoder with a multi-branch topology to improve the representation ability of the model. Experimentally, we collect dynamic trajectory data for random shape control task involving cloth or rope in a simulated environment. Then we apply it to train the proposed offline method to obtain latent dynamic models for shape control of deformable objects. In comparison with other baseline methods, our proposed method achieves substantial performance improvements.
Huimin Lu 0001, Yadong Teng, Yujie Li 0001
IEEE Trans. Intell. Transp. Syst.1
2023 Multidimensional Deformable Object Manipulation Based on DN-Transporter Networks
abstract
In the process of transportation, the handling and loading methods of rigid objects are becoming more and more perfect. However, whether in today’s transportation system or in daily life, such as packing objects or sorting cables before transportation, the manipulation of deformable objects has been always inevitable and has attracted more and more attention. Due to the super degrees of freedom and the unpredictable physical state of deformed objects. It is difficult for robots to complete tasks under the environment of the deformable object. Therefore, we present a method based on imitation learning. In the generated expert demonstration, the agent is offered to learn the state sequence, and then imitate the expert’s trajectory sequence which avoid the above-mentioned difficulties. In addition, compared with the baseline method, our proposed DN-Transporter Networks are more competitive in a simulation environment involving cloth, ropes or bags.
Yadong Teng, Huimin Lu 0001, Yujie Li 0001, Tohru Kamiya, Yoshihisa Nakatoh, Seiichi Serikawa, Pengxiang Gao
IEEE Trans. Intell. Transp. Syst.2
2023 Lightweight Real-Time Semantic Segmentation Network With Efficient Transformer and CNN
abstract
In the past decade, convolutional neural networks (CNNs) have shown prominence for semantic segmentation. Although CNN models have very impressive performance, the ability to capture global representation is still insufficient, which results in suboptimal results. Recently, Transformer achieved huge success in NLP tasks, demonstrating its advantages in modeling long-range dependency. Recently, Transformer has also attracted tremendous attention from computer vision researchers who reformulate the image processing tasks as a sequence-to-sequence prediction but resulted in deteriorating local feature details. In this work, we propose a lightweight real-time semantic segmentation network called LETNet. LETNet combines a U-shaped CNN with Transformer effectively in a capsule embedding style to compensate for respective deficiencies. Meanwhile, the elaborately designed Lightweight Dilated Bottleneck (LDB) module and Feature Enhancement (FE) module cultivate a positive impact on training from scratch simultaneously. Extensive experiments performed on challenging datasets demonstrate that LETNet achieves superior performances in accuracy and efficiency balance. Specifically, It only contains 0.95M parameters and 13.6G FLOPs but yields 72.8% mIoU at 120 FPS on the Cityscapes test set and 70.5% mIoU at 250 FPS on the CamVid test dataset using a single RTX 3090 GPU. Source code will be available athttps://github.com/IVIPLab/LETNet.
Guoan Xu, Juncheng Li 0003, Guangwei Gao, Huimin Lu 0001, Jian Yang 0003, Dong Yue 0001
IEEE Trans. Intell. Transp. Syst.4
2023 Multifeature Fusion-Based Object Detection for Intelligent Transportation Systems
abstract
The detection of 3D objects with high precision from point cloud data has become a crucial research topic in intelligent transportation systems. By effectively modeling global and local features, it can be acquired the state-of-the-art detector for 3D object detection. Nevertheless, regarding the previous work on feature representations, volumetric generation or point learning methods have difficulty building the relationships between local features and global features. Thus, we propose a multi-feature fusion network (MFFNet) to improve detection precision for 3D point cloud data by combining the global features from 3D voxel convolutions with the local features from the point learning network. Our algorithm is an end-to-end detection framework that contains a voxel convolutional module, a local point feature module and a detection head. Significantly, MFFNet constructs the local point feature set with point learning and sampling and the global feature map through 3D voxel convolution from raw point clouds. The detection head can use the obtained fusion feature to predict the position and category of the examined 3D object, so the proposed method can obtain higher precision than existing approaches. An experimental evaluation on the KITTI 3D object detection dataset obtain 97% MAP (Mean Average Precision) and Waymo Open dataset obtain 80% MAP, which proves the efficiency of the developed feature fusion representation method for 3D objects, and it can achieve satisfactory location accuracy.
Shuo Yang 0013, Huimin Lu 0001, Jianru Li
IEEE Trans. Intell. Transp. Syst.2
2023 Generalized Label Enhancement With Sample Correlations
abstract
Recently, label distribution learning (LDL) has drawn much attention in machine learning, where LDL model is learned from labelel instances. Different from single-label and multi-label annotations, label distributions describe the instance by multiple labels with different intensities and accommodate to more general scenes. Since most existing machine learning datasets merely provide logical labels, label distributions are unavailable in many real-world applications. To handle this problem, we propose two novel label enhancement methods, i.e., Label Enhancement with Sample Correlations (LESC) and generalized Label Enhancement with Sample Correlations (gLESC). More specifically, LESC employs a low-rank representation of samples in the feature space, and gLESC leverages a tensor multi-rank minimization to further investigate the sample correlations in both the feature space and label space. Benefitting from the sample correlations, the proposed methods can boost the performance of label enhancement. Extensive experiments on 14 benchmark datasets demonstrate the effectiveness and superiority of our methods.
Qinghai Zheng, Jihua Zhu, Haoyu Tang 0002, Xinyuan Liu 0001, Zhongyu Li 0002, Huimin Lu 0001
IEEE Trans. Knowl. Data Eng.6
2023 Context-Patch Representation Learning With Adaptive Neighbor Embedding for Robust Face Image Super-Resolution
abstract
Representation learning steered robust face image super-resolution (FSR) methods have attracted extensive attention in the past few decades. Most previous methods were devoted to exploiting the local position patches in the training set for FSR. However, they usually overlooked the sufficient usage of the contextual information around the testing patches, which are useful for stable representation learning. In this article, we attempt to utilize the context-patch around the testing patch and propose a method named context-patch representation learning with adaptive neighbor embedding (CRL-ANE) for FSR. On one hand, we simultaneously use the testing position patch and its adjacent ones for stable representation weight learning. This contextual information can compensate for recovering missing details in the target patch. On the other hand, for each input patch set, due to its inherent facial structural properties, we design an adaptive neighbor embedding strategy to elaborately and adaptively choose primary candidates for more accurate reconstruction. These two improvements enable the proposed method to achieve better SR performance than some of the other methods. Qualitative and quantitative experiments on some benchmarks have validated the superiority of the proposed method over some state-of-the-art methods.
Guangwei Gao, Yi Yu 0001, Huimin Lu 0001, Jian Yang 0003, Dong Yue 0001
IEEE Trans. Multim.3
2023 JDSR-GAN: Constructing an Efficient Joint Learning Network for Masked Face Super-Resolution
abstract
With the growing importance of preventing the COVID-19 virus in cyber-manufacturing security, face images obtained in most video surveillance scenarios are usually low resolution together with mask occlusion. However, most of the previous face super-resolution solutions can not efficiently handle both tasks in one model. In this work, we consider both tasks simultaneously and construct an efficient joint learning network, called JDSR-GAN, for masked face super-resolution tasks. Given a low-quality face image with mask as input, the role of the generator composed of a denoising module and super-resolution module is to acquire a high-quality high-resolution face image. The discriminator utilizes some carefully designed loss functions to ensure the quality of the recovered face images. Moreover, we incorporate the identity information and attention mechanism into our network for feasible correlated feature expression and informative feature learning. By jointly performing denoising and face super-resolution, the two tasks can complement each other and attain promising performance. Extensive qualitative and quantitative results show the superiority of our proposed JDSR-GAN over some competitive methods.
Guangwei Gao, Fei Wu 0004, Huimin Lu 0001, Jian Yang 0003
IEEE Trans. Multim.4
2023 FBSNet: A Fast Bilateral Symmetrical Network for Real-Time Semantic Segmentation
abstract
Real-time semantic segmentation, which can be visually understood as the pixel-level classification task on the input image, currently has broad application prospects, especially in the fast-developing fields of autonomous driving and drone navigation. However, the huge burden of calculation together with redundant parameters are still the obstacles to its technological development. In this article, we propose a Fast Bilateral Symmetrical Network (FBSNet) to alleviate the above challenges. Specifically, FBSNet employs a symmetrical encoder-decoder structure with two branches, semantic information branch and spatial detail branch. The Semantic Information Branch (SIB) is the main branch with semantic architecture to acquire the contextual information of the input image and meanwhile acquire sufficient receptive field. While the Spatial Detail Branch (SDB) is a shallow and simple network used to establish local dependencies of each pixel for preserving details, which is essential for restoring the original resolution during the decoding phase. Meanwhile, a Feature Aggregation Module (FAM) is designed to effectively combine the output of these two branches. Experimental results of Cityscapes and CamVid show that the proposed FBSNet can strike a good balance between accuracy and efficiency. Specifically, it obtains 70.9% and 68.9% mIoU along with the inference speed of 90 fps and 120 fps on these two test datasets, respectively, with only 0.62 million parameters on a single RTX 2080Ti GPU. The code is available athttps://github.com/IVIPLab/FBSNet.
Guangwei Gao, Guoan Xu, Juncheng Li 0003, Yi Yu 0001, Huimin Lu 0001, Jian Yang 0003
IEEE Trans. Multim.5
2023 Depth-Distilled Multi-Focus Image Fusion
abstract
Homogeneous regions, which are smooth areas that lack blur clues to discriminate if they are focused or non-focused. Therefore, they bring a great challenge to achieve high accurate multi-focus image fusion (MFIF). Fortunately, we observe that depth maps are highly related to focus and defocus, containing a preponderance of discriminative power to locate homogeneous regions. This offers the potential to provide additional depth cues to assist MFIF task. Taking depth cues into consideration, in this paper, we propose a new depth-distilled multi-focus image fusion framework, namely D2MFIF. In D2MFIF, depth-distilled model (DDM) is designed for adaptively transferring the depth knowledge into MFIF task, gradually improving MFIF performance. Moreover, multi-level fusion mechanism is designed to integrate multi-level decision maps from intermediate outputs for improving the final prediction. Visually and quantitatively experimental results demonstrate the superiority of our method over several state-of-the-art methods.
Fan Zhao 0005, Wenda Zhao 0003, Huimin Lu 0001, Yong Liu 0017, Libo Yao, Yu Liu 0005
IEEE Trans. Multim.3
2023 Lightweight Feature De-redundancy and Self-calibration Network for Efficient Image Super-resolution
abstract
In recent years, thanks to the inherent powerful feature representation and learning abilities of the convolutional neural network (CNN), deep CNN-steered single image super-resolution approaches have achieved remarkable performance improvements. However, these methods are often accompanied by large consumption of computing and memory resources, which is difficult to be adopted in real-world application scenes. To handle this issue, we design an efficient Feature De-redundancy and Self-calibration Super-resolution network (FDSCSR). In particular, a Feature De-redundancy and Self-calibration Block (FDSCB) is proposed to reduce the repetitive feature information extracted by the model and further enhance the efficiency of the model. Then, based on FDSCB, a Local Feature Fusion Module is presented to elaborately utilize and fuse the feature information extracted by each FDSCB. Abundant experiments on benchmarks have demonstrated that our FDSCSR achieves superior performance with relatively less computational consumption and storage resource than other state-of-the-art approaches. The code is available at https://github.com/IVIPLab/FDSCSR .
Zhengxue Wang, Guangwei Gao, Juncheng Li 0003, Huimin Lu 0001
ACM Trans. Multim. Comput. Commun. Appl.6
2022 Feature Distillation Interaction Weighting Network for Lightweight Image Super-resolution
abstract
Convolutional neural networks based single-image superresolution (SISR) has made great progress in recent years. However, it is difficult to apply these methods to real-world scenarios due to the computational and memory cost. Meanwhile, how to take full advantage of the intermediate features under the constraints of limited parameters and calculations is also a huge challenge. To alleviate these issues, we propose a lightweight yet efficient Feature Distillation Interaction Weighted Network (FDIWN). Specifically, FDIWN utilizes a series of specially designed Feature Shuffle Weighted Groups (FSWG) as the backbone, and several novel mutual Wide-residual Distillation Interaction Blocks (WDIB) form an FSWG. In addition, Wide Identical Residual Weighting (WIRW) units and Wide Convolutional Residual Weighting (WCRW) units are introduced into WDIB for better feature distillation. Moreover, a Wide-Residual Distillation Connection (WRDC) framework and a Self-Calibration Fusion (SCF) unit are proposed to interact features with different scales more flexibly and efficiently. Extensive experiments show that our FDIWN is superior to other models to strike a good balance between model performance and efficiency. The code is available at https://github.com/IVIPLab/FDIWN.
Guangwei Gao, Juncheng Li 0003, Fei Wu 0004, Huimin Lu 0001, Yi Yu 0001
AAAI5
2022 Grasp Position Estimation from Depth Image Using Stacked Hourglass Network Structure
abstract
In recent years, robots have been used not only in factories. However, most robots currently used in such places can only perform the actions programmed to perform in a predefined space. For robots to become widespread in the future, not only in factories, distribution warehouses, and other places but also in homes and other environments where robots receive complex commands and their surroundings are constantly being updated, it is necessary to make robots intelligent. Therefore, this study proposed a deep learning grasp position estimation model using depth images to achieve intelligence in pick-and-place. This study used only depth images as the training data to build the deep learning model. Some previous studies have used RGB images and depth images. However, in this study, we used only depth images as training data because we expect the inference to be based on the object's shape, independent of the color information of the object. By performing inference based on the target object's shape, the deep learning model is expected to minimize the need for re-training when the target object package changes in the production line since it is not dependent on the RGB image. In this study, we propose a deep learning model that focuses on the stacked encoder-decoder structure of the Stacked Hourglass Network. We compared the proposed method with the baseline method in the same evaluation metrics and a real robot, which shows higher accuracy than other methods in previous studies.
Keisuke Hamamoto, Huimin Lu 0001, Yujie Li 0001, Tohru Kamiya, Yoshihisa Nakatoh, Seiichi Serikawa
COMPSAC2
2022 Robotic Grasp Detection for Parallel Grippers: A Review
abstract
With the continuous progress of robot grasping technology, the application of robots in industrial applications is promoted. However, reliable grasping of any object is still a difficult problem for robot grasping tasks. In this paper, the parallel grabber is studied as the grabber used in robot grabber detection. The grab detection includes the two-dimensional plane grab method and six-degree-of-freedom grab method, in which the former is constrained to grab from one direction. This paper summarizes the development trend of the two methods and analyzes their advantages and disadvantages.
Zhiyun Yin, Yujie Li 0001, Jintong Cai, Huimin Lu 0001
COMPSAC4
2022 DHHN: Dual Hierarchical Hybrid Network for Weakly-Supervised Audio-Visual Video Parsing
abstract
The Weakly-Supervised Audio-Visual Video Parsing (AVVP) task aims to parse a video into temporal segments and predict their event categories in terms of modalities, labeling them as either audible, visible, or both. Since the temporal boundaries and modalities annotations are not provided, only video-level event labels are available, this task is more challenging than conventional video understanding tasks.Most previous works attempt to analyze videos by jointly modeling the audio and video data and then learning information from the segment-level features with fixed lengths. However, such a design exist two defects: 1) The various semantic information hidden in temporal lengths is neglected, which may lead the models to learn incorrect information; 2) Due to the joint context modeling, the unique features of different modalities are not fully explored. In this paper, we propose a novel AVVP framework termedDual Hierarchical Hybrid Network (DHHN) to tackle the above two problems. Our DHHN method consists of three components: 1) A hierarchical context modeling network for extracting different semantics in multiple temporal lengths; 2) A modality-wise guiding network for learning unique information from different modalities; 3) A dual-stream framework generating audio and visual predictions separately. It maintains the best adaptions on different modalities, further boosting the video parsing performance. Extensive quantitative and qualitative experiments demonstrate that our proposed method establishes the new state-of-the-art performance on the AVVP task.
Xun Jiang 0001, Xing Xu 0001, Jingran Zhang, Jingkuan Song, Fumin Shen, Huimin Lu 0001, Heng Tao Shen
ACM Multimedia7
2022 Guest Editorial Special Section on Learning With Multimodal Data for Biomedical Informatics
abstract
In this Special Section of the IEEE Transactions on Circuits and Systems for Video Technology, it is our honor to present emerging advanced machine learning and data analytics algorithms aiming at catalyzing synergies among image/video processing, text/speech understanding, and multimodal learning in biomedical informatics. Our goals are to 1) introduce novel data-driven models to accelerate knowledge discovery in biomedicine through the seamless integration of medical data collected from imaging systems, laboratory and wearable devices, as well as other related medical devices; 2) promote the development of new multi-modal learning systems to enhance the healthcare quality and patient safety; and 3) promote new applications in biomedical informatics that can leverage or benefits from the integration of multi-modal data and machine learning.
Zhangyang Wang, Vishal M. Patel, Steve B. Jiang, Huimin Lu 0001, Yang Shen 0001
IEEE Trans. Circuits Syst. Video Technol.5
2022 Image-Scale-Symmetric Cooperative Network for Defocus Blur Detection
abstract
Defocus blur detection (DBD) for natural images is a challenging vision task especially in the presence of homogeneous regions and gradual boundaries. In this paper, we propose a novel image-scale-symmetric cooperative network (IS2CNet) for DBD. On one hand, in the process of image scales from large to small, IS2CNet gradually spreads the recept of image content. Thus, the homogeneous region detection map can be optimized gradually. On the other hand, in the process of image scales from small to large, IS2CNet gradually feels the high-resolution image content, thereby gradually refining transition region detection. In addition, we propose a hierarchical feature integration and bi-directional delivering mechanism to transfer the hierarchical feature of previous image scale network to the input and tail of the current image scale network for guiding the current image scale network to better learn the residual. The proposed approach achieves state-of-the-art performance on existing datasets.Codes and results are available at:https://github.com/wdzhao123/IS2CNet.
Fan Zhao 0006, Huimin Lu 0001, Wenda Zhao 0003, Libo Yao
IEEE Trans. Circuits Syst. Video Technol.2
2022 Generalizable Crowd Counting via Diverse Context Style Learning
abstract
Existing crowd counting approaches predominantly perform well on the training-testing protocol. However, due to large style discrepancies not only among images but also within a single image, they suffer from obvious performance degradation when applied to unseen domains. In this paper, we aim to design a generalizable crowd counting framework which is trained on a source domain but can generalize well on the other domains. To reach this, we propose a gated ensemble learning framework. Specifically, we first propose a diverse fine-grained style attention model to help learn discriminative content feature representations, allowing for exploiting diverse features to improve generalization. We then introduce a channel-level binary gating ensemble model, where diverse feature prior, input-dependent guidance and density grade classification constraint are implemented, to optimally select diverse content features to participate in the ensemble, taking advantage of their complementary while avoiding redundancy. Extensive experiments show that our gating ensemble approach achieves superior generalization performance among four public datasets. Codes are publicly available athttps://github.com/wdzhao123/DCSL.
Wenda Zhao 0003, Yu Liu 0005, Huimin Lu 0001, Cong'an Xu, Libo Yao
IEEE Trans. Circuits Syst. Video Technol.4
2022 Learning Cross-Modal Common Representations by Private-Shared Subspaces Separation
abstract
Due to the inconsistent distributions and representations of different modalities (e.g., images and texts), it is very challenging to correlate such heterogeneous data. A standard solution is to construct one common subspace, where the common representations of different modalities are generated to bridge the heterogeneity gap. Existing methods based on common representation learning mostly adopt a less effective two-stage paradigm: first, generating separate representations for each modality by exploiting the modality-specific properties as the complementary information, and then capturing the cross-modal correlation in the separate representations for common representation learning. Moreover, these methods usually neglect that there may exist interference in the modality-specific properties, that is, the unrelated objects and background regions in images or the noisy words and incorrect sentences in the text. In this article, we hypothesize that explicitly modeling the interference within each modality can improve the quality of common representation learning. To this end, we propose a novel model private-shared subspaces separation (P3S) to explicitly learn different representations that are partitioned into two kinds of subspaces: 1) the common representations that capture the cross-modal correlation in a shared subspace and 2) the private representations that model the interference within each modality in two private subspaces. By employing the orthogonality constraints between the shared subspace and the private subspaces during the one-stage joint learning procedure, our model is able to learn more effective common representations for different modalities in the shared subspace by fully excluding the interference within each modality. Extensive experiments conducted on cross-modal retrieval verify the advantages of our P3S method compared with 15 state-of-the-art methods on four widely used cross-modal datasets.
Xing Xu 0001, Kaiyi Lin, Lianli Gao, Huimin Lu 0001, Heng Tao Shen, Xuelong Li 0001
IEEE Trans. Cybern.4
2022 Cognitive Memory-Guided AutoEncoder for Effective Intrusion Detection in Internet of Things
abstract
With the development of the Internet of Things (IoT) technology, intrusion detection has become a key technology that provides solid protection for IoT devices from network intrusion. At present, artificial intelligence technologies have been widely used in the intrusion detection task in previous methods. However, unknown attacks may also occur with the development of the network and the attack samples are difficult to collect, resulting in unbalanced sample categories. In this case, the previous intrusion detection methods have the problem of high false positive rates and low detection accuracy, which restricts the application of these methods in a real situation. In this article, we propose a novel method based on deep neural networks to tackle the intrusion detection task, which is termedCognitive Memory-guided AutoEncoder(CMAE). The CMAE method leverages a memory module to enhance the ability to store normal feature patterns while inheriting the advantages of autoencoder. Therefore, it is robust to the imbalanced samples. Besides, using the reconstruction error as an evaluation criterion to detect attacks effectively detects unknown attacks. To obtain superior intrusion detection performance, we propose feature reconstruction loss and feature sparsity loss to constrain the proposed memory module, promoting the discriminative of memory items and the ability of representation for normal data. Compared to previous state-of-the-art methods, sufficient experimental results reveal that the proposed CMAE method achieves excellent performance and effectiveness for intrusion detection.
Huimin Lu 0001, Xing Xu 0001, Ting Wang 0013
IEEE Trans. Ind. Informatics1
2022 Editorial Introduction to Responsible Artificial Intelligence for Autonomous Driving
abstract
Artificial Intelligence is in transition as the fast convergence of digital technologies and data science holds the promise to liberate consumer data and provide a faster and more cost-effective way of improving human initiatives. Particularly, artificial intelligence (AI) is heavily influencing autonomous vehicles nowadays. The data driven-based AI autonomous vehicles have the potential to reshape the expectations of human’s actions, the way that companies’ stakeholders collaborate, and revamp business models in the various industries.
Huimin Lu 0001, Mohsen Guizani, Pin-Han Ho
IEEE Trans. Intell. Transp. Syst.1
2022 Action Recognition Framework in Traffic Scene for Autonomous Driving System
abstract
For the autonomous driving system, accurately recognizing the actions of different roles in the traffic scene is the prerequisite for realizing this kind of human-vehicle information interaction. In this paper, we propose a complete framework based on 3D human pose estimation to recognize the actions of different roles on the road. The main objects recognized include traffic police, cyclists, and some passersby in need. We perform action recognition based on a dynamic adaptive graph convolutional network, which can realize the action recognition of objects based on 3D human pose. In addition to the action recognition module, we have optimized both the object detection module and the human pose estimation module in the framework so that the framework can handle multiple objects at the same time, which can be closer to the real traffic scene. To realize complex and changeable human action recognition, we built a multi-view camera system to collect responsible 3D human pose datasets containing traffic police gestures, cyclist gestures, and pedestrians’ body movements. In the experiments, compared to other state-of-the-art researches, the proposed framework can achieve comparable results with the same dataset. Satisfactory performance has also been obtained on the real data we collected, which can handle a variety of different action recognition tasks at the same time.
Feiyi Xu, Feng Xu 0005, Jiucheng Xie, Chi-Man Pun, Huimin Lu 0001, Hao Gao 0005
IEEE Trans. Intell. Transp. Syst.5
2022 Global-PBNet: A Novel Point Cloud Registration for Autonomous Driving
abstract
Registration performs an individual and deciding role in multiple intelligent transport systems. The advancement of deep-learning-based methods enhances the robustness and effectiveness of the preliminary registration stage, although the algorithm will effortlessly fall into local optima when improving the ultimate exactitude. Similarly, traditional method based on optimization has a more reliable performance in terms of precision. However, its performance still counts on the quality of initialization. In order to solve the above problems, we propose a PBNet that combines a point cloud network with a global optimization method. This framework uses the feature information of objects to perform high-precision rough registration and then searches the entire 3D motion space to implement branch-and-bound and iterative nearest point methods. The evaluation results show that PBNet significantly reduce the influence of initial values on registration and has good robustness against noise and outliers.
Yuchao Zheng 0001, Yujie Li 0001, Shuo Yang 0013, Huimin Lu 0001
IEEE Trans. Intell. Transp. Syst.4
2022 Answer Again: Improving VQA With Cascaded-Answering Model
abstract
Visual Question Answering (VQA) is a very challenging task, which requires to understand visual images and natural language questions simultaneously. In the open-ended VQA task, most previous solutions focus on understanding the question and image contents, as well as their correlations. However, they mostly reason the answers in a one-stage way, which results in that the generated answers are significantly ignored. In this paper, we propose a novel approach, termed Cascaded-Answering Model (CAM), which extends the conventional one-stage VQA model to a two-stage model. Hence, the proposed model can fully explore the semantics embedded in the predicted answers. Specifically, CAM is composed of two cascaded answering modules: Candidate Answer Generation (CAG) module and Final Answer Prediction (FAP) module. In CAG module, we select multiple relevant candidates from the generated answers using a typical VQA approach with Co-Attention. While in FAP module, we integrate the information of question and image, together with the semantics explored from the selected candidate answers to predict the final answer. Experimental results demonstrate that the proposed model produces high-quality candidate answers and achieves the state-of-the-art performance on five large benchmark datasets, VQA-1.0, VQA-2.0, VQA-CP v2, TDIUC and COCO-QA.
Yang Yang 0002, Xiaopeng Zhang 0008, Yanli Ji, Huimin Lu 0001, Heng Tao Shen
IEEE Trans. Knowl. Data Eng.5
2022 Cross-Modal Dynamic Networks for Video Moment Retrieval With Text Query
abstract
Video moment retrieval with text query aims to retrieve the most relevant segment from the whole video based on the given text query. It is a challenging cross-modal alignment task due to the huge gap between visual and linguistic modalities and the noise generated by manual labeling of time segments. Most of the existing works only use language information in the cross-modal fusion stage, neglecting that language information plays an important role in the retrieval stage. Besides, these works roughly compress the visual information in the video clips to reduce the computation cost which loses subtle video information in the long video. In this paper, we propose a novel model termed Cross-modal Dynamic Networks (CDN) which dynamically generates convolution kernel by visual and language features. In the feature extraction stage, we also propose a frame selection module to capture the subtle video information in the video segment. By this approach, the CDN can reduce the impact of the visual noise without significantly increasing the computation cost and leads to a better video moment retrieval result. The experiments on two challenge datasets,i.e., Charades-STA and TACoS, show that our proposed CDN method outperforms a bundle of state-of-the-art methods with more accurately retrieved moment video clips. The implementation code and extensive instruction of our proposed CDN method are provided athttps://github.com/CFM-MSG/Code_CDN.
Gongmian Wang, Xing Xu 0001, Fumin Shen, Huimin Lu 0001, Yanli Ji, Heng Tao Shen
IEEE Trans. Multim.4
2022 Discrete-Time Predictive Sliding Mode Control for a Constrained Parallel Micropositioning Piezostage
abstract
This article proposes a new discrete-time predictive sliding mode control (DPSMC) for a parallel micropositioning piezostage to improve the motion accuracy in the presence of cross-coupling hysteresis nonlinearities and input constraints. Unlike the traditional linear discrete-time sliding mode control (DSMC), the proposed DPSMC is chattering free and has a faster convergence rate thanks to the design of a nonlinear discrete-time fast integral terminal sliding mode surface. Moreover, by combining with the receding horizon optimization, the sliding mode state is predicted to follow the expected trajectory of a predefined continuous sliding mode reaching law, which also allows the proposed controller to explicitly deal with constraints. The stability of the closed-loop system is analyzed under the model disturbances and constraints, and proves that the proposed DPSMC can offer a smaller quasi-sliding mode bandwidth than the traditional DSMC. The effectiveness of the proposed controller is validated by a series of numerical simulations and experiments. Results demonstrate the advantages of proposed DPSMC over the traditional DSMC method.
Shengzheng Kang, Hongtao Wu 0001, Yao Li 0020, Jiafeng Yao, Bai Chen 0002, Huimin Lu 0001
IEEE Trans. Syst. Man Cybern. Syst.7
2022 Guest Editorial: Special issue on Cognitive computing for web applications
Huimin Lu 0001
Wirel. Networks1
2022 Special Issue on Synthetic Media on the Web
Huimin Lu 0001, Xing Xu 0001, Joze Guna, Gautam Srivastava 0001
World Wide Web1
2021 Enhancing Audio-Visual Association with Self-Supervised Curriculum Learning
abstract
The recent success of audio-visual representations learning can be largely attributed to their pervasive concurrency property, which can be used as a self-supervision signal and extract correlation information. While most recent works focus on capturing the shared associations between the audio and visual modalities, they rarely consider multiple audio and video pairs at once and pay little attention to exploiting the valuable information within each modality. To tackle this problem, we propose a novel audio-visual representation learning method dubbed self-supervised curriculum learning (SSCL) under the teacher-student learning manner. Specifically, taking advantage of contrastive learning, a two-stage scheme is exploited, which transfers the cross-modal information between teacher and student model as a phased process. The proposed SSCL approach regards the pervasive property of audiovisual concurrency as latent supervision and mutually distills the structure knowledge of visual to audio data. Notably, the SSCL method can learn discriminative audio and visual representations for various downstream applications. Extensive experiments conducted on both action video recognition and audio sound recognition tasks show the remarkably improved performance of the SSCL method compared with the state-of-the-art self-supervised audio-visual representation learning methods.
Jingran Zhang, Xing Xu 0001, Fumin Shen, Huimin Lu 0001, Xin Liu 0011, Heng Tao Shen
AAAI4
2021 Partial Feature Selection and Alignment for Multi-Source Domain Adaptation
abstract
Multi-Source Domain Adaptation (MSDA), which dedicates to transfer the knowledge learned from multiple source domains to an unlabeled target domain, has drawn increasing attention in the research community. By assuming that the source and target domains share consistent key feature representations and identical label space, existing studies on MSDA typically utilize the entire union set of features from both the source and target domains to obtain the feature map and align the map for each category and domain. However, the default setting of MSDA may neglect the issue of "partialness", i.e., 1) a part of the features contained in the union set of multiple source domains may not present in the target domain; 2) the label space of the target domain may not completely overlap with the multiple source domains. In this paper, we unify the above two cases to a more generalized MSDA task as Multi-Source Partial Domain Adaptation (MSPDA). We propose a novel model termed Partial Feature Selection and Alignment (PFSA) to jointly cope with both MSDA and MSPDA tasks. Specifically, we firstly employ a feature selection vector based on the correlation among the features of multiple sources and target domains. We then design three effective feature alignment losses to jointly align the selected features by preserving the domain information of the data sample clusters in the same category and the discrimination between different classes. Extensive experiments on various benchmark datasets for both MSDA and MSPDA tasks demonstrate that our proposed PFSA approach remarkably outperforms the state-of-the-art MSDA and unimodal PDA methods.
Yangye Fu, Xing Xu 0001, Zuo Cao, Yanli Ji, Kai Zuo, Huimin Lu 0001
CVPR8
2021 Multimodal Transformer Networks with Latent Interaction for Audio-Visual Event Localization
abstract
The task of audio-visual event localization (AVEL) aims to localize a visible and audible event in a video. Previous methods first divide a video into segments and then fuse visual and acoustic features at the segment level via a co-attention mechanism. However, existing methods mostly model relations between individual visual and audio segments in a limitedly short period, which may not cover a longer video duration for better high-level event information modeling. In this paper, we proposed a novel model termed Multimodal Transformer Network with Latent Interaction (MTNLI) to tackle this problem. The proposed MTNLI model employs a multimodal Transformer structure to learn the cross-modality relationships between latent visual and audio summarizations in long segment sequences, which summarize the visual and audio segments into a small number of latent representations to avoid modeling uninformative individual visual-audio relations. The cross-modality information between the latent summarizations is propagated to fuse valuable information from both modalities, which can effectively handle large temporal inconsistent between vision and audio. Our MTNLI method achieves state-of-the-art performance on the benchmark AVE (Audio-Visual Event) dataset for the event localization task.
Xing Xu 0001, Xin Liu 0011, Weihua Ou, Huimin Lu 0001
ICME5
2021 Lightweight Image Super-Resolution with Multi-Scale Feature Interaction Network
abstract
Recently, the single image super-resolution (SISR) approaches with deep and complex convolutional neural network structures have achieved promising performance. However, those methods improve the performance at the cost of higher memory consumption, which is difficult to be applied for some mobile devices with limited storage and computing resources. To solve this problem, we present a lightweight multi-scale feature interaction network (MSFIN). For lightweight SISR, MSFIN expands the receptive field and adequately exploits the informative features of the low-resolution observed images from various scales and interactive connections. In addition, we design a lightweight recurrent residual channel attention block (RRCAB) so that the network can benefit from the channel attention mechanism while being sufficiently lightweight. Extensive experiments on some benchmarks have confirmed that our proposed MSFIN can achieve comparable performance against the state-of-the-arts with a more lightweight model.
Zhengxue Wang, Guangwei Gao, Juncheng Li 0003, Yi Yu 0001, Huimin Lu 0001
ICME5
2021 Graph Convolutional Hourglass Networks for Skeleton-Based Action Recognition
abstract
Graph convolution networks (GCNs) have become the mainstream framework for the skeleton-based action recognition task, since the skeleton representation of human action can be naturally modeled by the graph structure. Generally, most of the existing GCN based models extract and aggregate skeleton features by exploiting single-scale joint information, while neglecting the valuable multi-scale information such as part and body features in the skeleton. To address this issue, we propose a novel Graph Convolutional Hourglass Network (GCHN) model, which is scalable by stacking several basic modules of Graph Convolutional Hourglass Block (GCHB). Each GCHB module consists of the sequential operations of graph convolution, graph pooling and graph unpooling, which can promote the interaction of multi-scale information in the skeleton and effectively improve the recognition performance. Extensive experiments on the challenging NTU-RGB+D and Kinetics-Skeleton datasets demonstrate that the proposed GCHN model achieves state-of-the-art performance.
Yiran Zhu, Xing Xu 0001, Yanli Ji, Fumin Shen, Heng Tao Shen, Huimin Lu 0001
ICME6
2021 Robust Motion Averaging under Maximum Correntropy Criterion
abstract
Recently, the motion averaging method has been introduced as an effective means to solve the multi-view registration problem. This method aims to recover global motions from a set of relative motions, where the original method is sensitive to outliers due to using the Frobenius norm error in the optimization. Accordingly, this paper proposes a novel robust motion averaging method based on the maximum correntropy criterion (MCC). Specifically, the correntropy measure is used instead of utilizing Frobenius norm error to improve the robustness of motion averaging against outliers. According to the half-quadratic technique, the correntropy measure based optimization problem can be solved by the alternating minimization procedure, which includes operations of weight assignment and weighted motion averaging. Further, we design a selection strategy of adaptive kernel width to take advantage of correntropy. Experimental results on benchmark data sets illustrate that our method has superior performance on accuracy and robustness for multi-view registration. What’s more, it can be applied to robot mapping.
Jihua Zhu, Huimin Lu 0001, Badong Chen, Zhongyu Li 0002, Yaochen Li
ICRA3
2021 CAA: Candidate-Aware Aggregation for Temporal Action Detection
abstract
Temporal action detection aims to locate specific segments of action instances in an untrimmed video. Most existing approaches commonly extract the features of all candidate video segments and then classify them separately. However, they may neglect the underlying relationship among candidates unconsciously. In this paper, we propose a novel model termed Candidate-Aware Aggregation (CAA) to tackle this problem. In CAA, we design the Global Awareness (GA) module to exploit long-range relations among all candidates from a global perspective, which enhances the features of action instances. The GA module is then embedded into a multi-level hierarchical network named FENet, to aggregate local features in adjacent candidates to suppress background noise. As a result, the relationship among candidates is explicitly captured from both local and global perspectives, which ensures more accurate prediction results for the candidates. Extensive experiments conducted on two popular benchmarks ActivityNet-1.3 and THUMOS-14 demonstrate the superiority of CAA comparing to the state-of-the-art methods.
Yifan Ren, Xing Xu 0001, Fumin Shen, Yazhou Yao, Huimin Lu 0001
ACM Multimedia5
2021 Special issue on cognitive computing for robotic vision
abstract
Cognitive computing breaks the boundary between two separate fields, neuroscience and computer science. It paves the way for machines to have reasoning abilities which is analogous to human. The research field of cognitive computing is interdisciplinary, and uses knowledge and methods from many areas such as psychology, biology, signal processing, physics, information theory, mathematics, and statistics. The development of cognitive computing will keep cross-fertilizing these research areas. However, in robotic vision applications there still remain many open problems for cognitive computing. Technologies like cloud computing and deep learning are essential to upgrade the robotic vision systems with near human intelligence by using new capabilities. From the total papers submitted to this special issue, including the selected best paper from the 4th International Symposium on Artificial Intelligence and Robotics 2019 (https://isair.site/). Finally, the high-quality articles were selected. Each paper was peer reviewed by two or more experts during the assessment process. The selected articles have exceptional diversity in terms of cognitive computing and computer vision techniques and applications. They represent the most recent development in both theory and practice.
Huimin Lu 0001, Joze Guna
Concurr. Comput. Pract. Exp.1
2021 Ontology negotiation: Knowledge interchange between distributed ontologies through agent negotiation
abstract
Summary With the proliferation of knowledge source on the internet as well as the widely professional agents, the knowledge interchange is drawing much attention. Ontology is recognized as the crucial technology due to their nature of sharing, formalization, and conceptualization to integrate and share the knowledge. In this paper, by interpreting and negotiating the communication content, a unified understanding of knowledge is formed; then, we can realize the interoperability between ontologies. We have developed the ontology automatic negotiation by agent elect protocol (AEP) to elect optimal participants and encourage agents to obey the protocol, concept mapping protocol (CMP) to find the corresponding concept mappings with the highest relevancy, and in addition, agent negotiation protocol (ANP). In ANP, we define the simultaneous negotiation protocol and agents' strategies to combine distributed ontologies interchange with agent negotiation. Finally, the implementation and preliminary results are given to verify the validity of the proposed ontology negotiation.
Junwu Zhu, Ling Teng, Huimin Lu 0001, Jieke Shi, Bin Li 0006
Concurr. Comput. Pract. Exp.3
2021 A Security Awareness and Protection System for 5G Smart Healthcare Based on Zero-Trust Architecture
abstract
The key features of 5G network (i.e., high bandwidth, low latency, and high concurrency) along with the capability of supporting big data platforms with high mobility make it valuable in coping with emerging medical needs, such as COVID-19 and future healthcare challenges. However, enforcing the security aspect of a 5G-based smart healthcare system that hosts critical data and services is becoming more urgent and critical. Passive security mechanisms (e.g., data encryption and isolation) used in legacy medical platforms cannot provide sufficient protection for a healthcare system that is deployed in a distributed manner and fail to meet the need for data/service sharing across "cloud-edge-terminal" in the 5G era. In this article, we propose a security awareness and protection system that leverages zero-trust architecture for a 5G-based smart medical platform. Driven by the four key dimensions of 5G smart healthcare including "subject" (i.e., users, terminals, and applications), "object" (i.e., data, platforms, and services), "behavior," and "environment," our system constructs trustable dynamic access control models and achieves real-time network security situational awareness, continuous identity authentication, analysis of access behavior, and fine-grained access control. The proposed security system is implemented and tested thoroughly at industrial-grade, which proves that it satisfies the needs of active defense and end-to-end security enforcement of data, users, and services involved in a 5G-based smart medical system.
BaoZhan Chen, Siyuan Qiao, Dongqing Liu, Xiaobing Shi, Minzhao Lyu, Huimin Lu 0001, Yunkai Zhai
IEEE Internet Things J.8
2021 A Model for Joint Planning of Production and Distribution of Fresh Produce in Agricultural Internet of Things
abstract
The production and distribution planning of fresh produce is a complex optimization problem, which is affected by many factors, including its perishable characteristics. Farmers cannot guarantee the efficiency and accuracy of production and distribution decisions. Given the close relationship between the production and distribution of annual fresh produce, the intention of our research is to solve the two-stage joint planning problem and maximize the revenue of farmers ultimately. The internal relationship matrix between the two links of production and distribution is established. On this basis, we propose a mixed-integer programming (MIP) model, which covers the constraints of labor and capital. The decisions obtained are not only based on price estimation and resource availability but also on the impact of the agricultural Internet-of-Things technology and the special requirements of each distribution channel. Numerical experiments demonstrate that when the planting area is 1, 4, and 6 ha, the proposed joint planning model can improve the distribution revenue of farmers by 7.92%, 4.15%, and 4.94%, respectively, compared with the traditional separate decision-making approach of distribution. According to different decision scenarios, management insights have been obtained. For example, farmers should carefully sort and package products as well as choose a timely and safe third-party express delivery company. Additionally, the proposed strategy can evaluate the impact of distribution channels on farmers' revenue.
Jiliang Han, Na Lin 0003, Junhu Ruan, Xuping Wang, Wei Wei 0006, Huimin Lu 0001
IEEE Internet Things J.6
2021 Guest Editorial: Special Issue on Internet of Things for Industrial Security for Smart Cities
abstract
More than half of the world’s current population resides in urban areas to compare to just 30% in the 1950s. The process of urbanization leads to exurban sprawl, the formation of slums, scattered workplaces, and aging infrastructure. These may cause huge inefficiencies in energy use, traffic, governance, waste management, and pollution, among others. To overcome these social, economic, and environmental challenges, public and private sectors invest heavily in smart city technologies. However, the risks of using smart technologies due to security breaches and cyberattacks in critical sectors should be well addressed.
Huimin Lu 0001, Pin-Han Ho, Mohsen Guizani
IEEE Internet Things J.1
2021 ATTDC: An Active and Traceable Trust Data Collection Scheme for Industrial Security in Smart Cities
abstract
With billions of sensing devices are deployed in smart cities to monitor regions of interests and collect large sensing data, the Internet-of-Things (IoT) applications are being widely used in various fields and empower the intelligent smart cities. Due to the smart decision made by IoT applications depends on the reliability of data collection, it is pivotal to collect data from the trust sensing devices. However, how to identify the credibility of sensor nodes to ensure the credibility of data collection is a challenge issue. In this article, an active and traceable trust-based data collection (ATTDC) scheme is proposed to collect trust data in Internet of Thing. The main contribution of this article are as follows: 1) an active trust framework is proposed to quickly obtain the trustworthiness of sensor nodes by using unmanned aerial vehicles (UAVs) with a piggybacking method; 2) in order to accurately obtain the trust degree of the sensor nodes, a traceable trust method of obtaining is proposed in which nodes in the network send data packets by digital signature, Tracing suspicious nodes according to data routing paths to obtain active trust, therefore, the acquisition cost of the network can be effectively reduced; and 3) In order to reduce the acquisition cost of UAV, an ant colony algorithm-based flight path algorithm is designed to reduce the flight path of UAV, and obtain the credibility evaluation of as many nodes as possible. The experimental results show that the ATTDC scheme proposed in this article can identify the trust of the sensing nodes faster and more accurately, ensuring the credibility of data collection.
Mengqiu Shen, Anfeng Liu, Guosheng Huang, Naixue Xiong, Huimin Lu 0001
IEEE Internet Things J.5
2021 SCCGAN: Style and Characters Inpainting Based on CGAN
Xiangshang Wang, Huimin Lu 0001, Shanxi Li, Xin Jin 0015
Mob. Networks Appl.3
2021 MADNet: A Fast and Lightweight Network for Single-Image Super Resolution
abstract
Recently, deep convolutional neural networks (CNNs) have been successfully applied to the single-image super-resolution (SISR) task with great improvement in terms of both peak signal-to-noise ratio (PSNR) and structural similarity (SSIM). However, most of the existing CNN-based SR models require high computing power, which considerably limits their real-world applications. In addition, most CNN-based methods rarely explore the intermediate features that are helpful for final image recovery. To address these issues, in this article, we propose a dense lightweight network, called MADNet, for stronger multiscale feature expression and feature correlation learning. Specifically, a residual multiscale module with an attention mechanism (RMAM) is developed to enhance the informative multiscale feature representation ability. Furthermore, we present a dual residual-path block (DRPB) that utilizes the hierarchical features from original low-resolution images. To take advantage of the multilevel features, dense connections are employed among blocks. The comparative results demonstrate the superior performance of our MADNet model while employing considerably fewer multiadds and parameters.
Rushi Lan, Zhenbing Liu, Huimin Lu 0001
IEEE Trans. Cybern.4
2021 Cascading and Enhanced Residual Networks for Accurate Single-Image Super-Resolution
abstract
Deep convolutional neural networks (CNNs) have contributed to the significant progress of the single-image super-resolution (SISR) field. However, the majority of existing CNN-based models maintain high performance with massive parameters and exceedingly deeper structures. Moreover, several algorithms essentially have underused the low-level features, thus causing relatively low performance. In this article, we address these problems by exploring two strategies based on novel local wider residual blocks (LWRBs) to effectively extract the image features for SISR. We propose a cascading residual network (CRN) that contains several locally sharing groups (LSGs), in which the cascading mechanism not only promotes the propagation of features and the gradient but also eases the model training. Besides, we present another enhanced residual network (ERN) for image resolution enhancement. ERN employs a dual global pathway structure that incorporates nonlocal operations to catch long-distance spatial features from the the original low-resolution (LR) input. To obtain the feature representation of the input at different scales, we further introduce a multiscale block (MSB) to directly detect low-level features from the LR image. The experimental results on four benchmark datasets have demonstrated that our models outperform most of the advanced methods while still retaining a reasonable number of parameters.
Rushi Lan, Zhenbing Liu, Huimin Lu 0001, Zhixun Su
IEEE Trans. Cybern.4
2021 A Two-Phase Learning-Based Swarm Optimizer for Large-Scale Optimization
abstract
In this article, a simple yet effective method, called a two-phase learning-based swarm optimizer (TPLSO), is proposed for large-scale optimization. Inspired by the cooperative learning behavior in human society, mass learning and elite learning are involved in TPLSO. In the mass learning phase, TPLSO randomly selects three particles to form a study group and then adopts a competitive mechanism to update the members of the study group. Then, we sort all of the particles in the swarm and pick out the elite particles that have better fitness values. In the elite learning phase, the elite particles learn from each other to further search for more promising areas. The theoretical analysis of TPLSO exploration and exploitation abilities is performed and compared with several popular particle swarm optimizers. Comparative experiments on two widely used large-scale benchmark datasets demonstrate that the proposed TPLSO achieves better performance on diverse large-scale problems than several state-of-the-art algorithms.
Rushi Lan, Yu Zhu 0004, Huimin Lu 0001, Zhenbing Liu
IEEE Trans. Cybern.3
2021 Deep Fuzzy Hashing Network for Efficient Image Retrieval
abstract
Hashing methods for efficient image retrieval aim at learning hash functions that map similar images to semantically correlated binary codes in the Hamming space with similarity well preserved. The traditional hashing methods usually represent image content by hand-crafted features. Deep hashing methods based on deep neural network (DNN) architectures can generate more effective image features and obtain better retrieval performance. However, the underlying data structure is hardly captured by existing DNN models. Moreover, the similarity (either visually or semantically) between pairwise images is ambiguous, even uncertain, to be measured in the existing deep hashing methods. In this article, we propose a novel hashing method termed deep fuzzy hashing network (DFHN) to overcome the shortcomings of existing deep hashing approaches. Our DFHN method combines the fuzzy logic technique and the DNN to learn more effective binary codes, which can leverage fuzzy rules to model the uncertainties underlying the data. Derived from fuzzy logic theory, the generalized hamming distance is devised in the convolutional layers and fully connected layers in our DFHN to model their outputs, which come from an efficientxoroperation on given inputs and weights. Extensive experiments show that our DFHN method obtains competitive retrieval accuracy with highly efficient training speed compared with several state-of-the-art deep hashing approaches on two large-scale image datasets: CIFAR-10 and NUS-WIDE.
Huimin Lu 0001, Xing Xu 0001, Yujie Li 0001, Heng Tao Shen
IEEE Trans. Fuzzy Syst.1
2021 Accurate 3-D Reconstruction Under IoT Environments and Its Applications to Augmented Reality
abstract
With the remarkable development of sensor devices and the Internet of Things (IoT), today's researchers can easily know what changes have taken place in the real world by acquiring a 3-D model. Conversely, a large amount of image data promotes the development of perceptual computing technology. In this article, we focus on modeling 3-D scenes from the multisource image data obtained from the IoT with cameras. Although great progress has been made in 3-D reconstruction, it is still challenging to recover the 3-D model from IoT data because the captured images are usually noisy, incomplete, varying scale, and with repetitive structures or features. In this article, we propose an accurate 3-D reconstruction method under IoT environments for perceptual computing of the scene. This method consists of sparse, dense, and surface reconstruction processes, which can gradually recover high-quality geometric models from the image data and efficiently deal with various repetitive structures. By analyzing the reconstructed model, we can detect the changes of scenes. We evaluate the proposed method on the benchmark data sets (i.e., tanks and temples) and publicly available data sets(in which samples usually contain repeated structures, lighting change, and different scales). Experimental results show that the proposed method outperforms the state-of-the-art methods according to the standard evaluation metric. We also use our method to enhance the real scenes with virtual objects, thus producing promising results.
Mingwei Cao, Liping Zheng, Wei Jia 0001, Huimin Lu 0001, Xiaoping Liu 0003
IEEE Trans. Ind. Informatics4
2021 Construction of a Hierarchical Feature Enhancement Network and Its Application in Fault Recognition
abstract
Industrial Internet of Things (IIoT) provide significant support for observing and controlling industrial machinery. In this article, a novel hierarchical feature enhancement network (HFEN) is proposed by combining signal processing and representation learning. The signal processing block extracts features with definite physical significance. Then, the representability of the physical features is improved by connecting stacked denoising autoencoders and squeeze-and-excitation networks. A novel two-stream architecture is designed for HFEN to fuse two types of features. Consequently, HFEN can extract features that can be analyzed for physical significance and that are also representative in terms of recognizable patterns. The experimental results prove that the performance of HFEN is satisfactory in terms of accuracy and efficiency when compared to other methods. Finally, this article also aims to demonstrate the potential of a new pairing that fuses the model- and data-driven strategies for IIoT.
Zhe Chen 0004, Huimin Lu 0001, Shiqing Tian, Junlin Qiu, Tohru Kamiya, Seiichi Serikawa
IEEE Trans. Ind. Informatics2
2021 Synergic Adversarial Label Learning for Grading Retinal Diseases via Knowledge Distillation and Multi-Task Learning
abstract
The need for comprehensive and automated screening methods for retinal image classification has long been recognized. Well-qualified doctors annotated images are very expensive and only a limited amount of data is available for various retinal diseases such as diabetic retinopathy (DR) and age-related macular degeneration (AMD). Some studies show that some retinal diseases such as DR and AMD share some common features like haemorrhages and exudation but most classification algorithms only train those disease models independently when the only single label for one image is available. Inspired by multi-task learning where additional monitoring signals from various sources is beneficial to train a robust model. We propose a method called synergic adversarial label learning (SALL) which leverages relevant retinal disease labels in both semantic and feature space as additional signals and train the model in a collaborative manner using knowledge distillation. Our experiments on DR and AMD fundus image classification task demonstrate that the proposed method can significantly improve the accuracy of the model for grading diseases by 5.91% and 3.69% respectively. In addition, we conduct additional experiments to show the effectiveness of SALL from the aspects of reliability and interpretability in the context of medical imaging application.
Lie Ju, Xin Wang 0094, Huimin Lu 0001, Dwarikanath Mahapatra, C. Paul Bonnington, ZongYuan Ge
IEEE J. Biomed. Health Informatics4
2021 User-Oriented Virtual Mobile Network Resource Management for Vehicle Communications
abstract
Currently, advanced communications and networks greatly enhance user experiences and have a major impact on all aspects of people's lifestyles in terms of work, society, and the economy. However improving competitiveness and sustainable vehicle network services, such as higher user experience, considerable resource utilization and effective personalized services, is a great challenge. Addressing these issues, this paper proposes a virtual network resource management based on user behavior to further optimize the existing vehicle communications. In particular, ensemble learning is implemented in the proposed scheme to predict the user's voice call duration and traffic usage for supporting user-centric mobile services optimization. Sufficient experiments show that the proposed scheme can significantly improve the quality of services and experiences and that it provides a novel idea for optimizing vehicle networks.
Huimin Lu 0001, Yin Zhang 0002, Yujie Li 0001, Haider Abbas
IEEE Trans. Intell. Transp. Syst.1
2021 Multi-Aspect Aware Session-Based Recommendation for Intelligent Transportation Services
abstract
In the intelligent transportation system, the session data usually represents the users' demand. However, the traditional approaches only focus on the sequence information or the last item clicked by the user, which cannot fully represent user preferences. To address this issue, this paper proposes an Multi-aspect Aware Session-based Recommendation (MASR) model for intelligent transportation services, which comprehensively considers the user's personalized behavior from multiple aspects. In addition, it developed a concise and efficient transformer-style self-attention to analyze the sequence information of the current session, for accurately grasping the user's intention. Finally, the experimental results show that MASR is available to improve user satisfaction with more accurate and rapid recommendations, and reduce the number of user operations to decrease the safety risk during the transportation service.
Yin Zhang 0002, Yujie Li 0001, Ranran Wang 0001, M. Shamim Hossain, Huimin Lu 0001
IEEE Trans. Intell. Transp. Syst.5
2021 Robust Facial Image Super-Resolution by Kernel Locality-Constrained Coupled-Layer Regression
abstract
Super-resolution methods for facial image via representation learning scheme have become very effective methods due to their efficiency. The key problem for the super-resolution of facial image is to reveal the latent relationship between the low-resolution ( LR ) and the corresponding high-resolution ( HR ) training patch pairs. To simultaneously utilize the contextual information of the target position and the manifold structure of the primitive HR space, in this work, we design a robust context-patch facial image super-resolution scheme via a kernel locality-constrained coupled-layer regression (KLC2LR) scheme to obtain the desired HR version from the acquired LR image. Here, KLC2LR proposes to acquire contextual surrounding patches to represent the target patch and adds an HR layer constraint to compensate the detail information. Additionally, KLC2LR desires to acquire more high-frequency information by searching for nearest neighbors in the HR sample space. We also utilize kernel function to map features in original low-dimensional space into a high-dimensional one to obtain potential nonlinear characteristics. Our compared experiments in the noisy and noiseless cases have verified that our suggested methodology performs better than many existing predominant facial image super-resolution methods.
Guangwei Gao, Huimin Lu 0001, Yi Yu 0001, Heyou Chang, Dong Yue 0001
ACM Trans. Internet Techn.3
2021 A Hybrid Feature Selection Algorithm Based on a Discrete Artificial Bee Colony for Parkinson's Diagnosis
abstract
Parkinson's disease is a neurodegenerative disease that affects millions of people around the world and cannot be cured fundamentally. Automatic identification of early Parkinson's disease on feature data sets is one of the most challenging medical tasks today. Many features in these datasets are useless or suffering from problems like noise, which affect the learning process and increase the computational burden. To ensure the optimal classification performance, this article proposes a hybrid feature selection algorithm based on an improved discrete artificial bee colony algorithm to improve the efficiency of feature selection. The algorithm combines the advantages of filters and wrappers to eliminate most of the uncorrelated or noisy features and determine the optimal subset of features. In the filter, three different variable ranking methods are employed to pre-rank the candidate features, then the population of artificial bee colony is initialized based on the significance degree of the re-rank features. In the wrapper part, the artificial bee colony algorithm evaluates individuals (feature subsets) based on the classification accuracy of the classifier to achieve the optimal feature subset. In addition, for the first time, we introduce a strategy that can automatically select the best classifier in the search framework more quickly. By comparing with several publicly available datasets, the proposed method achieves better performance than other state-of-the-art algorithms and can extract fewer effective features.
Haolun Li 0001, Chi-Man Pun, Feng Xu 0005, Longsheng Pan, Rui Zong, Hao Gao 0005, Huimin Lu 0001
ACM Trans. Internet Techn.7
2021 Introduction to the Special Section on Cognitive Robotics on 5G/6G Networks
abstract
introduction Share on Introduction to the Special Section on Cognitive Robotics on 5G/6G Networks Authors: Huimin Lu Kyushu Institute of Technology, Japan Kyushu Institute of Technology, JapanView Profile , Liao Wu University of New South Wales, Australia University of New South Wales, AustraliaView Profile , Giancarlo Fortino University of Calabria (Unical), Italy University of Calabria (Unical), ItalyView Profile , Schahram Dustdar Vienna University of Technology, Austria Vienna University of Technology, AustriaView Profile Authors Info & Claims ACM Transactions on Internet TechnologyVolume 21Issue 4November 2021 Article No.: 91epp 1–3https://doi.org/10.1145/3476466Published:28 September 2021Publication History 1citation36DownloadsMetricsTotal Citations1Total Downloads36Last 12 Months26Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Huimin Lu 0001, Liao Wu, Giancarlo Fortino, Schahram Dustdar
ACM Trans. Internet Techn.1
2021 Zero-shot Cross-modal Retrieval by Assembling AutoEncoder and Generative Adversarial Network
abstract
Conventional cross-modal retrieval models mainly assume the same scope of the classes for both the training set and the testing set. This assumption limits their extensibility on zero-shot cross-modal retrieval (ZS-CMR), where the testing set consists of unseen classes that are disjoint with seen classes in the training set. The ZS-CMR task is more challenging due to the heterogeneous distributions of different modalities and the semantic inconsistency between seen and unseen classes. A few of recently proposed approaches are inspired by zero-shot learning to estimate the distribution underlying multimodal data by generative models and make the knowledge transfer from seen classes to unseen classes by leveraging class embeddings. However, directly borrowing the idea from zero-shot learning (ZSL) is not fully adaptive to the retrieval task, since the core of the retrieval task is learning the common space. To address the above issues, we propose a novel approach named Assembling AutoEncoder and Generative Adversarial Network (AAEGAN), which combines the strength of AutoEncoder (AE) and Generative Adversarial Network (GAN), to jointly incorporate common latent space learning, knowledge transfer, and feature synthesis for ZS-CMR. Besides, instead of utilizing class embeddings as common space, the AAEGAN approach maps all multimodal data into a learned latent space with the distribution alignment via three coupled AEs. We empirically show the remarkable improvement for ZS-CMR task and establish the state-of-the-art or competitive performance on four image-text retrieval datasets.
Xing Xu 0001, Kaiyi Lin, Huimin Lu 0001, Jie Shao 0001, Heng Tao Shen
ACM Trans. Multim. Comput. Commun. Appl.4
2021 Output-Bounded and RBFNN-Based Position Tracking and Adaptive Force Control for Security Tele-Surgery
abstract
In security e-health brain neurosurgery, one of the important processes is to move the electrocoagulation to the appropriate position in order to excavate the diseased tissue. 1 However, it has been problematic for surgeons to freely operate the electrocoagulation, as the workspace is very narrow in the brain. Due to the precision, vulnerability, and important function of brain tissues, it is essential to ensure the precision and safety of brain tissues surrounding the diseased part. The present study proposes the use of a robot-assisted tele-surgery system to accomplish the process. With the aim to achieve accuracy, an output-bounded and RBF neural network–based bilateral position control method was designed to guarantee the stability and accuracy of the operation process. For the purpose of accomplishing a minimal amount of bleeding and damage, an adaptive force control of the slave manipulator was proposed, allowing it to be appropriate to contact the susceptible vessels, nerves, and brain tissues. The stability was analyzed, and the numerical simulation results revealed the high performance of the proposed controls.
Ting Wang 0013, Xiangjun Ji, Aiguo Song, Kurosh Madani, Amine Chohra, Huimin Lu 0001, Ramon Monero
ACM Trans. Multim. Comput. Commun. Appl.6
2021 Chinese Image Captioning via Fuzzy Attention-based DenseNet-BiLSTM
abstract
Chinese image description generation tasks usually have some challenges, such as single-feature extraction, lack of global information, and lack of detailed description of the image content. To address these limitations, we propose a fuzzy attention-based DenseNet-BiLSTM Chinese image captioning method in this article. In the proposed method, we first improve the densely connected network to extract features of the image at different scales and to enhance the model’s ability to capture the weak features. At the same time, a bidirectional LSTM is used as the decoder to enhance the use of context information. The introduction of an improved fuzzy attention mechanism effectively improves the problem of correspondence between image features and contextual information. We conduct experiments on the AI Challenger dataset to evaluate the performance of the model. The results show that compared with other models, our proposed model achieves higher scores in objective quantitative evaluation indicators, including BLEU , BLEU , METEOR, ROUGEl, and CIDEr. The generated description sentence can accurately express the image content.
Huimin Lu 0001, Rui Yang 0018, Zhenrong Deng, Yonglin Zhang, Guangwei Gao, Rushi Lan
ACM Trans. Multim. Comput. Commun. Appl.1
2020 Deep Learning for Visual Segmentation: A Review
abstract
Big data-driven deep learning methods have been widely used in image or video segmentation. The main challenge is that a large amount of labeled data is required in training deep learning models, which is important in real-world applications. To the best of our knowledge, there exist few researches in the deep learning-based visual segmentation. To this end, this paper summarizes the algorithms and current situation of image or video segmentation technologies based on deep learning and point out the future trends. The characteristics of segmentation that based on semi-supervised or unsupervised learning, all of the recent novel methods are summarized in this paper. The principle, advantages and disadvantages of each algorithms are also compared and analyzed.
Yujie Li 0001, Huimin Lu 0001, Tohru Kamiya, Seiichi Serikawa
COMPSAC3
2020 Data Analytics for the COVID-19 Epidemic
abstract
With the spread of COVID-19 worldwide, people¡¯s production and life have been significantly affected. Artificial intelligence and big data technologies have been vigorously developed in recent years. It is very significant to use data science and technology to help humans in a timely and accurate manner to prevent and control the development of the epidemic, maintain social stability and assess the impact of the epidemic. This paper explores how data science can play a role from the perspectives of epidemiology, social networking, and economics. In particular, for the existing epidemic model SIR, we present a parameter learning method using particle swarm optimization (PSO) and the least squares method, and use it to predict the trend of the epidemic. Aiming at the social network data, we provide a specific method to realize sentiment analysis during the epidemic and propose an explainable fake news detection technique based on a variety of data mining methods.
Ranran Wang 0001, Huimin Lu 0001, Yin Zhang 0002
COMPSAC4
2020 Temporal Denoising Mask Synthesis Network for Learning Blind Video Temporal Consistency
abstract
Recently, developing temporally consistent video-based processing techniques has drawn increasing attention due to the defective extend-ability of existing image-based processing algorithms (e.g., filtering, enhancement, colorization, etc). Generally, applying these image-based algorithms independently to each video frame typically leads to temporal flickering due to the global instability of these algorithms. In this paper, we consider enforcing temporal consistency in a video as a temporal denoising problem that removing the flickering effect in given unstable pre-processed frames. Specifically, we propose a novel model termed Temporal Denoising Mask Synthesis Network (TDMS-Net) that jointly predicts the motion mask, soft optical flow and the refining mask to synthesize the temporal consistent frames. The temporal consistency is learned from the original video and the learned temporal features are applied to reprocess the output frames that are agnostic (blind) to specific image-based processing algorithms. Experimental results on two datasets for 16 different applications demonstrate that the proposed TDMS-Net significantly outperforms two state-of-the-art blind temporal consistency approaches.
Xing Xu 0001, Fumin Shen, Lianli Gao, Huimin Lu 0001, Heng Tao Shen
ACM Multimedia5
2020 Correlated Features Synthesis and Alignment for Zero-shot Cross-modal Retrieval
abstract
The goal of cross-modal retrieval is to search for semantically similar instances in one modality by using a query from another modality. Existing approaches mainly consider the standard scenario that requires the source set for training and the target set for testing share the same scope of classes. However, they may not generalize well on zero-shot cross-modal retrieval (ZS-CMR) task, where the target set contains unseen classes that are disjoint with the seen classes in the source set. This task is more challenging due to 1) the absence of the unseen classes during training, 2) inconsistent semantics across seen and unseen classes, and 3) the heterogeneous multimodal distributions between the source and target set. To address these issues, we propose a novel Correlated Feature Synthesis and Alignment (CFSA) approach to integrate multimodal feature synthesis, common space learning and knowledge transfer for ZS-CMR. Our CFSA first utilizes class-level word embeddings to guide two coupled Wassertein generative adversarial networks (WGANs) to synthesize sufficient multimodal features with semantic correlation for stable training. Then the synthetic and true multimodal features are jointly mapped to a common semantic space via an effective distribution alignment scheme, where the cross-modal correlations of different semantic features are captured and the knowledge can be transferred to the unseen classes under the cycle-consistency constraint. Experiments on four benchmark datasets for image-text retrieval and two large-scale datasets for image-sketch retrieval show the remarkable improvements achieved by our CFAS method comparing with a bundle of state-of-the-art approaches.
Xing Xu 0001, Kaiyi Lin, Huimin Lu 0001, Lianli Gao, Heng Tao Shen
SIGIR3
2020 Wireless high-frequency NLOS monitoring system for heart disease combined with hospital and home
Jun Yang 0014, Wenjing Xiao, Huimin Lu 0001, Ahmed Barnawi
Future Gener. Comput. Syst.3
2020 Deep hierarchical encoding model for sentence semantic matching
Wenpeng Lu, Xu Zhang 0053, Huimin Lu 0001
J. Vis. Commun. Image Represent.3
2020 Endmember Extraction of Hyperspectral Remote Sensing Images Based on an Improved Discrete Artificial Bee Colony Algorithm and Genetic Algorithm
Zheng Fu, Chi-Man Pun, Hao Gao 0005, Huimin Lu 0001
Mob. Networks Appl.4
2020 Virtual Reality Sickness and Challenges Behind Different Technology and Content Settings
Joze Guna, Gregor Gersak, Iztok Humar, Maja Krebl, Marko Orel, Huimin Lu 0001, Matevz Pogacnik
Mob. Networks Appl.6
2020 Humanized Computing for Mass Customization Application in Curriculum Management
Yuqian Shi, Yi Bu 0001, Yang Xu 0022, Huimin Lu 0001, Xiangshang Wang, Weihua Lu, Changjiang Ji
Mob. Networks Appl.5
2020 Editorial: Cognitive Science and Artificial Intelligence for Human Cognition and Communication
Huimin Lu 0001, Yujie Li 0001
Mob. Networks Appl.1
2020 Cognitive Computing for Intelligence Systems
Huimin Lu 0001, Yujie Li 0001
Mob. Networks Appl.1
2020 Deep-Sea Organisms Tracking Using Dehazing and Deep Learning
Huimin Lu 0001, Tomoki Uemura, Dong Wang 0004, Jihua Zhu, Zi Huang, Hyoungseop Kim
Mob. Networks Appl.1
2020 Correction to: Deep-Sea Organisms Tracking Using Dehazing and Deep Learning
Huimin Lu 0001, Tomoki Uemura, Dong Wang 0004, Jihua Zhu, Zi Huang, Hyoungseop Kim
Mob. Networks Appl.1
2020 Human Emotion Recognition Using an EEG Cloud Computing Platform
Huimin Lu 0001, Mei Wang 0002, Arun Kumar Sangaiah
Mob. Networks Appl.1
2020 Detection of Circulating Tumor Cells in Fluorescence Microscopy Images Based on ANN Classifier
Kouki Tsuji, Huimin Lu 0001, Joo Kooi Tan, Hyoungseop Kim, Kazue Yoneda, Fumihiro Tanaka
Mob. Networks Appl.2
2020 Blind Image Deblurring Based on Local Rank
Li Zhu 0003, Jihua Zhu, Zhongyu Li 0002, Huimin Lu 0001
Mob. Networks Appl.6
2020 Locality-constrained feature space learning for cross-resolution sketch-photo face recognition
Guangwei Gao, Yannan Wang, Heyou Chang, Huimin Lu 0001, Dong Yue 0001
Multim. Tools Appl.5
2020 Improving multi-label chest X-ray disease diagnosis by exploiting disease and health labels dependencies
ZongYuan Ge, Dwarikanath Mahapatra, Xiaojun Chang, Zetao Chen, Lianhua Chi, Huimin Lu 0001
Multim. Tools Appl.6
2020 Effect of VR technology matureness on VR sickness
Gregor Gersak, Huimin Lu 0001, Joze Guna
Multim. Tools Appl.2
2020 Learning adaptive contrast combinations for visual saliency detection
Quan Zhou 0004, Huimin Lu 0001, Yawen Fan, Suofei Zhang, Xiaofu Wu, Baoyu Zheng, Weihua Ou, Longin Jan Latecki
Multim. Tools Appl.3
2020 High-quality-guided artificial bee colony algorithm for designing loudspeaker
Hao Gao 0005, Haolun Li 0001, Ye Liu 0005, Huimin Lu 0001, Hyoungseop Kim, Chi-Man Pun
Neural Comput. Appl.4
2020 Face image super-resolution with pose via nuclear norm regularized structural orthogonal Procrustes regression
Guangwei Gao, Meng Yang 0001, Huimin Lu 0001, Wankou Yang, Hao Gao 0005
Neural Comput. Appl.4
2020 An LBP encoding scheme jointly using quaternionic representation and angular information
Rushi Lan, Huimin Lu 0001, Yicong Zhou, Zhenbing Liu
Neural Comput. Appl.2
2020 A pricing method of online group-buying for continuous price function
Junwu Zhu, Ling Teng, Zhengnan Zhu, Huimin Lu 0001
Neural Comput. Appl.4
2020 Ternary Adversarial Networks With Self-Supervision for Zero-Shot Cross-Modal Retrieval
abstract
Given a query instance from one modality (e.g., image), cross-modal retrieval aims to find semantically similar instances from another modality (e.g., text). To perform cross-modal retrieval, existing approaches typically learn a common semantic space from a labeled source set and directly produce common representations in the learned space for the instances in a target set. These methods commonly require that the instances of both two sets share the same classes. Intuitively, they may not generalize well on a more practical scenario of zero-shot cross-modal retrieval, that is, the instances of the target set contain unseen classes that have inconsistent semantics with the seen classes in the source set. Inspired by zero-shot learning, we propose a novel model called ternary adversarial networks with self-supervision (TANSS) in this paper, to overcome the limitation of the existing methods on this challenging task. Our TANSS approach consists of three paralleled subnetworks: 1) two semantic feature learning subnetworks that capture the intrinsic data structures of different modalities and preserve the modality relationships via semantic features in the common semantic space; 2) a self-supervised semantic subnetwork that leverages the word vectors of both seen and unseen labels as guidance to supervise the semantic feature learning and enhances the knowledge transfer to unseen labels; and 3) we also utilize the adversarial learning scheme in our TANSS to maximize the consistency and correlation of the semantic features between different modalities. The three subnetworks are integrated in our TANSS to formulate an end-to-end network architecture which enables efficient iterative parameter optimization. Comprehensive experiments on three cross-modal datasets show the effectiveness of our TANSS approach compared with the state-of-the-art methods for zero-shot cross-modal retrieval.
Xing Xu 0001, Huimin Lu 0001, Jingkuan Song, Yang Yang 0002, Heng Tao Shen, Xuelong Li 0001
IEEE Trans. Cybern.2
2019 CAN: Contextual Aggregating Network for Semantic Segmentation
abstract
Fully convolutional neural networks (FCNs) have shown great success in dense estimation tasks. One key pillar of such progress is mining multi-scale context cues from features in different convolutional layers. This paper introduces contextual aggregating network(CAN), a generic convolutional feature ensembling framework for semantic segmentation. Our framework first captures multi-scale contextual clues by concatenating multi-level feature representation, which carries both coarse semantics and fine details. Then it adaptively integrates stacked features to perform dense pixel estimation. The proposed CAN is trainable end-to-end, and allows us to fully investigate multi-scale context information embedded in images. The experiments show the promising results of our method on PASCAL VOC 2012 and Cityscapes dataset.
Dechun Cong, Quan Zhou 0004, Xiaofu Wu, Suofei Zhang, Weihua Ou, Huimin Lu 0001
ICASSP7
2019 A 6-DOF Telexistence Drone Controlled by a Head Mounted Display
abstract
Recently, a new form of telexistence is achieved by recording images with cameras on an unmanned aerial vehicle (UAV) and displaying them to the user via a head mounted display (HMD). A key problem here is how to provide a free and natural mechanism for the user to control the viewpoint and watch a scene. To this end, we propose an improved rate-control method with an adaptive origin update (AOU) scheme. Without the aid of any auxiliary equipment, our scheme handles the self-centering problem. In addition, we present a full 6-DOF viewpoint control method to manipulate the motion of a stereo camera, and we build a real prototype to realize this by utilizing a pan-tilt-zoom (PTZ) which not only provides 2-DOF to the camera but also compensates the jittering motion of the UAV to record more stable image streams.
Xingyu Xia, Chi-Man Pun, Yang Yang 0002, Huimin Lu 0001, Hao Gao 0005, Feng Xu 0005
VR5
2019 Touch switch sensor for cognitive body sensor networks
Yujie Li 0001, Huimin Lu 0001, Hyoungseop Kim, Seiichi Serikawa
Comput. Commun.2
2019 Data offloading in cache-enabled cross-haul networks
Haoran Mei, Huimin Lu 0001, Limei Peng
Comput. Commun.2
2019 Recognition of surrounding environment from electric wheelchair videos based on modified YOLOv2
Yuki Sakai, Huimin Lu 0001, Joo Kooi Tan, Hyoungseop Kim
Future Gener. Comput. Syst.2
2019 Cognitive data science methods and models for engineering applications
Arun Kumar Sangaiah, Mu-Yen Chen, Huimin Lu 0001, Francesco Mercaldo
Soft Comput.4
2019 Deep adversarial metric learning for cross-modal retrieval
Xing Xu 0001, Li He 0001, Huimin Lu 0001, Lianli Gao, Yanli Ji
World Wide Web3
2019 Dilated-aware discriminative correlation filter for visual tracking
Guoxia Xu, Hu Zhu, Lizhen Deng, Lixin Han, Yujie Li 0001, Huimin Lu 0001
World Wide Web6
2019 Multi-scale deep context convolutional neural networks for semantic segmentation
Quan Zhou 0004, Guangwei Gao, Weihua Ou, Huimin Lu 0001, Longin Jan Latecki
World Wide Web5
2018 Dual Learning for Visual Question Generation
abstract
Recently, automatic answering of visually related questions (VQA) has gained a lot of attention in computer vision community. However, there is little work on automatically generating questions for images (VQG). Actually, VQG itself closes the loop to question-answering and diverse questions, which is useful to the research on VQA. Motivated by the assumption that learning to answer questions may boost the question generation, in this paper, we introduce the VQA task as the complementary of our primary VQG task, and propose a novel model that uses dual learning framework to jointly learn the dual tasks. In the framework, we devise an agent for VQG and VQA with pre-trained models respectively, and the learning tasks of the two agents form a closed loop, whose objectives are optimized together to guide each other via a reinforcement learning process. Specific rewards for each task are designed to update the models of the agents with policy gradient method. The relation of these two tasks can be exploited to further improve the performance of the primary VQG task. Extensive experiments conducted on two large-scale datasets show that the proposed method is capable to generate grounded visual questions of sufficient coverage and outperforms previous VQG methods on standard measures.
Xing Xu 0001, Jingkuan Song, Huimin Lu 0001, Li He 0001, Yang Yang 0002, Fumin Shen
ICME3
2018 BrainNets: Human Emotion Recognition Using an Internet of Brian Things Platform
abstract
Human wearable helmet is a useful tool for monitoring the status of miners in the mining industry. However, there is little research regarding human emotion recognition in an extreme environment. In this paper, an emotional state evoked paradigm is designed to identify the brain area where the emotion feature is most evident. Next, the correct electrode position is determined for the collection of the negative emotion by the electroencephalograph (EEG) based on the international 10-20 system of electrode placement. And then, a fusion algorithm of the anxiety level is proposed to evaluate the person's mental state using the θ, α, and β rhythms of an EEG. Experiments demonstrate that the position Fp2 is the best electrode position for obtaining the anxiety level parameter. The most visible EEG changes appear within the first two seconds following stimulation. The amplitudes of the θ rhythm increase most significantly in the negative emotional state.
Huimin Lu 0001, Hyoungseop Kim, Yujie Li 0001, Yin Zhang 0002
IWCMC1
2018 Modal-adversarial Semantic Learning Network for Extendable Cross-modal Retrieval
abstract
Cross-modal retrieval, e.g., using an image query to search related text and vice-versa, has become a highlighted research topic, to provide flexible retrieval experience across multi-modal data. Existing approaches usually consider the so-called non-extendable cross-modal retrieval task. In this task, they learn a common latent subspace from a source set containing labeled instances of image-text pairs and then generate common representation for the instances in a target set to perform cross-modal matching. However, these method may not generalize well when the instances of the target set contains unseen classes since the instances of both the source and target set are assumed to share the same range of classes in the non-extensive cross-modal retrieval task. In this paper, we consider a more practical issue of extendable cross-modal retrieval task where instances in source and target set have disjoint classes. We propose a novel framework, termed Modal-adversarial Semantic Learning Network (MASLN), to tackle the limitation of existing methods on this practical task. Specifically, the proposed MASLN consists two subnetworks of cross-modal reconstruction and modal-adversarial semantic learning. The former minimizes the cross-modal distribution discrepancy by reconstructing each modality data mutually, with the guidelines of class embeddings as side information in the reconstruction procedure. The latter generates semantic representation to be indiscriminative for modalities, while to distinguish the modalities from the common representation via an adversarial learning mechanism. The two subnetworks are jointly trained to enhance the cross-modal semantic consistency in the learned common subspace and the knowledge transfer to instances in the target set. Comprehensive experiment on three widely-used multi-modal datasets show its effectiveness and robustness on both non-extendable and extendable cross-modal retrieval task.
Xing Xu 0001, Jingkuan Song, Huimin Lu 0001, Yang Yang 0002, Fumin Shen, Zi Huang
ICMR3
2018 Domain Invariant Subspace Learning for Cross-Modal Retrieval
Chenlu Liu, Xing Xu 0001, Yang Yang 0002, Huimin Lu 0001, Fumin Shen, Yanli Ji
MMM (2)4
2018 Automatic road detection system for an air-land amphibious car drone
Yujie Li 0001, Huimin Lu 0001, Yoshiki Nakayama, Hyoungseop Kim, Seiichi Serikawa
Future Gener. Comput. Syst.2
2018 Low illumination underwater light field images reconstruction using deep convolutional neural networks
Huimin Lu 0001, Yujie Li 0001, Tomoki Uemura, Hyoungseop Kim, Seiichi Serikawa
Future Gener. Comput. Syst.1
2018 Loss-Tolerant Event Communications Within Industrial Internet of Things by Leveraging on Game Theoretic Intelligence
abstract
Internet of Things (IoT) is one of the key technologies paving the way for the next industrial revolution named as Industry 4.0, since it promises to realize smarter factories by optimizing costs and productivity. Traditionally, the adopted communication protocols among the sensors are required to manage the large scale of the infrastructure in terms on the high number of interconnected nodes and the massive volume of exchanged data. However, due to the key role of those Industrial IoT in exchanging business critical data, such protocols need to also provide high resiliency guarantees to the message exchange, with as few delivery misses as possible. The publish/subscribe interaction pattern and the protocol implementing it are a technically sound approach for achieving scalability and elasticity, thanks to their intrinsic decoupling among the interacting nodes. However, they are often unsuitable in their current form, because they provide only best-effort delivery guarantees, or they adopt naive solutions to achieve resilient communication, especially when wireless networks are used. This paper presents a clustered lightweight gossiping algorithm for resilient event based communications among the sensors, without requiring a pre-deployed brokering infrastructure supporting the adopted publish/subscribe protocol. A simulation-based assessment has been performed in order to empirically show the improvements in terms of successfully delivered notification without the excessive costs of the state-of-the-art solutions available in the literature.
Christian Esposito 0001, Massimo Ficco, Aniello Castiglione, Francesco Palmieri 0002, Huimin Lu 0001
IEEE Internet Things J.5
2018 Motor Anomaly Detection for Unmanned Aerial Vehicles Using Reinforcement Learning
abstract
Unmanned aerial vehicles (UAVs) are used in many fields including weather observation, farming, infrastructure inspection, and monitoring of disaster areas. However, the currently available UAVs are prone to crashing. The goal of this paper is the development of an anomaly detection system to prevent the motor of the drone from operating at abnormal temperatures. In this anomaly detection system, the temperature of the motor is recorded using DS18B20 sensors. Then, using reinforcement learning, the motor is judged to be operating abnormally by a Raspberry Pi processing unit. A specially built user interface allows the activity of the Raspberry Pi to be tracked on a Tablet for observation purposes. The proposed system provides the ability to land a drone when the motor temperature exceeds an automatically generated threshold. The experimental results confirm that the proposed system can safely control the drone using information obtained from temperature sensors attached to the motor.
Huimin Lu 0001, Yujie Li 0001, Shenglin Mu, Dong Wang 0004, Hyoungseop Kim, Seiichi Serikawa
IEEE Internet Things J.1
2018 PEA: Parallel electrocardiogram-based authentication for smart healthcare systems
Yin Zhang 0002, Raffaele Gravina, Huimin Lu 0001, Massimo Villari, Giancarlo Fortino
J. Netw. Comput. Appl.3
2018 Non-uniform de-Scattering and de-Blurring of Underwater Images
Yujie Li 0001, Huimin Lu 0001, Kuanching Li, Hyoungseop Kim, Seiichi Serikawa
Mob. Networks Appl.2
2018 Editorial: Artificial Intelligence for Mobile Robotic Networks
Huimin Lu 0001, Li He 0001, Quan Zhou 0004, ZongYuan Ge
Mob. Networks Appl.1
2018 Extraction of GGO Candidate Regions on Thoracic CT Images using SuperVoxel-Based Graph Cuts for Healthcare Systems
Huimin Lu 0001, Masashi Kondo, Yujie Li 0001, Joo Kooi Tan, Hyoungseop Kim, Seiichi Murakami, Takotoshi Aoki, Shoji Kido
Mob. Networks Appl.1
2018 Brain Intelligence: Go beyond Artificial Intelligence
Huimin Lu 0001, Yujie Li 0001, Min Chen 0003, Hyoungseop Kim, Seiichi Serikawa
Mob. Networks Appl.1
2018 Anxiety Level Detection Using BCI of Miner's Smart Helmet
Mei Wang 0002, Songzhi Zhang, Yuanjie Lv, Huimin Lu 0001
Mob. Networks Appl.4
2018 Editorial: Intelligent Industrial IoT Integration with Cognitive Computing
Yin Zhang 0002, Limei Peng, Yi Sun 0006, Huimin Lu 0001
Mob. Networks Appl.4
2018 Wavelet energy entropy and linear regression classifier for detecting abnormal breasts
Yi Chen 0023, Yin Zhang 0002, Huimin Lu 0001, Xian-Qing Chen, Jianwu Li, Shuihua Wang
Multim. Tools Appl.3
2018 FDCNet: filtering deep convolutional network for marine organism classification
Huimin Lu 0001, Yujie Li 0001, Tomoki Uemura, ZongYuan Ge, Xing Xu 0001, Li He 0001, Seiichi Serikawa, Hyoungseop Kim
Multim. Tools Appl.1
2018 Multisensor Image Fusion and Enhancement in Spectral Total Variation Domain
abstract
Most existing image fusion methods assume that at least one input image contains high-quality information at any place of an observed scene. Thus, these fusion methods will fail if every input image is degraded. To address this issue, this study proposes a novel fusion framework that integrates image fusion based on spectral total variation (TV) method and image enhancement. For spatially varying multiscale decompositions generated by the spectral TV framework, this study verifies that the decomposition components can be modeled efficiently by tailed α-stable-based random variable distribution (TRD) rather than the commonly used Gaussian distribution. Consequently, salience and match measures based on TRD are proposed to fuse each sub-band decomposition. The spatial intensity information is also adopted to fuse the remainder of the image decomposition components. A sub-band adaptive gain function family based on TV spectrum and space variation is constructed for fused multiscale decompositions to enhance fused image simultaneously. Finally, numerous experiments with various multisensor image pairs are conducted to evaluate the proposed method. Experimental results show that even if the input images are degraded, the fused image obtained by the proposed method achieves significant improvement in terms of edge details and contrast while extracting the main features of the input images, thereby achieving better performance compared with the state-of-the-art methods.
Wenda Zhao 0003, Huimin Lu 0001, Dong Wang 0004
IEEE Trans. Multim.2
2018 Mobile Intelligence Assisted by Data Analytics and Cognitive Computing
Yin Zhang 0002, Huimin Lu 0001, Haider Abbas
Wirel. Commun. Mob. Comput.2
2017 Unsupervised cross-modal retrieval through adversarial learning
abstract
The core of existing cross-modal retrieval approaches is to close the gap between different modalities either by finding a maximally correlated subspace or by jointly learning a set of dictionaries. However, the statistical characteristics of the transformed features were never considered. Inspired by recent advances in adversarial learning and domain adaptation, we propose a novel Unsupervised Cross-modal retrieval method based on Adversarial Learning, namely UCAL. In addition to maximizing the correlations between modalities, we add an additional regularization by introducing adversarial learning. In particular, we introduce a modality classifier to predict the modality of a transformed feature. This can be viewed as a regularization on the statistical aspect of the feature transforms, which ensures that the transformed features are also statistically indistinguishable. Experiments on popular multimodal datasets show that UCAL achieves competitive performance compared to state of the art supervised cross-modal retrieval methods.
Li He 0001, Xing Xu 0001, Huimin Lu 0001, Yang Yang 0002, Fumin Shen, Heng Tao Shen
ICME3
2017 Hearing Loss Detection in Medical Multimedia Data by Discrete Wavelet Packet Entropy and Single-Hidden Layer Neural Network Trained by Adaptive Learning-Rate Back Propagation
Shuihua Wang, Sidan Du, Yang Li 0063, Huimin Lu 0001, Ming Yang 0011, Bin Liu 0043, Yudong Zhang 0001
ISNN (2)4
2017 Wound intensity correction and segmentation with convolutional neural networks
abstract
Summary Wound area changes over multiple weeks are highly predictive of the wound healing process. A big data eHealth system would be very helpful in evaluating these changes. We usually analyze images of the wound bed for diagnosing injury. Unfortunately, accurate measurements of wound region changes from images are difficult. Many factors affect the quality of images, such as intensity inhomogeneity and color distortion. To this end, we propose a fast level set model‐based method for intensity inhomogeneity correction and a spectral properties‐based color correction method to overcome these obstacles. State‐of‐the‐art level set methods can segment objects well. However, such methods are time‐consuming and inefficient. In contrast to conventional approaches, the proposed model integrates a new signed energy force function that can detect contours at weak or blurred edges efficiently. It ensures the smoothness of the level set function and reduces the computational complexity of re‐initialization. To increase the speed of the algorithm further, we also include an additive operator‐splitting algorithm in our fast level set model. In addition, we consider using a camera, lighting, and spectral properties to recover the actual color. Numerical synthetic and real‐world images demonstrate the advantages of the proposed method over state‐of‐the‐art methods. Experimental results also show that the proposed model is at least twice as fast as methods used widely. Copyright © 2016 John Wiley & Sons, Ltd.
Huimin Lu 0001, Bin Li 0006, Junwu Zhu, Yujie Li 0001, Yun Li 0010, Xing Xu 0001, Li He 0001, Xin Li 0034, Jianru Li, Seiichi Serikawa
Concurr. Comput. Pract. Exp.1
2017 Underwater Optical Image Processing: a Comprehensive Review
Huimin Lu 0001, Yujie Li 0001, Yudong Zhang 0001, Min Chen 0003, Seiichi Serikawa, Hyoungseop Kim
Mob. Networks Appl.1
2016 Registration of Point Clouds Based on the Ratio of Bidirectional Distances
abstract
Despite the fact that original Iterative Closest Point(ICP) algorithm has been widely used for registration, itcannot tackle the problem when two point clouds are par-tially overlapping. Accordingly, this paper proposes a ro-bust approach for the registration of partially overlappingpoint clouds. Given two initially posed clouds, it firstlybuilds up bilateral correspondence and computes bidirec-tional distances for each point in the data shape. Based onthe ratio of bidirectional distances, the exponential functionis selected and utilized to calculate the probability value,which can indicate whether the point pair belongs to theoverlapping part or not. Subsequently, the probability val-ue can be embedded into the least square function for reg-istration of partially overlapping point clouds and a novelvariant of ICP algorithm is presented to obtain the optimalrigid transformation. The proposed approach can achievegood registration of point clouds, even when their overlappercentage is low. Experimental results tested on public da-ta sets illustrate its superiority over previous approaches onrobustness.
Jihua Zhu, Di Wang 0006, Xiuxiu Bai, Huimin Lu 0001, Congcong Jin, Zhongyu Li 0002
3DV4
2016 Underwater image descattering and quality assessment
abstract
Vision-based underwater navigation and object detection requires robust computer vision algorithms to operate in turbid water. Many conventional methods aimed at improving visibility in low turbid water. In this paper, we propose a novel contrast enhancement to enhance high turbid underwater images using descattering and color correction. The proposed enhancement method removes the scatter and preserves colors. In addition, as a rule to compare the performance of different image enhancement algorithms, a more comprehensive image quality assessment index Qu is proposed. The index combines the benefits of SSIM index and color distance index. Experimental results show that the proposed approach statistically outperforms state-of-the-art general purpose underwater image contrast enhancement algorithms. The experiment also demonstrated that the proposed method performs well for image classification.
Huimin Lu 0001, Yujie Li 0001, Xing Xu 0001, Li He 0001, Yun Li 0010, Donald G. Dansereau, Seiichi Serikawa
ICIP1
2016 Super Resolving of the Depth Map for 3D Reconstruction of Underwater Terrain Using Kinect
abstract
In recent years, sonar has been widely used for restoring the underwater terrain. Sonar imaging has the benefits such as long-range photographing, robust for turbidity water. However, it is not suitable for short-range imaging. Meanwhile, it also cannot meet the need of mining machine. Therefore, it is important to develop a 3D reconstruction method for short-range imaging. In this paper, we propose a Kinect-based underwater 3D image reconstruction method. To overcome the drawbacks of low accuracy of depth maps, we propose a novel super-resolution (SR) method, which uses the underwater dark channel prior dehazing, weight guided image SR, and inpainting. The proposed method considered the influence of mud sediments in water, it performs better than the traditional methods. The experimental results demonstrated that, after inpainting, dehazing and the super-resolution, it can obtain high accuracy depth maps.
Yu Nakagawa, Keita Kihara, Ryunosuke Tadoh, Seiichi Serikawa, Huimin Lu 0001, Yudong Zhang 0001, Yujie Li 0001
ICPADS5
2016 Learning unified binary codes for cross-modal retrieval via latent semantic hashing
Xing Xu 0001, Li He 0001, Atsushi Shimada 0001, Rin-Ichiro Taniguchi, Huimin Lu 0001
Neurocomputing5
2016 Underwater image enhancement method using weighted guided trigonometric filtering and artificial light correction
Huimin Lu 0001, Yujie Li 0001, Xing Xu 0001, Jian-Ru Lin, Zhifei Liu, Xin Li 0034, Jianmin Yang, Seiichi Serikawa
J. Vis. Commun. Image Represent.1
2016 Single image dehazing through improved atmospheric light estimation
Huimin Lu 0001, Yujie Li 0001, Shota Nakashima, Seiichi Serikawa
Multim. Tools Appl.1
2015 Single underwater image descattering and color correction
abstract
Absorption, scattering, and color distortion are three major issues in underwater optical imaging. Light rays traveling through water are scattered and absorbed according to their wavelength. Scattering is caused by large suspended particles that degrade optical images captured underwater. Color distortion occurs because different wavelengths are attenuated to different degrees in water; consequently, images of ambient underwater environments are dominated by a bluish tone. In the present paper, we propose a novel underwater imaging model that compensates for the attenuation discrepancy along the propagation path. In addition, we develop a fast weighted guided normalized convolution domain filtering algorithm for enhancing underwater optical images in shallow oceans. The enhanced images are characterized by a reduced noised level, better exposure in dark regions, and improved global contrast, by which the finest details and edges are enhanced significantly.
Huimin Lu 0001, Yujie Li 0001, Seiichi Serikawa
ICASSP1
2015 Underwater Image Devignetting and Colour Correction
Yujie Li 0001, Huimin Lu 0001, Seiichi Serikawa
ICIG (3)2
2015 Real-Time Underwater Image Contrast Enhancement Through Guided Filtering
Huimin Lu 0001, Yujie Li 0001, Xuelong Hu, Seiichi Serikawa
ICIG (3)1
2014 Underwater scene enhancement using weighted guided median filter
abstract
We present a novel method of enhancing shallow ocean optical images or videos using weighted guided median filter and wavelength properties. Absorption, scattering and color distortion are three major distortion issues for underwater optical imaging. Light rays traveling through water are scattered and absorbed depending on the wavelength. Scattering is caused by large suspended particles, as in turbid water that contains abundant particles, which causes the degradation of the image. Color distortion occurs because different wavelengths are attenuated to different degrees in water, causing ambient underwater environments to be dominated by a bluish tone. Our key contributions are proposed include a novel shallow water imaging model that compensates for the attenuation discrepancy along the propagation path and an effective underwater scene enhancement scheme. The recovered images are characterized by a reduced noised level, better exposure of the dark regions, and improved global contrast where the finest details and edges are enhanced significantly.
Huimin Lu 0001, Seiichi Serikawa
ICME1
2013 Underwater image enhancement using guided trigonometric bilateral filter and fast automatic color correction
abstract
This paper describes a novel method to enhance underwater optical images by guided trigonometric bilateral filters and color correction. Scattering and color distortion are two major problems of distortion for underwater optical imaging. Scattering is caused by large suspended particles, like fog or turbid water which contains abundant particles. Color distortion corresponds to the varying degrees of attenuation encountered by light traveling in the water with different wavelengths, rendering ambient underwater environments dominated by a bluish tone. Our key contributions are proposed a new underwater model to compensate the attenuation discrepancy along the propagation path, and to propose a fast guided trigonometric bilateral filtering enhancing algorithm and a novel fast automatic color enhancement algorithm. The enhanced images are characterized by reduced noised level, better exposedness of the dark regions, improved global contrast while the finest details and edges are enhance significantly. In addition, our enhancement method is comparable to higher quality than the state-of-the-art methods by assuming in the latest image evaluation systems.
Huimin Lu 0001, Yujie Li 0001, Seiichi Serikawa
ICIP1
2013 Underwater optical image dehazing using guided trigonometric bilateral filtering
abstract
This paper describes a novel method to enhance underwater optical images by dehazing. Scattering and color change are two major problems of distortion for underwater imaging. Scattering is caused by large suspended particles, like fog or turbid water which contains abundant particles, plankton etc. Color change corresponds to the varying degrees of attenuation encountered by light traveling in the water with different wavelengths, rendering ambient underwater environments dominated by a bluish tone. Our key contribution is to propose a fast image and video dehazing algorithm, to compensate the attenuation discrepancy along the propagation path, and to take the influence of the possible presence of an artificial lighting source into consideration. The enhanced images are characterized by reduced noised level, better exposedness of the dark regions, improved global contrast while the finest details and edges are enhance significantly. In addition, our enhancement method is comparable to higher quality than the state-of-the-art methods.
Huimin Lu 0001, Yujie Li 0001, Akira Yamawaki 0002, Seiichi Serikawa
ISCAS1
2013 Cross Depth Image Filter-Based Natural Image Matting
abstract
In this paper we propose a novel explicit image filter called guided depth image filter for natural image matting. Different from the traditional matting model, the guided image filter computes the filtering output by considering the content of a depth image. The guided depth image filter can be used as an edge-preserving smoothing operator like bilateral filter, but has better behaviors near edges. The proposed filter by using nonlocal neighborhoods, and contribute a simple and fast algorithm giving competitive results. Experimental results indicate that our matting results are comparable to the state of the art methods.
Yujie Li 0001, Huimin Lu 0001, Seiichi Serikawa
SNPD2
2013 Multiframe Medical Images Enhancement on Dual Tree Complex Wavelet Transform Domain
abstract
As a novel of multi-resolution analysis tool, dual tree complex wavelet transform (DTCWT) provides flexible multiresolution and directional expansion for medical image fusion. In this paper, a novel fusion method for multiframe medical images based on DTCWT is proposed. Contrary to present fusion methods, the proposed algorithm extends the 2 inputs model to multi inputs model. We take 8 input unclearly medical images as input. Then, we estimate the decomposed coefficients through the weighted soft threshold in each image, and choose the weighted coefficients. After DTCWT reconstruction, the clear image is gotten. During abundant experiments, we evaluate the proposed method both human visual and quantitative analysis. Compare with the-state-of-the-art methods, the new strategy for attaining image fusion with satisfactory performance.
Huimin Lu 0001, Yujie Li 0001, Shota Nakashima
SNPD1
2013 Distance Measurement with a General 3D Camera by Using a Modified Phase Only Correlation Method
abstract
This paper proposed a new approach of 3D measurement using a home use 3D camera. Stereo image measurement is different from the other active measurement method like using a laser range finder or an ultra-sonic sensor. It is a passive method, which acquire the object as two images, and then calculate the distance information form the two images according to the principle of triangulation. Ordinarily, such of two images is taken by a special use stereo camera which were settled with a precisely accuracy so that to keep a parallel optical axes, depth of focus and so on. But the accuracy settings of a home use 3D camera cannot satisfy such a requirement. In this paper, phase-only-correlation method which can yield sub-pixel accuracy is used, also with some modification and new approach. The simulation shows a good result.
Yujie Li 0001, Huimin Lu 0001, Seiichi Serikawa
SNPD3
2012 Multimodal Medical Image Fusion in Modified Sharp Frequency Localized Contourlet Domain
abstract
As a novel of multi-resolution analysis tool, the modified sharp frequency localized contour let transforms (MSFLCT) provides flexible multiresolution, anisotropy, and directional expansion for medical images. In this paper, we proposed a new fusion rule for multimodal medical images based on MSFLCT. The multimodal medical images are decomposed by MSFLCT. For the high-pass sub band, the weighted sum modified laplacian (WSML) method is used for choose the high frequency coefficients. For the low pass sub band, the maximum local energy (MLE) method is combined with "region" idea for low frequency coefficient selection. The final fusion image is obtained by applying inverse MSFLCT to fused low pass and high pass sub bands. Abundant experiments have been made on groups of multimodality datasets, both human visual and quantitative analysis show that the new strategy for attaining image fusion with satisfactory performance.
Seiichi Serikawa, Huimin Lu 0001, Yujie Li 0001
SNPD2
2012 Maximum Local Energy Based Multifocus Image Fusion in Mirror Extended Curvelet Transform Domain
abstract
In this paper, we firstly propose the maximum local energy (MLE) method to calculate the low frequency coefficients of images and compare the results with those of mirror extended curve let transform, which enhance the edge features and details of images. An image fusion step was performed as follows: First, we obtained the coefficients of two different types of images through mirror extended curve let transform. Second, we selected the low frequency coefficients by maximum local energy and obtaining the high-frequency coefficients using the absolute maximum value (AMV) method. Finally, the fused image was obtained by performing an inverse mirror extended curve let transform. In addition to human vision analysis, the images were also compared through quantitative analysis. multifocus images were used in the experiments to compare the results among the beyond wavelets. The numerical experiments reveal that maximum local energy is a new strategy for attaining image fusion with satisfactory performance.
Huimin Lu 0001, Yujie Li 0001, Seiichi Serikawa
SNPD2