Bin Song 0001

dblp:09/2085-1 · DBLP profile ↗
← Back
109ranked-venue papers
2as first author
63since 2021 · last 2027
0000-0002-8096-3370ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 34 · 1 first-author · 23 since 2021Computer networks · 29 · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 1 first-author · 7 since 2021Systems, architecture and hardware · 10 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 4 since 2021Databases, data management, data science and information retrieval · 7 · 7 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2027 Combating label noise in sequential recommendation: A symmetric cross-entropy and contrastive denoising framework
Yangcheng Huang, Bin Song 0001
Expert Syst. Appl.2
2026 A knowledge-driven model selection and resource management method with information entropy
Dan Wang 0002, Zhenshen Liang, Bin Song 0001
Sci. China Inf. Sci.3
2026 CDMR2F: Correlation-guided Denoising Multimodal Robust Fusion Framework for multimodal recommendation
Rong Cheng, Bin Song 0001
Expert Syst. Appl.2
2026 Uncertainty-aware neural-symbolic task planning with balanced exploration and execution for domestic service robot
Bin Song 0001
Expert Syst. Appl.3
2026 EAI-DMCU: Evolutionary algorithm-inspired diffusion model for concept unlearning
Yixin Yao, Bin Song 0001
Expert Syst. Appl.3
2026 Color-Shape Disentangled Representation Learning with channel augmentation in interactive image retrieval
Chen Chen 0128, Bin Song 0001
Neurocomputing2
2026 TRTMPM: Tight-coupling reasoning task message passing model for multi-modal knowledge graph completion
Bo Li 0034, Bin Song 0001
Neurocomputing2
2026 Dynamic LLM Node Selection in AIoT Networks: A Compatibility-Driven Diffusion Reinforcement Learning Approach
abstract
The growing complexity of Artificial Internet of Things (AIoT) applications, driven by massive data generation and intricate interaction demands, necessitates intelligent decision-making that surpasses basic automation. While traditional machine learning is already integrated into edge devices, the evolving AIoT paradigm still faces severe challenges in complex tasks. Large language models (LLMs) emerge as a promising solution for diverse AIoT tasks, yet their deployment on resource-constrained edge networks faces two major challenges: heterogeneous resource limitations and the complexity of dynamic model selection. To address these issues, we propose an edge computing framework that leverages a compatibility evaluation mechanism and a diffusion-based deep reinforcement learning (DRL) algorithm for dynamic LLM service node selection. This approach enhances deployment efficiency and adaptability in heterogeneous environments. Our diffusion-based DRL algorithm continuously learns from environmental states to make robust, real-time selection decisions, assigning each task to the most suitable LLM instance to maximize long-term task completion quality (TCQ). Experimental results demonstrate that our algorithm significantly outperforms baseline methods in TCQ, convergence speed, and stability. Furthermore, sensitivity and robustness analyses demonstrate the effectiveness of the proposed method under resource fluctuations, varying node scales, and different reward preferences.
Dan Wang 0002, Chi Cao, Bin Song 0001
IEEE Internet Things J.3
2026 Information entropy-guided knowledge graph recommendation
Jie Guo 0008, Yunfei Zhao 0004, Bin Song 0001
Knowl. Based Syst.5
2026 Fine-tuning large language models in federated learning with fairness-aware prompt selection
Yalan Jiang, Bin Song 0001
Neural Networks3
2026 A Structural Knowledge Enhanced Re-ranking method with large language models for temporal knowledge graph prediction
Bo Li 0001, Bin Song 0001, Tiantian He 0001, Yew-Soon Ong
Pattern Recognit.2
2026 Robust Multi-Stream Massive MIMO Satellite Systems Based on Statistical CSI
Hangsong Yan, Alexei E. Ashikhmin, Hong Yang 0001, Bin Song 0001, Shu Sun 0001
IEEE Trans. Commun.4
2026 Adaptive Graph Convolution With Diffusion Models for Multimodal Recommendation
Jie Guo 0008, Bin Song 0001
IEEE Trans. Knowl. Data Eng.3
2026 TextBridge: A Text-Centered Framework for Enhanced Multimodal Integration and Retrieval
abstract
Despite significant advancements in multimodal pre-training, effectively integrating and using latent semantic information across multiple modalities remains a challenge. In this paper, we introduce TextBridge, a text-centered framework that uses the text modality as a semantic anchor to guide cross-modal integration and alignment. TextBridge employs frozen encoders from state-of-the-art pre-trained models and introduces an innovative modality bridge module that enhances semantic alignment and reduces redundancy among different modal features. The framework also incorporates a multi-projection text feature fusion method, enhancing the alignment and integration of text features from diverse modalities into a cohesive semantic representation. To optimize the integration of multimodal information, we make the text encoder trainable and use a text-centered contrastive loss function to enhance the model's ability to capture complementary information across modalities. Extensive experiments on the M5Product dataset demonstrate that TextBridge significantly outperforms the SCALE model in mean average precision (mAP) and precision (Prec), underscoring its effectiveness in multimodal retrieval tasks.
Jie Guo 0008, Haiyang Jing, Bin Song 0001
IEEE Trans. Multim.4
2026 LETTER: Self-Harmonized Representation Learning for Multimodal Recommendation
abstract
Multimodal recommender systems try to integrate multimedia data (images, texts, etc.) with user-item historical records to better model user preference. However, most previous methods largely ignored the underlying fine-grained attribute features of items, which makes it difficult to fully explore users' nuanced attention across individual and combined attributes, resulting in low recommendation performance. To address these issues, this paper proposes a novel and effective self-harmonized representation learning network for multimodal recommendation, named LETTER. LETTER has the ability to effectively optimize the user and item representations for multimodal recommendation. Specifically, we design a factorized attribute interaction module that captures diverse combinations of item latent attributes using a bilinear pooling strategy. Then a dual graph convolution module is established to learn the modality-specific representations from user-item interactive and item semantic relations. Finally, we design a preference self-harmonization module that adaptively identifies the salient influencing factors of user preference, thus refining user and item representations to improve recommendation accuracy. We conduct extensive experiments on three real-world datasets, demonstrating that LETTER outperforms state-of-the-art multimodal recommendation methods.
Jie Guo 0008, Longyu Wen, Yunfei Zhao 0004, Bin Song 0001, Yuhao Chi
IEEE Trans. Multim.4
2026 RIS-Assisted AirComp FL: Joint Data Allocation and Power Optimization for Reduced Distortion
abstract
Federated learning (FL) ensures privacy by training models on multiple edge devices and aggregating updates at a central base station (BS), while preserving data locally on devices. Over-the-Air Computing (AirComp) enhances efficiency as it enables edge devices to transmit model parameters simultaneously to the BS, leveraging wireless channels’ superposition characteristics to minimize communication overhead and latency. However, challenges such as signal distortion from fading and noise in wireless channels, along with disparities in communication capabilities among edge devices, can hinder the model aggregation process, leading to increased communication errors at the BS and overall performance degradation. To address these issues, we propose a novel reconfigurable intelligent surface (RIS)-assisted AirComp FL framework, where we optimize the transmit power at edge devices and the received power at the BS to effectively mitigate signal distortion issues, and utilize RIS technology to enhance communication reliability and efficiency by addressing straggler effects in wireless channels. We formulate a joint optimization problem for the design of over-the-air transceiver and RIS configurations, which can be solved using Karush-Kuhn-Tucker (KKT) conditions and Sequential Convex Approximation (SCA) principles. Numerical results demonstrate that our proposed method enhances model accuracy by approximately 21.5%, surpassing state-of-the-art approaches and significantly bolstering performance in AirComp FL models.
Yalan Jiang, Bin Song 0001
IEEE Trans. Wirel. Commun.2
2026 Learning When and Where to Handover: A Hierarchical Reinforcement Learning Framework for Dense LEO Satellite Constellations
abstract
The rapid development of low Earth orbit (LEO) satellite constellations has enabled global broadband coverage and low-latency communication, but also brings new challenges to connection continuity and network stability. Due to the ultra-dense deployment and high orbital velocity of LEO satellites, users frequently experience mandatory handovers and are prone to triggering “ping-pong” effects in overlapping coverage areas. To address these issues, we propose a hierarchical reinforcement learning (HRL)-based inter-satellite handover framework that decouples decision-making into two subproblems: when to initiate a handover at the temporal level and which satellite to handover to at the spatial level. To achieve long-term planning, the temporal agent utilizes the proximal policy optimization (PPO) algorithm to determine the handover timing by predicting future satellite load and interference levels. For short-term adaptability, the spatial agent employs the deep Q-network (DQN) to select the target satellite based on a utility function incorporating satellite load, signal quality, and service duration. We implement the proposed method in a Starlink constellation environment. Simulation results show that the proposed HRL-based scheme outperforms traditional and non-hierarchical RL methods in terms of user handover frequency and handover success rate. Furthermore, we validate that the proposed approach significantly improves transmission rate and load balancing, demonstrating its effectiveness in highly dynamic LEO satellite environments.
Bin Song 0001, Yejun Zhou, Pengfei Qin
IEEE Trans. Wirel. Commun.3
2025 RMultiplex200K: Toward Reliable Multimodal Process Supervision for Visual Language Models on Telecommunications
Bin Song 0001
ICCV2
2025 Satellite-assisted 6G wide-area edge intelligence: dynamics-aware task offloading and resource allocation for remote IoT services
Rui Ding 0002, Bin Song 0001
Sci. China Inf. Sci.3
2025 Video-text retrieval based on multi-grained hierarchical aggregation and semantic similarity optimization
Jie Guo 0008, Shujie Lan, Bin Song 0001, Mengying Wang 0003
Neurocomputing3
2025 Frequency-domain keyframe interpolation denoising for text and image-guided video editing acceleration with diffusion models
Chenhao Pang, Chen Chen 0128, Bin Song 0001
Neurocomputing3
2025 Large Language Models and Artificial Intelligence Generated Content Technologies Meet Communication Networks
abstract
Artificial intelligence generated content (AIGC) technologies, with a predominance of large language models (LLMs), have demonstrated remarkable performance improvements in various applications, which have attracted great interests from both academia and industry. Although some noteworthy advancements have been made in this area, a comprehensive exploration of the intricate relationship between AIGC and communication networks remains relatively limited. To address this issue, this article conducts an exhaustive survey from dual standpoints: first, it scrutinizes the integration of LLMs and AIGC technologies within the domain of communication networks and second, it investigates how the communication networks can further bolster the capabilities of LLMs and AIGC. Additionally, this research explores the promising applications along with the challenges encountered during the incorporation of these AI technologies into communication networks. Through these detailed analyses, our work aims to deepen the understanding of how LLMs and AIGC can synergize with and enhance the development of advanced intelligent communication networks, contributing to a more profound comprehension of next-generation intelligent communication networks.
Jie Guo 0008, Meiting Wang, Hang Yin 0007, Bin Song 0001, Yuhao Chi, F. Richard Yu, Chau Yuen
IEEE Internet Things J.4
2025 TSFL-SS: Task-Specific Federated Learning for Resource-Limited Edge Devices
abstract
The widespread deployment of IoT devices in resource-constrained edge environmentsłcharacterized by limited computational power, storage, and network bandwidthłposes significant challenges for federated learning (FL). Traditional FL frameworks inadequately support frequent model updates and complex computations, while existing methods fail to balance computational efficiency and model performance on resource-limited devices, resulting in inefficient training and prohibitive communication costs. To address these limitations, we propose TSFL-SS, a task-specific federated learning framework that integrates three synergistic modules: (1) an Importance-Aware Sub-model Selection (IASS) module dynamically extracts lightweight task-relevant parameters via one-shot low-rank decomposition, enabling edge devices to bypass full-model computations; (2) a Task-Specific Model Partitioning (TSMP) module decouples models into shared encoders (fixed task-agnostic features) and personalized predictors, minimizing redundant transmissions; and (3) a Semantic-Aware Aggregation (SAP) module aligns client updates by clustering parameters with high semantic similarity, mitigating conflicts from non-IID data distributions. Experiments on ResNet18 demonstrate that TSFL-SS achieves a 26.84% average accuracy improvement over FedAvg across both IID and non-IID settings, while reducing communication overhead by 55.44%. These advancements establish TSFL-SS as an efficient, scalable solution for FL in resource-constrained edge environments.
Bin Song 0001
IEEE Internet Things J.3
2025 Transcoding-Enabled Edge Caching Strategy Optimization: A Dual-Timescale Meta-Learning-Based Stackelberg Game Approach
abstract
The explosive growth in video services has significantly strained current mobile network infrastructure, leading to spectrum scarcity, backhaul congestion, and degraded quality of experience. While edge caching has emerged as a promising solution to address these challenges and deliver seamless video playback experience, multiversion edge caching for heterogeneous clients remains challenging due to varying client requirements and network conditions. This article proposes a transcoding-enabled edge caching framework for mobile edge-cloud computing networks. Specifically, we combine video transcoding with edge caching to support both “direct cache hits” and “soft cache hits” for reducing transmission latency. We model this joint caching and resource allocation problem as a Stackelberg game to minimize video transmission latency. To solve this problem, we develop a novel dual timescale model agnostic meta-learning (MAML)-based Stackelberg game (DTMSG) optimization approach that determines the delay-optimal Stackelberg equilibrium (SE) and accelerates convergence. Simulation results demonstrate that our DTMSG optimization algorithm efficiently converges to the SE point, maximizing the utility function of the MVNO and BSs while reducing the average video transmission delay.
Dan Wang 0002, Keke Zhu, Bin Song 0001, F. Richard Yu
IEEE Internet Things J.3
2025 Rect-ViT: Rectified attention via feature attribution can improve the adversarial robustness of Vision Transformers
Xu Kang 0002, Bin Song 0001
Neural Networks2
2025 Service-Oriented Resource Allocation and Task Scheduling for Wi-Fi and Bluetooth Coexistence in Smart Home IoT Systems
abstract
In IoT-enabled smart home scenarios, heterogeneous communication devices such as Bluetooth (BT) and Wi-Fi are widely used in applications such as home automation, remote monitoring, and intelligent device interconnection. However, in such a multi-device coexistence environment, efficiently allocating limited time-frequency resources to mitigate communication interference and enhance system performance has become a critical challenge. To address these issues, this article proposes a comprehensive solution that integrates master selection, resource allocation, and task scheduling to optimize resource utilization and service quality in smart home IoT systems. For device management, we propose a hierarchical entropy weight method (HEWM), considering factors such as device parameters, sensing capabilities, communication performance, and device interoperability. This method ensures efficient and stable selection of the primary device, optimizing network topology and communication efficiency. For resource allocation, we introduce a proximal policy optimization (PPO) algorithm that dynamically adjusts time-frequency resource allocation based on the varying device usage, network load, and communication condition. This adaptive strategy reduces interference between devices and improves system throughput. For task scheduling, we develop a task urgency-based queueing (TUQ) mechanism that prioritizes tasks based on urgency. A task preemption mechanism ensures that high-urgency tasks are processed with minimal delay, enhancing scheduling efficiency and service responsiveness. Simulation results show that the proposed approach significantly outperforms traditional methods in smart home IoT scenarios, achieving higher primary device scores, a 3%–22% improvement in system throughput, and a 5%–36% reduction in task delay.
Tianxu Niu, Bin Song 0001, Xiaojiang Du
ACM Trans. Internet Things3
2025 Viewport Prediction With Unsupervised Multiscale Causal Representation Learning for Virtual Reality Video Streaming
abstract
The rise of the metaverse has driven the rapid development of various applications, such as Virtual Reality (VR) and Augmented Reality (AR). As a form of multimedia in the metaverse, VR video streaming (a.k.a., VR spherical video streaming and 360$^{\circ }$video streaming) can provide users with a 360$^{\circ }$immersive experience. Generally, transmitting VR video requires far more bandwidth than regular videos, which greatly strains existing network transmission. Predicting and selectively streaming VR video in the users' viewports in advance can reduce bandwidth consumption and system latency. However, existing methods either consider only historical viewport-based prediction methods or predict viewports by correlations between visual features of video frames, making it hard to adapt to the dynamics of users and video content. In the meantime, spurious correlations between visual features lead to inaccurate and unreliable prediction results. Hence, we propose an unsupervised multiscale causal representation learning (UMCRL)-based method to predict viewports in VR video streaming, including user preference-based and video content-based viewport prediction models. The former is designed by a position predictor to predict the future users' viewports based on their historical viewports in multiple video frames to adapt to users' dynamic preferences. The latter achieves unsupervised multiscale causal representation learning through an asymmetric causal regressor, used to infer the causalities between local and global-local visual features in video frames, thereby helping the model understand the contextual information in the videos. We embed the causalities in the transformer decoder via causal self-attention for predicting the users' viewports, adapting to the dynamic changes of video content. Finally, combining the results of the two aforementioned models yields the final prediction of the users' viewports. In addition, the QoE of users is satisfied by assigning different bitrates to the tiles in the viewport through a pyramid-based bitrate allocation. The experimental results verify the effectiveness of the method.
Dan Wang 0002, Bin Song 0001
IEEE Trans. Multim.3
2025 Multi-Scale Semantic Communication for Object Detection: Single and Cross-Domain Scenarios
abstract
With the rapid popularity of vision-driven communication applications, object detection has become one of the fundamental techniques for performing practical vision tasks. In traditional communication systems, images are compressed for transmission, reconstructed at the receiver, and then processed by existing object detection algorithms. However, transmitting large amounts of images consumes significant storage and communication resources. To address this challenge, a semantic communication-based image reconstruction scheme has been proposed for object detection, which transmits only the semantic information relevant to image reconstruction. However, this method is prone to losing key information, such as object position and texture details, leading to degraded object detection performance. Additionally, it is sensitive to environmental factors such as weather and lighting, resulting in poor adaptability across multiple scenarios. To address these issues, we propose a multi-scale semantic communication framework for object detection that transmits only multi-scale semantic features relevant to the task and employs decoupling at the receiver to separate positional and classification information of target objects without requiring image reconstruction. To improve adaptability across multiple scenarios, we introduce a cross-domain object detection technique that ensures reliable object detection in new scenarios by optimizing the framework’s multi-scale semantic encoder through domain adversarial learning. Numerical results demonstrate that the proposed framework achieves mean average precision improvements of$15.4\% \sim 38.5\%$over the traditional communication framework within low to medium signal-to-noise ratio regions in additive white Gaussian noise and Rayleigh fading channels.
Jie Guo 0008, Hang Yin 0007, Bin Song 0001, Yuhao Chi, Zhaoyang Zhang 0001, Chau Yuen, Dusit Niyato
IEEE Trans. Wirel. Commun.3
2024 Hardware Latency-Aware Differential Architecture Search: Search for Latency-Friendly Architectures on Different Hardware
abstract
As a result of its low search cost, Differentiable Architecture Search (DARTS) has recently received a lot of interest. Nowadays, most methods based on DARTS only focus on improving a single indicator (e.g., accuracy), making the search process more inclined to complex networks with more robust representational capabilities. Hence, the architectures searched by these methods tend to own high latency, leading to DARTS in some low-latency scenarios or edges with limited computing power, and on-device deployment becomes difficult. To deal with this challenge, we propose a Hardware Latency-Aware Differentiable Search (HL-DARTS) algorithm. This algorithm designs a multi-layer regression network that uses the soft attention mechanism to predict the latency on the corresponding hardware devices, thus adding a differentiable latency loss term based on the DARTS algorithm. We further propose an adaptive constraint amplitude—a mechanism for balancing accuracy and latency while searching for a latency-friendly architecture for a given hardware device. We conduct ablation experiments on different datasets and different hardware devices. The experimental results show that HL-DARTS can find the ideal architecture for different hardware devices and that this architecture is also broadly applicable to various datasets.
Dan Wang 0002, Bin Song 0001
TrustCom5
2024 A credible traffic prediction method based on self-supervised causal discovery
Dan Wang 0002, Bin Song 0001
Sci. China Inf. Sci.3
2024 A prototype-assisted clustered federated learning for big data security and privacy preservation
Yalan Jiang, Dan Wang 0002, Bin Song 0001, Xiaojiang Du
Future Gener. Comput. Syst.3
2024 HDHRFL: A hierarchical robust federated learning framework for dual-heterogeneous and noisy clients
Yalan Jiang, Dan Wang 0002, Bin Song 0001, Shengyang Luo
Future Gener. Comput. Syst.3
2024 RRA-FFSCIL: Inter-intra classes representation and relationship augmentation federated few-shot incremental learning
Yalan Jiang, Dan Wang 0002, Bin Song 0001
Neurocomputing4
2024 Multitask Fine-Grained Feature Mining for Multilabel Remote Sensing Image Classification
abstract
Multilabel remote sensing image classification can provide comprehensive object-level semantic descriptions of remote sensing images. However, most existing methods cannot fully mine the fine-grained features of images and labels, resulting in low classification accuracy. To address this issue, we propose a novel multitask framework for multilabel remote sensing image classification. The framework establishes the class-specific feature extraction as a binary classification auxiliary task to assist the main multilabel classification task, which can improve the model’s local and global feature extraction ability. Meanwhile, the framework updates the label correlation graph using the graph transformer layer to accurately identify label node pairs with potential correlation, which effectively mines the correlation of multiple labels to generate more accurate label co-occurrence embedding for image label prediction. Experimental results on UCM, AID, and DFC15 multilabel datasets show that the proposed method outperforms existing state-of-the-art methods.
Jie Guo 0008, Hao Sun 0033, Jinheng Han, Bin Song 0001, Yuhao Chi, Bingxi Song
IEEE Trans. Geosci. Remote. Sens.4
2024 HSMH: A Hierarchical Sequence Multi-Hop Reasoning Model With Reinforcement Learning
abstract
The incompleteness of knowledge graphs (KGs) negatively impacts the performance of KGs in downstream applications (e.g., recommendation systems and information retrieval). This phenomenon has brought an increasing rise in research related to knowledge graph reasoning. Recently, emerged reinforcement learning (RL)-based multi-hop reasoning methods can infer missing information through multi-hop reasoning according to the existing information in KGs, which has better reasoning performance and interpretability. However, these methods always use relation-entity pairs that have been pre-cropped as the action space of agents for path reasoning, which leads to two problems: 1) insufficient learning and reasoning ability of reasoning models and 2) the hard convergence of the training process of agents. To address these problems, we propose aHierarchicalSequenceMultiHop (HSMH) reasoning framework, which consists of the interactive search reasoning model, local-global knowledge fusion mechanism, and action optimization mechanism. We use interactive search reasoning models to select relations and entities independently, thus fully mining the semantic information of relations and entities and improving the learning and reasoning ability of reasoning models. In the HSMH framework, we design the local-global knowledge fusion and action optimization mechanisms for path reasoning, which can enhance agents' state information and action space. Specifically, the local-global knowledge fusion mechanism is designed to acquire the local knowledge of entities and neighboring relations and the global knowledge about KG structure. This local-global knowledge can improve the learning ability of reasoning models. In addition, the action optimization mechanism can combine the filtered action space and the additional action space for efficient path reasoning for agents. Experimental results on five benchmark datasets show that our proposed HSMH framework comprehensively outperforms the state-of-the-art multi-hop reasoning model.
Dan Wang 0002, Bo Li 0034, Bin Song 0001, Chen Chen 0128, F. Richard Yu
IEEE Trans. Knowl. Data Eng.3
2024 SPACE: Self-Supervised Dual Preference Enhancing Network for Multimodal Recommendation
abstract
Multimodal recommendation is an emerging task with the goal of improving the effectiveness of the recommendation system by utilizing multimodal data (images, texts, etc.). Most previous methods have struggled with the ability to mine item semantic relationships while guaranteeing accurate modeling of user modality preferences, resulting in low recommendation accuracy. To address this issue, this paper proposes a novel and effective Self-suPervised duAl preference enhanCing nEtwork for multimodal recommendation, named SPACE, which further mines user preferences towards historical interactions and multimodal features of items to obtain more precise user and item representation. Specifically, we design an interaction preference enhancing module to learn both interactive and latent semantic relationships between users and items. Then, a modality preference enhancing module is established by introducing self-supervised learning (SSL), which aims to strengthen the role of dominant modality-specific representation of items. Finally, the enhanced interaction and modality representations are fused, and the recommendation performance is largely improved by utilizing dual joint prediction. Extensive experiments are conducted on three real-world datasets, and the simulation results demonstrate that the proposed SPACE model outperforms the state-of-the-art multimodal recommendation methods.
Jie Guo 0008, Longyu Wen, Bin Song 0001, Yuhao Chi, F. Richard Yu
IEEE Trans. Multim.4
2023 Attention-guided Multi-step Fusion: A Hierarchical Fusion Network for Multimodal Recommendation
abstract
The main idea of multimodal recommendation is the rational utilization of the item's multimodal information to improve the recommendation performance. Previous works directly integrate item multimodal features with item ID embeddings, ignoring the inherent semantic relations contained in the multimodal features. In this paper, we propose a novel and effective aTtention-guided Multi-step FUsion Network for multimodal recommendation, named TMFUN. Specifically, our model first constructs modality feature graph and item feature graph to model the latent item-item semantic structures. Then, we use the attention module to identify inherent connections between user-item interaction data and multimodal data, evaluate the impact of multimodal data on different interactions, and achieve early-step fusion of item features. Furthermore, our model optimizes item representation through the attention-guided multi-step fusion strategy and contrastive learning to improve recommendation performance. The extensive experiments on three real-world datasets show that our model has superior performance compared to the state-of-the-art models.
Jie Guo 0008, Hao Sun 0033, Bin Song 0001, F. Richard Yu
SIGIR4
2023 Interpretability for reliable, efficient, and self-cognitive DNNs: From theories to applications
Xu Kang 0002, Jie Guo 0008, Bin Song 0001, Binghuang Cai, Hongyu Sun 0001, Zhebin Zhang
Neurocomputing3
2023 A Deep-Reinforcement-Learning-Based Social-Aware Cooperative Caching Scheme in D2D Communication Networks
abstract
Device-to-device (D2D) caching is becoming prevalent in relieving network congestion. However, there remain challenges in exploring efficient D2D caching strategies due to the diverse user requirements. In this article, we propose a social-aware D2D caching scheme that integrates the concept of social incentive and recommendation with D2D caching decision making. First, we investigate federated learning (FL)-based prediction method to achieve the social-aware in a privacy-preserving manner. Then, the predicted social relationship provides prior knowledge for deep reinforcement learning (DRL) to make optimal D2D caching decisions. The optimization problem of this article is to maximize the data offloading probability, which can be formulated as a Markov decision process. To solve it, we propose a double deep$Q$-learning network (DDQN)-based D2D caching algorithm. Finally, simulation results validate the prediction and convergence performance of the proposed scheme. Besides, the scheme also shows superior caching performance in reducing the average delay and improving overall offloading probability.
Yalu Bai, Dan Wang 0002, Gang Huang 0004, Bin Song 0001
IEEE Internet Things J.4
2023 Cache-Aided MEC for IoT: Resource Allocation Using Deep Graph Reinforcement Learning
abstract
With the growing demand for latency-sensitive and compute-intensive services in the Internet of Things (IoT), multiaccess edge computing (MEC)-enabled IoT is envisioned as a promising technique that allows network nodes to have computing and caching capabilities. In this article, we propose a cache-aided MEC (CA-MEC) offloading framework for joint optimization of communication, computing, and caching (3C) resources in the MEC-enabled IoT. Our goal is to optimize the offloading decision and resource allocation strategy to minimize the system latency subject to dynamic cache capacities and computing resource constraints. We first formulate this optimization problem as a multiagent decision problem, a partially observable Markov decision process (POMDP). Then, the deep graph convolution reinforcement learning (DGRL) method is applied to motivate the agents to learn optimal strategies cooperatively in a highly dynamic environment. Simulations show that our method is highly effective for computation offloading and resource allocation and performs superior results in a large-scale network.
Dan Wang 0002, Yalu Bai, Gang Huang 0004, Bin Song 0001, F. Richard Yu
IEEE Internet Things J.4
2023 Dual-Driven Resource Management for Sustainable Computing in the Blockchain-Supported Digital Twin IoT
abstract
Nowadays, emerging sixth-generation (6G) mobile networks, the Internet of Things (IoT), and mobile-edge computing (MEC) technologies have played significant roles in developing a sustainable computing network. In sustainable computing networks, with the increasing scale of data-driven applications, massive privacy-sensitive data are generated. How to effectively process such data on resource-limited IoT devices is challenging. Although edge intelligence (EI) is designed to maintain an appropriate level of ultradelay reliability, low-latency communication (URLLC), real-time data processing, and security and privacy are concerning. In this article, we propose a novel blockchain-supported hierarchical digital twin IoT (HDTIoT) framework, which combines the digital twin to edge network and adopts blockchain technology to achieve secure and reliable real-time computation. We first propose a data and knowledge dual-driven learning solution to ensure real-time interaction and efficient optimization between the physical and the digital worlds. To improve communication and computation efficiency with data and knowledge dual-driven learning, the optimization goal is to minimize the system delay and energy consumption and ensure system reliability and the learning accuracy of IoT devices. Moreover, we propose a proximal policy optimization (PPO)-based multiagent reinforcement learning (MARL) algorithm to solve the resource allocation (RA) problem. Experimental results show that the proposed RA scheme can improve the efficiency of the HDTIoT system, guarantee learning accuracy, reliability, and security, and make a balance between system delay and energy consumption.
Dan Wang 0002, Bo Li 0034, Bin Song 0001, Khan Muhammad 0001, Xiaokang Zhou
IEEE Internet Things J.3
2023 Black-box attacks on image classification model with advantage actor-critic algorithm in latent space
Xu Kang 0002, Bin Song 0001, Jie Guo 0008, Hao Qin 0001, Xiaojiang Du, Mohsen Guizani
Inf. Sci.2
2023 A deep learning-based approach for fault diagnosis of current-carrying ring in catenary system
Bin Song 0001, Xiaojiang Du, Mohsen Guizani
Neural Comput. Appl.2
2023 Inter-Intra Modal Representation Augmentation With Trimodal Collaborative Disentanglement Network for Multimodal Sentiment Analysis
abstract
Recently, Multimodal Sentiment Analysis (MSA) is a challenging research area given its complex nature, and humans express emotional cues across various modalities such as language, facial expressions, and speech. Representation and fusion of features are the most crucial tasks in multimodal sentiment analysis research. However, in the current research, most methods ignore the importance of eliminating potential irrelevant features in the original features of each modality and cross-modal common feature. Moreover, the features extracted from all the modalities contain cluttered background noise and different occlusions noise, which negatively affects feature alignment. Different from these methods, we propose a novel Trimodal Collaborative Disentanglement Network (TCDN) to solve these problems in this paper. This work can obtain effective sentiment results on two aspects: i) Trimodal collaborative uses L1-norm to eliminate irrelevant features and unify the characteristics of the three modals (inter-modal). ii) Disentanglement network introduces an adversary noise by combining the original features of various single modalities and the common representation, alleviating the background noises within each modality (intramodal). This inter-intra modal feature augmentation method is the first work to obtain the common representation by implementing data augmentation as far as we know. Extensive experiments are completed on two benchmark datasets, including MOSI and MOSEI, demonstrating the superiority of the TCDN model over the state-of-the-art methods.
Chen Chen 0128, Hansheng Hong, Jie Guo 0008, Bin Song 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 Flexible Resource Management in High-Throughput Satellite Communication Systems: A Two-Stage Machine Learning Framework
abstract
With digitization and globalization in the era of 5G and beyond, research on high-throughput satellites (HTS) to increase communication capacity and improve flexibility is becoming essential. To achieve efficient resource utilization and dynamic traffic demand matching, the multi-dimensional resource management (MDRM) problem of the HTS communication system has been studied in this paper. Since the MDRM problem is a non-convex mixed integer problem, we decompose it into two tractable sub-problems. First, the beam-domain resource configuration problem is formed to enable on-demand coverage. Next, the user-domain resource allocation problem is modeled to enable on-demand communication. Considering the two-domain optimization problem, a two-stage framework is developed based on the combination of self-supervised learning and deep reinforcement learning. Specifically, in the first stage, a maximum co-channel interference based self-supervised learning method is proposed to perform traffic demand matching through demand awareness. In the second stage, a soft frequency reuse based proximal policy optimization approach is presented to further increase the system capacity through interference coordination. The simulation results demonstrate that our proposed two-stage algorithm outperforms the benchmark schemes in terms of spectrum efficiency and demand satisfaction.
Hao Qin 0001, Ning Xin, Bin Song 0001
IEEE Trans. Commun.4
2023 A VAE-Based User Preference Learning and Transfer Framework for Cross-Domain Recommendation
abstract
The core idea of cross-domain recommendation is to alleviate the problem of data scarcity. Previous methods have made brilliant successes. However, many of them mainly focus on learning an ideal mapping function across-domains, ignoring the user preferences within a specific domain, which leads to suboptimal results. In this paper, we propose a Cross-Domain Recommendation Variational AutoEncoder framework (CDRVAE), a novel extension of a variational autoencoder on cross-domain recommendations for user behaviour distribution modeling. It applies a new hybrid architecture of VAE as the backbone and simultaneously constructs two information flows, within-domain and cross-domain modeling. For the former, an asymmetric codec structure is designed to reconstruct preference distribution from domain-specific latent factors. To relieve the posterior collapse dilemma, a combined prior is employed to increase the distribution complexity. The equivalent transition by a transformation matrix and the unobserved interaction generation by cross-domain reconstruction contribute to the latter. We combine all the above components for the more accurate and reliable user features. Extensive experiments are conducted on three public benchmark datasets to validate the effectiveness of the proposed CDRVAE. Experimental results demonstrate that CDRVAE is consistently superior to other state-of-the-art alternative baseline models.
Tong Zhang 0015, Chen Chen 0128, Dan Wang 0002, Jie Guo 0008, Bin Song 0001
IEEE Trans. Knowl. Data Eng.5
2023 Trust-Aware Multi-Task Knowledge Graph for Recommendation
abstract
Data sparsity and cold start problems are common in recommender systems. Adding some side information, such as knowledge graph and users' trust relationship, is an effective method to alleviate these problems. However, few work jointly explore the fine-grained implicit relationships between the external heterogeneous graphs to enhance the recommendation accuracy. To address this issue, in this paper, we propose a new method named Trust-aware Multi-task Knowledge Graph (TMKG), which uses multi-task learning to integrate two kinds of side information of trust graph and knowledge graph in an end-to-end manner. Firstly, we mine the intra-graph and inter-graph high-order connections through the node propagation and aggregation, and optimize the embedding of nodes through the implicit relationships obtained. Furthermore, through the shared cross unit, the connection relationships between each layer is mined, and the high-order interaction of nodes of different layers is obtained. We conduct extensive experiments on real-world datasets and prove that our model has the superior performance compared with the state-of-the-art models.
Jie Guo 0008, Bin Song 0001, Chen Chen 0128, Jianglong Chang, F. Richard Yu
IEEE Trans. Knowl. Data Eng.3
2023 Inter-Intra Modal Representation Augmentation With DCT-Transformer Adversarial Network for Image-Text Matching
abstract
Image-text matching has become a challenging task in the multimedia analysis field. Many advanced methods have been used to explore local and global cross-modal correspondence in matching. However, most methods ignore the importance of eliminating potential irrelevant features in the original features of each modality and cross-modal common feature. Moreover, the features extracted from regions in images and words in sentences contain cluttered background noise and different occlusion noise, which negatively affects alignment. Different from these methods, we propose a novel DCT-Transformer Adversarial Network (DTAN) for image-text matching in this paper. This work can obtain an effective metric based on two aspects: i) DCT-Transformer uses DCT (Discrete Cosine Transform) method based on a transformer mechanism to extract multi-domain common representations and eliminate irrelevant features from different modalities (inter-modal). Among them, DCT divides multi-modal content into chunks of different frequencies and quantifies them. ii) The adversarial network introduces an adversary idea by combining the original features of various single modalities and the multi-domain common representation, alleviating the background noise within each modality (intra-modal). The proposed adversarial feature augmentation method can easily obtain the common representation that is only useful for alignment. Extensive experiments are completed on the benchmark datasets Flickr30K and MS-COCO, demonstrating the superiority of the DTAN model over the state-of-the-art methods.
Chen Chen 0128, Dan Wang 0002, Bin Song 0001
IEEE Trans. Multim.3
2023 HGAN: Hierarchical Graph Alignment Network for Image-Text Retrieval
abstract
Image-text retrieval (ITR) is a challenging task in the field of multimodal information processing due to the semantic gap between different modalities. In recent years, researchers have made great progress in exploring the accurate alignment between image and text. However, existing works mainly focus on the fine-grained alignment between image regions and sentence fragments, which ignores the guiding significance of context background information. Actually, integrating the local fine-grained information and global context background information can provide more semantic clues for retrieval. In this paper, we propose a novel Hierarchical Graph Alignment Network (HGAN) for image-text retrieval. First, to capture the comprehensive multimodal features, we construct the feature graphs for the image and text modality respectively. Then, a multi-granularity shared space is established with a designed Multi-granularity Feature Aggregation and Rearrangement (MFAR) module, which enhances the semantic corresponding relations between the local and global information, and obtains more accurate feature representations for the image and text modalities. Finally, the ultimate image and text features are further refined through three-level similarity functions to achieve the hierarchical alignment. To justify the proposed model, we perform extensive experiments on MS-COCO and Flickr30 K datasets. Experimental results show that the proposed HGAN outperforms the state-of-the-art methods on both datasets, which demonstrates the effectiveness and superiority of our model.
Jie Guo 0008, Meiting Wang, Bin Song 0001, Yuhao Chi, Jianglong Chang
IEEE Trans. Multim.4
2022 A Knowledge Graph-based Cooperative Caching Scheme in MEC-enabled Heterogeneous Networks
abstract
To meet the user demand for high-speed and low-latency video services, this paper envisions a cooperative caching and video transcoding architecture in the multi-access edge computing (MEC)-enabled heterogeneous network. Under this architecture, we propose a novel knowledge graph (KG)-based video caching scheme. Specifically, KG reveals the relation between videos, which acts as an external knowledge to reflect user preferences thus guiding caching decisions. The goal of this paper is to minimize the service delay within the caching and computing resource constraints. To this end, we combine KG with the deep reinforcement learning (DRL) method and design a KG-deep Q network (DQN) based caching algorithm. KG is designed to select the related videos as candidate actions for DQN's caching decision. This way improves the convergence performance of DRL while providing rich external references for caching decisions. Numerous simulation results demonstrate that the proposed algorithm outperforms the traditional baselines on cache hit rate and delay performance.
Yalu Bai, Dan Wang 0002, Bin Song 0001
GLOBECOM3
2022 MR-DARTS: Restricted connectivity differentiable architecture search in multi-path search space
Bin Song 0001, Dan Wang 0002, Hao Qin 0001
Neurocomputing2
2022 Crafting universal adversarial perturbations with output vectors
Xu Kang 0002, Bin Song 0001, Dan Wang 0002, Xiaohui Cai
Neurocomputing2
2022 CAMA: Class activation mapping disruptive attack for deep neural networks
Sainan Sun, Bin Song 0001, Xiaohui Cai, Xiaojiang Du, Mohsen Guizani
Neurocomputing2
2022 Two-stream network with phase map for few-shot classification
Bin Song 0001, Dan Wang 0002, Hao Qin 0001
Neurocomputing2
2022 Resource Management for Edge Intelligence (EI)-Assisted IoV Using Quantum-Inspired Reinforcement Learning
abstract
Recent developments in the Internet of Vehicles (IoV) enable interconnected vehicles to support ubiquitous services. Various emerging service applications are promising to increase the Quality of Experience (QoE) of users. On-board computation tasks generated by these applications have heavily overloaded the resource-constrained vehicles, forcing it to offload on-board tasks to other edge intelligence (EI)-assisted servers. However, excessive task offloading can lead to severe competition for communication and computation resources among vehicles, thereby increasing the processing latency, energy consumption, and system cost. To address these problems, we investigate the transmission-awareness and computing-sense uplink resource management problem and formulate it as a time-varying Markov decision process. Considering the total delay, energy consumption, and cost, quantum-inspired reinforcement learning (QRL) is proposed to develop an intelligence-oriented edge offloading strategy. Specifically, the vehicle can flexibly choose the network access mode and offloading strategy through two different radio interfaces to offload tasks to multiaccess edge computing (MEC) servers through WiFi and cloud servers through 5G. The objective of this joint optimization is to maintain a self-adaptive balance between these two aspects. Simulation results show that the proposed algorithm can significantly reduce the transmission latency and computation delay.
Dan Wang 0002, Bin Song 0001, F. Richard Yu, Xiaojiang Du, Mohsen Guizani
IEEE Internet Things J.2
2022 Task-Oriented Image Transmission for Scene Classification in Unmanned Aerial Systems
abstract
The vigorous developments of the Internet of Things make it possible to extend its computing and storage capabilities to computing tasks in the aerial system with the collaboration of cloud and edge, especially for artificial intelligence (AI) tasks based on deep learning (DL). Collecting a large amount of image/video data, unmanned aerial vehicles (UAVs) can only hand over intelligent analysis tasks to the back-end mobile edge computing (MEC) server due to their limited storage and computing capabilities. How to efficiently transmit the most correlated information for the AI model is a challenging topic. Inspired by task-oriented communication in recent years, we propose a new aerial image transmission paradigm for the scene classification task. A lightweight model is developed on the front-end UAV for semantic block transmission with the perception of images and channel states. To achieve the tradeoff between transmission latency and classification accuracy, deep reinforcement learning (DRL) is applied to explore the semantic blocks which have the greatest contribution to the back-end classifier under various channel states. Experimental results show that the proposed method can significantly improve classification accuracy by more than 4% under the same conditions, compared to other semantic saliency learning methods.
Xu Kang 0002, Bin Song 0001, Jie Guo 0008, Zhijin Qin, F. Richard Yu
IEEE Trans. Commun.2
2022 Few-Shot Scale-Insensitive Object Detection for Edge Computing Platform
abstract
In the era of the Internet of Things, the construction of edge computing platform has become more and more important, which has led lots of object detection applications being deployed on embedded devices. However, traditional object detection algorithms require lots of engery and a large amount of well labeled samples for training. The time spent on model training and data labeling also slows down the upgrade iteration of applications. Therefore, an object detection algorithm that requires only few energy and a few samples to update parameters could help the long-term benign development of IoT technology. In this paper, we propose an effective object detection method based on the few-shot learning, which could achieve considerable performance with few data for novel(new) classes. Our well-designed strategies could alleviate the impact of scale variation in support set under few-shot setting. Through extensive experiments, we prove that our model is superior to well-recognized baselines on few-shot object detection task.
Bin Song 0001, Xiaojiang Du, Mohsen Guizani
IEEE Trans. Sustain. Comput.2
2021 A portable acoustofluidic device for multifunctional cell manipulation and reconstruction
abstract
Microbubble-induced acoustic microstreaming for efficient on-chip micromanipulation is widely developed in biological applications. However, it is still challenging to simultaneously transport, trap, and rotate single cells using one device in a biocompatible manner, while expensive and bulky traditional acoustic driving system also increases its limitation. This paper presents a portable acoustofluidic device for multifunctional cell manipulation and 3D reconstruction, using acoustically oscillating bottom bubble array. Based on the Arduino-based driving system, multiple bubble-induced microvortices were generated and utilized to achieve multifunctional manipulation in a noninvasive manner. Self-propelled transportation of single or multiple cells is first accomplished by bottom bubble array; Controllable trapping, 3D rotation (in the x-y or x-z plane) of DU145 cells are further performed by every single microbubble. Through experiments, rotation direction, speed and axis can be modulated by tuning the driving frequency and voltage. Finally, 3D cell reconstruction combining imaging processing algorithm with out-of-plane rotation enables a sufficient illustration of cell structures and surface morphology, providing an efficient properties measurement function. All these aspects of this device show great potentials in bioengineering, biophysics and biomedicine.
Wei Zhang 0049, Bin Song 0001, Jingli Guo, Lin Feng 0002, Fumihito Arai
ICRA2
2021 Dual Attention Transfer in Session-based Recommendation with Multi-dimensional Integration
abstract
Session-based recommendation (SBR) is widely used in e-commerce to predict the anonymous user's next click action according to a short sequence. Many previous studies have shown the potential advantages of applying Graph Neural Networks (GNN) to SBR tasks. However, the existing SBR models using GNN to solve user preference problems are only based on one single dataset to obtain one recommendation model during training. While the single dataset has the problems including the excessive sparse data source and the long-distance relationship of items. Therefore, introducing the dual transfer, which can enrich the data source, to SBR is absolutely necessary. To this end, a new method is proposed in this paper, which is called dual attention transfer based on multi-dimensional integration (DAT-MDI): (i) DAT uses a potential mapping method based on a slot attention mechanism to extract the user's representation information in different sessions between multiple domains. (ii) MDI combines the graph neural network for the graphs (session graph and global graph) and the gate recurrent unit (GRU) for the sequence to learn the item representation in each session. Then the multi-level session representation are combined by a soft-attention mechanism. We do a variety of experiments on four benchmark datasets which have shown that the superiority of the DAT-MDI model over the state-of-the-art methods.
Chen Chen 0128, Jie Guo 0008, Bin Song 0001
SIGIR3
2021 Image encryption based on a single-round dictionary and chaotic sequences in cloud computing
abstract
Summary With the increasing popularity of multimedia technology and the prevalence of various smart electronic devices, severe security problems have recently arisen in information and communication systems. The image maintenance can be treated as a typical example of cloud storage outsourcing as images require much more storage space than text documents. An image encryption method with high efficiency is crucial to preserve the privacy of sensitive and critical images in cloud‐edge communications. Compressive sensing (CS) is capable of reconstructing compressible or sparse signals using fewer measurements than traditional methods when the support of the signals is unobtainable in advance. Encryption methods based on CS decrease resources demands during signal acquisition and protect the image data to avoid theft of confidential information by malicious users. Moreover, the CS‐based encryption method can achieve the encryption and the compression of an image simultaneously. Recently, a considerable amount of work using CS theory has been reported. However, most of the existing methods result in recovered images with unsatisfactory quality and employ a massive measurement matrix as the encryption key. This paper proposes an encryption scheme for images on cloud based on a single‐round dictionary and chaotic sequences. Both the experimental and analytical results verify the proposed scheme's effectiveness, high level of security, and substantial image quality improvement after decryption.
Bin Song 0001, Jinjun Chen
Concurr. Comput. Pract. Exp.3
2021 Resource Management for Secure Computation Offloading in Softwarized Cyber-Physical Systems
abstract
The evolution of the Internet of Things (IoT) makes an increased emphasis on extending their computing and storage capabilities by relying particularly on the cloud/edge computing (EC) for cyber-physical systems (CPSs). Especially, in software-defined CPS (SD-CPS), different software-defined networking (SDN) controllers share information and cooperate to make global decisions. To further enhance system security during the information sharing process, we introduce blockchain technology into SD-CPS. However, because many security-related decisions are sensitive to latency, it is vital to minimize the system latency in blockchain-empowered SD-CPS. In this article, a blockchain-empowered distributed SD-CPS framework is proposed to realize consensus and distributed resource management by offloading data in a hybrid network paradigm that combines cloud computing and EC. Moreover, to adaptively implement offloading and control strategies while guaranteeing data security, we design a resource management scheme for reducing system latency and provide the flexibility of cooperation. To foster intelligence, we formulate the joint communication, computation, and consensus problems as a Markov decision process and use deep reinforcement learning to balance resource allocation, reduce latency, and guarantee data security. Compared with other schemes, simulation results verify the effectiveness of the proposed scheme, which performs better on self-adaptation decision making and system delay reduction.
Dan Wang 0002, Bin Song 0001, F. Richard Yu
IEEE Internet Things J.3
2021 Is 5G Handover Secure and Private? A Survey
abstract
The next-generation mobile cellular communication and networking system (5G) is highly flexible and heterogeneous. It integrates different types of networks, such as 4G legacy networks, Internet of Things (IoT), vehicular ad hoc network (VANET), and wireless local access network (WLAN) to form a heterogeneous network. This easily results in continual vertical handovers between different networks. On the other hand, substantial deployment of small/micro base stations (BSs) brings frequent horizontal handovers within a network. The continual handovers among BSs and various networks expose mobile equipment (ME) to the risk of security and privacy threats. So far, many security and privacy mechanisms have been proposed to ensure secure handover either vertically or horizontally in 5G networks. Nevertheless, there still lacks a thorough survey to summarize recent advances and explore open issues although handover security and privacy are crucial to 5G. In this article, we summarize security and privacy requirements in handovers to resist potential attacks. Following these requirements as the evaluation criteria, we review secure and privacy-preserving handover schemes by categorizing them into two scenarios, i.e., vertical handover and horizontal handover. As for the vertical handover, we review related work from three classes, i.e., handovers within Third-Generation Partnership Project (3GPP) networks, between 3GPP and non-3GPP networks, and between non-3GPP networks. Concerning horizontal handovers, we review related work from two classes, i.e., intramobile service controller (MSC) and inter-MSC handover. Meanwhile, we analyze and compare the technical means and performance of these works in order to uncover open issues and inspire future research directions.
Zheng Yan 0002, Peng Zhang 0004, Bin Song 0001
IEEE Internet Things J.5
2021 Hierarchical Deep Embedding for Aurora Image Retrieval
abstract
Retrieving informative images from the large-scale aurora data is of great significance in the field of space physics. In this article, we propose a hierarchical deep embedding (HDE) model to assist scientists for their aurora image retrieval. Other than conventional bag-of-words (BoW) models employing local cues individually, HDE performs visual matching in a hierarchical way, that is, only keypoints which are similar on local, regional, and global simultaneously can be treated as a true match. The added contextual evidences can effectively alleviate the occurrence of false matches and improve the precision of visual matching. Specifically, to complement the local SIFT feature, the convolutional neural network (CNN) is refined with a polar region pooling (PRP) layer to extract features from regional patches and global image, forming a group of hierarchical deep features with strong discriminative power. Also, an improved polar meshing (IPM) scheme is presented to determine the positions of keypoints, which is more suitable for images captured by circular fisheye lens and capable of reflecting the physical information in aurora images. Extensive experiments are conducted on the big aurora data, which indicate that the proposed HDE model greatly promotes the retrieval accuracy with acceptable memory cost and efficiency. In addition, the effectiveness of the IPM scheme and the superiority of the hierarchical deep feature integration are separately demonstrated.
Xi Yang 0011, Xinbo Gao 0001, Bin Song 0001, Bing Han 0003
IEEE Trans. Cybern.3
2020 ABP: Adaptive Body Partition Model For Visible Infrared Person Re-Identification
abstract
Person re-identification (Re-ID) aims to match pedestrian images cross multiple cameras. Most Re-ID studies focus on visible pedestrian images, without considering the images obtained by infrared cameras in the dark. To solve the cross-modality person Re-ID problem, current methods usually exploit global feature descriptors to obtain discriminative representations. However, they ignore the fine-grained information of heterogeneous images. In this paper, we propose an adaptive body partition (ABP) model to automatically detect and distinguish effective part representations. Instead of utilizing a two-stream convolutional neural network (CNN) to extract modality-specific information, we directly design an end-to-end one-stream CNN to simultaneously learn multi-modality sharable features and map them on a common space. Global loss, part losses and threefold triplet loss are integrated to enhance the feature discriminability and minimize the cross-modality gap. Extensive experimental results on two cross-modality Re-ID datasets exhibit the superiority of the proposed method compared with the state-of-the-art solutions.
Xi Yang 0011, Nannan Wang 0001, Bin Song 0001, Xinbo Gao 0001
ICME4
2020 A novel portable cell sonoporation device based on open-source acoustofluidics
abstract
Sonoporation, which typically employs acoustic cavitation microbubbles, can enhance the permeability of the cell membrane, allowing foreign matter to enter cells across the natural barriers. However, the diameter nonuniformity and random distribution of microbubbles make it difficult to achieve controllable and high-efficiency sonoporation, while complex extern acoustic driving system also limits its applicability. Herein, we demonstrate a low-cost, expandable, and portable acoustofluidic device for cell sonoporation using acoustic streaming generated by oscillating sharp edges. The streaming-induced high shear forces can (i) quickly trap target cells at the tip of sharp edges and (ii) transiently modulate the permeability of the cell membrane, which is utilized to perform cell sonoporation events. Using our device, sonoporation is successfully achieved in a microbubble-free manner, with a sonoporation efficiency of more than 90%. Furthermore, our acoustic driving system is designed around the open-source Arduino prototyping platform due to its extendibility and portability. In addition to these benefits, our acoustofluidic device is simple to fabricate and operate, and it can work at relatively low frequency (4.6 kHz). All these advantages make our novel cell sonoporation device invaluable for many biological and biomedical applications such as drug delivery and gene transfection.
Bin Song 0001, Wei Zhang 0049, Lin Feng 0002, Deyuan Zhang, Fumihito Arai
IROS1
2020 The enhancement of catenary image with low visibility based on multi-feature fusion network in railway industry
Bin Song 0001, Xiaojiang Du, Nadra Guizani
Comput. Commun.2
2020 Context-Aware Object Detection for Vehicular Networks Based on Edge-Cloud Cooperation
abstract
Due to high mobility and high dynamic environments, object detection for vehicular networks is one of the most challenging tasks. However, the development of integration techniques, such as software-defined networking (SDN) and network function visualization (NFV), in networking, caching, and computing provides us with new approaches. In this article, we propose a novel context-aware object detection method based on edge-cloud cooperation. Specifically, an object detection model based on deep learning is established in the cloud server. Different from other methods, to further explore the underlying inner spatial features of collected images, the visual objects of images are regarded as nodes and the spatial relations between objects as edges, then a type of message-passing method is employed to update the nodes' features. In the mobile edge computing (MEC) servers, the context information and captured images of the vehicular environments are extracted and then are used to adjust the object detection model from the cloud server. In this way, the cloud server cooperates with the MEC servers to realize context-aware object detection, which improves the adaptation and performance of the detection model under different scenarios. The simulation results also demonstrate that the proposed method is more accurate and faster than the previous methods.
Jie Guo 0008, Bin Song 0001, F. Richard Yu, Xiaojiang Du, Mohsen Guizani
IEEE Internet Things J.2
2020 Person Re-Identification with Feature Pyramid Optimization and Gradual Background Suppression
Yingzhi Tang, Xi Yang 0011, Nannan Wang 0001, Bin Song 0001, Xinbo Gao 0001
Neural Networks4
2020 HCNN-PSI: A hybrid CNN with partial semantic information for space target recognition
Xi Yang 0011, Tan Wu, Nannan Wang 0001, Yan Huang 0018, Bin Song 0001, Xinbo Gao 0001
Pattern Recognit.5
2020 A novel deformable body partition model for MMW suspicious object detection and dynamic tracking
Xi Yang 0011, Nannan Wang 0001, Bin Song 0001, Xinbo Gao 0001
Signal Process.4
2020 A Rotational Libra R-CNN Method for Ship Detection
abstract
Recently, ship detection methods based on deep learning have attracted significant attention due to their superior accuracy over traditional methods. However, there still exist two problems affecting its robustness in practical application. 1) The size of ships in one image varies greatly, i.e., different sizes; 2) Numerous ships gather in limited field-of-view, i.e., dense distribution. To address these problems, we propose a rotational Libra R-convolutional neural network (CNN) method. Our idea is to balance the three levels of neural networks for predicting the location of ships with rotational angle information, which refers to the feature level, sample level, and objective level. First, to extract a discriminative feature and improve its robustness against the impact of different sizes of ships, the concept of balanced feature pyramid is introduced. Second, to generate reliable proposals for feature pyramid and efficiently mine hard negative samples, we employ intersection over union (IoU)-balanced sampling. Finally, to eliminate the redundant background and detect densely distributed ships, we bring in a rotational region detection branch with balanced L1 loss. In general, we develop the balanced learning with rotational region detection to achieve consistent improvement on accuracy and visualization. Experimental results on DOTA data set show that the proposed method achieves the state-of-the-art accuracy.
Haoyuan Guo, Xi Yang 0011, Nannan Wang 0001, Bin Song 0001, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.4
2020 D2N4: A Discriminative Deep Nearest Neighbor Neural Network for Few-Shot Space Target Recognition
abstract
With the rapid development of space exploration worldwide, there is a sudden increase in the type and number of spacecraft, thus leading to a more complex space environment. To enhance the ability of space situational awareness, the most important step is to effectively recognize space targets of interests from various spacecraft and debris. Traditional space target recognition approaches adopt manual feature extraction with limited data, resulting in a semantic gap between low-level visual features and high-level semantic representation. Although deep learning models alleviate this problem with a unified framework for combined learning feature extraction and classification simultaneously, it is easy to overfit and leads to poor generalization results when faced with a situation of small examples. To address these issues, we present an end-to-end few-shot deep learning framework for space target recognition, i.e., discriminative deep nearest neighbor neural network (D2N4). Our D2N4 aims to improve the discriminability of the deeply learned features with mainly two strategies. On the one hand, we add an intraclass compactness principle by introducing center loss to efficiently pull deep features of the same classes to their centers and, thus overcoming significant intraclass variation of space target. On the other hand, we introduce the global pooling information for each deep local descriptor to reduce interference from local background noise, thus enhancing the model robustness. In practice, under the joint supervision of soft-max loss and center loss, the deep embedding module and image-to-class metric module are trained in an end-to-end way. Extensive experiments on the space target data set BUAA-SID-share1.0 demonstrate that our simple and effective approach outperforms previous space target recognition methods and is more efficient than recent few-shot approaches. In addition, the proposed framework is equally applicable to natural images and achieves state-of-the-art performance on data sets CUB-200-2010, Stanford Dogs, and Stanford Cars.
Xi Yang 0011, Xiaoting Nan, Bin Song 0001
IEEE Trans. Geosci. Remote. Sens.3
2020 Aurora Image Search With a Saliency-Weighted Region Network
abstract
On account of the remarkable performance of convolutional neural network (CNN) features for natural image searches, utilizing it for other images collected with the anamorphic lens has become a research hotspot. This article selects the aurora images generated from a circular fisheye lens as a typical example. By considering the imaging principle and geomagnetic information, a saliency-weighted region network (SWRN) is presented and introduced into the Mask R-CNN pipeline. Our SWRN selects salient regions with important semantic information and weights them both hierarchically and spatially. Hence, regions encompassing the search target are strengthened while uninformative regions are discarded, which benefits the suppression of background interference and reduction of computational complexity. In practice, by aggregating the outputs of SWRN with post-processing, a compact CNN feature is generated to represent the aurora image. Large-scale aurora image search experiments are conducted, and the results prove that our method performs better than the state-of-the-art methods on both accuracy and efficiency.
Xi Yang 0011, Nannan Wang 0001, Bin Song 0001, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.3
2020 CGAN-TM: A Novel Domain-to-Domain Transferring Method for Person Re-Identification
abstract
Person re-identification (re-ID) is a technique aiming to recognize person cross different cameras. Although some supervised methods have achieved favorable performance, they are far from practical application owing to the lack of labeled data. Thus, unsupervised person re-ID methods are in urgent need. Generally, the commonly used approach in existing unsupervised methods is to first utilize the source image dataset for generating a model in supervised manner, and then transfer the source image domain to the target image domain. However, images may lose their identity information after translation, and the distributions between different domains are far away. To solve these problems, we propose an image domain-to-domain translation method by keeping pedestrian's identity information and pulling closer the domains' distributions for unsupervised person re-ID tasks. Our work exploits the CycleGAN to transfer the existing labeled image domain to the unlabeled image domain. Specially, a Self-labeled Triplet Net is proposed to maintain the pedestrian identity information, and maximum mean discrepancy is introduced to pull the domain distribution closer. Extensive experiments have been conducted and the results demonstrate that the proposed method performs superiorly than the state-ofthe- art unsupervised methods on DukeMTMC-reID and Market- 1501.
Yingzhi Tang, Xi Yang 0011, Nannan Wang 0001, Bin Song 0001, Xinbo Gao 0001
IEEE Trans. Image Process.4
2020 A Novel Symmetry Driven Siamese Network for THz Concealed Object Verification
abstract
Security inspection aims to improve the high detection rate as well as reduce the false alarm rate. However, it still suffers from two challenges affecting its robustness. 1) Existing security inspection methods are mostly designed for natural images, which cannot reflect the uniqueness and imaging principle of THz images. 2) Existing methods is sensitive to noise interference and pose variations. This work revisits these challenges and presents a novel symmetry driven Siamese network (SDSN) for THz concealed object verification. Our idea is to employ a specially designed network architecture for THz concealed object verification. First, to reflect the uniqueness and the special property of THz images, Siamese network with Contrastive loss is used for feature extraction along with symmetrical prior information consideration, which can learn symmetrical metrics from the same person. Second, to alleviate the impact of noise interference and pose variations, the adaptive identity normalization (A-IDN) is proposed to normalize the symmetrical metrics each person. Finally, to enhance the generalization of network, an adaptive selective threshold based on Gaussian mixture model (AST-GMM) is designed, which serves as a classifier for the final classification results. Extensive experiments show that SDSN significantly improves the accuracy. Specially, SDSN outperforms the state-of-the-art methods without symmetrical prior information on THz security dataset.
Xi Yang 0011, Haoyuan Guo, Nannan Wang 0001, Bin Song 0001, Xinbo Gao 0001
IEEE Trans. Image Process.4
2019 Person re-Identification with Gradual Background Suppression
abstract
Person re-identification plays an important role in public security. However, owing to the interference of background clutters, its performance still needs to be improved. Several mask-based methods aim to solve this problem by totally removing the background clutters, but the promotion is limited because of the mask sharpening effect. In this paper, we propose a novel person re-identification method with Gradual Background Suppression (GBS). The GBS adopts several CNN branches to extract deep features of images with different weight distributions between background and human body. Thus, it can not only reduce the background clutters but also keep the smoothness of target pedestrians. Afterwards, deep features from different CNN branches are integrated with a fusion scheme, and the fused feature is capable of balancing the influence of background clutter and mask sharpening. Extensive experiments have been conducted and the results prove the superiority of the proposed GBS over the background removal approach. Comparing with the state-of-the-art methods, our method achieves remarkable performance with 6.6%, 7.58% and 8.26% improvement of mAP on dataset Market-1501, CUHK03-labeled and CUHK03-detected, respectively.
Yingzhi Tang, Xi Yang 0011, Nannan Wang 0001, Bin Song 0001, Xinbo Gao 0001
ICME5
2019 Cell Injection Microrobot Development and Evaluation in Microfluidic Chip
abstract
We propose an innovative design of microrobot, which can achieve donor cell suction, delivery and injection in a mammalian oocyte on microfluidic chip. The microrobot body contains a hollow space that produces suction and ejection forces for injection of cell nuclei using a nozzle at the tip of the robot. Specifically, a controller changes the hollow volume by balancing the magnetic and elastic forces of the membrane, and along with motion of stages in the XY plane. A glass capillary attached at the tip of the robot contains the nozzle is able to absorb and inject cell nuclei. The microrobot provides three degrees of freedom and generates micronewton forces. We demonstrate the effectiveness of the proposed microrobot through an experiment of absorption and ejection of 20 μm particles from the nozzle using magnetic control in a microfluidic chip.
Lin Feng 0002, Dixiao Chen, Bin Song 0001, Wei Zhang 0049
ICRA4
2019 T-SCNN: A Two-Stage Convolutional Neural Network for Space Target Recognition
abstract
Space target recognition plays an important role in the field of space security and exploration. With the rapid development of artificial intelligence technique and explosive increase of image dataset, object recognition based on deep learning has achieved favorable performance. However, the recognition of deep space targets in visible spectrum images still remains in the traditional manual interpretation approach, thus leading to low efficiency and inevitable subjective errors. In this paper, we propose an artificial intelligence method for space target recognition, called Two-Stage Convolutional Neural Network (T-SCNN). Our T-SCNN is composed of two stages, i.e., target locating and target recognition. In the stage of target locating, we first detect all suspected targets from the total image dataset by presenting a minimum bounding rectangle with threshold (MBRT) approach, then cut out all regions encompassing targets to generate target images for training. In the stage of target recognition, we send target images to the well-trained recognition network for identification. Additionally, data augmentation is conducted in the CNN training to satisfy its data quantity requirement. Extensive experiments are performed on our synthetic space target image dataset, and the result demonstrate that the proposed method achieves high accuracy within a short time.
Tan Wu, Xi Yang 0011, Bin Song 0001, Nannan Wang 0001, Xinbo Gao 0001, Liyang Kuang, Xiaoting Nan, Dong Yang 0012
IGARSS3
2019 Deep neural network-aided Gaussian message passing detection for ultra-reliable low-latency communications
Jie Guo 0008, Bin Song 0001, Yuhao Chi, Lahiru Jayasinghe, Chau Yuen, Yong Liang Guan 0001, Xiaojiang Du, Mohsen Guizani
Future Gener. Comput. Syst.2
2019 BoSR: A CNN-based aurora image retrieval method
Xi Yang 0011, Nannan Wang 0001, Bin Song 0001, Xinbo Gao 0001
Neural Networks3
2019 Exemplar based regular texture synthesis using LSTM
Xiuxia Cai, Bin Song 0001, Zhiqian Fang
Pattern Recognit. Lett.2
2019 CNN with spatio-temporal information for fast suspicious object detection and recognition in THz security images
Xi Yang 0011, Tan Wu, Lei Zhang 0019, Dong Yang 0012, Nannan Wang 0001, Bin Song 0001, Xinbo Gao 0001
Signal Process.6
2018 Saliency Deep Embedding for Aurora Image Search
abstract
Deep neural networks have achieved remarkable success in the field of image search. However, the state-of-the-art algorithms are trained and tested for natural images captured with ordinary cameras. In this paper, we aim to explore a new search method for images captured with circular fisheye lens, especially the aurora images. To reduce the interference from uninformative regions and focus on the most interested regions, we propose a saliency proposal network (SPN) to replace the region proposal network (RPN) in the recent Mask R-CNN. In our SPN, the centers of the anchors are not distributed in a rectangular meshing manner, but exhibit spherical distortion. Additionally, the directions of the anchors are along the deformation lines perpendicular to the magnetic meridian, which perfectly accords with the imaging principle of circular fisheye lens. Extensive experiments are performed on the big aurora data, demonstrating the superiority of our method in both search accuracy and efficiency.
Xi Yang 0011, Xinbo Gao 0001, Bin Song 0001, Nannan Wang 0001, Dong Yang 0012
ICME3
2018 FPAN: Fine-grained and progressive attention localization network for data retrieval
Bin Song 0001, Jie Guo 0008, Yanling Zhang, Xiaojiang Du, Mohsen Guizani
Comput. Networks2
2018 Leveraging high-order statistics and classification in frame timing estimation for reliable vehicle-to-vehicle communications
abstract
In vehicle‐to‐vehicle (V2V) communications, achieving reliable physical layer performance is a challenging task due to the highly dynamic nature of V2V propagation channels. Frame timing estimation, as one of the most critical signal processing procedures that rely on channel statistics, has to be appropriately enhanced to tackle this challenge. This study presents a novel frame timing estimation scheme based on both the available periodical preambles in IEEE 802.11p standard. By designing the fourth‐order statistics‐based correlation and differential normalisation functions, the proposed timing metric not only is capable of possessing an extensible correlation length, but also achieves the robustness to multipath effect and large carrier frequency offset. From the standpoints of hypothesis testing and classification, the proposed approach can effectively increase the distinction between correct and wrong timing indexes in terms of the class‐separability criteria, and consequently has a significantly improved timing estimation performance compared with the existing methods. Simulation results consist with theoretical analysis under the typical V2V channel model, and demonstrate that the proposed method can significantly reduce both the probabilities of false alarm and missed detection, and make the selection of a suitable threshold for frame detection much easier.
Li Zhen, Hao Qin 0001, Bin Song 0001, Rui Ding 0002, Yanling Zhang
IET Commun.3
2018 Aurora image search with contextual CNN feature
Xi Yang 0011, Xinbo Gao 0001, Bin Song 0001, Dong Yang 0012
Neurocomputing3
2018 Random Access Preamble Design and Detection for Mobile Satellite Communication Systems
abstract
Reasonable design and effective detection of the random access preamble has become a challenging task due to the unique characteristics of mobile satellite communications. To tackle this challenge, we first design a universal long sequence structure by concatenating multiple short Zadoff-Chu sequences that are insensitive to carrier frequency offset (CFO), and then propose the new principles of parameter selection for short sequences to ensure the minimum utilization of root sequence and the independence of the cyclic shift offset on the beam radius. To further reduce the detection complexity and improve the multi-user access performance, a fast timing detection approach is also presented by leveraging the piecewise cumulative detection and the multi-peaks joint estimation to obtain an accurate timing advance for each access user. Simulation results and complexity analysis validate the effectiveness of the new preamble in a typical satellite communication environment, and reveal that the proposed timing detection can achieve the robustness to CFO and offer outstanding performance improvements especially in multi-user scenarios while having a notably reduced computational complexity.
Li Zhen, Hao Qin 0001, Bin Song 0001, Rui Ding 0002, Xiaojiang Du, Mohsen Guizani
IEEE J. Sel. Areas Commun.3
2018 Semantic object removal with convolutional neural network feature-based inpainting approach
Xiuxia Cai, Bin Song 0001
Multim. Syst.2
2018 Image-based pencil drawing synthesized using convolutional neural network feature maps
Xiuxia Cai, Bin Song 0001
Mach. Vis. Appl.2
2018 Vehicle Tracking Using Surveillance With Multimodal Data Fusion
abstract
Vehicle location prediction or vehicle tracking is a significant topic within connected vehicles. This task, however, is difficult if merely a single modal data is available, probably causing biases and impeding the accuracy. With the development of sensor networks in connected vehicles, multimodal data are becoming accessible. Therefore, we propose a framework for vehicle tracking with multimodal data fusion. Specifically, we fuse the results of two modalities, images and velocities, in our vehicle-tracking task. Images, being processed in the module of vehicle detection, provide visual information about the features of vehicles, whereas velocity estimation can further evaluate the possible locations of the target vehicles, which reduces the number of candidates being compared, decreasing the time consumption and computational cost. Our vehicle detection model is designed with a color-faster R-CNN, whose inputs are both the texture and color of the vehicles. Meanwhile, velocity estimation is achieved by the Kalman filter, which is a classical method for tracking. Finally, a multimodal data fusion method is applied to integrate these outcomes so that vehicle-tracking tasks can be achieved. Experimental results suggest the efficiency of our methods, which can track vehicles using a series of surveillance cameras in urban areas.
Yue Zhang 0022, Bin Song 0001, Xiaojiang Du, Mohsen Guizani
IEEE Trans. Intell. Transp. Syst.2
2017 An Unbalanced Data Hybrid-Sampling Algorithm Based on Multi-Information Fusion
abstract
The emergence of big data bringsnewissues and challenges for the data imbalance problem.Therefore, unbalanced data sampling technology has been a hot research topic in the field of big data.However, the existing sampling methods cannot accurately define the harmful and useless samplescontained in the originaldataset. That is, based on the single information of the dataset, a large number of actuallyharmful samples are being used for sampling, which results in a sharp decline in the identifiable performance of the sampled data. In order to overcome the problems caused by only using one kind of information, an unbalanced data hybrid-sampling algorithm based on multi-information fusion(MIFS)is presented in this paper. The MIFS combines the feature information learned by the boostingmodel with the position information of the data to define the sample, and then divides the samples into different subsets by the information contained. According to the definition of samples, the algorithm performs corresponding under-sampling and over-sampling on these subsets. Experiments show that the MIFS method can improve the performance of sampling operations and produce a high F-score and AUC against bothminority and majority classes in the classification of balanced data.
Bin Song 0001, Jie Guo 0008, Xiaojiang Du
GLOBECOM2
2017 An Advanced Random Forest Algorithm Targeting the Big Data with Redundant Features
Bin Song 0001, Yue Zhang 0022
ICA3PP2
2017 An effective DDoS defense scheme for SDN
abstract
In this paper, we propose a scheme to protect the Software Defined Network(SDN) controller from Distributed Denial-of-Service(DDoS) attacks. We first predict the amount of new requests for each openflow switch periodically based on Taylor series, and the requests will then be directed to the security gateway if the prediction value is beyond the threshold. The requests that caused the dramatic decrease of entropy will be filtered out and rules will be made in security gateway by our algorithm; the rules of these requests will be sent to the controller. The controller will send the rules to each switch to make them direct the flows matching with the rules to the honey pot. The simulation shows the averages of both false positive and false negative are less than 2%.
Xueli Huang, Xiaojiang Du, Bin Song 0001
ICC3
2017 Achieving Fair Spectrum Allocation for Co-Existing Heterogeneous Secondary User Networks
abstract
The rapid growth of mobile network traffic has posed a serious challenge to the limited spectrum. The United States Federal Communications Commission (FCC) allowed the utilization of unused TV White Space (TVWS) by unlicensed secondary users (SUs). Particularly, the IEEE 802.19.1 standard is proposed to regulate the coexistence of dissimilar or independently operated SU networks and devices on the TV band. In this paper, we propose a fair spectrum allocation scheme for co-existing SU networks under the IEEE 802.19.1 system architecture. The entire heterogeneous wireless system is divided into two levels, and the spectrum allocation is formulated into a four-stage problem. Unlike previous allocation schemes that maximize the aggregated throughput, the aim of our scheme is to maximize the end user satisfactions within each SU network while maintaining fairness among and within the SU networks. Extensive simulations demonstrate the effectiveness of our spectrum allocation scheme.
Longfei Wu, Xiaojiang Du, Jie Wu 0001, Bin Song 0001
ICCCN4
2017 Object detection among multimedia big data in the compressive measurement domain under mobile distributed architecture
Jie Guo 0008, Bin Song 0001, F. Richard Yu, Zheng Yan 0002, Laurence T. Yang
Future Gener. Comput. Syst.2
2017 Data-driven vs. model-driven: Fast face sketch synthesis
Nannan Wang 0001, Mingrui Zhu, Jie Li 0001, Bin Song 0001, Zan Li 0001
Neurocomputing4
2017 Social behavior study under pervasive social networking based on decentralized deep reinforcement learning
Yue Zhang 0022, Bin Song 0001, Peng Zhang 0004
J. Netw. Comput. Appl.2
2017 Unified framework for face sketch synthesis
Nannan Wang 0001, Shengchuan Zhang, Xinbo Gao 0001, Jie Li 0001, Bin Song 0001, Zan Li 0001
Signal Process.5
2016 Frame timing estimation based on statistical analysis for orthogonal frequency division multiplexing systems in multipath fading channels
abstract
This study investigates the problem of frame timing estimation in orthogonal frequency division multiplexing systems. Conventional timing estimation methods, which take advantage of the correlation property of a given preamble, always experience performance degradation in multipath fading channels with severe channel dispersion. To achieve accurate timing estimation, the authors propose a robust threshold‐based timing detection method independent of the preamble structure. Based on the autocorrelation and cross‐correlation, a novel timing metric with an extended correlation length is proposed to mitigate noise and resist large carrier frequency offsets. Due to the superior statistical property of the proposed timing metric, the threshold can be easily determined with no need for the process of noise variance estimation. Simulation results under different multipath fading channels demonstrate that the proposed method achieves a remarkably improved timing accuracy compared to the existing methods.
Li Zhen, Hao Qin 0001, Bin Song 0001, Rui Ding 0002
IET Commun.3
2016 Evaluation on synthesized face sketches
Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001, Bin Song 0001, Zan Li 0001
Neurocomputing4
2016 Combining inconsistent textures using convolutional neural networks
Xiuxia Cai, Bin Song 0001
J. Vis. Commun. Image Represent.2
2016 Significance Evaluation of Video Data Over Media Cloud Based on Compressed Sensing
abstract
Given the varying communication environment between the media cloud and users, there is a need to ensure the most significant part of a video will be successfully transmitted. Although there exist some techniques to evaluate the significance of video data in traditional video coding methods, such as H.264, the evaluation algorithms are often simple and inaccurate. This paper presents a novel significance evaluation method for video data based on compressed sensing. Specifically, we propose a method to obtain a trained dictionary directly by using the measurements of the video data, and then keep the sparse components and generate a saliency map. Since the sparse components can reflect the essential parts of videos, we discuss how to analyze the area and distribution of salient regions. At last, we present a computing method that gives the degree of significance of a frame. Experimental results show that the proposed saliency map reflects the focus points of humans. The method can be used in the distribution of video data over “wireless” transmissions and provide good video quality to mobile users.
Jie Guo 0008, Bin Song 0001, Xiaojiang Du
IEEE Trans. Multim.2
2016 Image Encryption Based on Compressive Sensing and Scrambled Index for Secure Multimedia Transmission
abstract
With the rapid growth of multimedia message exchange and digital communication, multimedia big data has become a research hotspot in various fields. The storage and transmission of multimedia big data have high requirements for security. Images, covering the highest proportion of multimedia data, should be processed and transmitted with high security. Compressive sensing (CS) has a beneficial property for the encryption that the image can be recovered with fewer samples than conventional approaches use. In recent years, CS has been studied not only to reduce the resource requirements for signal acquisition but also to ensure the security of data. It is still an open challenge to improve security and enhance the quality of the decrypted image simultaneously using the key with small size. In this article, a CS-based encryption method is presented that associates the quantization with random measurement permutation. An enormous number of experiments have been conducted on both standard test images and face images chosen from the big database LFW. Experimental results show that our proposal has dramatic improvements on ensuring the security, enhancing the quality of the decrypted image, and raising the efficiency. Additionally, this proposal remarkably reduces storage and transmission resources. Accordingly, this encryption scheme can be applied to ensure the security of multimedia transmission.
Bin Song 0001, Rong Cao, Yue Zhang 0022, Hao Qin 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2015 Optimal-correlation-based reconstruction for distributed compressed video sensing
Haixiao Liu, Bin Song 0001, Hao Qin 0001
J. Vis. Commun. Image Represent.2
2014 Estimation of measurements for block-based compressed video sensing: study of correlation noise in measurement domain
abstract
Compressed video sensing (CVS) is an application of compressed sensing theory which samples a signal below the Shannon–Nyquist rate. However, previous research about CVS has largely ignored the inter‐frame correlation analysis in the measurement domain, and then is not able to remove the time redundancy. In this study, the authors consider the estimation of the measurements of a block in any possible position in a frame by introducing a correlation noise (CN) between the actual and the estimated measurements. In this work, they first establish a correlation model (CM) in the pixel domain between a block which is in a random unknown position in a frame and the adjacent non‐overlapping blocks that they already have. Then, a novel measurement domain CM is presented to approximate the measurements for the random block. Lastly, they employ the CN to characterise the accuracy of the CM in the measurement domain. The simulation results show that the proposed model can make an accurate estimation to the actual measurements of an arbitrary block in a frame and that by using the proposed CN to perform motion estimation, they can improve the peak signal‐to‐noise ratio of the video sequences by 0.1–1.7 dB compared with the existing methods.
Bin Song 0001, Jie Guo 0008, Lingquan Li, Haixiao Liu
IET Image Process.1
2014 Compressed sensing with partial support information: coherence-based performance guarantees and alternative direction method of multiplier reconstruction algorithm
abstract
The recently introduced theory of compressed sensing (CS) enables the recovery of sparse or compressible signals from a small set of non‐adaptive measurements, and furthermore, it holds promise for substantially improving the performance by leveraging more signal structures that go beyond simple sparsity. In this study, the authors study the weighted l 1 minimisation problem for CS reconstruction when partial support information is available. Firstly, they focus on the coherence‐based performance guarantees and show that if an estimated support can be obtained with its accuracy and relative size satisfying certain coherence‐related conditions, the weighted l 1 minimisation is then stable and robust under weaker sufficient conditions than that of the analogous standard l 1 optimisation. Meanwhile, better upper bounds on the reconstruction error could also be achieved. Besides, a novel adaptive alternating direction method of multipliers with iterative support detection is outlined to solve the weighted l 1 minimisation problem. Simulation results show that the authors’ method achieves good convergence, and obtains improved reconstruction performance in comparison with the conventional methods.
Haixiao Liu, Bin Song 0001, Hao Qin 0001
IET Signal Process.2
2014 Joint Sampling Rate and Bit-Depth Optimization in Compressive Video Sampling
abstract
Compressed sensing is a novel technology that exploits sparsity of a signal to perform sampling below the Nyquist rate, and thus has great potential in low-complexity video sampling and compression applications, due to the significant reduction of the sampling rate ( SR) and computational complexity. However, most current work about compressive video sampling (CVS) has focused on real-valued measurements without being quantized, and thus is not applicable to engineering practices. Moreover, in many circumstances, the total number of bits is often constrained. Therefore, how to achieve a compromise between the number of measurements and the number of bits per measurement to maximize the visual quality is a great challenge for CVS, which has still not been addressed in literature. In this paper, we first present a novel distortion model that reveals the relationship between distortion, SR, and quantization bit-depth ( B). Then, using this model, we propose a joint SR - B optimization algorithm, by which we are able to easily derive the values of SR and B. Finally, we present an adaptive and unidirectional CVS framework with rate-distortion (RD) optimized rate allocation, wherein we use video characteristics extracted from partial sampling to allocate the required bits for each block, and then implement “optimized” video sampling and measurement quantization with the estimated SR and B, respectively. Simulation results show that our proposal offers comparable RD performance to the conventional method, with a 4.6 dB improvement in the average PSNR.
Haixiao Liu, Bin Song 0001, Hao Qin 0001
IEEE Trans. Multim.2
2013 Dictionary learning based reconstruction for distributed compressed video sensing
Haixiao Liu, Bin Song 0001, Hao Qin 0001, Zhiliang Qiu
J. Vis. Commun. Image Represent.2
2013 An Adaptive-ADMM Algorithm With Support and Signal Value Detection for Compressed Sensing
abstract
This letter presents a novel adaptive alternating direction method of multipliers with support/signal value detection for compressed sensing. The support/signal value detection in our algorithm can achieve an efficient reconstruction by leveraging more information that goes beyond simple sparsity. Especially for time-correlated signals in large-scale problems, our proposal performs better than conventional methods, since more accurate signal information could be estimated from prior knowledge during initialization. Simulation results show that our method can improve the average PSNR by 1.02-2.05 dB for undersampled video sequences.
Haixiao Liu, Bin Song 0001, Hao Qin 0001, Zhiliang Qiu
IEEE Signal Process. Lett.2