VLDB 2026 Research / reviewers in the wild / expert
Jigang Wu
dblp:21/1136 · also Wu Jigang
· DBLP profile ↗
215ranked-venue papers
25as first author
114since 2021 · last 2026
0000-0002-6470-9794ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 71 · 14 first-author · 29 since 2021Artificial intelligence and machine learning · 40 · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 28 since 2021Applied, interdisciplinary, general and emerging computing · 26 · 6 first-author · 13 since 2021Computer networks · 19 · 12 since 2021Databases, data management, data science and information retrieval · 8 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 7 · 4 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Theory of computation · 3 · 3 first-authorSecurity and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Progressive Part-Token Expansion and Slot-Wise Momentum Contrast for Generalized Zero-Shot Learning
Zhenghui Luo, Min Meng 0001, Jigang Wu |
ICIC (5) | 4 |
| 2026 | Text-aided alignment for multi-view clustering
Min Meng 0001, Jigang Liu, Jigang Wu |
Expert Syst. Appl. | 4 |
| 2026 | Adaptive Pseudo-labeling for Causal-driven Cross-Modal hashing
Qintai Hu, Min Meng 0001, Zetao Ma, Jigang Wu |
Expert Syst. Appl. | 6 |
| 2026 | AnomalyLVM:Vision-language models for zero-shot anomaly detection
Min Meng 0001, Jigang Wu |
Expert Syst. Appl. | 3 |
| 2026 | Dynamic patch selection and dual-granularity alignment for cross-modal retrieval
Zhenghui Luo, Min Meng 0001, Jigang Wu |
Neurocomputing | 3 |
| 2026 | Pro-CLIP: Residual learning and object-agnostic prompts for few-shot anomaly detection
Min Meng 0001, Jigang Wu |
Neurocomputing | 3 |
| 2026 | CLIP-guided sample selection for active domain adaptation
Zengmao Li, Min Meng 0001, Jigang Liu, Jigang Wu |
Knowl. Based Syst. | 4 |
| 2026 | ONE-FOR-ALL: Towards unified zero-shot anomaly detection via adaptive prompt learning
Min Meng 0001, Jigang Wu |
Pattern Recognit. | 3 |
| 2026 | GaitPR: A high-confidence multi-granular residual attention network for gait recognition
Xiaona Zheng, Qintai Hu, Shuping Zhao, Jigang Wu |
Pattern Recognit. Lett. | 4 |
| 2026 | Adaptive neighbor-aware alignment for multi-view clustering
You Xiang, Min Meng 0001, Jigang Liu, Jigang Wu |
Signal Process. | 4 |
| 2026 | DAGSIS: A DAG-Aware MAGIC-Based Synthesis Framework for In-Memory ComputingabstractThis paper presents a comprehensive synthesis framework, named DAGSIS, for memristor-aided logic (MAGIC)-based in-memory computing system. DAGSIS addresses the limitations of prior works, such as overlooking the benefits of MAGIC’s high fan-in capability and the impact of global properties of netlists on the scheduling of computation sequence (CS). DAGSIS achieves the optimization in two synthesis stages. In the technology-independent optimization stage, DAGSIS encourages the merging of nodes in the network to reduce circuit size, by utilizing equivalent transformation of multiplexer (MUX). In the CS scheduling stage, DAGSIS introduces two schemes for optimizing area overhead and latency, respectively. For area optimization, DAGSIS maximizes the utilization of memristive cells by erasing the expired data as early as possible. For latency optimization, DAGSIS aims to minimize erasing operations, by maximizing the number of erased cells in each epoch of filling the memory. To achieve better CS scheduling, DAGSIS introduces two design rules to guide CS scheduling, which fully considers the global attributes of circuit design, such as critical path and high fan-out nodes. Experiment results show that DAGSIS reduces the circuit size by 6.69% on ISCAS’85 benchmarks compared to ABC tool, an open-source logic synthesis framework. Compared to the state-of-the-art works, DAGSIS achieves a reduction of 40.68% and 12.67% in area overhead and erasing operations respectively, on ISCAS’85 and EPFL benchmarks. The improvements are further translated into the reduction in energy consumption by up to 13.7%. Lian Yao, Jigang Wu, Peng Liu 0045, Siew-Kei Lam |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2026 | A Novel Memristive Combinational Logic for Accelerating N-bit AddersabstractMemristors are anticipated to replace CMOS technology due to their low power consumption and high-speed in-memory processing capabilities. However, most existing in-memory technologies implement traditional Boolean functions to achieve complex functions, which lead to increased logical depth and long delays due to the repeated iterations. In this paper, we propose a novel combinational logic, namely the AND-OR gate, which integrates the functionality of AND and OR logic into a single function. The proposed gate is able to implement multiple commonly-used logic functions within a single cycle. To highlight the advantages of the proposed AND-OR gate, twoN-bit adders, based on parallel prefix algorithms, are designed by integrating the AND-OR gate into a memristive crossbar array. Benefiting from the proposed AND-OR gate that supports prefix computation within a single cycle, the latency of the adders is significantly reduced to$O(log(N))$. Compared with the fastest reported adder that uses Majority gate for the implementation, our proposed adder achieves notable performance improvements of$1.2\times $and$2.7\times $in terms of latency and area, respectively. Moreover, the Figure of Merits (FoMs) are employed for fair comparison, in which both latency and area are considered simultaneously. Simulation results demonstrate a remarkable improvement of$40\times $over state-of-the-art circuit (i.e., the carry-select adder). Lian Yao, Jigang Wu, Peng Liu 0045, Siew-Kei Lam |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2026 | DKGZSL: Leveraging Dynamic Visual-Semantic Knowledge for Generative Zero-Shot LearningabstractGenerative Zero-Shot Learning (GZSL) methods address the challenge of recognizing unseen classes by synthesizing visual features, thereby converting ZSL into a supervised learning task. However, existing approaches are predominantly constrained to two multi-stage strategies: pre-generation prior knowledge enhancement and post-generation feature refinement. These paradigms often suffer from error propagation across stages, ultimately limiting generation quality and representational fidelity. To overcome these limitations, we propose DKGZSL, a novel generative framework that injects dynamic visual-semantic knowledge directly into the feature synthesis process, effectively unifying generation and refinement into a single cohesive stage. Specifically, a Knowledge Transfer Network (KTN) is introduced to convert semantic information into hierarchical visual knowledge representations. To ensure accurate semanticvisual alignment, we further design a Semantic-Oriented Visual Refinement (SOVR) module that reshapes real visual features into semantically aligned and noise-suppressed representations, providing precise guidance for the KTN. Moreover, hierarchical knowledge extracted from each KTN layer is progressively transmitted to the generator via Meta-Fusion Units (MFUs), enabling dynamic semantic guidance and improving generation quality. Extensive experiments on three benchmark datasets demonstrate that DKGZSL achieves consistent state-of-the-art performance with both ResNet-101 and ViT-B/16 feature extractors. Comprehensive ablation studies further confirm the effectiveness and complementarity of each proposed component. The code is available at https://github.com/JingHu-gdut/DKGZSL. Min Meng 0001, Jigang Liu, Jun Yu 0002, Jigang Wu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | UAD: A Unified Model for Zero-Shot Anomaly DetectionabstractZero-shot anomaly detection (ZSAD) aims to identify and localize anomalies in previously unseen target domains without accessing any target-domain training data, which is crucial under privacy, security, or proprietary constraints. However, existing ZSAD methods often struggle to generalize across domains, as they are tightly coupled to specific object categories or rely on fragmented designs that fail to capture both semantic consistency and structural abnormality. In this paper, we propose UAD, a unified framework that addresses ZSAD from a holistic perspective by jointly modeling semantic regularity and anomaly-aware representations. The key insight of UAD is that effective ZSAD requires aligning multi-level semantic understanding with fine-grained structural cues, rather than relying solely on object-centric semantics or local appearance statistics. To this end, UAD organizes image representations into coherent semantic contexts and identifies anomalies as deviations from both local structural patterns and high-level semantic consistency. Furthermore, we enhance cross-domain robustness by improving semantic supervision and diversity through prompt concatenation and intensity-guided anomaly synthesis, enabling UAD to better generalize to unseen anomaly types and domains. Extensive experiments on 17 real-world anomaly detection datasets show that UAD achieves superior zero-shot performance of detecting and segmenting anomalies in datasets of highly diverse class semantics from various defect inspection and medical imaging domains. Our works are available at https://github.com/hanli6688/UAD. Min Meng 0001, Jigang Wu, Ruiwei Xie |
IEEE Trans. Image Process. | 3 |
| 2026 | Vehicle Coalition-Based Incentive Algorithm for Model Deployment and Task Offloading
Yalan Wu, Zhibing Fang, Longkun Guo, Jigang Wu |
IEEE Trans. Netw. Serv. Manag. | 5 |
| 2026 | Maximizing Edge Throughput in Collaborative Multi-Task Inference With Shareable Model StructuresabstractRecent studies in collaborative edge computing fail to take advantage of shareable structures in multi-task learning (MTL) models and the potential of MTL models sharing at the edge. This leads to resource under-utilization at the edge. Thus, this paper focuses on shareable-structure-aware model deployment and task scheduling in collaborative multi-task inference, so as to fully utilize the low-latency potential of edge computing. Specifically, we formally define the problem with an objective to maximize edge throughput under multiple constraints (e.g., resource constraint, model integrity, etc.), and prove that it is NP-hard. To solve the problem, we first propose an approximation algorithm based on randomized rounding to generate sub-optimal solutions. We then present an adjustment strategy that generates a feasible solution when the approximation algorithm violates any of the given constraints. To evaluate the proposed algorithms, we conduct comprehensive simulations based on state-of-the-art MTL models, Google cluster-usage trace, and four kinds of computing units. Extensive experiments show that the proposed approximation algorithm coupled with the adjustment strategy, outperforms state-of-the-art methods for all cases, in terms of edge throughput. Yalan Wu, Jigang Wu, Longkun Guo, Siew-Kei Lam |
IEEE Trans. Serv. Comput. | 4 |
| 2026 | Recurrent semantic disentanglement: enhancing zero-shot learning through dynamic feature refinement
Min Meng 0001, Jigang Wu |
Vis. Comput. | 3 |
| 2025 | Enhanced Discriminant Sparse Feature Extraction for Image Classification
Zhuojie Huang, Jigang Wu |
ADMA (1) | 4 |
| 2025 | GOLD: Guiding Contrastive Learning with Out-of-Distribution Detection for Universal Domain AdaptationabstractUniversal domain adaptation seeks to extend knowledge from a labeled source domain to an unlabeled target domain, unconstrained by label space alignment. It is challenging to align shared categories and separate private ones without prior category overlap information. Existing methods often rely heavily on source domain information to learn a transferable classifier, neglecting the relationships between the manifold structures inherent within the two domains. This paper introduces a novel framework, called Guiding cOntrastive Learning with out-of-distribution Detection (GOLD) for universal domain adaptation. GOLD employs an instance-prototype hybrid contrastive learning approach with self-attention to reveal domain structures. Subsequently, it constructs a residual subspace from source prototypes to filter unknown categories and refines neighborhood structures through instance-level virtual adversarial training to reduce noise. Experiments on three datasets show GOLD surpasses current methods in various UniDA settings. Jigang Wu, Jigang Liu, Min Meng 0001 |
CSCWD | 2 |
| 2025 | Coalition Formation-Based Auction for Deep Neural Network Inference in Vehicular Edge Computing
Zhibing Fang, Yalan Wu, Jigang Wu |
ICA3PP (2) | 4 |
| 2025 | DAGait: Enhancing Gait Recognition with Dynamic Adversarial TrainingabstractGait recognition systems are increasingly deployed in real-world scenarios where adversarial robustness is critical but remains underexplored. To this end, we propose DAGait, a novel adversarial defense framework for gait recognition, which operates in the gait embedding space to address the incompatibility of conventional image-space defenses with metric-based gait models. Specifically, a feature-level progressive adversarial training strategy (FPA) is proposed, which applies bounded perturbations and gradually integrates adversarial examples into the training phase, enabling a smooth transition from clean to robust representations. We further propose TriWeight, a novel triplet loss that emphasizes adversarially hard examples while employing three dynamic weights to mitigate view variations and sample interference. Extensive experiments conducted on the CASIA-B* dataset demonstrate that DAGait significantly enhances adversarial robustness, offering a practical and effective solution for secure gait-based biometric systems. Xiaona Zheng, Qintai Hu, Shuping Zhao, Jigang Wu |
IJCB | 4 |
| 2025 | Dependency-driven spectral embedding based multi-view clustering
Zien Liang, Zhuojie Huang, Shuping Zhao, Jigang Wu |
Appl. Intell. | 4 |
| 2025 | High-confidence alignment and clustering for multi-view clustering
You Xiang, Min Meng 0001, Jigang Liu, Jigang Wu |
Appl. Intell. | 4 |
| 2025 | Incentive-Based Two-Level Scheduling Algorithms for Load Balance in Vehicular Edge ComputingabstractIn vehicular edge computing (VEC), two-level scheduling both at intra-vehicle and inter-vehicle offers great potential to improve quality of services for deep neural network (DNN) inference. However, existing works on two-level scheduling failed to jointly consider load balance among vehicles and the selfishness of vehicles, which results in the absence of guarantee in quality of services. This paper seeks to fill this gap by formulating an incentive problem associated with two-level scheduling aimed at load balance for DNN inference in VEC, with the objective of maximize the system utility in VEC under the constraints of per task response time, per vehicle energy consumption, per vehicle utility guarantee, etc. Then, we prove the problem is NP-complete. A coalition based incentive algorithm, called CBA, is proposed. CBA makes intra-vehicle scheduling decisions by a heuristic strategy and it makes inter-vehicle scheduling decisions by a coalition game based strategy. The Nash-stable and convergence for CBA are proved. In addition, a deep reinforcement learning based algorithm, called DRL, is proposed to solve the formulated problem. DRL introduces a heuristic strategy to generate the intra-vehicle scheduling decisions, and it exploits deep reinforcement learning method to generate the inter-vehicle scheduling decisions. The proposed algorithms are evaluated on a platform with CPUs, SCALE-Sim, OSM and SUMO. Simulation results show that two proposed algorithms outperform the state-of-the-art methods for all cases, in terms of system utility. Compared with two baseline algorithms, CBA and DRL improve system utility by an average of 0.56× and 1.24×, respectively, for different numbers of vehicles. Yalan Wu, Rongtian Zhang, Longkun Guo, Jigang Wu |
IEEE Internet Things J. | 5 |
| 2025 | HARBOR: Harnessing Bandwidth, Computation, and Batch for Fair QoE Having Collaborative Edge-AI Services in Industrial CPSabstractInadequate resource coordination and control can result in poor quality of experience (QoE) for user devices in heterogeneous edge-enabled cyber-physical systems. Unfortunately, in a cooperative edge network, existing studies have rarely jointly optimized communication, computing resources, and batch size for QoE guarantee when controlling task offloading. To this end, we investigate the problem of harnessing bandwidth, computation, and batch size for fair quality of experience (HARBOR) in a practical collaborative edge-AI environment, where UEs have different accuracy requirements of inference services and edge devices possess different batch processing capabilities. Specifically, we introduce the task completion efficiency as the task-completion-time-to-deadline ratio to quantify individual QoE. Then, we formulate the problem HARBOR as a mixed integer nonlinear programming with constraints of accuracy, bandwidth, computation, task hard deadlines and so on. The objective is to minimize the maximum task completion efficiency among all tasks to achieve task-level fairness. After providing the NP-hardness proof for HARBOR, we then devise an efficient scheme named e-HARBOR with a competitive ratio guarantee, to solve the decoupled sub-problems of HARBOR with calibrated long short-term memory network for resource prediction. Both testbed and simulation experiments evidently demonstrate that the proposed scheme works efficiently and scales well compared to baselines. Long Chen 0006, Shaojie Zheng, Jigang Wu, Hongning Dai, Dusit Niyato, Jiafu Wan |
IEEE J. Sel. Areas Commun. | 3 |
| 2025 | Deep multi-view clustering with diverse and discriminative feature learning
Junpeng Xu, Min Meng 0001, Jigang Liu, Jigang Wu |
Pattern Recognit. | 4 |
| 2025 | CoDi: Contrastive Disentanglement Generative Adversarial Networks for Zero-Shot Sketch-Based 3D Shape RetrievalabstractSketch-based 3D shape retrieval has attracted increasing attention in recent years. Most existing methods fail to address the zero-shot scenario, and the few dedicated to zero-shot learning encounter the following two issues: 1) the features learned by these methods lack informativeness and generalization, rendering them ineffective in identifying unseen samples; 2) the generation of low-quality samples, aimed at facilitating the recognition of unseen categories, paradoxically diminishes their ability to identify these unseen classes. This paper introduces a novel contrastive disentanglement generative adversarial networks (CoDi) tailored for zero-shot sketch-based 3D shape retrieval. Initially, we introduce a paradoxical feature construction approach designed to assist the networks in capturing certain low-level features. Despite their weak semantic relevance, these features play a crucial role in sample recognition. Subsequently, a SemContrast fusion module is employed to align the semantic space with the prototype embedding space of categories. This alignment facilitates knowledge transfer to unseen classes and promotes the generation of high-quality samples. The networks are jointly trained on real and generated samples to achieve retrieval for unseen categories. Extensive experiments demonstrate a significant improvement in retrieval performance for unseen categories using our method. Min Meng 0001, Wenhang Chen, Jigang Liu, Jun Yu 0002, Jigang Wu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Cross-Scatter Sparse Dictionary Pair Learning for Cross-Domain ClassificationabstractIn cross-domain recognition tasks, the divergent distributions of data acquired from various domains degrade the effectiveness of knowledge transfer. Additionally, in practice, cross-domain data also contain a massive amount of redundant information, usually disturbing the training processes of cross-domain classifiers. Seeking to address these issues and obtain efficient domain-invariant knowledge, this paper proposes a novel cross-domain classification method, named cross-scatter sparse dictionary pair learning (CSSDL). Firstly, a pair of dictionaries is learned in a common subspace, in which the marginal distribution divergence between the cross-domain data is mitigated, and domain-invariant information can be efficiently extracted. Then, a cross-scatter discriminant term is proposed to decrease the distance between cross-domain data belonging to the same class. As such, this term guarantees that the data derived from same class can be aligned and that the conditional distribution divergence is mitigated. In addition, a flexible label regression method is introduced to match the feature representation and label information in the label space. Thereafter, a discriminative and transferable feature representation can be obtained. Moreover, two sparse constraints are introduced to maintain the sparse characteristics of the feature representation. Extensive experimental results obtained on public datasets demonstrate the effectiveness of the proposed CSSDL approach. Jigang Wu, Shuping Zhao, Jiaxing Li 0009 |
IEEE Trans. Multim. | 2 |
| 2024 | Lotus: Loading Cost-Aware Joint Mining Service Caching, Request Routing, and Bandwidth Orchestration in Cooperative MEC Networks
Long Chen 0006, Yalan Wu, Jigang Wu |
ADMA (1) | 4 |
| 2024 | Quantization-aware Optimization Approach for CNNs Inference on CPUsabstractData movements through the memory hierarchy are a fundamental bottleneck in the majority of convolutional neural network (CNN) deployments on CPUs. Loop-level optimization and hybrid bitwidth quantization are two representative optimization approaches for memory access reduction. However, they were carried out independently because of the significantly increased complexity of design space exploration. We present QAOpt, a quantization-aware optimization approach that can reduce the high complexity when combining both for CNN deployments on CPUs. We develop a bitwidth-sensitive quantization strategy that can perform the trade-off between model accuracy and data movements when deploying both loop-level optimization and mixed precision quantization. Also, we provide a quantization-aware pruning process that can reduce the design space for high efficiency. Evaluation results demonstrate that our work can achieve better energy efficiency under acceptable accuracy loss. Jiasong Chen, Zeming Xie, Weipeng Liang, Bosheng Liu, Xin Zheng 0001, Jigang Wu, Xiaoming Xiong |
ASPDAC | 6 |
| 2024 | Multi-scale Similarity Information Fusion Hashing for Unsupervised Cross-Modal Retrieval
Jianxi He, Min Meng 0001, Jigang Liu, Jigang Wu |
CGI (2) | 4 |
| 2024 | Distributed Incentive Algorithm for Fine-Grained Offloading in Vehicular Ad Hoc Networks
Junhong Wu, Yalan Wu, Jigang Wu |
ICA3PP (4) | 4 |
| 2024 | Sparse Discriminant Graph Embedding for Feature Extraction
Shuping Zhao, Xinpeng Zhang 0003, Jigang Wu |
ICIC (5) | 5 |
| 2024 | Accelerating Frequency-domain Convolutional Neural Networks Inference using FPGAsabstractLow-end field programmable gate arrays (FPGAs) are difficult to deploy typical convolutional neural networks (C- NNs) owing to the limited hardware resources and the increasing model computational complexity. Fast Fourier transform (FFT) is a promising solution for saving both computation and memory footprint by convolving in the frequency domain. However, few FPGA accelerators can take full advantage at the computation level, because of the distinct element-wise complex calculation in the frequency domain. In this work, we present an FPGA-based 8-bit inference accelerator (called FAF) that packs frequency-domain calculations into digital signal processing (DSP) blocks to fully utilize DSPs for performance boost. We then provide a mapping dataflow to maximize the reduction of redundant packing operations by frequency-domain data reuse. Evaluations based on representative CNN benchmarks show that our work can achieve 1.5-6.9× better power efficiency compared with representative FPGA baselines. Bosheng Liu, Yongqi Xu, Jigang Wu, Xiaoming Chen 0003, Peng Liu 0045, Qingguo Zhou, Yinhe Han 0001 |
ISCAS | 4 |
| 2024 | Deep Reinforcement Learning Helps: Making Loading Cost-aware Joint Cooperative Edge-AI Service Deployment and Request Routing PracticalabstractCooperative edge artificial intelligence (AI) has shown its advantages via edge-edge collaboration. By deploying deep neural network (DNN) inference service models at the edge server, the lifetime of user devices (UDs) can be prolonged through computation offloading. In practise, the service configuration delay or loading cost can potentially degrade the performance of cooperative edge-AI services. Although there have been recent studies on service caching and request routing having loading cost in mind, there exists performance gap between theory and practise, especially when UD applications have stringent deadlines, for example, running big data inference applications. This paper thus resolves the flaw of algorithm running time violates the task deadline using deep reinforcement learning. The original loading cost-aware joint cooperative edge-AI service deployment and computation offloading problem is reformulated with Markov decision process. The state, action spaces and the reward function have been well defined and the objective is to minimize the difference between the target network and evaluation network. Extensive simulation results demonstrate that compared with benchmark algorithms, the proposed algorithm can achieve more than 200 times performance gain on the algorithm running time, while obtains over 15% throughput enhancement than the benchmarks on various indices. Long Chen 0006, Jigang Wu |
ISPA | 4 |
| 2024 | Envision: Application Level Fairness for Cooperative Edge-AI Services with Deep Reinforcement LearningabstractEdge artificial intelligence (Edge-AI) is emerging with the proliferation of both multi-access edge computing (MEC) and AI. Cooperative Edge-AI can not only increase the computing resource utilization ratio with edge-edge collaboration, but also improve the big data processing efficiency of mobile end devices through computation offloading to a group of edge servers. Existing paradigms for cooperative edge-AI applications are not tailored for heterogeneous types of applications, thus harming the quality of experience (QoE) of different application users or operators in the network. This paper thus fills the gap by firstly defining the fairness index as service completion ratio, and then formulating the max-min fairness problem subject to edge server’s storage, computation, deadline constraints and so on. The problem is proven to be NP-hard through reduction from a well-known NP-complete problem, the multi-knapsack problem. To tackle the dynamics of both computing resources and channel fading conditions, a deep reinforcement learning algorithm is invented on the basis of buffer replay and evaluation-target networks, to derive the joint service deployment and computation offloading strategy. Extensive experimental results demonstrate that the proposed scheme named as Envision is at least 17× faster than the existing ORA algorithm. Long Chen 0006, Shaojie Zheng, Jigang Wu |
ISPA | 4 |
| 2024 | Energy Balanced Cooperative Edge-AI Services for Service Quality GuaranteeabstractCooperative edge computing has shown its advantage to expedite the computing speed and enhance resource utilization ratio when offering edge-AI services. Under such setting, existing works have studied the joint service deployment and request routing problem with cooperative edge servers, however, they have seldom considered the energy balance of edge servers, especially for those battery limited edge devices. To guarantee the edge-AI service quality and achieve a balanced energy consumption, this work thus addresses the joint service deployment and request offloading problem by optimizing the maximum energy consumption of an edge node in each small cell base station in the heterogeneous network. The problem is proven to be NP-hard with a randomized approximation solution. Experimental results demonstrate that the proposed algorithm RRMME can well guarantee the service quality and achieve energy balance. Compared to the designed benchmark algorithm without service quality guarantee, but has energy budget constraint, RRMME algorithm can significantly reduce the average energy consumption by about 28.7%, while has only a slightly service completion rate reduction of less than 8% averagely. Kongyang Li, Long Chen 0006, Tianwen Peng, Jigang Wu |
ISPA | 4 |
| 2024 | Contract-based Service Fairness Guarantee in Vehicular Fog NetworksabstractIn vehicular ad-hoc networks, vehicle-to-vehicle fog computing (VFC) can not only alleviate the computing delay of inference tasks from vehicles, but also reduce the computational overheads of RSUs. Existing studies on cooperative vehicle task computation offloading assume that RSUs can obtain global computing capability information of vehicles and the service-providing vehicles are always willing to offer services, while overlooking the privacy and selfishness of vehicles. Motivated by contract theory, we propose a joint service caching and task offloading NP-hard problem for vehicular fog computing, aiming to maximize the minimum service completion rate and to offer incentives for both service vehicles and RSUs. By designing contract-based joint service caching and task offloading algorithms, vehicles are encouraged to provide fog computing resources while protecting privacy. Extensive simulation results show that the proposed greedy algorithm CGA can improve the minimum service completion rate by over 10.8% and 14.7%, given fixed number of edge servers, when compared to a benchmark algorithm without contract, and a designed contract-based algorithm that maximizes the total throughput. Moreover, the designed contract-based approximation algorithm CRA can achieve the performance that are close to the benchmark algorithm without contract. Long Chen 0006, Jigang Wu, Ming Tao 0001 |
ISPA | 3 |
| 2024 | Having Energy Depletion in Mind to Make Service Fairness PracticalabstractThe energy consumption of edge devices or nodes is critical to ensure a long lifetime of cooperative edge-AI service network, which has been somehow overlooked in the literature. Failure to accommodating the energy depletion can not only bring quality of service degradation of mobile terminal devices, but can also harm the connectivity of the multi-access edge computing network. This paper thus addresses the application service fairness problem under energy depletion constraints, to make the service fairness paradigm practical to suit for energy-limited edge nodes, e.g., the unmanned areal vehicles (UAVs), solar energy powered road side units (RSUs). The problem is formulated as a non-convex integer linear programming, which is NP-hard. Then a randomized rounding algorithm as well as a greedy algorithm are designed to maximize the minimum service type’s completion rate. Extensive simulation results have shown that compared to the algorithms without energy constraints, the proposed randomized rounding algorithm and greedy algorithm with energy constraints can reduce the average energy consumption by about 39.87% and 40.31% respectively, at the cost of a mild average system throughput degradation. Tianwen Peng, Long Chen 0006, Jigang Wu |
ISPA | 4 |
| 2024 | A Complementary Resistive Switch-Based Balanced Ternary LogicabstractMemristors offer advantages in terms of high speed, high integration density, and non-volatility, making them a promising option for efficient logic applications. Recent works have explored the design methodology for ternary logic in memristor-based computing-in-memory (CIM) systems. However, existing methods require a large number of devices and are susceptible to noise interference. To address these issues, this work proposes a reliable in-memory computing paradigm for balanced ternary logic based on complementary resistive switch (CRS), which can be considered as two anti-serially connected memristors. Six balanced ternary logic gates are designed based on the proposed method, which support parallel operations when integrated into the CRS crossbar array. To demonstrate the efficiency of the proposed method, a 1-tri full adder is designed by using the proposed logic gates. The feasibility of the design is verified by Cadence Virtuoso using the Voltage Threshold Adaptive Memristor (VTEAM) model. The Monte Carlo simulation of the full adder verifies the reliability of the proposed method. Compared to existing methods, both the operation steps and area overhead are reduced using the proposed approach. Zhijian Peng, Peng Liu 0045, Lian Yao, Zhiqiang You, Bosheng Liu, Jigang Wu |
ITC-Asia | 7 |
| 2024 | Balanced Clustering with Discretely Weighted Pseudo-label
Zien Liang, Shuping Zhao, Zhuojie Huang, Jigang Wu |
PRCV (1) | 4 |
| 2024 | Coalitional Double Auction For Ridesharing With Desired Benefit And QoE ConstraintsabstractAbstract Ridesharing is an effective approach to alleviate traffic congestion. In most existing works, drivers and passengers are assigned prices without considering the constraints of desired benefits. This paper investigates ridesharing by formulating a matching and pricing problem to maximize the total payoff of drivers, with the constraints of desired benefit and quality of experience. An efficient algorithm is proposed to solve the formulated problem based on coalitional double auction. Secondary pricing based strategy and sacrificed minimum bid based strategy are proposed to support the algorithm. This paper also proves that the proposed algorithm can achieve a Nash-stable coalition partition in finite steps, and the proposed two strategies guarantee truthfulness, individually rational and budget balance. Extensive simulation results on the real-world dataset of taxi trajectory in Beijing city show that the proposed algorithm outperforms the existing ones, in terms of average total payoff of drivers while meeting the benefits of passengers. Jigang Wu, Long Chen 0006, Yalan Wu, Yidong Li |
Comput. J. | 2 |
| 2024 | Discriminative Subspace Learning With Adaptive Graph RegularizationabstractAbstract Many subspace learning methods based on low-rank representation employ the nearest neighborhood graph to preserve the local structure. However, in these methods, the nearest neighborhood graph is a binary matrix, which fails to precisely capture the similarity between distinct samples. Additionally, these methods need to manually select an appropriate number of neighbors, and they cannot adaptively update the similarity graph during projection learning. To tackle these issues, we introduce Discriminative Subspace Learning with Adaptive Graph Regularization (DSL_AGR), an innovative unsupervised subspace learning method that integrates low-rank representation, adaptive graph learning and nonnegative representation into a framework. DSL_AGR introduces a low-rank constraint to capture the global structure of the data and extract more discriminative information. Furthermore, a novel graph regularization term in DSL_AGR is guided by nonnegative representations to enhance the capability of capturing the local structure. Since closed-form solutions for the proposed method are not easily obtained, we devise an iterative optimization algorithm for its resolution. We also analyze the computational complexity and convergence of DSL_AGR. Extensive experiments on real-world datasets demonstrate that the proposed method achieves competitive performance compared with other state-of-the-art methods. Zhuojie Huang, Shuping Zhao, Zien Liang, Jigang Wu |
Comput. J. | 4 |
| 2024 | Mutual dimensionless improved bearing fault diagnosis based on Bp-increment broad learning system in computer vision
Qintai Hu, Shuping Zhao, Jigang Wu, Jianbin Xiong |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | Blockchain-based public auditing with deep reinforcement learning for cloud storage
Jiaxing Li 0009, Jigang Wu, Jin Li 0002 |
Expert Syst. Appl. | 2 |
| 2024 | Data Collection Algorithms for Model Training in Internet of VehiclesabstractIn Internet of Vehicles (IoV), it is critical to collect sufficient data for model training, to support vehicular intelligent applications. However, the environment of IoV is highly dynamic due to the mobility of vehicles, making it challenging to efficiently allocate resources for data collection. Additionally, timely training of machine learning models with collected data is important for accurate representation in a constantly changing environment. This article aims to improve the performance of model training by collecting sufficient data from vehicles in IoV. A system throughput maximization problem is formulated under the limited bandwidth, storage, and computing resources, which is an NP-hard problem. To solve the problem, an iterative algorithm, namely, the iterative algorithm based on approaching minimum bandwidth (IAMB), is proposed to preferentially collect data from the vehicles with sufficient data and reliable communication quality. Besides, a genetic algorithm, namely, the genetic algorithm based on approaching minimum bandwidth (GAMB), is proposed to further improve the probability of superior individual by replacing operation. We also customize three greedy strategy-based algorithms as the baselines. Extensive experimental results show that our proposed algorithms outperform the baseline algorithms for all cases. Specifically, IAMB and GAMB can improve the throughput by up to 5% and 8%, respectively, compared with baseline algorithms. In addition, the customized genetic algorithm is also superior to the iterative algorithm on performance of system throughput. Moreover, the customized genetic algorithm is more stable than the proposed iterative algorithm in terms of system throughput for model training in the dynamic network environment. Yifei Sun 0017, Jigang Wu, Yalan Wu, Long Chen 0006, Weijun Sun, Yidong Li |
IEEE Internet Things J. | 2 |
| 2024 | Domain-invariant feature learning with label information integration for cross-domain classification
Jigang Wu, Shuping Zhao, Jiaxing Li 0009 |
Neural Comput. Appl. | 2 |
| 2024 | Coding self-representative and label-relaxed hashing for cross-modal retrieval
Jigang Wu, Shuping Zhao, Jiaxing Li 0009 |
Pattern Recognit. Lett. | 2 |
| 2024 | Semantic Disentanglement Adversarial Hashing for Cross-Modal RetrievalabstractCross-modal hashing has gained considerable attention in cross-modal retrieval due to its low storage cost and prominent computational efficiency. However, preserving more semantic information in the compact hash codes to bridge the modality gap still remains challenging. Most existing methods unconsciously neglect the influence of modality-private information on semantic embedding discrimination, leading to unsatisfactory retrieval performance. In this paper, we propose a novel deep cross-modal hashing method, called Semantic Disentanglement Adversarial Hashing (SDAH), to tackle these challenges for cross-modal retrieval. Specifically, SDAH is designed to decouple the original features of each modality into modality-common features with semantic information and modality-private features with disturbing information. After the preliminary decoupling, the modality-private features are shuffled and treated as positive interactions to enhance the learning of modality-common features, which can significantly boost the discriminative and robustness of semantic embeddings. Moreover, the variational information bottleneck is introduced in the hash feature learning process, which can avoid the loss of a large amount of semantic information caused by the high-dimensional feature compression. Finally, the discriminative and compact hash codes can be computed directly from the hash features. A large number of comparative and ablation experiments show that SDAH achieves superior performance than other state-ofthe- art methods. Min Meng 0001, Jiaxuan Sun, Jigang Liu, Jun Yu 0002, Jigang Wu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Learning to Discover Knowledge: A Weakly-Supervised Partial Domain Adaptation ApproachabstractDomain adaptation has shown appealing performance by leveraging knowledge from a source domain with rich annotations. However, for a specific target task, it is cumbersome to collect related and high-quality source domains. In real-world scenarios, large-scale datasets corrupted with noisy labels are easy to collect, stimulating a great demand for automatic recognition in a generalized setting, i.e., weakly-supervised partial domain adaptation (WS-PDA), which transfers a classifier from a large source domain with noises in labels to a small unlabeled target domain. As such, the key issues of WS-PDA are: 1) how to sufficiently discover the knowledge from the noisy labeled source domain and the unlabeled target domain, and 2) how to successfully adapt the knowledge across domains. In this paper, we propose a simple yet effective domain adaptation approach, termed as self-paced transfer classifier learning (SP-TCL), to address the above issues, which could be regarded as a well-performing baseline for several generalized domain adaptation tasks. The proposed model is established upon the self-paced learning scheme, seeking a preferable classifier for the target domain. Specifically, SP-TCL learns to discover faithful knowledge via a carefully designed prudent loss function and simultaneously adapts the learned knowledge to the target domain by iteratively excluding source examples from training under the self-paced fashion. Extensive evaluations on several benchmark datasets demonstrate that SP-TCL significantly outperforms state-of-the-art approaches on several generalized domain adaptation tasks. Code is available at https://github.com/mc-lan/SP-TCL. Mengcheng Lan, Min Meng 0001, Jun Yu 0002, Jigang Wu |
IEEE Trans. Image Process. | 4 |
| 2024 | D-SPAC: Double-Sided Preference-Aware Carpooling of Private Cars for Maximizing Passenger UtilityabstractPrivate car-based carpooling (PCC) has become an important transportation mode in our daily life. Unlike ride-hailing or taxi-based carpooling, PCC has two unique features that have yet to be fully explored: (i) A private-car driver has more bargaining space than a non-private car driver; (ii) There exists unfriendly congestion in private car-based carpooling if not handled well. Existing carpooling schemes are not tailored for PCC services with an oversimplified assumption that passengers pay detour fees and there is no guarantee on the passenger’s travel time. Consequently, such limitations not only harm the passenger’s carpooling incentive but also hurt the passenger’s quality of experience as well as the driver’s utility. We propose a novel framework for the double-sided preference-aware carpooling (D-SPAC) problem, after comprehensively addressing the above two unique features. We formulate the D-SPAC problem as a mixed-integer non-linear programming problem, which is proved to be NP-hard, to maximize the total utility of passengers while meeting the driver’s buyout asking price, traversal radius, passenger’s waiting time, budget and both sides’ detour length constraints. We design a coalitional double auction-based scheme that can better motivate both sides with guaranteed economic properties. We further design a deep reinforcement learning algorithm to cope with the position dynamics and the changing user requests. Extensive experimental results based on real-world data sets demonstrate the effectiveness of proposed algorithms over three benchmark algorithms. Long Chen 0006, Hongning Dai, Xingyi Yuan, Yalan Wu, Jigang Wu |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2024 | Share-Aware Joint Model Deployment and Task Offloading for Multi-Task InferenceabstractIn vehicular edge computing, efficient strategies for model deployment and task offloading offer tremendous potential to reduce response time for machine learning inference. However, existing works do not pay much attention to that there are shared structures among different types of inference tasks. This limits the improvement in response time. This paper aims to fill this gap by investigating a share-aware joint model deployment and task offloading problem for multi-task inference in vehicular edge computing. We formulate the problem with an objective to minimize the total response time of all inference requests, under constraints of per task response time, per roadside unit storage capacity, etc. We prove that the formulated problem is NP-hard. To solve the problem, a time period aware algorithm, called TPA, is proposed with guaranteed approximation ratio. In TPA, an iterative approach is designed to solve the problem of maximizing system throughput during a certain time period. Then, the certain time period approximates to the minimum time period of completing all requests. The algorithms are evaluated in the environment comprising two CPUs, two GPUs, state-of-the-art multi-task learning models and the dataset of Google cluster-usage trace. Simulation results derived from this environment show that, the proposed TPA outperforms the state-of-the-art methods for all cases, in terms of the total response time of all requests. For example, TPA can significantly reduce the total response time by at least$73.72\%$for different numbers of RSUs considered, compared with state-of-the-art methods. Yalan Wu, Jigang Wu, Long Chen 0006, Bosheng Liu, Mianyang Yao, Siew-Kei Lam |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | Efficient approaches for task offloading in point-of-interest based vehicular fog computing
Yifei Sun 0017, Jigang Wu, Yalan Wu, Long Chen 0006, Weijun Sun |
J. Supercomput. | 2 |
| 2024 | Joint Dataset Reconstruction and Power Control for Distributed Training in D2D Edge NetworkabstractThe intrinsic nature of non-independent and identically distributed datasets on heterogeneous devices slows down the distributed model training process and reduces the training accuracy. To settle this problem, we propose a dataset reconstruction scheme to transform the data distribution of training device’s dataset into independent and identically distributed dataset via data exchange among trusted devices. For energy efficiency, we further consider power control for the devices. We then formulate an optimization problem, which is a mixed integer non-linear programming problem, to minimize the total energy consumption for each round of distributed training. Due to the NP-hardness and coupling property of the optimization problem, we decompose it into two subproblems for dataset reconstruction and power control, respectively. An approximation algorithm is designed to obtain a near-optimal auxiliary devices set for dataset reconstruction with minimum energy consumption, while meeting the variance constraint of the optimization problem. We prove that approximation algorithm has a worst-case approximation ratio of$1+\ln |\boldsymbol{\Omega }_{i}(t)|$, where$|\boldsymbol{\Omega }_{i}(t)|$is the required data samples for dataset reconstruction of each training device. For power control, we design a dynamic programming algorithm to further reduce the energy consumption. For comparison, we propose three benchmark schemes that adopt either one of the algorithms or neither. We also customize three baseline algorithms based on the state-of-the-arts to compare with our proposed algorithm. Numerical results show that, our proposed algorithm outperforms three benchmarks on the average energy consumption for one round for different cases. When varying the labels that each device owns, our proposed algorithm outperforms the other three baseline algorithms on training accuracy. Besides, when setting a target accuracy, our proposed algorithm always has the lowest energy consumption. Jiaxin Wu 0004, Jigang Wu, Long Chen 0006, Yifei Sun 0017, Yalan Wu |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2023 | Accelerating Convolutional Neural Networks in Frequency Domain via Kernel-Sharing ApproachabstractConvolutional neural networks (CNNs) are typically computationally heavy. Fast algorithms such as fast Fourier transforms (FFTs), are promising in significantly reducing computation complexity by replacing convolutions with frequency-domain element-wise multiplication. However, the increased high memory access overhead of complex weights counteracts the computing benefit, because frequency-domain convolutions not only pad weights to the same size as input maps, but also have no sharable complex kernel weights. In this work, we propose an FFT-based kernel-sharing technique called FS-Conv to reduce memory access. Based on FS-Conv, we derive the sharable complex weights in frequency-domain convolutions, which has never been solved. FS-Conv includes a hybrid padding approach, which utilizes the inherent periodic characteristic of FFT transformation to provide sharable complex weights for different blocks of complex input maps. We in addition build a frequency-domain inference accelerator (called Yixin) that can utilize the sharable complex weights for CNN accelerations. Evaluation results demonstrate the significant performance and energy efficiency benefits compared with the state-of-the-art baseline. Bosheng Liu, Hongyi Liang, Jigang Wu, Xiaoming Chen 0003, Peng Liu 0045, Yinhe Han 0001 |
ASP-DAC | 3 |
| 2023 | Relation-Aware Graph Attention Network for Multi-Behavior RecommendationabstractIn practical recommendation scenarios, the types of user behaviors are usually diverse (e.g., click, add-to-cart, purchase), and different types of user behaviors can provide different aspects of user preference information. However, most existing methods only consider a single type of user behavior for modeling, which is not sufficient to fully learn complex user preferences. Besides multiple behavior types, the heterogeneous preference strength of users for items under the same behavior is also a factor that is often overlooked by most methods. Moreover, different types of behaviors may be correlated due to various factors, inadequate exploration of the implicit relationships between different types of behaviors may lead to the loss of potential information across behaviors. To solve the above problems, we propose a novel multi-behavior model with relation-aware graph attention network (RGAN), which is built on a graph-based neural architecture to explore high-order user-item relations. Specifically, we design a relation-aware attention propagation layer and an inter-behavior dependency encoder to capture heterogeneous collaborative signals from type-specific and inter-type behavior relations, respectively. During behavior integration, our proposed model automatically learns which types of behaviors are more important for assisting target behavior prediction. Extensive experiments conducted on three real-world datasets demonstrate that the RGAN model consistently outper-forms many state-of-the-art baselines, in terms of HR@n and NDCG@n. Qiufen Ni, Jigang Wu |
IJCNN | 3 |
| 2023 | Robust Subspace Learning with Double Graph Embedding
Zhuojie Huang, Shuping Zhao, Zien Liang, Jigang Wu |
PRCV (7) | 4 |
| 2023 | Inter-class Sparsity Based Non-negative Transition Sub-space Learning
Miaojun Li, Shuping Zhao, Jigang Wu |
PRCV (3) | 3 |
| 2023 | Hierarchical Triple-Level Alignment for Multiple Source and Target Domain Adaptation
Zhuanghui Wu, Min Meng 0001, Tianyou Liang, Jigang Wu |
Appl. Intell. | 4 |
| 2023 | Efficient Parameter Server Placement for Distributed Deep Learning in Edge ComputingabstractAbstract Parameter servers (PSs) placement is one of the most important factors for global model training on distributed deep learning. This paper formulates a novel problem for placement strategy of PSs in the dynamic available storage capacity, with the objective of minimizing the training time of the distributed deep learning under the constraints of storage capacity and the number of local PSs. Then, we provide the proof for the NP-hardness of the proposed problem. The whole training epochs are divided into two parts, i.e. the first epoch and the other epochs. For the first epoch, an approximation algorithm and a rounding algorithm are proposed in this paper, to solve the proposed problem. For the other epochs, an adjustment algorithm is proposed, by continuously adjusting the decisions for placement strategy of PSs to decrease the training time of the global model. Simulation results show that the proposed approximation algorithm and rounding algorithm perform better than existing works for all cases, in terms of the training time of global model. Meanwhile, the training time of global model for the proposed approximation algorithm is very close to that for optimal solution generated by the brute-force approach for all cases. Besides, the integrated algorithm outperforms the existing works when the available storage capacity varies during the training. Yalan Wu, Jiaquan Yan, Long Chen 0006, Jigang Wu, Yidong Li |
Comput. J. | 4 |
| 2023 | Loading Cost-Aware Model Caching And Request Routing In Edge-enabled Wireless Sensor NetworksabstractAbstract Existing works on caching in multi-access edge computing focus on service caching and request routing. However, loading cost and execution time influenced by resource sharing have not been well exploited. To fill this gap, we investigate the joint optimization problem over deep neural network (DNN) model caching and DNN request routing with edge collaboration in edge-enabled wireless sensor networks. A problem is formulated, with the objective of maximizing throughput, under constraints of budget, accuracy and latency etc. The proof of NP-hardness for the formulated problem is provided. To solve the problem, an approximation algorithm based on randomized rounding is presented. In addition, the approximation ratio for the presented algorithm is proved to be $1/(1-\sqrt{4\ln S/\xi^\dagger})$, where $S$ is the number of edge servers and $\xi^\dagger$ is the objective value from linear relaxation. Extensive experiments demonstrate that the system throughput for the presented algorithm can be improved by 58.8% on average, compared with that of the baseline algorithm. Mianyang Yao, Long Chen 0006, Yalan Wu, Jigang Wu |
Comput. J. | 4 |
| 2023 | Intra-cluster aggregation aware routing for distributed training in wireless sensor networksabstractAbstract In wireless sensor networks (WSNs), wireless sensor nodes can be equipped with deep neural network accelerators to deal with the computation challenges in distributed training. However, the communication overhead of distributed training and the limited battery capacity of sensor nodes still impedes the broad deployment of distributed training applications. This article investigates the distributed training in WSNs by formulating an aggregation‐aware routing problem into a non‐linear integer programming problem. The objective of the formulated problem is to reduce the training time using data aggregation‐aware routing under the constraints of memory size and energy cost. Meanwhile, the NP‐Hardness of the formulated problem is proved in this article. Then, an intra‐cluster aggregation‐aware routing algorithm is proposed. The proposed algorithm accelerates the transmission of the data packet by integrating the K‐Means clustering and shortest path routing to choose the aggregators and the route paths. Extensive experiments demonstrate that the proposed algorithm outperforms two classical clustering routing algorithms UC‐LEACH and K‐Means by 29% and 37% in terms of average training time, and reducing the energy consumption by 21% and 15%, respectively. Zhaohong Chen, Long Chen 0006, Yalan Wu, Jigang Wu, Shuangyin Liu |
Concurr. Comput. Pract. Exp. | 5 |
| 2023 | Efficient decentralized access control for secure data sharing in cloud computingabstractSummary Access control is an important technique in information security that allows legitimate users to gain access to and prevent unauthorized users from getting access to resources in a system. The restriction between access from a user and a shared file of the data owner can be determined by the access policy. In most existing access control models, it is assumed that all entities, including users, the third party, and cloud service provider (CSP), are in the same trust domain. However, in cloud computing environments, it is usually assumed that the CSP cannot be fully trusted, and the data owner (DO) is desired to have the absolute initiative to control data access. This article proposes a blockchain‐based access control scheme for cloud computing, in which the DO maintains an access matrix to describe the access policy. Then, the public keys of all nodes and the access matrix are stored in the blockchain, to ensure the security of the proposed scheme. The DOs can encrypt the large shared files once using a symmetric key in a long time. And they also can encrypt the symmetric key in parallel using the public key of authorized users in a short time. Security analysis proves that the proposed scheme is able to prevent outsourced files from unauthorized access and collusion attack. Experimental results show that the proposed scheme outperforms the existing baselines in terms of overheads on computation and storage. On average, the computation overhead of the proposed scheme is lower than that of the scheme SVPAC, PpBAC, and Timely CP‐ABE by 25.37%, 45.46%, and 36.44%, respectively. The communication overhead of the proposed scheme is lower than that of the scheme Timely CP‐ABE by 17.16%, and it is more secure, although it is higher than that of the scheme SVPAC and PpBAC by 5.88% and 39.05%. And the storage overhead of the proposed scheme is lower than that of the scheme SVPAC, PpBAC, and Timely CP‐ABE by 59.36%, 20.25%, and 61.88%, respectively. Tonglai Liu, Jigang Wu, Jiaxing Li 0009, Yidong Li |
Concurr. Comput. Pract. Exp. | 2 |
| 2023 | STADIA: Photonic Stochastic Gradient Descent for Neural Network AcceleratorsabstractDeep Neural Networks (DNNs) have demonstrated great success in many fields such as image recognition and text analysis. However, the ever-increasing sizes of both DNN models and training datasets make deep leaning extremely computation- and memory-intensive. Recently, photonic computing has emerged as a promising technology for accelerating DNNs. While the design of photonic accelerators for DNN inference and forward propagation of DNN training has been widely investigated, the architectural acceleration for equally important backpropagation of DNN training has not been well studied. In this paper, we propose a novel silicon photonic-based backpropagation accelerator for high performance DNN training. Specifically, a general-purpose photonic gradient descent unit named STADIA is designed to implement the multiplication, accumulation, and subtraction operations required for computing gradients using mature optical devices including Mach-Zehnder Interferometer (MZI) and Mircoring Resonator (MRR), which can significantly reduce the training latency and improve the energy efficiency of backpropagation. To demonstrate efficient parallel computing, we propose a STADIA-based backpropagation acceleration architecture and design a dataflow by using wavelength-division multiplexing (WDM). We analyze the precision of STADIA by quantifying the precision limitations imposed by losses and noises. Furthermore, we evaluate STADIA with different element sizes by analyzing the power, area and time delay for photonic accelerators based on DNN models such as AlexNet, VGG19 and ResNet. Simulation results show that the proposed architecture STADIA can achieve significant improvement by 9.7× in time efficiency and 147.2× in energy efficiency, compared with the most advanced optical-memristor based backpropagation accelerator. Chengpeng Xia, Yawen Chen 0001, Haibo Zhang 0001, Jigang Wu |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2023 | Two-Level Scheduling Algorithms for Deep Neural Network Inference in Vehicular NetworksabstractIn vehicular networks, task scheduling at the microarchitecture-level and network-level offers tremendous potential to improve the quality of computing services for deep neural network (DNN) inference. However, existing task scheduling works only focus on either one of the two levels, which results in inefficient utilization of computing resources. This paper aims to fill this gap by formulating a two-level scheduling problem for DNN inference tasks in a vehicular network, with an objective of minimizing total weighted sum of response time and energy consumption for all tasks under the following constraints: per task response time, per vehicle energy consumption, per vehicle storage capacity. We first formulate the problem and prove that it is NP-hard. A group transformation based algorithm, called GTA, is proposed. GTA makes scheduling decisions at the network-level using the group transformation based approach, and at the microarchitecture-level using a greedy strategy. In addition, an algorithm, denoted as DRL, is proposed to decrease total weighted sum of response time and energy consumption for all tasks. DRL trains two models with deep reinforcement learning to achieve two-level scheduling. The proposed algorithms are evaluated on a platform consisting of a desktop, Raspberry Pi, Eyeriss, OSM, SUMO, NS-3. Simulation results show that DRL outperforms the state-of-the-art methods for all cases, while the proposed GTA outperforms the state-of-the-art methods for most cases, in terms of total weighted sum of response time and energy consumption. Compared with four baseline algorithms, GTA and DRL reduce the total weighted sum of response time and energy consumption by 41.49% and 62.38%, on average respectively, for different numbers of tasks. Yalan Wu, Jigang Wu, Mianyang Yao, Bosheng Liu, Long Chen 0006, Siew-Kei Lam |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Algorithms for tree-shaped task partition and allocation on heterogeneous multiprocessors
Suna He, Jigang Wu, Bing Wei 0002, Jiaxin Wu 0004 |
J. Supercomput. | 2 |
| 2023 | Blockchain-Based Secure Key Management for Mobile Edge ComputingabstractMobile edge computing (MEC) is a promising edge technology to provide high bandwidth and low latency shared services and resources to mobile users. However, the MEC infrastructure raises major security concerns when the shared resources involve sensitive and private data of users. This paper proposes a novel blockchain-based key management scheme for MEC that is essential for ensuring secure group communication among the mobile devices as they dynamically move from one subnetwork to another. In the proposed scheme, when a mobile device joins a subnetwork, it first generates lightweight key pairs for digital signature and communication, and broadcasts its public key to neighbouring peer users in the subnetwork blockchain. The blockchain miner in the subnetwork packs all the public key of mobile devices into a block that will be sent to other users in the subnetwork. This enables the mobile device to communicate with its peers in the subnetwork by encrypting the data with the public key stored in the blockchain. When the mobile device moves to another subnetwork in the tree network, all the mobile devices of the new subnetwork can quickly verify its identity by checking its record in the local or higher hierarchy subnetwork blockchain. Furthermore, when the mobile device leaves the subnetwork, it does not need to do anything and its records will remain in the blockchain which is an append-only database. Theoretical security analysis shows that the proposed scheme can defend against the 51 percent attack and malicious entities in the blockchain network utilizing Proof-of-Work consensus mechanism. Moreover, the backward and forward secrecy is also preserved. Experimental results demonstrate that the proposed scheme outperforms two baselines in terms of computation, communication and storage. Jiaxing Li 0009, Jigang Wu, Long Chen 0006, Jin Li 0002, Siew-Kei Lam |
IEEE Trans. Mob. Comput. | 2 |
| 2023 | Dual-Level Adaptive and Discriminative Knowledge Transfer for Cross-Domain RecognitionabstractUnsupervised domain adaptation is an appealing technique to learn robust classifiers for unlabeled target domain by borrowing knowledge from well-established source domain. However, previous works mainly suffer from two limitations: 1) the classifier trained on labeled source data may be prone to overfitting the source distribution, lowering its performance on the target domain; 2) the adaptation process will be misled by conditional distribution matching using hard pseudo labels of target samples. This paper presents a Dual-Level Adaptive and Discriminative (DLAD) classifier learning framework, in which transfer classifier and distribution adaptation can be mutually beneficial for effective knowledge transfer. Specifically, we aim to achieve a domain-level adaptive classifier by considering structural risk minimization (SRM) on both domains and performing weighted distribution adaptation, which facilitates joint classifier learning in a semi-supervised manner. To further achieve a class-level discriminative classifier, we explicitly leverage unlabeled target data to promote classifier learning based on class probabilities, which refines the decision boundary to be more discriminative for unlabeled target data. To the best of our knowledge, DLAD is the first attempt to consider the principle of SRM on the target domain, which significantly boosts the discriminative power of transfer classifier and yields a tighter generalization bound. Experimental evaluations on several standard cross-domain datasets show that DLAD significantly outperforms other competitive methods. Min Meng 0001, Mengcheng Lan, Jun Yu 0002, Jigang Wu, Ligang Liu 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Intrinsic and Complete Structure Learning Based Incomplete Multiview ClusteringabstractIn the real-world, some views of samples are often missing for the collected multiview data. Faced with the incomplete multiview data, most of the existing clustering methods tended to learn a common graph from the available views, where the hidden information of the absent views was ignored. Furthermore, some methods filled the absent instances with the average vector of the available samples for each view, which could not reflect a real distribution of the data. To solve these problems, in this paper an intrinsic and complete structure learning based incomplete multiview clustering method (ICSL_IMC) is proposed. Firstly, we calculate the initial complete graphs for all views by exploring the available incomplete graphs, which are further taken as the constraints for the reconstruction of the absent data integrating the self-representation method. Afterwards, encouraged by the complete multiview data, a complete structure inferring strategy is proposed to learn the intrinsic and complete structures for all views, such that the real distribution of the absent instances can be reflected in the completed structure of each view. We integrate these three learning phases into a joint optimization model, which can promote each other in the iterative learning procedure, simultaneously. Comparing with the other state-of-the-art methods, the proposed ICSL_IMC can achieve the best performances on different databases. Shuping Zhao, Lunke Fei, Jie Wen 0001, Jigang Wu, Bob Zhang 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Frequency-Domain Inference Acceleration for Convolutional Neural Networks Using ReRAMsabstractConvolutional neural networks (CNNs) (including 2D and 3D convolutions) are popular in video analysis tasks such as action recognition and activity understanding. Fast algorithms such as fast Fourier transforms (FFTs) are promising in significantly reducing computation complexity by transforming convolution into frequency domain. In frequency space, conventional spatial convolutions are replaced with simpler element-wise complex multiplications. Conventional application-specific-integrated-circuit (ASIC) based frequency-domain accelerators can achieve effective performance boost but come at the cost of significant energy consumption, owing to the hierarchical memory organization. We propose a frequency-domain resistive random access memory (ReRAM) based inference accelerator called FDA that can process element-wise complex multiplication in memory for both 2D and 3D CNNs. Each ReRAM-based frequency-domain process element (PE) with two ReRAM cells can perform an element-wise complex multiplication in two continuous execution cycles. We then provide a flexible dataflow to alleviate the redundant data movements by frequency-domain data reuse and inherent symmetrical characteristic for both 2D and 3D convolutions. Evaluation results based on representative both 2D and 3D CNN benchmarks demonstrate that FDA outperforms state-of-the-art baselines with better performance and energy efficiency. Bosheng Liu, Zhuoshen Jiang, Yalan Wu, Jigang Wu, Xiaoming Chen 0003, Peng Liu 0045, Qingguo Zhou, Yinhe Han 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2023 | Adaptive Graph Embedded Preserving Projection Learning for Feature Extraction and SelectionabstractPreserving projection learning has been widely used in feature extraction and selection for unsupervised image classification. Generally, some related methods constructed a graph to represent the nearest neighbor relationships of the data based on the Euclidean distances among different samples, which used 0 or 1 to predefine whether two samples are from the same class. Since a simple Euclidean distance is sensitive to noise, the predefined graph cannot produce exact correlations between the two samples. What is more, the predefined graph cannot reflect the structure of the projected data on a latent subspace when the projection matrix is learned. To solve these problems, in this article a novel adaptive graph embedded preserving projection learning (AGE_PPL) method is proposed, first combining the sparsity-based graph learning and the projection learning as an integral framework for feature extraction and feature selection. In particular, a sparse representation term with$l_{1}$-norm is exploited in AGE_PPL to achieve the adaptive graph of the data to preserve the local structures among different samples while the projection matrix is learned. Meanwhile, a global-scale constraint is imposed to preserve the global structure of the data on a latent subspace. Therefore, the transformed samples will be more discriminative, allowing margins of the same class to be reduced, and margins among different classes to be enlarged. Experimental results proved the effectiveness of the proposed algorithm by obtaining competitive performances over other baseline and state-of-the-art methods. In addition, the proposed method is very flexible for feature selection and dimensionality reduction. Shuping Zhao, Jigang Wu, Bob Zhang 0001, Lunke Fei, Shuyi Li 0003, Pengyang Zhao |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2023 | Adequate alignment and interaction for cross-modal retrievalabstractCross-modal retrieval has attracted widespread attention in many cross-media similarity search applications, especially image-text retrieval in the fields of computer vision and natural language processing. Recently, visual and semantic embedding (VSE) learning has shown promising improvements on image-text retrieval tasks. Most existing VSE models employ two unrelated encoders to extract features, then use complex methods to contextualize and aggregate those features into holistic embeddings. Despite recent advances, existing approaches still suffer from two limitations: 1) without considering intermediate interaction and adequate alignment between different modalities, these models cannot guarantee the discriminative ability of representations; 2) existing feature aggregators are susceptible to certain noisy regions, which may lead to unreasonable pooling coefficients and affect the quality of the final aggregated features. To address these challenges, we propose a novel cross-modal retrieval model containing a well-designed alignment module and a novel multimodal fusion encoder, which aims to learn adequate alignment and interaction on aggregated features for effectively bridging the modality gap. Experiments on Microsoft COCO and Flickr30k datasets demonstrates the superiority of our model over the state-of-the-art methods. Mingkang Wang, Min Meng 0001, Jigang Liu, Jigang Wu |
Virtual Real. Intell. Hardw. | 4 |
| 2023 | Energy-efficient cooperative offloading for mobile edge computing
Wenjun Shi, Jigang Wu, Long Chen 0006, Xinxiang Zhang, Huaiguang Wu |
Wirel. Networks | 2 |
| 2022 | Joint Matrix Factorization and Structure Preserving for Domain Adaptation
Wenhao Shao, Min Meng 0001, Jigang Wu |
CGI | 4 |
| 2022 | Weighted Graph Embedded Low-Rank Projection Learning for Feature ExtractionabstractLow-rank based methods have been widely adopted to structure preserving, when the projection matrix is learned for feature extraction. However, some dilemmas still exist that degrade the classification performance: 1) The local structure of the data is ignored; 2) the reconstructed data is not consistent with the original data. To solve those problems, in this paper a weighted graph embedded low-rank projection (WGE_LRP) method is proposed. In WGE_LRP, a novel weighted graph regularization term is proposed, which can learn the local structure of the data based on the similarity of different samples. Meanwhile, an extra global information term is introduced to keep the reconstructed data consistent with the original data. Experimental results show that the proposed method can obtain competitive performance in comparison to the state-of-the-arts. Zhuojie Huang, Shuping Zhao, Lunke Fei, Jigang Wu |
ICASSP | 4 |
| 2022 | Loading Cost-Aware Model Caching and Request Routing for Cooperative Edge InferenceabstractMost existing works on edge service caching and request routing fail to consider the influence of the service loading time. Meanwhile, the requests generated by end devices will change dynamically, which means that the caching strategy should adapt accordingly. In this paper, we investigate loading cost-aware joint model caching and request routing with cooperative edge computing, considering both the service loading time and the dynamic user requests. A system throughput maximization problem is formulated, which is proved to be NP-hard. Then, a randomized rounding-based online algorithm with M/(M − 2 ln N)-approximation ratio is proposed to solve it, where M and N are the numbers of end devices and deep neural network (DNN) models, respectively. Extensive experimental results demonstrate that our algorithm achieves more than 42.7% throughput gain than baseline algorithms. Mianyang Yao, Long Chen 0006, Jun Zhang 0004, Jigang Wu |
ICC | 5 |
| 2022 | Group Correspondence: A Statistical Perspective for Incomplete Multi-View Clustering AugmentationabstractCross-view consistency is the fundamental property of multiview clustering. However, in incomplete multi-view scenarios, existing methods can only pursue consistency through the paired data while ignoring the information in unpaired data. In this paper, we show a new insight from the data pattern and provide a novel perspective to incorporate unpaired data for consistency maximization by mining group correspondence. We first formulate cross-view consistency in a statistical perspective to by-pass the strict demand of instance correspondence, and then propose a technique to construct corresponding groups across views to enhance the objective of consistency maximization. Our proposal can be used as a universal plug-in to augment existing approaches. We test the efficacy and generality of our proposal by adapting it to two base methods as augmentations and comparing the augmented models against the original ones and other baselines. Experiment results demonstrate the effectiveness of our proposal and validate the value of our insight. Tianyou Liang, Min Meng 0001, Mengcheng Lan, Jun Yu 0002, Jigang Wu |
ICME | 5 |
| 2022 | Online Robust Specific and Consistent HashingabstractMost of the existing cross-modal hashing (CMH) methods are trained in a batch-based manner, which is time-consuming and unable to handle streaming data. Recently, online CMH methods have attracted increasing attention. However, existing online CMH methods face two limitations: 1) they usu-ally excavate the common semantic information by learning modality-specific projection matrix for each modality, while ignoring the intrinsic relationship among different modali-ties; 2) they often suffer from learning less discriminative hash codes because of insufficiently exploiting the pairwise similarity. To tackle these challenges, we propose a novel Online Robust Specific and Consistent Hashing (ORSCH) method. Specifically, ORSCH decomposes the projection ma-trices into consistent and modality-specific ones, which ef-fectively exploits intrinsic semantic information of streaming data in different modalities. Furthermore, we utilize both dis-crete and continuous labels to construct affinity matrices to improve the discrimination of hash codes. Experiments on three benchmark datasets show the superiority of ORSCH. Min Meng 0001, Jigang Wu |
ICME | 3 |
| 2022 | Triple Disentangling Network for Unsupervised Domain AdaptationabstractMost existing unsupervised domain adaptation methods learn domain-invariant representations with entangled domain in-formation, semantic information, and instance information. Differently, in this paper, we propose a Triple Disentangling Network (TDN), to disentangle these three types of information and then predict the target labels merely using semantic information. Specifically, TDN consists of a reconstruction module and a disentanglement module. In the reconstruction module, TDN utilizes a variational auto-encoder to re-construct the domain, semantic, and instance latent variables behind the data. In the disentanglement module, adversar-ial learning, discriminative clustering, and instance separation are seamlessly integrated to disentangle these three sets of re-constructed latent variables. Significantly, TDN can not only effectively alleviate the negative transfer of outliers through disentangling instance information, but also disentangle se-mantic information more thoroughly by exploring discriminative structure knowledge. Experimental studies on two bench-mark datasets demonstrate the superiority of TDN. Zhuanghui Wu, Tianyou Liang, Min Meng 0001, Jigang Liu, Jun Yu 0002, Jigang Wu |
ICME | 6 |
| 2022 | Efficient Algorithms For Storage Load Balancing Of Outsourced Data In Blockchain NetworkabstractAbstract Decentralized storage of data is one of the typical applications in the blockchain network. However, most of the existing works neglected the storage balancing problem in the blockchain network, which has an immediate impact on the availability and stability of the network. Therefore, this paper proposes a storage balancing problem for non-local data storage in the blockchain network and proves that the problem is non-deterministic polynomial (NP)-hard. The criterion of the storage balance is established by a balanced coefficient in the proposed scheme. A heuristic matching algorithm (HMA), a genetic algorithm (GA) and a tabu search algorithm (TSA) are customized to solve the problem of imbalanced storage formalized in this paper. Compared with our previous algorithm fast matching algorithm (FMA), experimental results demonstrate that HMA achieves better performance in terms of accuracy, computation overhead and storage overhead. Specifically, the computation overhead of HMA is lower than that of FMA by 84.45% on average, whereas the storage overhead of HMA is lower than that of FMA by 32.26% on average. By using the initial solution of HMA, TSA achieves the highest accuracy among GA, TSA and moth-flame optimization (MFO). Meanwhile, by using the initial solution of FMA, TSA achieves the highest accuracy among GA, TSA and MFO. Tonglai Liu, Jigang Wu, Jiaxing Li 0009, Zikai Zhang 0004 |
Comput. J. | 2 |
| 2022 | Small Target Recognition Using Dynamic Time Warping and Visual AttentionabstractAbstract Microaneurysm is a kind of small targets in color retinal image, and it is an essential work to recognize the small target for the early diagnosis of diabetic retinopathy. This paper proposes an efficient method to accurately recognize microaneurysm. A symmetric extended curvature Gabor wavelet is presented to generate candidate objects, where some novel features are extracted for classification. A kind of statistic features is generated to distinguish between microaneurysm and thin vessels, in terms of the shape similarity of cross-section profiles. Furthermore, the visual attention-based features are proposed to compute local contrast of small targets in complex background. Random undersampling with AdaBoost (RUSBoost) classifier is employed to discriminate true microaneurysm from an overwhelming amount of candidate objects. Experimental results demonstrate that the proposed method achieves significant sensitivity and accuracy on the public datasets, in comparison to the state-of-the-arts. Xinpeng Zhang 0003, Jigang Wu, Min Meng 0001 |
Comput. J. | 2 |
| 2022 | Three-stage auction scheme for computation offloading on mobile blockchain with edge computingabstractSummary Blockchain has been applied in wide range of fields to guarantee security. However, it has been very challenging for blockchain to flourish in mobile environment with limited resources. Existing studies mainly assume that single mobile user can buy the whole resources from edge servers in mobile blockchain. This paper formulates the problem of maximizing the social welfare for computation offloading in mobile blockchain. A three‐stage auction scheme with approximation ratio of based on group‐buying mechanism is proposed to allocate edge server resources for mobile blockchain applications. In the first stage, the miners are divided into groups, and a Vickrey–Clarke–Groves based auction is proposed to determine the bid of each group for each edge server. In the second stage, a matching algorithm is proposed to match edge servers and Access Points for maximizing the profit of edge servers. In the third stage, the edge server resources are allocated to mobile users for mining base on the results in the above stages. We prove that our auction scheme guarantees truthfulness, individual rationality and budget balance. Simulation results show that, the social welfare of our scheme is improved by 33.78%, 21.84%, 19.69%, and 6.69% for 1000 miners, compared with the existing works. Chengpeng Xia, Yalan Wu, Long Chen 0006, Yawen Chen 0001, Jigang Wu |
Concurr. Comput. Pract. Exp. | 5 |
| 2022 | An improved reconfigurable logic in resistive random access memory
Peng Liu 0045, Jigang Wu, Dongxiang Luo |
Integr. | 4 |
| 2022 | Reconfiguration algorithms for synchronous communication on switch based degradable arrays
Yalan Wu, Jigang Wu, Peng Liu 0045, Yinhe Han 0001, Thambipillai Srikanthan |
Parallel Comput. | 2 |
| 2022 | Low-rank inter-class sparsity based semi-flexible target least squares regression for feature representation
Shuping Zhao, Jigang Wu, Bob Zhang 0001, Lunke Fei |
Pattern Recognit. | 2 |
| 2022 | Search-Free Inference Acceleration for Sparse Convolutional Neural NetworksabstractSparse convolution neural networks (CNNs) are promising in reducing both memory usage and computational complexity while still preserving high inference accuracy. State-of-the-art sparse CNN accelerators can deliver high throughput by skipping zero weights and/or activations. To operate on only nonzero weights and activations, sparse accelerators typically search pairs of nonzero weights and activations for multiplication-accumulation (MAC) operations. However, the conventional search operation results in a severe limitation in the processing element (PE) array scale because of the enormous demands of internal interconnection and memory bandwidth. In this article, we first provide a design principle to free the search process of sparse CNN accelerations. Specifically, the indexes of the static compressed weights access the dynamic activations directly to avoid the search process for MAC operations. We then develop two search-free inference accelerators, called Swan and Swan-flexible, for sparse CNN accelerations. Swan supports search-free sparse convolution accelerations for interconnection and bandwidth saving. Compared with Swan, Swan-flexible not only has the search-free capability but also comprises a configurable architecture for optimum throughput. We formulate a mathematical optimization problem by combining the configurable characterization with the compressive dataflow to optimize the overall throughput. Evaluations based on a place-and-route process show that the proposed designs, in a compact factor of 4096 PEs, achieve 1.5–$2.7\times $higher speedup and 6.0–$13.6\times $better energy efficiency than representative accelerator baselines with the same PE array scale. Bosheng Liu, Xiaoming Chen 0003, Yinhe Han 0001, Jigang Wu, Liang Chang 0003, Peng Liu 0045 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | Dependency-Aware Computation Offloading for Mobile Edge Computing With Edge-Cloud CooperationabstractMost of existing Multi-access edge computing (MEC) studies consider the remote cloud server as a special edge server, the opportunity of edge-cloud collaboration has not been well exploited. We propose a dependency-aware offloading scheme in MEC with edge-cloud cooperation under task dependency constraints. Each mobile device has a limited budget and has to determine which sub-task should be computed locally or should be sent to the edge or remote cloud. To address this issue, we divide the offloading problem into two application finishing time minimization sub-problems with two different cooperation modes, both of which are proved to be NP-hard. We then devise one greedy algorithm with approximation ratio of$1+\epsilon$for the first mode with edge-cloud cooperation but no edge-edge cooperation. Then we design an efficient greedy algorithm for the second mode, considering both edge-cloud and edge-edge co-operations. Extensive simulation results show that for the first mode, the proposed greedy algorithm achieves near optimal performance for typical task topologies. On average, it outperforms the modified Hermes benchmark algorithm by about$23\%\sim 43.6\%$in terms of application finishing time with given budgets. By further exploiting collaborations among edge servers in the second cooperation mode, the proposed algorithm helps to achieve over 20.3 percent average performance gain on the application finishing time over the first mode under various scenarios. Real-world experiments comply with simulation results. Long Chen 0006, Jigang Wu, Jun Zhang 0004, Hongning Dai, Mianyang Yao |
IEEE Trans. Cloud Comput. | 2 |
| 2022 | Generalized Multi-View Collaborative Subspace ClusteringabstractIn real-world applications, complete or incomplete multi-view data are common, which leads to the problem of generalized multi-view clustering. Recently, researchers attempt to learn the latent representation in the common subspace from heterogeneous data, which usually suffers from feature degeneration. Moreover, there are limited efforts on simultaneously revealing the underlying subspace structure and exploring the complementary information from incomplete multiple views. In this paper, we introduce a novel Generalized Multi-view Collaborative Subspace Clustering (GMCSC) framework to address the above issues, in which consensus subspace structure of all views and embedding subspaces for each view are jointly learned to benefit each other. Specifically, we develop a novel collaborative subspace learning strategy based on self-representation learning, which provides a brand-new way of pursuing the complete subspace structure directly from multi-view data. Furthermore, we explore complementary information by enforcing the consistency across different views and preserving the view-specific information of each view, which can alleviate the problem of feature degeneration and enhance the reasonability of using a consensus representation for multiple views. Experimental results on six benchmark datasets demonstrate that the proposed method can significantly outperform the state-of-the-art algorithms. Mengcheng Lan, Min Meng 0001, Jun Yu 0002, Jigang Wu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Exploring Fine-Grained Cluster Structure Knowledge for Unsupervised Domain AdaptationabstractUnsupervised domain adaptation aims to leverage knowledge from a labeled source domain to learn an accurate model in an unlabeled target domain. However, many previous approaches propose to learn domain agnostic feature representations using a global distribution alignment objective, which does not consider the fine-grained cluster structures in the source and target domains. As such, the goal of this paper is to address two challenging problems:1) how to thoroughly explore fine-grained cluster structure knowledge in the source and target domains, 2) how to effectively incorporate these structure knowledge for adaptation.Regarding the first point, we are motivated by structural domain similarity assumption and propose structural representation learning, which is achieved by enforcing structural consistency between the source and target domains while retaining their individual discriminative properties. Regarding the second point, we firstly devise a novel structural centroid-based label prediction method, which explicitly models structural representations to form discriminative source and target cluster centroids, and estimates the label distribution of each target sample through the cosine similarity between its corresponding target cluster centroid and all the other source cluster centroids. Then, we adopt clustering learning to incorporate these discriminative structure knowledge for adaptation by minimizing the KL divergence between the predictive target label distribution and an introduced auxiliary one. Comprehensive experiments and analyses on four benchmark datasets demonstrate the superiority of the proposed discriminative clustering framework. Min Meng 0001, Zhuanghui Wu, Tianyou Liang, Jun Yu 0002, Jigang Wu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Multiview Consensus Structure DiscoveryabstractMultiview subspace learning has attracted much attention due to the efficacy of exploring the information on multiview features. Most existing methods perform data reconstruction on the original feature space and thus are vulnerable to noisy data. In this article, we propose a novel multiview subspace learning method, called multiview consensus structure discovery (MvCSD). Specifically, we learn the low-dimensional subspaces corresponding to different views and simultaneously pursue the structure consensus over subspace clustering for multiple views. In such a way, latent subspaces from different views regularize each other toward a common consensus that reveals the underlying cluster structure. Compared to existing methods, MvCSD leverages the consensus structure derived from the subspaces of diverse views to better exploit the intrinsic complementary information that well reflects the essence of data. Accordingly, the proposed MvCSD is capable of producing a more robust and accurate representation structure which is crucial for multiview subspace learning. The proposed method can be optimized effectively, with theoretical convergence guarantee, by alternatively iterating the argument Lagrangian multiplier algorithm and the eigendecomposition. Extensive experiments on diverse datasets demonstrate the advantages of our method over the state-of-the-art methods. Min Meng 0001, Mengcheng Lan, Jun Yu 0002, Jigang Wu |
IEEE Trans. Cybern. | 4 |
| 2022 | Load Balance Guaranteed Vehicle-to-Vehicle Computation Offloading for Min-Max Fairness in VANETsabstractLoad balance in vehicular ad hoc networks (VANETs) is a challenge in vehicle-to-vehicle computation offloading, due to stochastic requests of users, heterogeneous service capabilities and high mobility of vehicles, etc. This paper aims to fill this gap by formulating a problem for load balance in a VANET, with the objective of minimizing the maximum load under transmit power, storage capacity, per task completion time and energy consumption constraints. The formulated problem is proved to be NP-hard, then it is investigated by decomposing it into two subproblems, i.e., how to offload tasks for the case of fixed transmit power and how to adjust transmit power for the given offloading decision. For the first subproblem, an approximation algorithm is proposed by offloading the tasks in the vehicle with the maximum load to the vehicle with minimum load. Meanwhile, a deep reinforcement learning algorithm is proposed, in order to focus on the network dynamics. A coalition based algorithm, a distributed coalition based algorithm, as well as an incentive algorithm based on deep reinforcement learning, are proposed to maximize the total payoff for the selfishness of vehicles. For the second subproblem, an adjustment strategy for transmit power is customized to further reduce the computing load. The algorithms are evaluated on an integrated simulation platform with open street map, SUMO, NS-3 and dataset of Google cluster-usage traces. Simulation results show that, the proposed algorithms outperform three state-of-the-art works for most cases, in terms of the maximum load. The proposed distributed algorithm can significantly accelerate the proposed centralized algorithm with acceptable increase in maximum load. Besides, the load can be further reduced by the proposed adjustment strategy. Yalan Wu, Jigang Wu, Long Chen 0006, Jiaquan Yan, Yinhe Han 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | EECDN: Energy-efficient Cooperative DNN Edge Inference in Wireless Sensor NetworksabstractMulti-access edge computing (MEC) is emerging to improve the quality of experience of mobile devices including internet of things sensors by offloading computing intensive tasks to MEC servers. Existing MEC-enabled cooperative computation offloading works focus on the optimization of total energy consumption but fail to exploit multi-relay diversity and min-max fairness of energy consumption on participated sensors. We explore a typical wireless sensor network with multi-source, multi-relay, and one edge server, where relay nodes can provide both cooperative communication and computation services. We divide the energy efficiency optimization problem into two sub-problems: One is to minimize the weighted average total energy consumption per time slot, and the other is to minimize the maximum weighted energy consumption. For the first sub-problem, we propose an optimal algorithm named as optimal weighted average total energy consumption algorithm (OTCA) based on bipartite matching. For the second sub-problem, greedy algorithm for fairness guarantee (GAF) is proposed with an approximation ratio of (1 + ε), where ε is a small positive constant. Extensive numerical results show that OTCA outperforms the baseline algorithms by 26.7–77.4% on the average total weighted energy consumption while GAF outperforms benchmark algorithms by 30.7–84.4%. NS-3 simulation experiments comply with numerical results. Long Chen 0006, Mianyang Yao, Yalan Wu, Jigang Wu |
ACM Trans. Internet Techn. | 4 |
| 2022 | Dual-level contrastive learning network for generalized zero-shot learning
Jiaqi Guan, Min Meng 0001, Tianyou Liang, Jigang Liu, Jigang Wu |
Vis. Comput. | 5 |
| 2021 | Task Offloading Algorithms for Novel Load Balancing in Homogeneous Fog NetworkabstractFog computing has become an emerging distributed computing paradigm to provide services with low latency and high throughput. However, load unbalance is serious due to the difference in geography, which results in performance deterioration and low utilization of resources in the fog network. In this paper, the load is the tradeoff between the delay and energy consumption for fog nodes. Meanwhile, the problem of minimizing the maximum load in the homogeneous fog network is formulated and its NP-hardness is proved. Then, a greedy algorithm is proposed for solving the problem by giving the preference to offloading the task in the fog node with the maximum load to the fog node with the minimum load in the network. Moreover, for solving the problem with consideration of selfishness of fog nodes, a coalition based algorithm is proposed to encourage the fog nodes with a light load to share their resources to reduce the maximum load. We evaluate the performance of the proposed algorithms on NS-3 and simulation results show that the proposed algorithms outperform the existing algorithm about 40% in terms of the maximum load. Jiaquan Yan, Jigang Wu, Yalan Wu, Long Chen 0006, Shuangyin Liu |
CSCWD | 2 |
| 2021 | F3D: Accelerating 3D Convolutional Neural Networks in Frequency Space Using ReRAMabstract3D convolutional neural networks (CNNs) are widely deployed in video analysis. Fast algorithms such as fast Fourier transforms (FFTs) are gaining popularity in reducing computation complexity for their superior capability of replacing convolutions with simpler element-wise multiplications. Conventional frequency-domain dedicated accelerators employ memory hierarchy organization for high throughput but at the expensive costs of a significant amount of data movements and energy consumptions. This paper presents F3D, a processingin-memory frequency-domain accelerator using resistive random access memory (ReRAM). F3D supports frequency-domain complex number multiplications directly in ReRAM-based crossbar architecture. We alleviate the overheads of redundant data movements in ReRAM-based complex number multiplications by data reuse and the inherent symmetry of inputs in the frequency space. Evaluation results demonstrate that F3D outperforms state-of-the-art accelerators with significant improvements in performance and energy efficiency. Bosheng Liu, Zhuoshen Jiang, Jigang Wu, Xiaoming Chen 0003, Yinhe Han 0001, Peng Liu 0045 |
DAC | 3 |
| 2021 | Erasure-Coded Multi-Block Updates Based on Hybrid Writes and Common XORs FirstabstractErasure code is widely used in storage systems since it can offer higher reliability at lower redundancy than data replication. However, erasure coding based storage systems have to perform multi-block updates for partial writes of an erasure coding group, which leads to a large number of XOR operations. This paper presents an efficient approach, named ECMU, for erasure-coded multi-block update under a stringent latency by scheduling update sequences. ECMU takes a hybrid of reconstructed-write and read-modify-write for parity blocks of an erasure coding group, it dynamically selects the write scheme with the fewer XORs for each parity block to be updated, in order to reduce the number of XORs. ECMU iteratively retrieves the unmodified parity blocks to calculate the minimum XORs for each write scheme. For all parity blocks to be updated, after the write schemes are determined, ECMU performs the common XORs first, then it reuses the computational results to further reduce the number of XORs. ECMU caches a certain number of scheduling schemes to reduce the construction count of the scheduling schemes. Experimental results on real-world trace replaying show that the number of XORs and update time can be reduced significantly, compared with the state-of-the-art. Bing Wei 0002, Jigang Wu, Limin Xiao 0001 |
ICCD | 3 |
| 2021 | Learning Controlled Semantic Embedding for Cross-Modal RetrievalabstractCross-modal retrieval has caught appealing attentions as it supports querying across different modalities. However, most existing methods have emphasized on directly mapping heterogeneous features into the common subspace, which inevitably results in highly entangled representations, thereby preventing them from bridging the modality gap. This paper presents a novel deep framework called Controlled Semantic Embedding (CSE), which is the first attempt to learn disentangled representations with controlled semantic structure for cross-modal retrieval. Specifically, we design two generative networks based on variational autoencoder, which incorporate semantic discriminators for effective prediction of structured semantics. Meanwhile, a self-supervised semantic network is seamlessly integrated into the generative networks to supervise the semantic embedding process, which is further coupled with a quantizer for controlling the quantizability of semantic representations. Extensive experiments show the superiority of CSE over other state-of-the-art methods in cross-modal retrieval. Min Meng 0001, Jun Yu 0002, Jigang Wu |
ICME | 4 |
| 2021 | Stack-VAE Network for Zero-Shot Learning
Jinghao Xie, Jigang Wu, Tianyou Liang, Min Meng 0001 |
ICONIP (4) | 2 |
| 2021 | Data Delta Based Hybrid Writes for Erasure-Coded Storage Systems
Bing Wei 0002, Jigang Wu, Limin Xiao 0001 |
NPC | 4 |
| 2021 | Adaptive Updates for Erasure-Coded Storage Systems Based on Data Delta and Logging
Bing Wei 0002, Jigang Wu, Xiaosong Su |
PDCAT | 2 |
| 2021 | Photonic Computing and Communication for Neural Network Accelerators
Chengpeng Xia, Yawen Chen 0001, Haibo Zhang 0001, Hao Zhang 0058, Jigang Wu |
PDCAT | 5 |
| 2021 | Available Time Aware Offloading for Dependent Tasks with Cooperative Edge Servers
Bingyan Zhou, Long Chen 0006, Jigang Wu |
WASA (1) | 3 |
| 2021 | Long-term optimization for MEC-enabled HetNets with device-edge-cloud collaboration
Long Chen 0006, Jigang Wu, Jun Zhang 0004 |
Comput. Commun. | 2 |
| 2021 | Structure preservation adversarial network for visual domain adaptation
Min Meng 0001, Qiguang Chen, Jigang Wu |
Inf. Sci. | 3 |
| 2021 | Feature-transfer network and local background suppression for microaneurysm detection
Xinpeng Zhang 0003, Jigang Wu, Min Meng 0001, Yifei Sun 0017, Weijun Sun |
Mach. Vis. Appl. | 2 |
| 2021 | Robust Discriminant Projection Via Joint Margin and Locality Structure Preservation
Min Meng 0001, Jigang Wu |
Neural Process. Lett. | 3 |
| 2021 | Context switch cost aware joint task merging and scheduling for deep learning applications
Jigang Wu, Yalan Wu, Long Chen 0006, Yidong Li |
Parallel Comput. | 2 |
| 2021 | Fault Modeling and Efficient Testing of Memristor-Based MemoryabstractMemristor-based memory technology is one of the emerging memory technologies, which is a potential candidate to replace traditional memories. Efficient test solutions are required to enable the quality and reliability of such products. In previous works, fault models are caused by open, short and bridge defects and parametric variations during the fabrication. However, these fault models cannot describe the bridge defects that cause the state of the faulty cell to an undefined state. In this paper, we analyze the different effects of bridge defects and aggregate their faulty behavior into new fault models, undefined coupling fault and dynamic undefined coupling fault. In addition, an enhanced March algorithm is designed to detect all the modeled faults. In one resistor crossbar with$N$memristors, the enhanced March algorithm requires$8N$write and$7N$read operations with negligible hardware overhead. To reduce the test time, a March RC algorithm is proposed based on read operations with new reference currents, which requires$4N+2$write and$6N$read operations. Analytical results show that the proposed test algorithms can detect all the modeled faults outperforming all the previous methods. Subsequently, a Design-for-Testability scheme is proposed to implement March RC algorithm with a little area overhead. Peng Liu 0045, Zhiqiang You, Jigang Wu, Bosheng Liu, Yinhe Han 0001, Krishnendu Chakrabarty |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2021 | Coupled Knowledge Transfer for Visual Data RecognitionabstractTransfer learning aims to learn an effective classifier for unlabeled target data by borrowing knowledge from well-labeled source data. However, most existing work has emphasized on learning domain invariant features to reduce the distribution discrepancy, which may suffer from the negative transfer problem caused by structure inconsistencies or distribution outliers. To address this challenge, in this paper, we propose a novel transfer learning approach, which seamlessly integrates domain invariant feature learning, discriminative structure preservation and sample reweighting into a unified learning model. Specifically, we attempt to learn domain invariant features by jointly adapting the marginal and conditional distributions. To transfer discriminative knowledge inferred from data, we enforce the structure consistency between the original feature space and the latent feature space. Furthermore, to enhance the robustness of our model, an efficient and more generalized sample reweighting strategy is developed to assign target predictions with different levels of confidence. The key advantage over previous methods is that our model can adaptively select pivot samples in target domain and retain the properties of discriminative structures underlying data domains, which enables coupled knowledge transfer during the learning process. Experimental results on several benchmark datasets have verified the superiority of the proposed method over other state-of-the-art algorithms. Min Meng 0001, Mengcheng Lan, Jun Yu 0002, Jigang Wu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Asymmetric Supervised Consistent and Specific Hashing for Cross-Modal RetrievalabstractHashing-based techniques have provided attractive solutions to cross-modal similarity search when addressing vast quantities of multimedia data. However, existing cross-modal hashing (CMH) methods face two critical limitations: 1) there is no previous work that simultaneously exploits the consistent or modality-specific information of multi-modal data; 2) the discriminative capabilities of pairwise similarity is usually neglected due to the computational cost and storage overhead. Moreover, to tackle the discrete constraints, relaxation-based strategy is typically adopted to relax the discrete problem to the continuous one, which severely suffers from large quantization errors and leads to sub-optimal solutions. To overcome the above limitations, in this article, we present a novel supervised CMH method, namely Asymmetric Supervised Consistent and Specific Hashing (ASCSH). Specifically, we explicitly decompose the mapping matrices into the consistent and modality-specific ones to sufficiently exploit the intrinsic correlation between different modalities. Meanwhile, a novel discrete asymmetric framework is proposed to fully explore the supervised information, in which the pairwise similarity and semantic labels are jointly formulated to guide the hash code learning process. Unlike existing asymmetric methods, the discrete asymmetric structure developed is capable of solving the binary constraint problem discretely and efficiently without any relaxation. To validate the effectiveness of the proposed approach, extensive experiments on three widely used datasets are conducted and encouraging results demonstrate the superiority of ASCSH over other state-of-the-art CMH methods. Min Meng 0001, Haitao Wang 0026, Jun Yu 0002, Jigang Wu |
IEEE Trans. Image Process. | 5 |
| 2021 | Fog Computing Model and Efficient Algorithms for Directional Vehicle Mobility in Vehicular NetworkabstractVehicular fog computing (VFC) has become an appealing paradigm to provide services for vehicles and traffic systems. However, high mobility is one of the great challenges to the communication and computation service qualities in VFC. A network model for directional vehicle mobility is proposed in this paper to guarantee the service qualities of vehicles in VFC. In the model, vehicles are configured into three vehicular subnetworks according to their turning directions at the next crossing. For each subnetwork, vehicles communicate with each other via vehicle-to-vehicle communication, and with roadside units via vehicle-to-infrastructure communication. The aim is to minimize the average response time of the tasks originated from vehicles. By carefully choosing neighboring vehicles as task processing helpers, a greedy algorithm is proposed to solve the mentioned optimization problem. Besides, two bipartite matching based algorithms, named BMA1and BMA2, are proposed by exploiting Kuhn-Munkras approach and minimum-cost maximum-flow approach, respectively. Performance of the proposed model and the offloading algorithms are evaluated on the combined simulation platform by open street map, SUMO and NS-3. Simulation results show that, the proposed model outperforms four existing models in terms of average response time, when the five models have similar number of unsuccessful tasks. Moreover, the proposed BMA1and BMA2are superior to the existing greedy algorithm in terms of the average response time of tasks, and the proposed greedy algorithm significantly accelerates the generation of offloading decisions in comparison to BMA1, BMA2and the existing greedy algorithm. Yalan Wu, Jigang Wu, Long Chen 0006, Gangqiang Zhou, Jiaquan Yan |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Learning Compact Multifeature Codes for Palmprint Recognition From a Single Training Image per PalmabstractIn this article, we propose a multifeature learning method to jointly learn compact multifeature codes (LCMFCs) for palmprint recognition with a single training sample per palm. Unlike most existing hand-crafted methods that extract single-type features from raw pixels, we first form the multi-type data vectors such as the direction-data, and texture-data to completely sample the multiple information of a palmprint image. Then, we learn the discriminative multifeatures from multi-type data vectors by maximizing the inter-palm distance, and minimizing the energy loss between the learned codes, and the original data. Moreover, our LCMFC method adaptively learns the optimal weights of multi-type features to jointly learn the compact multifeature codes. Finally, we cluster the nonoverlapping blockwise histograms of the compact multifeature codes into a feature vector for palmprint representation. Extensive experimental results on six benchmark palmprint databases are presented to show the effectiveness of the proposed method. Lunke Fei, Bob Zhang 0001, Lin Zhang 0014, Wei Jia 0001, Jie Wen 0001, Jigang Wu |
IEEE Trans. Multim. | 6 |
| 2021 | TARCO: Two-Stage Auction for D2D Relay Aided Computation Resource Allocation in HetNetabstractIn heterogeneous cellular network, task scheduling for computation offloading is one of the biggest challenges. Most works focus on alleviating heavy burden of macro base stations by moving the computation tasks on macro cell user equipment (MUE) to remote cloud or small cell base stations. But the selfishness of network users is seldom considered. Motivated by the multiple access mobile edge computing, this paper provides incentive for task transfer from macro cell users to small cell base stations. The proposed incentive scheme utilizes small cell user equipments to provide relay services. The problem of computation offloading is modeled as a two-stage auction, in which the remote MUEs with common social character can form a group and then buy the computation resource of small cell base stations with the relaying of small cell user equipment. A two-stage auction scheme named TARCO is contributed to maximize utilities for both sellers and buyers in the network. The truthfulness, individual rational and budget balance properties of TARCO are also proved in this paper. In addition, two algorithms are proposed to further refine TARCO on the social welfare of the network. One can achieve higher utility of MUEs and the other can obtain higher total social welfare. Extensive simulation results demonstrate that, TARCO is better than random algorithm by 104.90 percent in terms of average utility of MUEs, while the performance of TARCO is further improved up to 28.75 percent and 17.06 percent by the proposed two algorithms, respectively. Long Chen 0006, Jigang Wu, Xinxiang Zhang, Gangqiang Zhou |
IEEE Trans. Serv. Comput. | 2 |
| 2021 | Combinatorial Double Auction for Resource Allocation in Mobile Blockchain Network
Xuelian Liu, Jigang Wu, Long Chen 0006, Chengpeng Xia, Yidong Li |
Wirel. Networks | 2 |
| 2020 | Data Aggregation Aware Routing for Distributed Training
Zhaohong Chen, Yalan Wu, Long Chen 0006, Jigang Wu, Shuangyin Liu |
PDCAT | 5 |
| 2020 | Multiple-Choice Hardware/Software Partitioning for Tree Task-Graph on MPSoCabstractAbstract Hardware/software (HW/SW) partitioning, that decides which components of an application are implemented in hardware and which ones in software, is a crucial step in embedded system design. On modern heterogeneous embedded system platform, each component of application can typically have multiple feasible configurations/implementations, trading off quality aspects (e.g. energy consumption, completion time) with usage for various types of resources. This provides new opportunities for further improving the overall system performance, but few works explore the potential opportunity by incorporating the multiple choices of hardware implementation in the partitioning process. This paper proposes three algorithms for multiple-choice HW/SW partitioning of tree-shape task graph on multiple processors system on chip (MPSoC) with the objective of minimizing execution time, while meeting area constraint. Firstly, an efficient heuristic algorithm is proposed to rapidly generate an approximate solution. The obtained solution produced by the first algorithm is then further refined by a customized Tabu search algorithm. We also propose a dynamic programming algorithm to calculate the exact solutions for relatively smaller scale instances. Simulation results show that the proposed heuristic algorithm is able to quickly generate good approximate solutions, and the solutions become very close to the exact solutions after refined by the proposed Tabu search algorithm, in comparison to the exact solutions produced by the dynamic programming algorithm. Wenjun Shi, Jigang Wu, Guiyuan Jiang, Siew-Kei Lam |
Comput. J. | 2 |
| 2020 | Efficient task scheduling for servers with dynamic states in vehicular edge computing
Yalan Wu, Jigang Wu, Long Chen 0006, Jiaquan Yan, Yuchong Luo |
Comput. Commun. | 2 |
| 2020 | SODNet: small object detection using deconvolutional neural networkabstractConvolution neural network (CNN) is an efficient technique to detect objects in various kinds of images, especially for microaneurysm (MA) of diabetic retinopathy in retinal fundus image. This study proposes a deconvolutional neural network to accurately discriminate MA from non‐MA. The deconvolution, instead of pooling operation, is embedded into the CNN to recover the erased details of feature maps of convolutional layers. Three types of images are collected for training and predicting. Furthermore, the extracted features are fed into the fully‐connected layers to classify using a softmax layer. Experimental results demonstrate that the proposed method can achieve significant sensitivity and accuracy on multiple public datasets, in comparison to the state‐of‐the‐art. For Retinopathy Online Challenge dataset, the sensitivity and accuracy are improved up to 0.798 and 0.986, respectively. Xinpeng Zhang 0003, Jigang Wu, Zhihao Peng 0004, Min Meng 0001 |
IET Image Process. | 2 |
| 2020 | Joint discriminative attributes and similarity embeddings modeling for zero-shot recognition
Min Meng 0001, Xiaoyu Zhan, Jigang Wu |
Neurocomputing | 3 |
| 2020 | Blockchain-based public auditing for big data in cloud storage
Jiaxing Li 0009, Jigang Wu, Guiyuan Jiang, Thambipillai Srikanthan |
Inf. Process. Manag. | 2 |
| 2020 | Visual Sentiment Prediction with Attribute Augmentation and Multi-attention Mechanism
Zhuanghui Wu, Min Meng 0001, Jigang Wu |
Neural Process. Lett. | 3 |
| 2020 | Double Relaxed Regression for Image ClassificationabstractThis paper addresses two fundamental problems: 1) learning discriminative model parameters and 2) avoiding over-fitting, which often occurs in regression-based classification tasks. We formulate these two problems in terms of relaxing both the strict binary label matrix and graph regularization term into more flexible forms so that the margins between different classes are enlarged as much as possible and the problem of over-fitting is avoided to some extent. This task is accomplished by the proposed double relaxed regression (DRR) method. The convex problem of DRR is solved efficiently with an iterative procedure. Extensive experiments on synthetic and real world image data sets demonstrate the effectiveness of the proposed method in terms of both classification accuracy and running time. Na Han, Jigang Wu, Xiaozhao Fang, Wai Keung Wong, Yong Xu 0001, Jian Yang 0003, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Group Low-Rank Representation-Based Discriminant Linear RegressionabstractIn this paper, a novel least square regression method, named group low-rank representation-based discriminant linear regression (GLRRDLR), is proposed for multi-class classification. Unlike the conventional linear regression methods, the proposed method aims to learn a more discriminative projection. Specially, two main techniques are adopted to improve the discriminability of the projection. The first approach is to make the transformed samples locate in their own subspace by introducing a group low-rank constraint to the model, such that the distance between samples from the same class can be decreased greatly. The second approach is to simultaneously learn a discriminative target matrix for regression. The extensive experimental results show that the proposed method performs much better than the state-of-the-art methods, which proves the effectiveness of the above two approaches in improving the discriminability of the projection. Shanhua Zhan, Jigang Wu, Na Han, Jie Wen 0001, Xiaozhao Fang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Constrained Discriminative Projection Learning for Image ClassificationabstractProjection learning is widely used in extracting discriminative features for classification. Although numerous methods have already been proposed for this goal, they barely explore the label information during projection learning and fail to obtain satisfactory performance. Besides, many existing methods can learn only a limited number of projections for feature extraction which may degrade the performance in recognition. To address these problems, we propose a novel constrained discriminative projection learning (CDPL) method for image classification. Specifically, CDPL can be formulated as a joint optimization problem over subspace learning and classification. The proposed method incorporates the low-rank constraint to learn a robust subspace which can be used as a bridge to seamlessly connect the original visual features and objective outputs. A regression function is adopted to explicitly exploit the class label information so as to enhance the discriminability of subspace. Unlike existing methods, we use two matrices to perform feature learning and regression, respectively, such that the proposed approach can obtain more projections and achieve superior performance in classification tasks. The experiments on several datasets show clearly the advantages of our method against other state-of-the-art methods. Min Meng 0001, Mengcheng Lan, Jun Yu 0002, Jigang Wu, Dapeng Tao |
IEEE Trans. Image Process. | 4 |
| 2020 | Projective Double Reconstructions Based Dictionary Learning Algorithm for Cross-Domain RecognitionabstractDictionary learning plays a significant role in the field of machine learning. Existing works mainly focus on learning dictionary from a single domain. In this paper, we propose a novel projective double reconstructions (PDR) based dictionary learning algorithm for cross-domain recognition. Owing the distribution discrepancy between different domains, the label information is hard utilized for improving discriminability of dictionary fully. Thus, we propose a more flexible label consistent term and associate it with each dictionary item, which makes the reconstruction coefficients have more discriminability as much as possible. Due to the intrinsic correlation between cross-domain data, the data should be reconstructed with each other. Based on this consideration, we further propose a projective double reconstructions scheme to guarantee that the learned dictionary has the abilities of data itself reconstruction and data crossreconstruction. This also guarantees that the data from different domains can be boosted mutually for obtaining a good data alignment, making the learned dictionary have more transferability. We integrate the double reconstructions, label consistency constraint and classifier learning into a unified objective and its solution can be obtained by proposed optimization algorithm that is more efficient than the conventional l1 optimization based dictionary learning methods. The experiments show that the proposed PDR not only greatly reduces the time complexity for both training and testing, but also outperforms over the stateof- the-art methods. Na Han, Jigang Wu, Xiaozhao Fang, Shaohua Teng, Guoxu Zhou, Shengli Xie 0001, Xuelong Li 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | Latent Elastic-Net Transfer LearningabstractSubspace learning based transfer learning methods commonly find a common subspace where the discrepancy of the source and target domains is reduced. The final classification is also performed in such subspace. However, the minimum discrepancy does not guarantee the best classification performance and thus the common subspace may be not the best discriminative. In this paper, we propose a latent elastic-net transfer learning (LET) method by simultaneously learning a latent subspace and a discriminative subspace. Specifically, the data from different domains can be well interlaced in the latent subspace by minimizing Maximum Mean Discrepancy (MMD). Since the latent subspace decouples inputs and outputs and, thus a more compact data representation is obtained for discriminative subspace learning. Based on the latent subspace, we further propose a low-rank constraint based matrix elastic-net regression to learn another subspace in which the intrinsic intra-class structure correlations of data from different domains is well captured. In doing so, a better discriminative alignment is guaranteed and thus LET finally learns another discriminative subspace for classification. Experiments on visual domains adaptation tasks show the superiority of the proposed LET method. Na Han, Jigang Wu, Xiaozhao Fang, Shengli Xie 0001, Shanhua Zhan, Kan Xie 0002, Xuelong Li 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | Transferable Linear Discriminant AnalysisabstractLinear discriminant analysis (LDA) has been widely used as the technique of feature exaction. However, LDA may be invalid to address the data from different domains. The reasons are as follows: 1) the distribution discrepancy of data may disturb the linear transformation matrix so that it cannot extract the most discriminative feature and 2) the original design of LDA does not consider the unlabeled data so that the unlabeled data cannot take part in the training process for further improving the performance of LDA. To address these problems, in this brief, we propose a novel transferable LDA (TLDA) method to extend LDA into the scenario in which the data have different probability distributions. The whole learning process of TLDA is driven by the philosophy that the data from the same subspace have a low-rank structure. The matrix rank in TLDA is the key learning criterion to conduct local and global linear transformations for restoring the low-rank structure of data from different distributions and enlarging the distances among different subspaces. In doing so, the variations of distribution discrepancy within the same subspace can be reduced, i.e., data can be aligned well and the maximally separated structure can be achieved for the data from different subspaces. A simple projected subgradient-based method is proposed to optimize the objective of TLDA, and a strict theory proof is provided to guarantee a quick convergence. The experimental evaluation on public data sets demonstrates that our TLDA can achieve better classification performance and outperform the state-of-the-art methods. Na Han, Jigang Wu, Xiaozhao Fang, Jie Wen 0001, Shanhua Zhan, Shengli Xie 0001, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | Robust Multi-View Hashing for Cross-Modal RetrievalabstractExisting hashing methods barely explore the information loss problem during learning the common semantic subspace, thus retrieval performance may be degraded. Besides, these methods mainly rely on the inter-modality or intra-modality correlations separately and fail to exploit the full structure reflected by these correlations. To address these problems, we present a novel cross-modal hashing method, namely Robust Multi-View Hashing (RMVH). To learn a robust latent semantic subspace, we enforce the learnt representations to well reconstruct original features such that more important information can be retained. To comprehensively exploit the relationship between representations of multiple modalities, we utilize Multi-View Learning to construct an affinity matrix to guide the learning of common latent semantic subspace, which can preserve both inter-modality and intra-modality similarities. Instead of relaxing the binary constraints, we leverage the label information to learn hash codes discretely which can avoid the large quantization error and preserve the semantic similarity. Experimental results on three benchmark datasets show that the proposed RMVH achieves superior performance compared with other state-of-the-art methods. Haitao Wang 0026, Min Meng 0001, Jigang Wu |
ICME | 4 |
| 2019 | Supervised Consistent and Specific HashingabstractMost existing methods seek for the common semantics using different projections for different modalities, which isolates the intrinsic relationships among different modalities. Besides, to avoid the large quantization error, some of them adopt the discrete cyclic coordinate descent schemes which are usually time-consuming. To address these issues, we present a novel hashing method, namely Supervised Consistent and Specific Hashing (SCSH), for cross-modal retrieval. We explicitly decompose the mapping matrices into consistent part and modality-specific ones. Specifically, consistency excavates the semantic shared by different modalities, whereas specificity captures private properties for each modality. Different from prior works, SCSH can discover the intrinsic semantic shared among different modalities more accurately. Moreover, by regressing the semantic labels to hash codes, SCSH can further promote the discriminative power of hash codes and significantly accelerate the hashing learning process. Extensive experiments on three widely used datasets demonstrate that the proposed SCSH outperforms other state-of-the-art methods. Haitao Wang 0026, Min Meng 0001, Jigang Wu |
ICME | 4 |
| 2019 | Task Merging and Scheduling for Parallel Deep Learning Applications in Mobile Edge ComputingabstractMobile edge computing enables the execution of compute-intensive applications, e.g. deep learning applications, on the end devices with limited computation resources. However, the deep learning applications bring the performance bottleneck in mobile edge computing, due to the movements of a large amount of data incurred by the large number of layers and millions of weights. In this paper, the computing model for parallel deep learning applications in mobile edge computing is proposed, by considering the occupancy allocation of processors, cost of context switch, and multi-processors in edge server and remote cloud. The problem of minimizing the completion time for deep learning applications is formulated, and the NP-hardness of the problem is proved. To solve the problem, an integrated algorithm by merging and scheduling is proposed. Moreover, a real-world distributed platform is developed for evaluating the proposed algorithm. Experimental results show that, the completion time of deep learning application for the proposed algorithm is decreased by 63% and 75%, respectively, without extra control costs, compared with the existing algorithms. Jigang Wu, Yalan Wu, Long Chen 0006 |
PDCAT | 2 |
| 2019 | Collaborative Task Offloading with Computation Result Reusing for Mobile Edge ComputingabstractAbstract The task offloading problem, which aims to balance the energy consumption and latency for Mobile Edge Computing (MEC), is still a challenging problem due to the dynamic changing system environment. To reduce energy while guaranteeing delay constraint for mobile applications, we propose an access control management architecture for 5G heterogeneous network by making full use of Base Station’s storage capability and reusing repetitive computational resource for tasks. For applications that rely on real-time information, we propose two algorithms to offload tasks with consideration of both energy efficiency and computation time constraint. For the first scenario, i.e. the rarely changing system environment, an optimal static algorithm is proposed based on dynamic programming technique to get the exact solution. For the second scenario, i.e. the frequently changing system environment, a two-stage online algorithm is proposed to adaptively obtain the current optimal solution in real time. Simulation results demonstrate that the exact algorithm in the first scenario runs 4 times faster than the enumeration method. In the second scenario, the proposed online algorithm can reduce the energy consumption and computation time violation rate by 16.3% and 25% in comparison with existing methods. Zikai Zhang 0004, Jigang Wu, Long Chen 0006, Guiyuan Jiang, Siew-Kei Lam |
Comput. J. | 2 |
| 2019 | Unsupervised feature extraction by low-rank and sparsity preserving embedding
Shanhua Zhan, Jigang Wu, Na Han, Jie Wen 0001, Xiaozhao Fang |
Neural Networks | 2 |
| 2019 | Precision direction and compact surface type representation for 3D palmprint identification
Lunke Fei, Bob Zhang 0001, Yong Xu 0001, Wei Jia 0001, Jie Wen 0001, Jigang Wu |
Pattern Recognit. | 6 |
| 2019 | Flexible Affinity Matrix Learning for Unsupervised and Semisupervised ClassificationabstractIn this paper, we propose a unified model called flexible affinity matrix learning (FAML) for unsupervised and semisupervised classification by exploiting both the relationship among data and the clustering structure simultaneously. To capture the relationship among data, we exploit the self-expressiveness property of data to learn a structured matrix in which the structures are induced by different norms. A rank constraint is imposed on the Laplacian matrix of the desired affinity matrix, so that the connected components of data are exactly equal to the cluster number. Thus, the clustering structure is explicit in the learned affinity matrix. By making the estimated affinity matrix approximate the structured matrix during the learning procedure, FAML allows the affinity matrix itself to be adaptively adjusted such that the learned affinity matrix can well capture both the relationship among data and the clustering structure. Thus, FAML has the potential to perform better than other related methods. We derive optimization algorithms to solve the corresponding problems. Extensive unsupervised and semisupervised classification experiments on both synthetic data and real-world benchmark data sets show that the proposed FAML consistently outperforms the state-of-the-art methods. Xiaozhao Fang, Na Han, Wai Keung Wong, Shaohua Teng, Jigang Wu, Shengli Xie 0001, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2019 | Efficient three-stage auction schemes for cloudlets deployment in wireless access network
Gangqiang Zhou, Jigang Wu, Long Chen 0006, Guiyuan Jiang, Siew-Kei Lam |
Wirel. Networks | 2 |
| 2018 | Defect Analysis and Parallel March Test Algorithm for 3D Hybrid CMOS-Memristor MemoryabstractAs an attractive option of future non-volatile memories (NVM), resistive random access memory (RRAM) has attracted more attentions. CMOS Molecular (CMOL) architecture, which can alleviate the sneak path problem of one memristor (1R) crossbars and limit its power consumption in 1R crossbars, is used as a large-scale memory system. In this paper, we analyze the electrical defects in a CMOL circuit including open and bridge. A parallel March-like test algorithm is presented for the CMOL architecture, which covers defined faults caused by electrical defects. The test time of the proposed test algorithm is reduced significantly compared with previous test algorithms that are enhanced for CMOL architecture. Peng Liu 0045, Jigang Wu, Zhiqiang You, Michael Elimu, Weizheng Wang 0002, Shuo Cai |
ATS | 2 |
| 2018 | NESTLE: Incentive Mechanism Specialized for Computation Offloading in Local Edge Community
Jigang Wu, Long Chen 0006 |
ICA3PP (2) | 2 |
| 2018 | Blockchain-Based Secure and Reliable Distributed Deduplication Scheme
Jigang Wu, Long Chen 0006, Jiaxing Li 0009 |
ICA3PP (1) | 2 |
| 2018 | Energy-Efficient Offloading in Mobile Edge Computing with Edge-Cloud Collaboration
Jigang Wu, Long Chen 0006 |
ICA3PP (3) | 2 |
| 2018 | POEM: Pricing Longer for Edge Computing in the Device Cloud
Qiankun Yu, Jigang Wu, Long Chen 0006 |
ICA3PP (3) | 2 |
| 2018 | TAMSA: Two-Stage Auction Mechanism for Spectrum Allocation in Cooperative Cognitive Radio Networks
Xinxiang Zhang, Jigang Wu, Long Chen 0006 |
ICA3PP (3) | 2 |
| 2018 | Coalitional Game Based Carpooling Algorithms for Quality of ExperienceabstractTo motivate passengers to participate in carpooling, we focus on how to guarantee the quality of experience (QoE) of passengers in carpooling using coalition game. We formulate the QoE-guarantee problem as a benefit allocation problem. To solve the problem, we quantify the impatience of passengers due to detouring time delay. The algorithm named PCA is proposed to minimize the impatience of all passengers and calculate the compensation for them based on Shapley value. We prove that PCA can guarantee the fairness of passengers. Simulation results demonstrate that PCA can minimize the impatience of passengers and produce a win-win solution for both passengers and drivers in carpooling. Jigang Wu, Long Chen 0006 |
ICPADS | 2 |
| 2018 | ETRA: Efficient Three-Stage Resource Allocation Auction for Mobile Blockchain in Edge ComputingabstractBlockchain technology is emerging in various fields, to guarantee security of digital currency and internet of things. In this paper, we provide incentive to encourage edge servers to serve mobile users for the mobile blockchain application. We formulate the problem as a resource allocation problem, then we propose a three-stage auction to implement resource allocation specially designed for mobile blockchain, and introduce the group-buying mechanism to motivate mobile users. We prove that our auction scheme is truthful, individual rationality, and computational efficiency. We compare proposed scheme with TACD and HAF mechanisms, and simulation results show that the social welfare achieved by our scheme is higher than that of TACD and HAF mechanisms. Chengpeng Xia, Xuelian Liu, Jigang Wu, Long Chen 0006 |
ICPADS | 4 |
| 2018 | Algorithms for Replica Placement and Update in Tree NetworkabstractA critical issue in data replication is to wisely place data replicas which involves identifying the best possible nodes to duplicate data. Facing dynamics of data requests, this paper investigates the problem of replica placement and update in tree networks, where part of nodes have pre-existing replicas. We aim to develop efficient algorithms to accelerate the replica placement and update without causing obvious degradation in solution quality via reusing pre-existing replicas. Firstly, an efficient heuristic algorithm GRP is proposed to quickly place replicas when users change their requests dynamically, under the Closest policy where a client must be served by the closest server. Then, a Tabu search algorithm TSRP is customized to further refine the solution obtained by GRP. Furthermore, we propose a heuristic algorithm MPFSF for the replica placement and update problem, under the Multiple policy where requests of a client are served by multiple servers. Simulation results show that, GRP and TSRP can accelerate existing dynamic programming algorithm by 87.97% while quality degradation is bounded by 2.49%. MPFSF can achieve the best improvement for about 84.6% than existing heuristic algorithm. Jigang Wu, Long Chen 0006, Guiyuan Jiang, Siew-Kei Lam, Thambipillai Srikanthan |
Comput. J. | 1 |
| 2018 | Efficient hybrid multicast approach in wireless data center network
Longting Zhu, Jigang Wu, Guiyuan Jiang, Long Chen 0006, Siew-Kei Lam |
Future Gener. Comput. Syst. | 2 |
| 2018 | Generalized passivity of coupled neural networks with directed and undirected topologies
Shun-Yan Ren, Jin-Liang Wang 0001, Jigang Wu |
Neurocomputing | 3 |
| 2018 | Block-secure: Blockchain based scheme for secure P2P cloud storage
Jiaxing Li 0009, Jigang Wu, Long Chen 0006 |
Inf. Sci. | 2 |
| 2018 | Low-rank and sparse embedding for dimensionality reduction
Na Han, Jigang Wu, Yingyi Liang, Xiaozhao Fang, Wai Keung Wong, Shaohua Teng |
Neural Networks | 2 |
| 2018 | QUICK: QoS-guaranteed efficient cloudlet placement in wireless metropolitan area networks
Long Chen 0006, Jigang Wu, Gangqiang Zhou, Longjie Ma |
J. Supercomput. | 2 |
| 2018 | Approximate Low-Rank Projection Learning for Feature ExtractionabstractFeature extraction plays a significant role in pattern recognition. Recently, many representation-based feature extraction methods have been proposed and achieved successes in many applications. As an excellent unsupervised feature extraction method, latent low-rank representation (LatLRR) has shown its power in extracting salient features. However, LatLRR has the following three disadvantages: 1) the dimension of features obtained using LatLRR cannot be reduced, which is not preferred in feature extraction; 2) two low-rank matrices are separately learned so that the overall optimality may not be guaranteed; and 3) LatLRR is an unsupervised method, which by far has not been extended to the supervised scenario. To this end, in this paper, we first propose to use two different matrices to approximate the low-rank projection in LatLRR so that the dimension of obtained features can be reduced, which is more flexible than original LatLRR. Then, we treat the two low-rank matrices in LatLRR as a whole in the process of learning. In this way, they can be boosted mutually so that the obtained projection can extract more discriminative features. Finally, we extend LatLRR to the supervised scenario by integrating feature extraction with the ridge regression. Thus, the process of feature extraction is closely related to the classification so that the extracted features are discriminative. Extensive experiments are conducted on different databases for unsupervised and supervised feature extraction, and very encouraging results are achieved in comparison with many state-of-the-arts methods. Xiaozhao Fang, Na Han, Jigang Wu, Yong Xu 0001, Jian Yang 0003, Wai Keung Wong, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2018 | Passivity and Output Synchronization of Complex Dynamical Networks With Fixed and Adaptive Coupling StrengthabstractThis paper considers a complex dynamical network model, in which the input and output vectors have different dimensions. We, respectively, investigate the passivity and the relationship between output strict passivity and output synchronization of the complex dynamical network with fixed and adaptive coupling strength. First, two new passivity definitions are proposed, which generalize some existing concepts of passivity. By constructing appropriate Lyapunov functional, some sufficient conditions ensuring the passivity, input strict passivity and output strict passivity are derived for the complex dynamical network with fixed coupling strength. In addition, we also reveal the relationship between output strict passivity and output synchronization of the complex dynamical network with fixed coupling strength. By employing the relationship between output strict passivity and output synchronization, a sufficient condition for output synchronization of the complex dynamical network with fixed coupling strength is established. Then, we extend these results to the case when the coupling strength is adaptively adjusted. Finally, two examples with numerical simulations are provided to demonstrate the effectiveness of the proposed criteria. Jin-Liang Wang 0001, Huai-Ning Wu, Tingwen Huang, Shun-Yan Ren, Jigang Wu |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2018 | Analysis and Control of Output Synchronization in Directed and Undirected Complex Dynamical NetworksabstractThis research focuses on the problem of output synchronization in undirected and directed complex dynamical networks, respectively, by applying Barbalat's lemma. First, to ensure the output synchronization, several sufficient criteria are established for these network models based on some mathematical techniques, such as the Lyapunov functional method and matrix theory. Furthermore, some adaptive schemes to adjust the coupling weights among network nodes are developed to achieve the output synchronization. By applying the designed adaptive laws, several criteria for output synchronization are deduced for the network models. In addition, a design procedure of the adaptive law is shown. Finally, two simulation examples are used to show the effectiveness of the previous results. Jin-Liang Wang 0001, Huai-Ning Wu, Tingwen Huang, Shun-Yan Ren, Jigang Wu, Xiao-Xiao Zhang |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2017 | DOTA: Delay Bounded Optimal Cloudlet Deployment and User Association in WMANsabstractIn the large-scale Wireless Metropolitan Area Network (WMAN) consisting of many wireless Access Points (APs),choosing the appropriate position to place cloudlet is very important for reducing the user's access delay. For service provider, it isalways very costly to deployment cloudlets. How many cloudletsshould be placed in a WMAN and how much resource eachcloudlet should have is very important for the service provider. In this paper, we study the cloudlet placement and resourceallocation problem in a large-scale Wireless WMAN, we formulatethe problem as an novel cloudlet placement problem that givenan average access delay between mobile users and the cloudlets, place K cloudlets to some strategic locations in the WMAN withthe objective to minimize the number of use cloudlet K. Wethen propose an exact solution to the problem by formulatingit as an Integer Linear Programming (ILP). Due to the poorscalability of the ILP, we devise a clustering algorithm K-Medoids(KM) for the problem. For a special case of the problem whereall cloudlets computing capabilities have been given, we proposean efficient heuristic for it. We finally evaluate the performanceof the proposed algorithms through experimental simulations. Simulation result demonstrates that the proposed algorithms areeffective. Longjie Ma, Jigang Wu, Long Chen 0006 |
CCGrid | 2 |
| 2017 | Fast algorithms for capacitated cloudlet placementsabstractMobile cloud computing addresses resource scarcity problem of mobile devices by offloading computation data from mobile devices into the cloud. However, remote server may be far from mobile users. Cloudlet could be used to deal with the long access delay problem. In the large-scale Wireless Metropolitan Area Network (WMAN) consisting of many wireless Access Points (APs), choosing the appropriate position of cloudlet is very important to reducing access delay. Recently, a heuristic algorithm has been proposed. However, it has so many repeated sorting process of APs that the algorithm efficiency is poor. In this paper, we propose a New Heuristic Algorithm (NHA) and a Particle Swarm Optimization (PSO) algorithm for the delaying problem. We evaluate the performance of the proposed algorithms through extensive simulations. Simulation results demonstrate NHA is more efficient then existing algorithm. For the PSO algorithm, in the case of parallelized execution, it is more efficient than the new heuristic algorithm within a bounded delay. Longjie Ma, Jigang Wu, Long Chen 0006, Zhusong Liu |
CSCWD | 2 |
| 2017 | QoE-Aware Task Offloading for Time Constraint Mobile ApplicationsabstractIn this paper, we develop an access controller management model which provides new opportunities for further reducing the computation repetition and data transmission redundancy for Mobile Edge Computing (MEC) in 5G network. We propose novel algorithms for solving the offloading problem with consideration of tradeoff between energy consumption and the amount of offloaded data under constraint of overall task computation time. For sequential topology applications, we develop a dynamic programming algorithm to produce optimal solutions. For general topology applications, a critical-path based heuristic algorithm is proposed by repeatedly identifying partial critical path (PCP) from the application task graph and calculating optimal solution for the PCP by performing the proposed dynamic programming algorithm. In addition, the interference of parallel data transmission between tasks (one-to-many, manyto-one and many-to-many) using single channel is taken into consideration. Experimental results demonstrate the effectiveness of our proposed method. Zikai Zhang 0004, Jigang Wu, Guiyuan Jiang, Long Chen 0006, Siew-Kei Lam |
LCN | 2 |
| 2017 | TACD: A Three-Stage Auction Scheme for Cloudlet Deployment in Wireless Access Network
Gangqiang Zhou, Jigang Wu, Long Chen 0006 |
WASA | 2 |
| 2017 | Passivity and pinning passivity of complex dynamical networks with spatial diffusion coupling
Shun-Yan Ren, Jigang Wu, Beibei Xu |
Neurocomputing | 2 |
| 2017 | Passivity and Pinning Passivity of Coupled Delayed Reaction-Diffusion Neural Networks with Dirichlet Boundary Conditions
Shun-Yan Ren, Jigang Wu, Pu-Chong Wei |
Neural Process. Lett. | 2 |
| 2017 | Passivity of Directed and Undirected Complex Dynamical Networks With Adaptive Coupling WeightsabstractA complex dynamical network consisting of N identical neural networks with reaction-diffusion terms is considered in this paper. First, several passivity definitions for the systems with different dimensions of input and output are given. By utilizing some inequality techniques, several criteria are presented, ensuring the passivity of the complex dynamical network under the designed adaptive law. Then, we discuss the relationship between the synchronization and output strict passivity of the proposed network model. Furthermore, these results are extended to the case when the topological structure of the network is undirected. Finally, two examples with numerical simulations are provided to illustrate the correctness and effectiveness of the proposed results. Jin-Liang Wang 0001, Huai-Ning Wu, Tingwen Huang, Shun-Yan Ren, Jigang Wu |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2017 | Joint Charging Tour Planning and Depot Positioning for Wireless Sensor Networks Using Mobile ChargersabstractRecent breakthrough in wireless energy transfer technology has enabled wireless sensor networks (WSNs) to operate with zero-downtime through the use of mobile energy chargers (MCs), that periodically replenish the energy supply of the sensor nodes. Due to the limited battery capacity of the MCs, a significant number of MCs and charging depots are required to guarantee perpetual operations in large scale networks. Existing methods for reducing the number of MCs and charging depots treat the charging tour planning and depot positioning problems separately even though they are inter-dependent. This paper is the first to jointly consider charging tour planning and MC depot positioning for large-scale WSNs. The proposed method solves the problem through the following three stages: charging tour planning, candidate depot identification and reduction, and depot deployment and charging tour assignment. The proposed charging scheme also considers the association between the MC charging cycle and the operational lifetime of the sensor nodes, in order to maximize the energy efficiency of the MCs. This overcomes the limitations of existing approaches, wherein MCs with small battery capacity ends up charging sensor nodes more frequently than necessary, while MCs with large battery capacity return to the depots to replenish themselves before they have fully transferred their energy to the sensor nodes. Compared with existing approaches, the proposed method leads to an average reduction in the number of MCs by 64%, and an average increase of 19.7 times on the ratio of total charging time over total traveling time. Guiyuan Jiang, Siew-Kei Lam, Lijia Tu, Jigang Wu |
IEEE/ACM Trans. Netw. | 5 |
| 2017 | Passivity Analysis of Coupled Reaction-Diffusion Neural Networks With Dirichlet Boundary ConditionsabstractTwo coupled reaction-diffusion neural networks (CRDNNs) with different dimensions of input and output are considered in this paper. The only difference between them is whether time-varying delay is incorporated in the mathematical model of network. We respectively analyze dissipativity and passivity of these CRDNNs. First, for the systems with different dimensions of input and output vectors, two new passivity definitions are proposed. Then, by exploiting some inequality techniques, several dissipativity and passivity criteria for these CRDNNs are established. Furthermore, we analyze stability of passive CRDNNs. Finally, two examples with simulation results are presented to verify the effectiveness of the proposed criteria. Jin-Liang Wang 0001, Huai-Ning Wu, Tingwen Huang, Shun-Yan Ren, Jigang Wu |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2016 | Efficient Scheduling Strategy for Mobile Charger in Wireless Rechargeable Sensor NetworksabstractWith the development of wireless sensor networks, the limited battery capacity of sensor nodes has become one of energy bottleneck problem that dominates the wide application of wireless sensor networks. In recent years, wireless rechargeable sensor networks have attracted much attention due to their potential in solving the energy bottleneck problem. In this paper, we study the scheduling strategy of mobile charger in the on-demand mobile charging wireless sensor networks. The proposed strategy divides the sensor nodes of service pool into two categories, such that the mobile charger can provide charging service in some priority according to the degree of charging request urgency during the charging tour. The proposed algorithm successfully reduces charging missing ratio by 89 percent, and it can keep the charging throughput decline rate less than 9.91 percent. Shanhua Zhan, Jigang Wu, Lijun Qu, Dang Xin |
PDCAT | 2 |
| 2016 | Note on Edge-Colored Graphs for Networks with Homogeneous FaultsabstractThe failure on all homogeneous devices due to the same reason is called homogeneous fault in networks. In contrast, heterogeneous platforms deployed simultaneously in the network are more robust against homogeneous faults. One of the challenging problems is how to design survivable networks that against homogeneous faults. This paper utilizes edge-colored graphs to investigate the network topology with homogeneous faults, in order to guarantee network connectivity using minimum number of links. Two types of network topologies are proposed on the edge-colored graph. One type of networks is characterized by the fact that all the edges of the same color form a Hamiltonian path or a Hamiltonian cycle. An upper bound on the number of colors used in the proposed network topologies is obtained. The network topologies of the second type have edges colored with at most five colors. Additionally, the subnetworks induced by the edges of two colors contain a Hamiltonian path, or a Hamilton cycle in some cases. Rui Hou 0006, Jigang Wu, Yawen Chen 0001, Haibo Zhang 0001 |
Comput. J. | 2 |
| 2016 | Pinning Control for Synchronization of Coupled Reaction-Diffusion Neural Networks With Directed TopologiesabstractThis paper proposes a directed complex dynamical network consisting of N linearly and diffusively coupled identical reaction-diffusion neural networks. Based on the Lyapunov functional method and the pinning control technique, some sufficient conditions are obtained to guarantee the synchronization of the proposed network model. In addition, an adaptive strategy is proposed to obtain appropriate coupling strength for achieving network synchronization. Furthermore, the pinning adaptive synchronization problem is also investigated in this paper, and a general criterion for ensuring network synchronization is established. Finally, a numerical example is provided to illustrate the effectiveness of the proposed criteria. Jin-Liang Wang 0001, Huai-Ning Wu, Tingwen Huang, Shun-Yan Ren, Jigang Wu |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2015 | Fast Replica Placement and Update Strategies in Tree NetworksabstractData replication enhances data availability and thereby improves the system reliability and efficiency while reduces access latency and communication cost. A critical issue in data replication is to wisely place data replicas which involves identifying the best possible nodes to duplicate data according to user requests. In this paper, we address the problem of replica placement and update in tree networks, where some nodes of the network contain pre-existing replicas. It is obvious that reusing a pre-existing replica leads to smaller cost than creating a new replica, thus it is necessary to take full advantage of the pre-existing replicas. The only previous work that consider the same problem tries to find the optimal solution by developing a dynamic programming algorithm which runs in O(N5). However, this approach is not suitable for practical situation where the user requests change frequently. In this paper, we develop efficient algorithms to accelerate the replica placement without causing obvious degradation in solution quality. We first propose an efficient heuristic algorithm (named GreedyRP) for quickly placing replicas when users change their requests. Then a tabu search algorithm (named TSRP) is customized to further refine the solution obtained by algorithm GreedyRP. Experimental results show that, on tree networks with 600 nodes and 150 pre-existing replicas, the proposed algorithms can accelerate the previous work by 87.97% while the quality degradation is bounded by 2.49% in comparison to the optimal solution. Jigang Wu, Guiyuan Jiang, Siew-Kei Lam, Thambipillai Srikanthan |
CCGRID | 2 |
| 2015 | Reconfigurations for Processor Arrays with Faulty Switches and LinksabstractLarge scale multiprocessor array suffers from frequent hardware defects or soft faults due to overheating, overload or occupancy by other running applications. To obtain fault-free logical array, reconfiguration techniques are proposed to reuse the fault-free PEs by changing the interconnection among PEs. Previous research has worked on this topic but assume that switches and links are fault-free. In this paper, we consider faults not only on the processing elements (PEs) but also on the switches and links, and develop efficient algorithms to construct as large as possible logical arrays with optimized networks length. To deal with the faults on switches and links, an efficient pre-processing procedure is designed, in which switch faults are transformed into link faults, and then faulty links are classified into several categories to handle. Then, we propose an efficient algorithm, A-MLA, to produce as many as possible logical columns which are then combined to form a two dimensional processor array. After that, we propose an algorithm A-TMLA to reduce the interconnection length of the logical array obtained by algorithm A-MLA, as short interconnect leads to small communication latency and power consumption. Extensive experimental results show that, even with switch faults and link faults, our approach can produce larger logical fault-free arrays with shorter interconnection length, compared to the state-of-the-art. Jigang Wu, Longting Zhu, Peilan He, Guiyuan Jiang |
CCGRID | 1 |
| 2015 | Constructing Edge-Colored Graph for Heterogeneous Networks
Rui Hou 0006, Jigang Wu, Yawen Chen 0001, Haibo Zhang 0001, Xiufeng Sui |
J. Comput. Sci. Technol. | 2 |
| 2015 | Reconfiguring Three-Dimensional Processor Arrays for Fault-Tolerance: Hardness and Heuristic AlgorithmsabstractWith the increased density of three-dimensional (3D) processor arrays, faults can potentially occur quite often due to power overheating during massively parallel computing. In order to achieve fault-tolerance under such a scenario, an effective way is to find an as large as possible logical fault-free subarray of m' × n' × h' from a faulty array of m × n × h (m' ≤ m, n' ≤ n, h' ≤ h), such that an original application can still work on the m' × n' × h' subarray. This paper investigates the problem of constructing maximum fault-free subarrays with minimum interconnection length from 3D arrays with faults. First, we prove that constructing maximum logical array (MLA) is NP-complete. We propose a linear-time algorithm which is capable of producing an MLA for the problem with the constraint of selected indexes. Second, we prove that minimizing the interconnection length (inter-length) of the MLA is NP-hard. We propose an efficient heuristic which significantly reduces the inter-length by revising each logical plane of the MLA. This leads to the reduction of communication cost, capacitance and dynamic power dissipation. In addition, we propose a lower bound for the inter-length of the MLA to evaluate the proposed algorithms. Simulation results show that, the size of logical array can be improved up to 62.6 percent in average, and the inter-length redundancy can be reduced by 22.7 percent in average, compared to the state-of-the-art, for all cases considered. Guiyuan Jiang, Jigang Wu, Yajun Ha, Yi Estelle Wang |
IEEE Trans. Computers | 2 |
| 2015 | Algorithmic aspects of graph reduction for hardware/software partitioning
Guiyuan Jiang, Jigang Wu, Siew-Kei Lam, Thambipillai Srikanthan |
J. Supercomput. | 2 |
| 2014 | Reducing the Interconnection Length for 3D Fault-Tolerant Processor Arrays
Guiyuan Jiang, Jigang Wu, Longting Zhu |
ICA3PP (1) | 2 |
| 2014 | Interconnection Network Reconstruction for Fault-Tolerance of Torus-Connected VLSI Array
Longting Zhu, Jigang Wu, Guiyuan Jiang |
ICA3PP (1) | 2 |
| 2014 | Downsampling sparse representation and discriminant information aided occluded face recognition
Jufu Feng, Jigang Wu |
Sci. China Inf. Sci. | 4 |
| 2014 | Flexible rerouting schemes for reconfiguration of multiprocessor arrays
Guiyuan Jiang, Jigang Wu, Yiyi Gao |
J. Parallel Distributed Comput. | 2 |
| 2014 | Parallel reconfiguration algorithms for mesh-connected processor arrays
Jigang Wu, Guiyuan Jiang, Yuze Shen, Siew-Kei Lam, Thambipillai Srikanthan |
J. Supercomput. | 1 |
| 2014 | Constructing Sub-Arrays with ShortInterconnects from Degradable VLSI ArraysabstractReducing the interconnection length of VLSI arrays leads to less capacitance, power dissipation and dynamic communication cost between the processing elements (PEs). This paper develops efficient algorithms for constructing tightly-coupled subarrays from the mesh-connected VLSI arrays with faulty PEs. For a given size r·s of the target (logical) array, the proposed algorithm searches and reroutes a physical r×s subarray that has the least number of faults, resulting in an approximate target array, which is subsequently extended to the desired target array. Experimental results show that over 65 percent redundant interconnects can be reduced for a 64×64 target array on the 512×512 host array with no more than 1 percent faults. In addition, we propose a recursive divide-and-conquer algorithm for constructing the maximum target array (MTA). The lower bound of the total interconnection length of the MTA has been established. Experimental results show that the proposed algorithm is capable of reducing the long interconnects by over 33 percent for the MTA derived from the 512×512 host array with no more than 1 percent faults. Moreover, the proposed total interconnection length of target array is close to the lower bound for the cases with relatively fewer number of faults. Jigang Wu, Thambipillai Srikanthan, Guiyuan Jiang |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2013 | Preprocessing technique for accelerating reconfiguration of degradable VLSI arraysabstractThis paper presents a heuristic approach to accelerate the reconfiguration of two-dimensional degradable VLSI arrays linked by 4-port switches in presence of faulty processing elements (PEs). In particular, we proposed a technique to preprocess the host array by 1) identifying fault-free PEs that cannot form the target array due to their proximity to faulty PEs, and 2) labeling these fault-free PEs as faults. The proposed preprocessing method minimizes the number of PEs that will be considered for reconfiguration, thus accelerating the reconfiguration process. Simulation results show that the runtime of two well-known algorithms are significantly reduced by employing the preprocessing technique. In addition, we demonstrate the scalability of the proposed technique by showing that the runtime reduction rate increases with increasing fault density. Yuanbo Zhu, Jigang Wu, Siew-Kei Lam, Thambipillai Srikanthan |
ISCAS | 2 |
| 2013 | Efficiency of Flexible Rerouting Scheme for Maximizing Logical Arrays
Guiyuan Jiang, Jigang Wu |
NPC | 2 |
| 2013 | Efficient localization for mobile sensor networks based on constraint rules optimized Monte Carlo method
Ze Wang 0016, Maode Ma, Jigang Wu |
Comput. Networks | 4 |
| 2013 | CADSE: communication aware design space exploration for efficient run-time MPSoC management
Amit Kumar Singh 0002, Akash Kumar 0001, Jigang Wu, Thambipillai Srikanthan |
Frontiers Comput. Sci. | 3 |
| 2013 | Performance comparisons between cellular-only and cellular/WLAN integrated systems based on analytical models
Guozhi Song, Jigang Wu, John A. Schormans |
Frontiers Comput. Sci. | 3 |
| 2013 | Efficient semi-supervised feature selection with noise insensitive trace ratio criterion
Yun Liu 0021, Feiping Nie 0001, Jigang Wu, Lihui Chen 0001 |
Neurocomputing | 3 |
| 2013 | Efficient reconfiguration algorithms for communication-aware three-dimensional processor arrays
Guiyuan Jiang, Jigang Wu |
Parallel Comput. | 2 |
| 2013 | Efficient heuristic and tabu search for hardware/software partitioning
Jigang Wu, Siew-Kei Lam, Thambipillai Srikanthan |
J. Supercomput. | 1 |
| 2012 | Non-Backtracking Reconfiguration Algorithm for Three-dimensional VLSI ArraysabstractFast reconfiguration is one of the main challenges in fault tolerant VLSI arrays. In these arrays, there are some invalid processing elements (PEs) that are fault-free but cannot be used to form a target array. These invalid PEs lead to backtracking in reconfiguration. This paper proposes a non-backtracking reconfiguration (NBR) algorithm for three-dimensional degradable VLSI array with faults. The proposed algorithm accelerates the reconfiguration without loss of harvest, by eliminating the backtracking operation that frequently occurs in the existing algorithm (named as BGPR) cited in this paper. Initially, the invalid PEs are identified in the preprocessing for the host array. Then NBR algorithm constructs each logical plane from bottom to top in the host array, and updates the set of the invalid PEs in the host array after a logical plane is constructed. Experimental results show that the NBR algorithm is more scalable than the BGPR algorithm, and thus it can reconfigure large host arrays much faster. In addition, the runtime of NBR algorithm tends to decrease, rather than increase as did in BGPR algorithm, with the increasing fault density. Guiyuan Jiang, Jigang Wu |
ICPADS | 2 |
| 2012 | Reconfiguration Algorithms for Degradable VLSI Arrays with Switch FaultsabstractThe problem of reconfiguring two-dimensional VLSI arrays with faults is to find a maximum logical array without faults. The existing algorithms only consider faults associated with processing elements, and all switches and links are assumed to be fault-free. But switch faults may often occur in the network-on-chips with high density. In this paper, two novel approaches are proposed to tackle the reconfiguration problem of degradable VLSI arrays with switch faults. The first approach extends the well-known existing algorithm with simple pre-processing and row bypass scheme. The second one employs a novel row and column rerouting scheme to maximize the size of the logical array. Simulation results show that the proposed two approaches can effectively generate the logical arrays on the given host array with switch faults, and the second algorithm performs more favorably with the increasing number of the switch faults. Yuanbo Zhu, Jigang Wu, Siew-Kei Lam, Thambipillai Srikanthan |
ICPADS | 2 |
| 2012 | A New Fault-Tolerant Routing Scheme for N-dimesional MeshabstractFault tolerance is one of the most important issues for the design of cost-effective and high performance interconnection networks. In this paper, a new fault tolerance routing algorithm for n-dimensional meshes is presented. The presented algorithm is based on a planer fault model which only disables minimum fault-free nodes to form rectangular fault regions. The algorithm uses three virtual channels per physical channel and only employs a very simple deadlock avoidance scheme. In spit the variety fault regions in n-dimensional mesh, the presented algorithm is always connected as long as fault regions do not disconnect the network. The result of simulation shows that the proposed routing algorithm is of feasibility of gracefully degraded operation. Xinming Duan, Jigang Wu |
PDCAT | 2 |
| 2012 | Integrated Heuristic for Hardware/Software Co-design on Reconfigurable DevicesabstractHardware/Software (HW/SW) partitioning and scheduling are essential to the embedded systems. In this paper, a hybrid algorithm derived from Tabu Search and Simulated Annealing is proposed for solving the HW/SW partitioning problem. The virtual hardware resource is set to implement the customized Tabu Search. Earliest-Deadline-First strategy is introduced to describe the reconfiguration of FPGA. Moreover, an algorithm combining the Breadth-First-Search with Depth-First-Search is proposed for HW/SW tasks scheduling to fit for the feature of reconfigurable systems. Experimental results show that, the proposed algorithms produce better performance than the previoPus methods cited in this paper. Peng Liu 0045, Jigang Wu, Yongji Wang 0002 |
PDCAT | 2 |
| 2012 | Multithread Reconfiguration Algorithm for Mesh-Connected Processor ArraysabstractMesh-connected processor array is a popular architecture used in parallel processing. Extensive studies have been conducted on reconfiguration algorithms for the processor arrays with faults, but few work is on parallel algorithm to accelerate the reconfiguration. This paper presents a fast algorithm to reconfigure two dimensional mesh-connected processor arrays with faults. A traditional algorithm is successfully accelerated in the manner of multithread, without loss of harvest. The proposed algorithm reconfigures the processor array with the mechanics of route distance in order to avoid the routing errors. Simulation results show that the proposed algorithm can accelerate the reconfiguration nearly by 15 times on a 64 × 64 array in comparison to the traditional algorithm cited in this paper. Yuze Shen, Jigang Wu, Guiyuan Jiang |
PDCAT | 2 |
| 2012 | Securing wireless mesh networks in a unified security framework with corruption-resilience
Ze Wang 0016, Maode Ma, Jigang Wu |
Comput. Networks | 3 |
| 2010 | Selecting profitable custom instructions for reconfigurable processors
Tao Li 0008, Jigang Wu, Siew-Kei Lam, Thambipillai Srikanthan, Xicheng Lu |
J. Syst. Archit. | 2 |
| 2010 | Communication-aware heuristics for run-time task mapping on NoC-based MPSoC platforms
Amit Kumar Singh 0002, Thambipillai Srikanthan, Akash Kumar 0001, Jigang Wu |
J. Syst. Archit. | 4 |
| 2010 | Algorithmic Aspects of Hardware/Software Partitioning: 1D Search AlgorithmsabstractHardware/software (HW/SW) partitioning is one of the key challenges in HW/SW codesign. This paper presents efficient algorithms for the HW/SW partitioning problem, which has been proved to be NP-hard. We reduce the HW/SW partitioning problem to a variation of knapsack problem that is approximately solved by searching 1D solution space, instead of searching 2D solution space in the latest work cited in this paper, to reduce time complexity. Three heuristic algorithms are proposed to determine suitable partitions to satisfy HW/SW partitioning constraints. We have shown that the time complexity for partitioning a graph with n nodes and m edges is significantly reduced from O(dx· dy· n3) to O(n log n + d · (n + m)), where d and dx· dyare the number of the fragments of the searched 1D solution space and the searched 2D solution space, respectively. The lower bound on the solution quality is also proposed based on the new computing model to show that it is comparable to that reported in the literature. Moreover, empirical results show that the proposed algorithms produce comparable and often better solutions when compared to the latest algorithm while reducing the time complexity significantly. Jigang Wu, Thambipillai Srikanthan |
IEEE Trans. Computers | 1 |
| 2010 | Preprocessing and Partial Rerouting Techniques for Accelerating Reconfiguration of Degradable VLSI ArraysabstractThis paper presents novel techniques to accelerate the reconfiguration of degradable very large scale integration arrays. A preprocessing step is used to derive the upper and lower size bounds of the maximum logical array (MLA) such that only those subarrays that possibly contain the MLA are reconfigured, thereby reducing the reconfiguration time and also obtaining a same-sized logical array. In addition, the partial rerouting approach is generalized so that as many as possible previous routing results can be reused in the current rerouting step. The reconfiguration time is reduced from$ O((1-\rho)\cdot \beta \cdot m \cdot n)$to its lower bound$ O((1-\rho)\cdot m\cdot n)$for$ m\times n$host arrays with small fault density$ \rho $, where$ \beta $is the expected routing length required per logical column. Jigang Wu, Thambipillai Srikanthan, Xiaogang Han |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2009 | Fast enumeration of maximal valid subgraphs for custom-instruction identificationabstractExtensible processors are increasingly becoming popular as they allow for incorporating custom instructions to meet design constraints. However, identifying custom instructions under architectural input/output ports constraint is a time consuming process particularly when large applications are considered. To rapidly identify the most profitable custom instructions with large inputs and outputs, this paper proposes a novel identification algorithm for enumerating maximal convex subgraphs containing no invalid node (i.e., maximal valid subgraphs). The proposed enumerating strategy is based on divide-and-conquer with a top-down manner, rather than the bottom-up manner utilized in the state-of-the-art. The division operation only considers invalid inner nodes of the given DFG, rather than taking all the invalid nodes into account, and thus accelerates enumeration of the maximal valid subgraphs. Experimental results show that, the improvement over the latest work is more than 90% for 60% DFG instances of the acknowledged benchmarks. Tao Li 0008, Zhigang Sun 0002, Jigang Wu, Xicheng Lu |
CASES | 3 |
| 2009 | Mapping Algorithms for NoC-Based Heterogeneous MPSoC PlatformsabstractMapping of applications onto multiprocessor system-on-chip (MPSoC) can be realized either at design-time or run-time. At any time the number of tasks executing in MPSoC platform can exceed the available resources, requiring efficient run-time mapping techniques to meet the real-time constraints of the applications. This paper presents two run-time mapping heuristics for mapping the tasks of an application in close proximity so as to minimize the communication overhead. In particular, the communication overhead between two adjacent hardware tasks is eliminated by mapping them onto the same reconfigurable processing node. We show that the proposed approach is capable of alleviating network-on-chip (NoC) congestion bottlenecks to optimize the overall performance. Based on our investigations to map the tasks of applications' at run-time onto an 8×8 NoC-based heterogeneous MPSoC, our mapping heuristics are capable of reducing total execution time and average channel load of applications when compared to state-of-the-art runtime mapping heuristics. Amit Kumar Singh 0002, Jigang Wu, Alok Prakash, Thambipillai Srikanthan |
DSD | 2 |
| 2009 | Minimizing interconnect length on reconfigurable meshes
Jigang Wu, Thambipillai Srikanthan |
Frontiers Comput. Sci. China | 1 |
| 2008 | Finding minimum interconnect sub-arrays in reconfigurable VLSI arraysabstractShorter total interconnect and fewer switches in a VLSI array definitely lead to less capacitance, power dissipation and dynamic communication cost between the processing elements (PEs). This paper presents techniques to find a logical (target) array that has shorter interconnect and fewer switches in a reconfigurable VLSI array with faulty PEs. The proposed algorithm initially searches for a sub-array on the host array, which contains the minimum number of the faults. Then it reroutes the sub-array, rather than the whole host array as was done in previous algorithm, to an approximate target array whose size is less then but close to the size of the target array. Finally, the target array is obtained by simple extension of the approximate target array. Experimental results show that the proposed algorithm can construct a target array with much shorter total interconnect. The improvement over the previous work is up to 68% in terms of the interconnect redundancy for the case of the cluster faults. Jigang Wu, Thambipillai Srikanthan |
ISCAS | 1 |
| 2008 | A temperature-aware virtual submesh allocation scheme for noc-based manycore chipsabstractVarious continuous and non-continuous submesh allocation schemes have been proposed for traditional mesh-connected multiprocessor systems. Discussions of these schemes are limited to issues such as fragmentation, system performance and algorithmic complexity. NoC-based manycore chips with mesh topology are envisaged to enter the domain of parallel computing and cores of them are allocated to applications in the forms of submeshes. However, unlike the processors in traditional systems, these cores may face severe runtime thermal non-uniformities and a promising strategy to avoid these thermal crises is keeping good heat balance throughout chips at runtime. One opportunity for good heat balance is to include thermally favorable cores when submeshes are allocated. This requires a suitable temperature sensitive submesh allocation scheme for manycore chips. Xiongfei Liao, Jigang Wu, Thambipillai Srikanthan |
SPAA | 2 |
| 2008 | New Model and Algorithm for Hardware/Software Partitioning
Jigang Wu, Thambipillai Srikanthan, Guang-Wei Zou |
J. Comput. Sci. Technol. | 1 |
| 2007 | Temperature-Aware Submesh Allocation Scheme for Heat Balancing on Chip-MultiprocessorsabstractThis paper explores the thermal problems in future CMPs in multiprogrammed environment for heat balancing. We first give the observation of the temperature variation of cores in this scenario. Then we propose a temperature-aware submesh allocation scheme to manage cores with submeshes and allocate submeshes of cores to jobs under temperature-aware policies to balance heat chip-wide. Several scheduling policies are suggested and a HotSpot-based thermal simulator is used to evaluate the scheme and its policies under the workloads of benchmark programs. Simulation results show that our proposed scheme with global coolest policy and global neighbor-aware policy can lead to lower peak temperatures and effectively reduce the temporal variance and spatial variance of temperatures of cores to achieve better heat balance. Xiongfei Liao, Jigang Wu, Thambipillai Srikanthan |
ASAP | 2 |
| 2007 | One-dimensional Search Algorithms for Hardware/Software PartitioningabstractHardware/software (HW/SW) partitioning is one of the key challenges in HW/SW co-design. This paper presents a new formulation to handle the HW/SW partitioning problem, which has been proved to be NP-hard. The proposed formulation transforms the partitioning problem into an extended 0-1 knapsack problem that is approximately solved in this paper by scanning a one-dimensional search space, instead of scanning a two-dimensional search space as presented in the literature cited in this paper. Two heuristic algorithms are proposed to explore the feasible partitions to meet the given constraints. The time complexity of the latest heuristic algorithm is significantly reduced from O(dxldr dyldr n3) to O(n log n + d ldr (n + m)) for the given graphs with n nodes and m edges, where dxldr dyis the number of the fragments of the scanned two-dimensional search space, and d is that of the scanned one-dimensional search space. Empirical results show that the proposed algorithms run extremely fast and still produce better or similar solutions in comparison with the latest algorithm. Jigang Wu, Thambipillai Srikanthan |
MEMOCODE | 1 |
| 2007 | Integrated Row and Column Rerouting for Reconfiguration of VLSI Arrays with Four-Port SwitchesabstractThis paper deals with the issue of developing efficient algorithms for reconfiguring two-dimensional VLSI arrays linked by four-port switches in the presence of faulty processing elements (PEs). The proposed algorithm reroutes the arrays with faults in both row and column directions at the same time. Unlike previous work, the compensation technique to replace the faulty PE is not restricted to the adjacent rows of the excluded row. Instead, we consider the neighbor rows of any faulty PE for compensation purposes. The nonfaulty PEs lying in the excluded rows are also effectively utilized to form the maximal target arrays, making the proposed algorithm more efficient in terms of both the percentages of harvest and the degradation of VLSI arrays for random and clustered faults. Empirical study shows that the improvement in harvest increases with increasing fault size and is more notable for maximal square target arrays than for maximal target arrays. Our investigations show that the improvement can be up to 8 percent and 23 percent for a 256 times 256 VLSI array with random faults of size 25 percent for maximal target arrays and for maximal square target arrays, respectively. Jigang Wu, Thambipillai Srikanthan |
IEEE Trans. Computers | 1 |
| 2006 | Efficient algorithm for functional scheduling in hardware/software co-designabstractTask scheduling is one of the crucial steps during functional hardware/software co-design. Due to the possibly concurrent execution of the tasks implemented in hardware, the NP-hard scheduling problem becomes more difficult to solve optimally. In this paper an efficient algorithm is proposed for task scheduling in functional hadware/software co-design. The proposed algorithm assigns the priority for each task combining the information both in the communication penalty and the hardware-only critical path, to enhance the parallelism of the tasks. A large body of experimental results confirm that the proposed algorithm is superior to the most widely used approaches first-come first-schedule(FCFS) and level-by-level schedule (LBLS) in hardware/software scheduling, both for random graphs and some realistic application graphs, without large increase in running time. The improvement over FCFS and LBLS is up to 10% for some random graphs, and it is more significant for FFT application graphs, according to the simulation results on the same types of graphs (under the same assumptions) as in the literature where LBLS is employed Jigang Wu, Thambipillai Srikanthan, Tao Jiao |
FPT | 1 |
| 2006 | Low-complex dynamic programming algorithm for hardware/software partitioning
Jigang Wu, Thambipillai Srikanthan |
Inf. Process. Lett. | 1 |
| 2006 | An efficient algorithm for the collapsing knapsack problem
Jigang Wu, Thambipillai Srikanthan |
Inf. Sci. | 1 |
| 2006 | Reconfiguration Algorithms for Power Efficient VLSI Subarrays with Four-Port SwitchesabstractTechniques to determine subarrays when processing elements of VLSI arrays become faulty have been investigated extensively. These tend to identify the largest subarray that is possible without concentrating on the power efficiency of the resulting subarray. In this paper, we propose new techniques, based on heuristic strategy and dynamic programming, to minimize the interconnect length in an attempt to reduce power dissipation without performance penalty. Our algorithms show that notable improvements in the reduction of the number of long interconnects could be realized in linear time and without sacrificing the size of the subarray. Our evaluations show that, for a VLSI array of size 256/spl times/256, the number of long interconnects in the subarray can be reduced by up to 95 percent for clustered faults and up to 50 percent and 73 percent for a random fault with density of 10 percent and 0.1 percent, respectively, when compared with the most efficient implementation cited in the literature. The interconnect power saving for a VLSI array of size 512/spl times/512 is by up to 11 percent for a random fault. We have also shown that interconnect power savings of up to 14 percent are possible for the cases investigated. Simulations based on several random and clustered fault scenarios clearly reveal the superiority of the proposed techniques for power efficient realizations. In addition, the lower bound of the performance has been proposed to demonstrate that the proposed algorithms are nearly optimal for the cases considered in this paper. Jigang Wu, Thambipillai Srikanthan |
IEEE Trans. Computers | 1 |
| 2006 | Algorithmic aspects of area-efficient hardware/software partitioning
Jigang Wu, Thambipillai Srikanthan |
J. Supercomput. | 1 |
| 2005 | Efficient Techniques and Hardware Analysis for Mesh-Connected Processors
Jigang Wu, Thambipillai Srikanthan |
ICA3PP | 1 |
| 2005 | Power Efficient Sub-Array in Reconfigurable VLSI Meshes
Jigang Wu, Thambipillai Srikanthan |
J. Comput. Sci. Technol. | 1 |
| 2005 | Efficient reconfigurable techniques for VLSI arrays with 6-port switchesabstractThis paper proposes an efficient techniques to reconfigure a two-dimensional degradable very large scale integration/wafer scale integration (VLSI/WSI) array under the row and column routing constraints, which has been shown to be NP-complete. The proposed VLSI/WSI array consists of identical processing elements such as processors or memory cells embedded in a 6-port switch lattice in the form of a rectangular grid. It has been shown that the proposed VLSI structure with 6-port switches eliminates the need to incorporate internal bypass within processing elements and leads to notable increase in the harvest when compared with the one using 4-port switches. A new greedy rerouting algorithm and compensation approaches are also proposed to maximize harvest through reconfiguration. Experimental results show that the proposed VLSI array with 6-port switches consistently outperforms the most efficient alternative proposed in literature, toward maximizing the harvest in the presence of fault processing elements. Jigang Wu, Thambipillai Srikanthan |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2004 | An efficient data structure for branch-and-bound algorithm
Jigang Wu, Thambipillai Srikanthan |
Inf. Sci. | 1 |
| 2003 | An improved reconfiguration algorithm for degradable VLSI/WSI arrays
Jigang Wu, Thambipillai Srikanthan |
J. Syst. Archit. | 1 |
| 2002 | New Architecture and Algorithms for Degradable VLSI/WSI Arrays
Jigang Wu, Thambipillai Srikanthan |
COCOON | 1 |
| 2000 | An Optimal Online Algorithm for Halfplane Intersection
Jigang Wu, Yongchang Ji, Guoliang Chen 0001 |
J. Comput. Sci. Technol. | 1 |
| 1994 | The least basic operations on heap and improved heapsort
Jigang Wu, Hong Zhu 0004 |
J. Comput. Sci. Technol. | 1 |