EDBT 2026 Demo / reviewers in the wild / expert
Quan Wang 0006
dblp:86/5728-6
· DBLP profile ↗
97ranked-venue papers
0as first author
72since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 20 since 2021Computer networks · 14 · 12 since 2021Systems, architecture and hardware · 11 · 7 since 2021Databases, data management, data science and information retrieval · 10 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 6 since 2021Security and privacy · 5 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RSVG-ZeroOV: Exploring a Training-Free Framework for Zero-Shot Open-Vocabulary Visual Grounding in Remote Sensing ImagesabstractRemote sensing visual grounding (RSVG) aims to localize objects in remote sensing images based on free-form natural language expressions. Existing approaches are typically constrained to closed-set vocabularies, limiting their applicability in open-world scenarios. While recent attempts to leverage generic foundation models for open-vocabulary RSVG, they overly rely on expensive high-quality datasets and time-consuming fine-tuning. To address these limitations, we propose RSVG-ZeroOV, a training-free framework that aims to explore the potential of frozen generic foundation models for zero-shot open-vocabulary RSVG. Specifically, RSVG-ZeroOV comprises three key stages: (i) Overview: We utilize a vision-language model (VLM) to obtain cross-attention maps that capture semantic correlations between text queries and visual regions. (ii) Focus: By leveraging the fine-grained modeling priors of a diffusion model (DM), we fill in gaps in structural and shape information of objects, which are often overlooked by VLM. (iii) Evolve: A simple yet effective attention evolution module is introduced to suppress irrelevant activations, yielding purified segmentation masks over the referred objects. Without cumbersome task-specific training, RSVG-ZeroOV offers an efficient and scalable solution. Extensive experiments demonstrate that the proposed framework consistently outperforms existing weakly-supervised and zero-shot methods. Ke Li 0024, Di Wang 0011, Fuyu Dong, Quan Wang 0006 |
AAAI | 9 |
| 2026 | Anatomical Region-Guided Contrastive Decoding: A Plug-and-Play Strategy for Mitigating Hallucinations in Medical VLMsabstractMedical Vision-Language Models (MedVLMs) show immense promise in clinical applicability. However, their reliability is hindered by hallucinations, where models often fail to derive answers from visual evidence, instead relying on learned textual priors. Existing mitigation strategies for MedVLMs have distinct limitations: training-based methods rely on costly expert annotations, limiting scalability, while training-free interventions like contrastive decoding, though data-efficient, apply a global, untargeted correction whose effects in complex real-world clinical settings can be unreliable. To address these challenges, we introduce Anatomical Region-Guided Contrastive Decoding (ARCD), a plug-and-play strategy that mitigates hallucinations by providing targeted, region-specific guidance. Our module leverages an anatomical mask to direct a three-tiered contrastive decoding process. By dynamically re-weighting at the token, attention, and logits levels, it verifiably steers the model's focus onto specified regions, reinforcing anatomical understanding and suppressing factually incorrect outputs. Extensive experiments across diverse datasets, including chest X-ray, CT, brain MRI, and ocular ultrasound, demonstrate our method's effectiveness in improving regional understanding, reducing hallucinations, and enhancing overall diagnostic accuracy. Di Wang 0011, Bin Jing, Quan Wang 0006 |
AAAI | 6 |
| 2026 | VideoAgent: Personalized Synthesis of Scientific VideosabstractThe technical complexity of research papers often limits their reach, necessitating more accessible formats like scientific videos to disseminate key insights through engaging narration. However, existing automated methods primarily focus on static posters or slide presentations that remain template-bound and linear. Shifting to audience-adaptive video synthesis requires addressing non-linear narrative orchestration and the joint synchronization of disparate multimodal assets. We introduce VideoAgent, a modular framework that redefines scientific video synthesis as an intent-driven planning problem. By decoupling content understanding from multimodal synthesis, VideoAgent adaptively interleaves static slides with dynamic animations to match the semantic density of the narration. We further propose SciVidEval, a benchmark evaluating multimodal quality and pedagogical utility through automated metrics and human knowledge transfer studies. Extensive experiments demonstrate that VideoAgent effectively conveys complex technical logic with high narrative fidelity and communicative impact. Bangxin Li, Hanyue Zheng, Di Wang 0011, Cong Tian 0001, Quan Wang 0006 |
ICMR | 8 |
| 2026 | Riemannian spatio-temporal graph neural network for enhanced cognitive load detection using EEG
Jiayang Huang, Dingnan Li, Pengfei Yang 0001, Quan Wang 0006, Zhiqiang Zhang 0001 |
Neurocomputing | 6 |
| 2026 | PUF-D2PB: A Low-Latency and Low-Cost PUFs-Based Framework for Direct Mutual Authentication Between IIoT Devices and Peer Nodes in Consortium BlockchainsabstractIndustrial Internet of Things (IIoT) devices often rely on public communication channels for authentication, which makes them particularly susceptible to a variety of security threats. Physically Unclonable Function (PUF) provides a lightweight cryptographic method to address authentication challenges in resource-constrained devices. However, some PUF-based authentication frameworks store challenge-response pairs (CRPs) centrally on a server, posing risks of single point failure and performance bottlenecks in the authentication process. Although blockchain (BC) can mitigate this through decentralized storage, they still suffer from high computational overhead, large CRPs storage requirements, limited authentication efficiency, and ledger synchronization latency challenges. Therefore, we propose a novel direct mutual authentication framework for IIoT devices and peer nodes in consortium BCs, called PUF-D2PB. It enables peer nodes and IIoT devices to authenticate each other without involving BC-clients, and simultaneously supports CRP updates without requiring extra operations. In addition, we introduce a competition-based mechanism and a distributed encryption method to enhance CRP synchronization and prevent CRP leakage. Security analysis and a prototype implementation using a real PUF circuit demonstrate that our framework provides strong security guarantees and efficient authentication. Compared with existing BC-assisted PUF schemes, PUF-D2PB effectively reduces authentication latency and resource overhead while preventing CRP leakage. Yin Chen 0001, Xiaohong Jiang 0001, Quan Wang 0006 |
IEEE Internet Things J. | 6 |
| 2026 | CHIME: Cost-Constrained Hybrid Popularity-Aware Intelligent Service Caching Framework for MEC
Tianyang Zheng, Pengfei Yang 0001, Chenlu Zhai, Wenkai Lv, Yueli Ding, Quan Wang 0006 |
IEEE Internet Things J. | 9 |
| 2026 | I2ID: Disentangling identity features via synchronized masking for zero-shot composed person retrieval
Di Wang 0011, Chengwei Yan, Nan Luo, Yifeng Wang 0004, Quan Wang 0006 |
Pattern Recognit. | 7 |
| 2026 | Adaptive Task Offloading Strategy in Vehicle-Assisted Mobile Edge ComputingabstractTraditional Mobile Edge Computing (MEC) is over whelmed by time-sensitive applications in the Internet of Vehicles (IoV), leading to significant task completion delays because existing methods fail to leverage vehicle-to-infrastructure collaboration in dynamic network topologies. This paper proposes a vehicle-assisted adaptive task offloading strategy to minimize completion time through a dual-mode framework that adapts based on a vehicle's position relative to an edge server. When a vehicle is outside a server's range, the Best Service Vehicle Selection Algorithm (BSVSA) offloads tasks to the most suitable nearby vehicle while ensuring communication stability. When within server coverage, our novel Hybrid Differential Teaching Optimization Algorithm (HDTOA) determines the optimal offloading ratio and schedules tasks across edge servers to balance the computational load. Simulation results validate that our integrated approach (HDTOA+BSVSA) outperforms benchmarks like Differential Evolution (DE) and Particle Swarm Optimization (PSO), demonstrating faster convergence and lower average task execution times under heavy load. Under scenarios with a large task data size, the HDTOA reduces the average task execution time by 69.99% compared to the PSO algorithm. The strategy also provides a more balanced workload across servers, thus enhancing overall system efficiency Hui Zhao 0003, Jing Wang 0028, Quan Wang 0006 |
IEEE Trans. Serv. Comput. | 5 |
| 2025 | FD2-Net: Frequency-Driven Feature Decomposition Network for Infrared-Visible Object DetectionabstractInfrared-visible object detection (IVOD) seeks to harness the complementary information in infrared and visible images, thereby enhancing the performance of detectors in complex environments. However, existing methods often neglect the frequency characteristics of complementary information, such as the abundant high-frequency details in visible images and the valuable low-frequency thermal information in infrared images, thus constraining detection performance. To solve this problem, we introduce a novel Frequency-Driven Feature Decomposition Network for IVOD, called FD2-Net, which effectively captures the unique frequency representations of complementary information across multimodal visual spaces. Specifically, we propose a feature decomposition encoder, wherein the high-frequency unit (HFU) utilizes discrete cosine transform to capture representative high-frequency features, while the low-frequency unit (LFU) employs dynamic receptive fields to model the multi-scale context of diverse objects. Next, we adopt a parameter-free complementary strengths strategy to enhance multimodal features through seamless inter-frequency recoupling. Furthermore, we innovatively design a multimodal reconstruction mechanism that recovers image details lost during feature extraction, further leveraging the complementary information from infrared and visible images to enhance overall representational capacity. Extensive experiments demonstrate that FD2-Net outperforms state-of-the-art (SoTA) models across various IVOD benchmarks, i.e. LLVIP (96.2% mAP), FLIR (82.9% mAP), and M3FD (83.5% mAP). Ke Li 0024, Di Wang 0011, Zhangyuan Hu, Weiping Ni, Lin Zhao 0003, Quan Wang 0006 |
AAAI | 7 |
| 2025 | EvoFormer: Learning Dynamic Graph-Level Representations with Structural and Temporal Bias CorrectionabstractDynamic graph-level embedding aims to capture structural evolution in networks, which is essential for modeling real-world scenarios. However, existing methods face two critical yet under-explored issues: Structural Visit Bias, where random walk sampling disproportionately emphasizes high-degree nodes, leading to redundant and noisy structural representations; and Abrupt Evolution Blindness, the failure to effectively detect sudden structural changes due to rigid or overly simplistic temporal modeling strategies, resulting in inconsistent temporal embeddings. To overcome these challenges, we propose EvoFormer, an evolution-aware Transformer framework tailored for dynamic graph-level representation learning. To mitigate Structural Visit Bias, EvoFormer introduces a Structure-Aware Transformer Module that incorporates positional encoding based on node structural roles, allowing the model to globally differentiate and accurately represent node structures. To overcome Abrupt Evolution Blindness, EvoFormer employs an Evolution-Sensitive Temporal Module, which explicitly models temporal evolution through a sequential three-step strategy: (I) Random Walk Timestamp Classification, generating initial timestamp-aware graph-level embeddings; (II) Graph-Level Temporal Segmentation, partitioning the graph stream into segments reflecting structurally coherent periods; and (III) Segment-Aware Temporal Self-Attention combined with an Edge Evolution Prediction task, enabling the model to precisely capture segment boundaries and perceive structural evolution trends, effectively adapting to rapid temporal shifts. Extensive evaluations on five benchmark datasets confirm that EvoFormer achieves state-of-the-art performance in graph similarity ranking, temporal anomaly detection, and temporal segmentation tasks, validating its effectiveness in correcting structural and temporal biases. Code is available at https://github.com/zlx0823/EvoFormerCode. Haodi Zhong, Liuxin Zou, Di Wang 0011, Bo Wan 0002, Zhenxing Niu, Quan Wang 0006 |
CIKM | 6 |
| 2025 | Uncertainty-Driven Expert Control: Enhancing the Reliability of Medical Vision-Language ModelsabstractThe rapid advancements in Vision Language Models (VLMs) have prompted the development of multi-modal medical assistant systems. Despite this progress, current models still have inherent probabilistic uncertainties, often producing erroneous or unverified responses-an issue with serious implications in medical applications. Existing methods aim to enhance the performance of Medical Vision Language Model (MedVLM) by adjusting model structure, fine-tuning with high-quality data, or through preference fine-tuning. However, these training-dependent strategies are costly and still lack sufficient alignment with clinical expertise. To address these issues, we propose an expert-in-the-loop framework named Expert-Controlled Classifier-Free Guidance (Expert-CFG) to align MedVLM with clinical expertise without additional training. This framework introduces an uncertainty estimation strategy to identify unreliable outputs. It then retrieves relevant references to assist experts in highlighting key terms and applies classifier-free guidance to refine the token embeddings of MedVLM, ensuring that the adjusted outputs are correct and align with expert highlights. Evaluations across three medical visual question answering benchmarks demonstrate that the proposed Expert-CFG, with 4.2B parameters and limited expert annotations, outperforms state-of-the-art models with 13B parameters. The results demonstrate the feasibility of deploying such a system in resource-limited settings for clinical use. Di Wang 0011, Zhicheng Jiao, Ronghan Li, Pengfei Yang 0001, Quan Wang 0006, Tat-Seng Chua |
ICCV | 6 |
| 2025 | Text-Guided Attribute Enhancement Framework for Composed Image RetrievalabstractComposed image retrieval is a challenging multimodal task that refers to the process of retrieving target image by taking advantage of both complementary and synergistic image and text input. Existing efforts often focus on designing interaction models to fuse global query image and text features. However, these approaches struggle to capture fine-grained semantic association information between query image and text, especially when it comes to identifying specific objects or attributes in the query image that need to be modified in the text. In addition, these methods fail to adequately model cross-modal attention when dealing with composed query and target image, resulting in the model's inability to accurately map the semantic information from the composed query to the corresponding regions in the target image. To address these challenges, we propose a Text-Guided Attribute Enhancement Framework for Composed Image Retrieval (TAE-CIR). Our approach consists of three key modules: (a) Multi-granularity vision aggregation module, which extracts multi-granularity visual features and captures fine-grained object-level features related to the query text, refining object and attribute representations for more precise retrieval; (b) Multi-level fusion interaction module, which facilitates deep cross-modal interactions between the composed query and target image features, effectively capturing complex semantic relationships from the composed query to target image; (c) Composed feature alignment, which fuses multi-granularity visual features with the text using a text-guided Q-Former and contrastive learning to ensure accurate alignment between the composed query and the target image. Our extensive experiments on benchmark datasets FashionIQ and CIRR demonstrate the superiority of our proposed method. Yizi Huang, Di Wang 0011, Bo Wan 0002, Lin Zhao 0003, Quan Wang 0006 |
ICMR | 6 |
| 2025 | MST-Distill: Mixture of Specialized Teachers for Cross-Modal Knowledge Distillation
Pengfei Yang 0001, Juanyang Chen, Yanxin Chen, Quan Wang 0006 |
ACM Multimedia | 6 |
| 2025 | CheXPO: Preference Optimization for Chest X-ray VLMs with Counterfactual RationaleabstractVision-language models (VLMs) are prone to hallucinations that critically compromise reliability in medical applications. While preference optimization can mitigate these hallucinations through clinical feedback, its implementation faces challenges such as clinically irrelevant training samples, imbalanced data distributions, and prohibitive expert annotation costs. To address these challenges, we introduce CheXPO, a Chest X-ray Preference Optimization strategy that combines confidence-similarity joint mining with counterfactual rationale. Our approach begins by synthesizing a unified, fine-grained multi-task chest X-ray visual instruction dataset across different question types for supervised fine-tuning (SFT). We then identify hard examples through token-level confidence analysis of SFT failures and use similarity-based retrieval to expand hard examples for balancing preference sample distributions, while synthetic counterfactual rationales provide fine-grained clinical preferences, eliminating the need for additional expert input. Experiments show that CheXPO achieves 8.93% relative performance gain using only 5% of SFT samples, reaching state-of-the-art performance across diverse clinical tasks. 1Code: https://github.com/ResearchGroup-MedVLLM/CheX-Phi35V Di Wang 0011, Lin Zhao 0003, Ronghan Li, Bo Wan 0002, Quan Wang 0006 |
ACM Multimedia | 8 |
| 2025 | Alternating optimization for energy consumption-oriented task offloading in SAGIN
Pengfei Yang 0001, Tianyang Zheng, Weidi Su, Bijie Yi, Wenkai Lv, Quan Wang 0006 |
Comput. Networks | 7 |
| 2025 | Bayesian deep multi-instance learning for student performance prediction based on campus big data
Jiayang Huang, Keyi Yang, Quan Wang 0006, Pengfei Yang 0001, Ziling Ruan, Zhiqiang Zhang 0001 |
Neurocomputing | 3 |
| 2025 | A Reinforcement Learning Framework for Efficient Task Allocation Among AGVs in Smart WarehouseabstractIn smart warehouses that use automated guided vehicles (AGVs) for goods transportation, task allocation has a great impact on operational efficiency. Currently, warehouse task allocation is typically modeled as a pickup and delivery problem (PDP), which requires vehicles to start and return from the same depot to construct several closed-loop routes. This approach increases the vehicle travel distance without load in high-throughput warehouses and results in resource wastage. Thus, we remodel the task allocation problem as an open-loop routing problem with heterogeneous starting points and name it capacitied multiagent open PDP (CMOPDP), which has more complex solution space and constraints than PDP. The solving speed of existing heuristic methods cannot meet the real-time processing demands of large-scale warehouses. And deep reinforcement learning (DRL)-based methods typically satisfy constraints through the output mask of decoders, which leads to unsatisfactory quality of solutions under complex constraints. To address these limitations, we design an DRL-based model with encoder-decoder architecture to solve the CMOPDP. Specifically, first, an encoder with heterogeneous attention is designed to fully explore constraint relationships between nodes. Second, we utilize dual decoders and information sharing to maximize vehicle-customer nodes matching. Finally, entropy rewards are introduced to enhance exploration during reinforcement learning, preventing the model from getting stuck in local optima. Extensive experiments on random datasets and various warehouse maps demonstrate that our method improves solution quality by at least 1.76% over baselines, while maintaining competitive solving time and exhibiting good generalization performance. Zejian Zhao, Di Wang 0011, Ke Li 0024, Gang Liu 0006, Quan Wang 0006 |
IEEE Internet Things J. | 6 |
| 2025 | Multihardware Adaptive Latency Prediction for Neural Architecture SearchabstractIn hardware-aware neural architecture search (NAS), accurately assessing a model’s inference efficiency is crucial for search optimization. Traditional approaches, which measure numerous samples to train proxy models, are impractical across varied platforms due to the extensive resources needed to remeasure and rebuild models for each platform. To address this challenge, we propose a multihardware-aware NAS method that enhances the generalizability of proxy models across different platforms while reducing the required sample size. Our method introduces a multihardware adaptive latency prediction (MHLP) model that leverages one-hot encoding for hardware parameters and multihead attention mechanisms to effectively capture the intricate interplay between hardware attributes and network architecture features. Additionally, we implement a two-stage sampling mechanism based on probability density weighting to ensure the representativeness and diversity of the sample set. By adopting a dynamic sample allocation mechanism, our method can adjust the adaptive sample size according to the initial model state, providing stronger data support for devices with significant deviations. Evaluations on NAS benchmarks demonstrate the MHLP predictor’s excellent generalization accuracy using only 10 samples, guiding the NAS search process to identify optimal network architectures. Chengmin Lin, Pengfei Yang 0001, Quan Wang 0006 |
IEEE Internet Things J. | 3 |
| 2025 | Energy and Makespan Bi-Objective Optimization for UAV-Assisted MEC Task OffloadingabstractIn the framework of unmanned aerial vehicle (UAV)-assisted mobile edge computing (MEC), the UAV’s trajectory design is crucial for the effectiveness of task offloading. However, existing studies suffer from some limitations such as reliance on static decision metrics, prediction-based strategies with poor generalization, or high computational complexity, making them unsuitable for large-scale emergency scenarios with surging task demands. To address this issue, this study introduces the Location Priority Level (LPL) and proposes LPL-based Bi-objective Task Offloading Strategy (LBTOS), which tries to simultaneously minimize task completion time (makespan) and total energy consumption. By jointly considering task load, regional deadlines, and UAV energy constraints, LBTOS introduces the LPL to dynamically map task urgency to UAV trajectories. Leveraging LPL and an adaptive triggering mechanism, the proposed UAV Trajectory Design algorithm (UAVTD) achieves high responsiveness under computationally constrained emergency scenarios. Additionally, a task offloading algorithm named HWPSO-SA is designed, which integrates adaptive and time-varying inertia weights, and combines the Particle Swarm Optimization with Simulated Annealing, thereby enhancing the algorithm’s capability to obtain the global optimal solution. Lastly, simulation experiments are conducted to validate the significant effectiveness of LBTOS in the dual optimization objectives of minimizing task completion time and the total computing and transmission energy consumption of the system. Hui Zhao 0003, Xiaoqin Lu, Jinzhe Li, Jing Wang 0028, Quan Wang 0006 |
IEEE Internet Things J. | 5 |
| 2025 | A Task Scheduling Method for Minimizing Completion Time in Edge Collaboration EnvironmentabstractIn the edge computing environment, the uneven geographical distribution of tasks may lead to unbalanced load on the edge server. In addition, some larger tasks are difficult to completely offload to edge servers, which cannot fully utilize edge server resources. To solve the above problems, we propose a task scheduling method to minimize the completion time by combining the horizontal edge collaboration and fine-grained task partial offloading technology. First, combining horizontal edge collaboration and fine-grained task partial offloading technology, considering the location relationship between users and edge servers in multiuser multiedge server scenario, a task partial offloading optimization problem is established to minimize task completion time. Second, due to the nonconvex and variables coupling, we decompose the original problem into resource allocation, user-server association, and offloading strategy subproblems. A task scheduling algorithm based on improved teaching-learning-based optimization (ITLBO) is proposed to obtain the best task scheduling decision which includes task offloading location and offloading ratio. Simulation results show that the proposed method can effectively reduce the task completion time in edge collaboration environment. Hui Zhao 0003, Xiaoqin Lu, Jing Wang 0028, Pengfei Yang 0001, Bo Wan 0002, Quan Wang 0006 |
IEEE Internet Things J. | 7 |
| 2025 | Cortex: Enhancing Resource Utilization in Edge Clusters Through Efficient Co-Location of LC and BE WorkloadsabstractIn edge computing environments, the co-location of latency-critical (LC) services and best-effort (BE) jobs is a key strategy for enhancing resource utilization. However, existing analysis-based co-location strategies incur high analytical costs and struggle to rapidly adapt to the evolving fields of edge computing and microservice architectures, often failing to effectively meet the demands of edge computing environments. Feedback-based co-location strategies, while reducing analytical overhead, lack sufficient research in multi-node environments, resulting in overly coarse-grained deployment strategies. These strategies do not adequately consider the dynamic workloads and resource constraints inherent in edge computing, leading to improper resource allocation and degraded performance. This paper introduces Cortex, a Kubernetes-based co-location framework for edge device clusters that addresses these challenges by innovatively transforming the co-location deployment problem into a Minimum Cost Maximum Flow (MCMF) problem and employing the Network Simplex Algorithm (NSA) to optimize resource allocation and ensure QoS of LC services. Cortex also features a dynamic adjustment mechanism that adapts to changes in the request load of LC services, thereby minimizing the performance loss of BE jobs and reducing resource wastage. Our experiments in real edge device clusters demonstrate that Cortex significantly improves system resource utilization by 12.81%, increases the QoS satisfaction rate by 17.86%, and boosts the number of BE jobs by 51.96% compared to existing methods. Tianyang Zheng, Pengfei Yang 0001, Quan Wang 0006, Wenkai Lv |
IEEE Internet Things J. | 3 |
| 2025 | Multi-workflow fault-tolerance scheduling strategy considering resources supply delay in WaaS platforms
Hui Zhao 0003, Wentao Zhi, Xiaoqin Lu, Jing Wang 0028, Nan Luo, Bo Wan 0002, Quan Wang 0006 |
Parallel Comput. | 7 |
| 2025 | Adaptive spatial and scale label assignment for anchor-free object detection
Min Dang, Gang Liu 0006, Di Wang 0011, Xike Li, Quan Wang 0006 |
Pattern Recognit. | 6 |
| 2025 | An Ultrahigh-Throughput and FPGA-Compatible TRNG Based on Dynamic Hybrid Metastability and Jitter Entropy CellsabstractThe entropy source is the most critical component of a true random number generator (TRNG), which determines the quality of the random numbers. Current TRNGs mainly utilize a specific source of physical randomness as the entropy source, but it is difficult for this method to achieve a balance between low resource overhead and high throughput. This paper explores the self-feedback multiplexer (SFMUX) structure to obtain a novel dynamic hybrid entropy source for TRNGs. Unlike other MUX-based entropy source circuits, our SFMUX cross-connects the outputs of four independent high-frequency ring oscillators (ROs) as the input signals of four MUXs, and the output of each MUX is self-fed back to serve as a selection signal. Thus, the SFMUX can not only output jitter, but also update the selection signal rapidly and randomly, which increases the probability that the SFMUX outputs unstable signals. When using a D-flip-flop (DFF) to sample this signal, the DFF may become metastable. Modeling the entropy source shows that connecting 1-stage ROs and 2-stage ROs to each SFMUX can achieve higher minimum entropy than using ROs with other numbers of stages. The proposed TRNG design is implemented on Xilinx Virtex-6, Artix-7 and Kintex-7 FPGAs. The experimental results demonstrate that our TRNG achieves a maximum throughput of 550 Mbps while using only 6 slices, and it passes the NIST, AIS-31 and Dieharder tests without postprocessing. Yin Chen 0001, Lirong Zhou, Xiaohong Jiang 0001, Quan Wang 0006 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2025 | DFF-VIO: A General Dynamic Feature Fused Monocular Visual-Inertial OdometryabstractIntegrating dynamic effects has shown its significance in enhancing the accuracy and robustness of Visual-Inertial Odometry (VIO) systems in dynamic scenarios. Existing methods either prune dynamic features or rely heavily on prior semantic knowledge or kinetic models, proved unfriendly to scenes with a multitude of dynamic elements. This work proposes a novel dynamic feature fusion method for monocular VIO, named DFF-VIO, which requires no prior models or scene preference. By combining IMU-predicted poses with visual clues, it initially identifies dynamic features during the tracking stage by constraints of consistency and degree of motion. Then, we innovatively design a Dynamic Transformation Operation (DTO) to separate the effect of dynamic features on multiple frames into pairwise effects and construct a Dynamic Feature Cell (DFC) to preserve the eligible information. Subsequently, we reformulate the VIO nonlinear optimization problem and construct dynamic feature residuals with the transformed DFC as a unit. Based on the proposed inter-frame model of moving features, a so-called motion compensation is developed to resolve the reprojection issue of dynamic features, allowing their effects to be incorporated into the VIO’s tight coupling optimization, thereby realizing robust positioning in dynamic scenarios. We conduct accuracy evaluations on ADVIO and VIODE, degradation tests on EuRoC dataset, as well as ablation studies to highlight the joint optimization of dynamic residuals. Results reveal that DFF-VIO outperforms state-of-the-art methods in pose accuracy and robustness across various dynamic environments. Nan Luo, Zhexuan Hu, Hui Zhao 0003, Gang Liu 0006, Quan Wang 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | Physical Adversarial Patch Attack for Optical Fine-Grained Aircraft RecognitionabstractDeep neural networks (DNNs) have been widely used in remote sensing but demonstrated to be sensitive with adversarial examples. By introducing carefully designed perturbations to clean images, DNNs can be led to incorrect predictions. Adversarial patch is commonly used to conduct adversarial attack, where traditional methods optimize its content and position separately, neglecting the coupling relation of two factors. In this paper, we propose a black-box attack framework targeting fine-grained aircraft recognition, named PatchGen, simultaneously optimizing both content and position of physical adversarial patches. For the requirements of physical attack, we further constrain the patch in object region and utilize elaborate criteria to evaluate its naturalness to alleviate the distortion when applying the patch in real world. We comprehensively validate our method in fine-grained aircraft classification, extending to object detection subsequently. Extensive experiments demonstrate that the proposed method achieves superior attack performance efficiently for classification and detection tasks in digital domain. Moreover, we validate the effectiveness of the adversarial patch under diverse circumstances in the physical world and prove that our method can be applied to different models as well as various domains. Ke Li 0024, Di Wang 0011, Wenxuan Zhu, Quan Wang 0006, Xinbo Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | A Trusted Medical Image Zero-Watermarking Scheme Based on DCNN and Hyperchaotic SystemabstractThe zero-watermarking methods provide a means of lossless, which was adopted to protect medical image copyright requiring high integrity. However, most existing studies have only focused on robustness and there has been little discussion about the analysis and experiment on discriminability. Therefore, this paper proposes a trusted robust zero-watermarking scheme for medical images based on Deep convolution neural network (DCNN) and the hyperchaotic encryption system. Firstly, the medical image is converted into several feature map matrices by the specific convolution layer of DCNN. Then, a stable Gram matrix is obtained by calculating the colinear correlation between different channels in feature map matrices. Finally, the Gram matrixes of the medical image and the feature map matrixes of the watermark image are fused by the trained DCNN to generate the zero-watermark. Meanwhile, we propose two feature evaluation criteria for finding differentiated eigenvalues. The eigenvalue is used as the explicit key to encrypt the generated zero-watermark by Lorenz hyperchaotic encryption, which enhances security and discriminability. The experimental results show that the proposed scheme can resist common image attacks and geometric attacks, and is distinguishable in experiments, being applicable for the copyright protection of medical images. Ruotong Xiang, Gang Liu 0006, Min Dang, Quan Wang 0006, Rong Pan 0004 |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | PRA-Det: Anchor-Free Oriented Object Detection With Polar Radius RepresentationabstractOriented object detection typically adds an additional rotation angle to the regressed horizontal bounding box (HBB) for representing the oriented bounding box (OBB). However, existing oriented object detectors based on regression angles face inconsistency between metric and loss, boundary discontinuity or square-like problems. To solve the above problems, we propose an anchor-free oriented object detector named PRA-Det, which assigns the center region of the object to regress OBBs represented by the polar radius vectors. Specifically, the proposed PRA-Det introduces a diamond-shaped positive region of category-wise attention factor to assign positive sample points to regress polar radius vectors. PRA-Det regresses the polar radius vector of the edges from the assigned sample points as the regression target and suppresses the predicted low-quality polar radius vectors through the category-wise attention factor. The OBBs defined for different protocols are uniformly encoded by the polar radius encoding module into regression targets represented by polar radius vectors. Therefore, the regression target represented by the polar radius vector does not have angle parameters during training, thus solving the angle-sensitive boundary discontinuity and square-like problems. To optimize the predicted polar radius vector, we design a spatial geometry loss to improve the detection accuracy. Furthermore, in the inference stage, the center offset score of the polar radius vector is combined with the classification score as the confidence to alleviate the inconsistency between classification and regression. The extensive experiments on public benchmarks demonstrate that the PRA-Det is highly competitive with state-of-the-art oriented object detectors and outperforms other comparison methods. Min Dang, Gang Liu 0006, Hao Li 0095, Di Wang 0011, Rong Pan 0004, Quan Wang 0006 |
IEEE Trans. Multim. | 6 |
| 2024 | Unleashing Channel Potential: Space-Frequency Selection Convolution for SAR Object DetectionabstractDeep Convolutional Neural Networks (DCNNs) have achieved remarkable performance in synthetic aperture radar (SAR) object detection, but this comes at the cost of tremendous computational resources, partly due to extracting redundant features within a single convolutional layer. Recent works either delve into model compression methods or focus on the carefully-designed lightweight models, both of which result in performance degradation. In this paper, we propose an efficient convolution module for SAR object detection, called SFS-Conv, which increases feature diversity within each convolutional layer through a shunt-perceive-select strategy. Specifically, we shunt input feature maps into space and frequency aspects. The former perceives the context of various objects by dynamically adjusting receptive field, while the latter captures abundant frequency variations and textural features via fractional Gabor transformer. To adaptively fuse features from space and frequency aspects, a parameter-free feature selection module is proposed to ensure that the most representative and distinctive information are preserved. With SFS-Conv, we build a lightweight SAR object detection network, called SFS-CNet. Experimental results show that SFS-CNet outperforms state-of-the-art (SoTA) models on a series of SAR object detection benchmarks, while simultaneously reducing both the model size and computational cost. Ke Li 0024, Di Wang 0011, Zhangyuan Hu, Wenxuan Zhu, Quan Wang 0006 |
CVPR | 6 |
| 2024 | Leveraging Coarse-to-Fine Grained Representations in Contrastive Learning for Differential Medical Visual Question Answering
Di Wang 0011, Zhicheng Jiao, Haodi Zhong, Mengyu Yang, Quan Wang 0006 |
MICCAI (5) | 7 |
| 2024 | Divide and Conquer: Isolating Normal-Abnormal Attributes in Knowledge Graph-Enhanced Radiology Report GenerationabstractRadiology report generation aims to automatically generate clinical descriptions for radiology images, reducing the workload of radiologists. Compared to general image captioning tasks, the subtle differences in medical images and the specialized, complex nature of medical terminology limit the performance of data-driven radiology report generation. Previous research has attempted to leverage prior knowledge, such as organ-disease graphs, to enhance models' abilities to identify specific diseases and generate corresponding medical terminology. However, these methods cover only a limited number of disease types, focusing solely on disease terms mentioned in reports but ignoring their normal or abnormal attributes, which are critical to generating accurate reports. To address this issue, we propose a Divide-and-Conquer approach, named DCG, which separately constructs disease-free and disease-specific nodes within the knowledge graphs. Specifically, we extracted more comprehensive organ-disease entities from reports than previous methods and constructed disease-free and disease-specific nodes by rigorously distinguishing between normal conditions and specific diseases. This enables our model to consciously focus on abnormal information and mitigate the impact of excessively common diseases on report generation. Subsequently, the constructed graph is utilized to enhance the correlation between visual representations and disease terminology, thereby guiding the decoder in report generation. Extensive experiments conducted on benchmark datasets IU-Xray and MIMIC-CXR demonstrate the superiority of our proposed method. Code is available at https://github.com/ecoxial2007/DCG_Enhanced_distilGPT2. Yanlei Zhang, Di Wang 0011, Haodi Zhong, Ronghan Li, Quan Wang 0006 |
ACM Multimedia | 6 |
| 2024 | Explore Bayesian analysis in Cognitive-aware Key-Value Memory Networks for knowledge tracing in online learning
Juli Zhang, Ruoheng Xia, Qiguang Miao, Quan Wang 0006 |
Expert Syst. Appl. | 4 |
| 2024 | NIV-SSD: Neighbor IoU-voting single-stage object detector from point cloud
Shuai Liu 0009, Di Wang 0011, Quan Wang 0006, Kai Huang 0001 |
Neurocomputing | 3 |
| 2024 | Multimodal transformer with adaptive modality weighting for multimodal sentiment analysis
Yifeng Wang 0004, Di Wang 0011, Quan Wang 0006, Bo Wan 0002, Xuemei Luo |
Neurocomputing | 4 |
| 2024 | Ada-FA: A Comprehensive Framework for Adaptive Fault Tolerance and Aging Mitigation in FPGAsabstractCommercial SRAM-based field-programmable gate arrays (FPGAs) are extremely susceptible to failures caused by external ionizing radiation or prolonged internal overloading in harsh applications, such as single event effects (SEEs) and aging failure. Existing methods utilize the triple modular redundancy (TMR) architecture to shield against the effects of radiation on FPGA systems. However, these solutions are resource-costly and practically unnecessary. Additionally, hard faults caused by the aging effects of long-term usage of FPGA systems is not effectively alleviated. To address these issues, we present a comprehensive framework for ensuring adaptive fault tolerance and aging mitigation in FPGAs, i.e., Ada-FA. Ada-FA is a cross-layer-aware reliability framework that includes two phases: 1) offline and 2) online. (1) In the offline phase, a task criticality evaluation strategy supporting fine-grained fault tolerance is proposed to reduce the hardware resource overhead. Specifically, we improve the integer linear programming (ILP) formula, which considers both fault tolerance and aging mitigation, to obtain the optimal reliability-aware layout, thus maximizing the mean time to failure (MTTF). (2) In the online phase, we propose a runtime management architecture to further ensure the reliable operation of FPGA systems. The experimental results show that the resource usage (RU) of the proposed Ada-FA framework is reduced by 15.8% on average compared to that of existing fault-tolerant layout/scheduling methods. Moreover, our method provides a higher reliability and task accomplishment rate (TAR) than the state-of-the-art offline aging mitigation methods. Quan Wang 0006, Xiaohong Jiang 0001 |
IEEE Internet Things J. | 5 |
| 2024 | Graph-Reinforcement-Learning-Based Dependency-Aware Microservice Deployment in Edge ComputingabstractMicroservice architecture is a design philosophy that achieves decoupling by decomposing a monolithic application into multiple lightweight microservices. Meanwhile, edge computing can significantly reduce service latency and network congestion by extending computation and storage resources to the network edge. Therefore, in the microservice-oriented edge computing platform, a fundamental problem is how to efficiently deploy microservices with complex dependencies on the resource-constrained edge servers to satisfy the Quality of Service (QoS) constraints of users. Most of the existing studies ignore multiple call graphs with differentiated dependencies for an application, which often result in the violation of QoS. To address this issue, in this article, we first model the request response time of multiple instances and multiple call graphs scenario with service conflicts. Then, different from the existing heuristic or approximation algorithms which rely heavily on expert knowledge, we propose a graph-reinforcement-learning-based deployment (GRLD) framework. GRLD uses a graph convolutional network (GCN) to extract the graph data required for multiple call graphs with messages passing and aggregation, and the generated feature is fed into the underlying network of deep-reinforcement-learning (DRL). Experimental results show that GRLD outperforms counterparts in reducing service deployment overhead while satisfying QoS constraints of multiple call graphs. Wenkai Lv, Pengfei Yang 0001, Tianyang Zheng, Chengmin Lin, Minwen Deng, Quan Wang 0006 |
IEEE Internet Things J. | 7 |
| 2024 | Performance Prediction for Deep Learning Models With Pipeline Inference StrategyabstractFor Heterogeneous Multi-Processor System-on-Chips (HMPSoCs), a reasonable pipeline design can significantly improve the inference performance of Deep Learning (DL) models. The pipeline design optimization can be modeled as a search problem where an accurate prediction model can efficiently speed up the search process. However, the performance prediction of DL models for the pipeline inference strategy is challenging because of the inter-layer effect, inference details, and variety of model structures. In this paper, we propose TPPNet, a transformer-based model for predicting the inference performance of various DL models with the pipeline inference strategy. TPPNet represents the DL model as an execution sequence with operators and hardware details to extract the hidden factors between layers. Moreover, we apply the Multi-task Learning (MTL) method to accurately predict throughput and latency metrics by constructing a predictive model. To the best of our knowledge, this is the first study dedicated to pipeline inference performance prediction for the DL model on HMPSoCs. We evaluate TPPNet on six well-known DL models using RK3399. The experimental outcomes affirm the high accuracy of TPPNet and its capability to significantly reduce the time overhead associated with pipeline exploration. Pengfei Yang 0001, Linwei Hu, Wenkai Lv, Chengmin Lin, Quan Wang 0006 |
IEEE Internet Things J. | 8 |
| 2024 | Candidate-Heuristic In-Context Learning: A new framework for enhancing medical visual question answering with LLMs
Di Wang 0011, Haodi Zhong, Quan Wang 0006, Ronghan Li, Rui Jia, Bo Wan 0002 |
Inf. Process. Manag. | 4 |
| 2024 | Fine-grained complexity-driven latency predictor in hardware-aware neural architecture search using composite loss
Chengmin Lin, Pengfei Yang 0001, Wenkai Lv, Quan Wang 0006 |
Inf. Sci. | 7 |
| 2024 | Flexi-BOPI: Flexible granularity pipeline inference with Bayesian optimization for deep learning models on HMPSoC
Pengfei Yang 0001, Linwei Hu, Wenkai Lv, Chengmin Lin, Quan Wang 0006 |
Inf. Sci. | 7 |
| 2024 | DiagSWin: A multi-scale vision transformer with diagonal-shaped windows for object detection and segmentation
Ke Li 0024, Di Wang 0011, Gang Liu 0006, Wenxuan Zhu, Haodi Zhong, Quan Wang 0006 |
Neural Networks | 6 |
| 2024 | SLAPP: Subgraph-level attention-based performance prediction for deep learning models
Pengfei Yang 0001, Linwei Hu, Chengmin Lin, Wenkai Lv, Quan Wang 0006 |
Neural Networks | 7 |
| 2024 | Learning deep representation and discriminative features for clustering of multi-layer networks
Xiaoke Ma 0001, Quan Wang 0006, Maoguo Gong, Quanxue Gao |
Neural Networks | 3 |
| 2024 | Anonymity in Attribute-Based Access Control: Framework and MetricabstractAnonymous access is an effective method for preserving privacy in access control. This study assumes that anonymous access control requires both frameworks and policies. Numerous solutions have been proposed for anonymous access at the framework level. In this study, these solutions are analyzed and quantified using a unified attribute-based access control (ABAC) anonymous access reference framework. Anonymous access at the framework level is the first line of defense, and inappropriate policies may undermine subject anonymity. An anonymity metric is proposed at the policy level to prevent authorization authority from re-identification using specific attributes and policies. The anonymity metric evaluates the risk of re-identifying a subject due to inappropriate access requests, as well as subject attribute assignment schemes and policies. This study is the first to focus on anonymity at the policy level in ABAC. Furthermore, a formal definition of anonymity suitable for ABAC is proposed. The feasibility of the proposed anonymity metric is verified through simulations. Runnan Zhang, Gang Liu 0006, Hongzhaoning Kang, Quan Wang 0006, Bo Wan 0002, Nan Luo |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2024 | Gist, Content, Target-Oriented: A 3-Level Human-Like Framework for Video Moment RetrievalabstractVideo moment retrieval (VMR) aims to locate corresponding moments in an untrimmed video via a given natural language query. While most existing approaches treat this task as a cross-modal content matching or boundary prediction problem, recent studies have started to solve the VMR problem from a reading comprehension perspective. However, the cross-modal interaction processes of existing models are either insufficient or overly complex. Therefore, we reanalyze human behaviors in the document fragment location task of reading comprehension, and design a specific module for each behavior to propose a 3-level human-like moment retrieval framework (Tri-MRF). Specifically, we summarize human behaviors such as grasping the general structures of the document and the question separately, cross-scanning to mark the direct correspondences between keywords in the document and in the question, and summarizing to obtain the overall correspondences between document fragments and the question. Correspondingly, the proposed Tri-MRF model contains three modules: 1) a gist-oriented intra-modal comprehension module is used to establish contextual dependencies within each modality; 2) a content-oriented fine-grained comprehension module is used to explore direct correspondences between clips and words; and 3) a target-oriented integrated comprehension module is used to verify the overall correspondence between the candidate moments and the query. In addition, we introduce a biconnected GCN feature enhancement module to optimize query-guided moment representations. Extensive experiments conducted on three benchmarks, TACoS, ActivityNet Captions and Charades-STA demonstrate that the proposed framework outperforms State-of-the-Art methods. Di Wang 0011, Xiantao Lu, Quan Wang 0006, Yumin Tian, Bo Wan 0002, Lihuo He |
IEEE Trans. Multim. | 3 |
| 2024 | Dual-Perspective Fusion Network for Aspect-Based Multimodal Sentiment AnalysisabstractAspect-based multimodal sentiment analysis (ABMSA) is an important sentiment analysis task that analyses aspect-specific sentiment in data with different modalities (usually multimodal data with text and images). Previous works usually ignore the overall sentiment tendency when analyzing the sentiment of each aspect term. However, the overall sentiment tendency is highly correlated with aspect-specific sentiment. In addition, existing methods neglect to explore and make full use of the fine-grained multimodal information closely related to aspect terms. To address these limitations, we propose a dual-perspective fusion network (DPFN) that considers both global and local fine-grained sentiment information in multimodal data. From the global perspective, we use text-image caption pairs to obtain a global representation containing information about the overall sentiment tendencies. From the local fine-grained perspective, we construct two graph structures to explore the fine-grained information in texts and images. Finally, aspect-level sentiment polarities can be obtained by analyzing the combination of global and local fine-grained sentiment information. Experimental results on two multimodal Twitter datasets show that the proposed DPFN model outperforms state-of-the-art methods. Di Wang 0011, Changning Tian, Lin Zhao 0003, Lihuo He, Quan Wang 0006 |
IEEE Trans. Multim. | 6 |
| 2024 | Deep Hierarchical Multimodal Metric LearningabstractMultimodal metric learning aims to transform heterogeneous data into a common subspace where cross-modal similarity computing can be directly performed and has received much attention in recent years. Typically, the existing methods are designed for nonhierarchical labeled data. Such methods fail to exploit the intercategory correlations in the label hierarchy and, therefore, cannot achieve optimal performance on hierarchical labeled data. To address this problem, we propose a novel metric learning method for hierarchical labeled multimodal data, named deep hierarchical multimodal metric learning (DHMML). It learns the multilayer representations for each modality by establishing a layer-specific network corresponding to each layer in the label hierarchy. In particular, a multilayer classification mechanism is introduced to enable the layerwise representations to not only preserve the semantic similarities within each layer, but also retain the intercategory correlations across different layers. In addition, an adversarial learning mechanism is proposed to bridge the cross-modality gap by producing indistinguishable features for different modalities. Through integration of the multilayer classification and adversarial learning mechanisms, DHMML can obtain hierarchical discriminative modality-invariant representations for multimodal data. Experiments on two benchmark datasets are used to demonstrate the superiority of the proposed DHMML method over several state-of-the-art methods. Di Wang 0011, Aqiang Ding, Yumin Tian, Quan Wang 0006, Lihuo He, Xinbo Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | A Logic Encryption-Enhanced PUF Architecture to Deceive Machine Learning-Based Modeling AttacksabstractAs a low-cost hardware security primitive, physically unclonable function (PDF) has been widely utilized in secure key generation and identity authentication of physical devices due to the advantages of high reliability and randomness. However, hackers can model a PDF circuit by collecting a small number of challenge- response pairs (CRPs), making it potentially vulnerable to machine learning (ML) attacks. To effectively resist this risk, various ML-resistance methods have been proposed. However, most of them mainly enhance the anti-attack ability of PDFs by structure nonlinear and CRP obfuscation, reducing stability and reliability. Therefore, we present a logic encryption-enhanced PDF (LEE PDF) architecture to resist ML-based attacks without affecting PDF performance. A logic encryption unit is applied to protect the function of original PDF circuits, thus concealing the valid CRPs in a large number of useless ones. Since hackers can only obtain a sparse number of useful CRPs without knowing the correct key, making it impossible to model PDF using ML methods. We have implemented the proposed LEE PDF on FPGA micro-boards. The experimental results demonstrate that the LEE PDF with only 2-bit key can effectively resist various ML- based attacks, average prediction rate is close to 50 %. In addition, the PDF performance remains almost constant. Lirong Zhou, Quan Wang 0006, Bo Wan 0002 |
ATS | 5 |
| 2023 | Language-Guided Visual Aggregation Network for Video Question AnsweringabstractVideo Question Answering (VideoQA) aims to comprehend intricate relationships, actions, and events within video content, as well as the inherent links between objects and scenes, to answer text-based questions accurately. Transferring knowledge from the cross-modal pre-trained model CLIP is a natural approach, but its dual-tower structure hinders fine-grained modality interaction, posing challenges for direct application to VideoQA tasks. To address this issue, we introduce a Language-Guided Visual Aggregation (LGVA) network. It employs CLIP as an effective feature extractor to obtain language-aligned visual features with different granularities and avoids resource-intensive video pre-training. The LGVA network progressively aggregates visual information in a bottom-up manner, focusing on both regional and temporal levels, and ultimately facilitating accurate answer prediction. More specifically, it employs local cross-attention to combine pre-extracted question tokens and region embeddings, pinpointing the object of interest in the question. Then, graph attention is utilized to aggregate regions at the frame level and integrate additional captions for enhanced detail. Following this, global cross-attention is used to merge sentence and frame-level embeddings, identifying the video segment relevant to the question. Ultimately, contrastive learning is applied to optimize the similarities between aggregated visual and answer embeddings, unifying upstream and downstream tasks. Our method conserves resources by avoiding large-scale video pre-training and simultaneously demonstrates commendable performance on the NExT-QA, MSVD-QA, MSRVTT-QA, TGIF-QA, and ActivityNet-QA datasets, even outperforming some end-to-end trained models. Our code is available at https://github.com/ecoxial2007/LGVA_VideoQA. Di Wang 0011, Quan Wang 0006, Bo Wan 0002, Lingling An, Lihuo He |
ACM Multimedia | 3 |
| 2023 | An improved minimal noise role mining algorithm based on role interpretability
Hongzhaoning Kang, Gang Liu 0006, Quan Wang 0006, Jiamin Niu, Nan Luo |
Comput. Secur. | 3 |
| 2023 | Energy Consumption and QoS-Aware Co-Offloading for Vehicular Edge ComputingabstractBy deploying computing, storage, and bandwidth resources at the user side, vehicular edge computing (VEC) provides low-delay services for vehicle users. However, due to the limited resources of edge servers, how to efficiently meet the Quality-of-Service (QoS) requirements of multiple tasks and save the total energy consumption in a dynamic environment is an important issue in VEC. In this article, we first propose an energy consumption and QoS-aware co-offloading model. Unlike most previous studies, our goal is to minimize the total energy consumption while guaranteeing the QoS constraints of tasks, thus avoiding the overallocation of resources and high energy consumption caused by the one-sided pursuit of delay minimization. Then, without the requirements for domain experts, we propose Bayesian optimization-based computation offloading (BOCO) method to find the optimal offloading decision. To the best of our knowledge, this work is the first to apply Bayesian optimization to computation offloading in VEC. Furthermore, we conduct a series of experiments and comparisons with other offloading methods to analyze the effectiveness and performance of the proposed algorithm. Experimental results verify that our proposed BOCO outperforms counterparts. Wenkai Lv, Pengfei Yang 0001, Tianyang Zheng, Bijie Yi, Yunqing Ding, Quan Wang 0006, Minwen Deng |
IEEE Internet Things J. | 6 |
| 2023 | VM performance-aware virtual machine migration method based on ant colony optimization in cloud environment
Hui Zhao 0003, Nanzhi Feng, Guobin Zhang, Jing Wang 0028, Quan Wang 0006, Bo Wan 0002 |
J. Parallel Distributed Comput. | 6 |
| 2023 | Efficient and accurate compound scaling for convolutional neural networks
Chengmin Lin, Pengfei Yang 0001, Quan Wang 0006, Zeyu Qiu, Wenkai Lv |
Neural Networks | 3 |
| 2023 | QANS: Toward Quantized Neural Network Adversarial Noise SuppressionabstractNeural network quantization techniques play an important role in efficiently deploying deep learning models on the hardware with limited computing and storage resources. Numerous applications of this technology, such as autopilot, necessitate not just efficiency but also robustness. Research on the robustness of quantized networks against adversarial attacks is becoming one of the major points of interest. In this work, we rethink the impact of quantization on adversarial attacks and explore the boundary of the robustness of quantized neural networks. This study reveals that activation quantization can be used as a defense to weaken adversarial noise, but the robustness of quantized models is still limited by the amplification effect of network errors, including quantization errors and adversarial noise. To address this problem, we propose the quantization adversarial noise suppression (QANS) method that employs a Gaussian kernel regularization constraint to stabilize the model by restricting the perturbation error within two levels of tolerance. Extensive experiments are conducted with Wide ResNet and VGG-16 models on CIFAR-10 and street view house number datasets under different attack methods, including several white-box and black-box attacks. Experimental results show the proposed method achieves superior robustness to prior works. Chengmin Lin, Pengfei Yang 0001, Tianbing He, Jinpeng Liang, Quan Wang 0006 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2023 | Mixing Self-Attention and Convolution: A Unified Framework for Multisource Remote Sensing Data ClassificationabstractConvolution and self-attention are two powerful techniques for multi-source remote sensing (RS) data fusion that have been widely adopted in Earth observation tasks. However, Convolutional Neural Networks (CNNs) are inadequate for fully mining contextual information and representing the sequence attributes of spectral signatures. Additionally, the specific self-attention mechanism often comes with high computation costs, which hinders its application in the field of RS. To overcome the above limitations, this paper proposes a unified framework called “Mixing Self-Attention and Convolution Network" for comprehensive feature extraction and efficient feature fusion. First, the proposed MACN utilizes two adaptive CNN encoders (ACEs) to extract shallow convolutional features from multi-source RS data. Secondly, taking the complexity and varying scales of RS data into account, the proposed mixing self-attention and convolution transformer (MACT) layer achieves local and global multiscale perception through an elegant integration of self-attention and convolution. MACT can extract abundant spatial and high-dimensional information (e.g., spectral and elevation information) while maintaining minimal computational overhead compared to pure convolution or self-attention counterparts. Finally, a multi-source cross-guided fusion (MCGF) module is designed to achieve deep fusion of multi-source RS data features. MCGF utilizes a carefully designed cross-modal attention mechanism to capture the interaction between multi-source data and aggregate contextual information. Extensive tests on six public RS datasets have shown that our method outperforms other multi-source fusion models, delivering state-of-the-art results on multiple RS data fusion tasks without specific tuning. The source code of the proposed method will be available publicly at https://github.com/like413/MACN. Ke Li 0024, Di Wang 0011, Xu Wang 0057, Gang Liu 0006, Zili Wu, Quan Wang 0006 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Interpretable and Efficient Heterogeneous Graph Convolutional NetworkabstractGraph Convolutional Network (GCN) has achieved extraordinary success in learning representations of nodes in graphs. However, regarding Heterogeneous Information Network (HIN), existing HIN-oriented GCN methods still suffer from two deficiencies: (1) they cannot flexibly explore all possible meta-paths and extract the most useful ones for each target object, which hinders both effectiveness and interpretability; (2) before performing aggregation, they often require some additional time-consuming pre-processing operations, which increase the computational complexity. To address the above issues, we propose an interpretable and efficient Heterogeneous Graph Convolutional Network (ie-HGCN) to learn the representations of objects in HINs. It is designed as a hierarchical aggregation architecture, i.e., object-level aggregation and type-level aggregation. The new architecture can automatically evaluate all possible meta-paths within a length limit, and discover and exploit the most useful ones for each target object, i.e., at fine granularity. It also reduces the computational cost by avoiding additional time-consuming pre-processing operations. Theoretical analysis shows its ability to evaluate the usefulness of all possible meta-paths, its connection to the spectral graph convolution on HINs, and its quasi-linear time complexity. Extensive experiments on four real network datasets demonstrate its interpretability, efficiency as well as its superiority against thirteen baselines. Yaming Yang 0002, Ziyu Guan, Jianxin Li 0001, Wei Zhao 0019, Jiangtao Cui, Quan Wang 0006 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2023 | Cross-Modal Enhancement Network for Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis (MSA) plays an important role in many applications, such as intelligent question-answering, computer-assisted psychotherapy and video understanding, and has attracted considerable attention in recent years. It leverages multimodal signals including verbal language, facial gestures, and acoustic behaviors to identify sentiments in videos. Language modality typically outperforms nonverbal modalities in MSA. Therefore, strengthening the significance of language in MSA will be a vital way to promote recognition accuracy. Considering that the meaning of a sentence often varies in different nonverbal contexts, combining nonverbal information with text representations is conducive to understanding the exact emotion conveyed by an utterance. In this paper, we propose a Cross-modal Enhancement Network (CENet) model to enhance text representations by integrating visual and acoustic information into a language model. Specifically, it embeds a Cross-modal Enhancement (CE) module, which enhances each word representation according to long-range emotional cues implied in unaligned nonverbal data, into a transformer-based pre-trained language model. Moreover, a feature transformation strategy is introduced for acoustic and visual modalities to reduce the distribution differences between the initial representations of verbal and nonverbal modalities, thereby facilitating the fusion of distinct modalities. Extensive experiments on benchmark datasets demonstrate the significant gains of CENet over state-of-the-art methods. Di Wang 0011, Shuai Liu 0009, Quan Wang 0006, Yumin Tian, Lihuo He, Xinbo Gao 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | Hierarchical Semantic Structure Preserving Hashing for Cross-Modal RetrievalabstractCross-modal hashing has become a vital technique in cross-modal retrieval due to its fast query speed and low storage cost in recent years. Generally, most of the priors supervised cross-modal hashing methods are flat methods which are designed for non-hierarchical labeled data. They treat different categories independently and ignore the inter-category correlations. In practical applications, many instances are labeled with hierarchical categories. The hierarchical label structure provides rich information among different categories. To rationally take use of category correlations, hierarchical cross-modal hashing is proposed. However, existing methods intend to preserve instance-pairwise or class-pairwise similarities, which cannot fully explore the semantic correlations among different categories and make the learned hash codes less discriminative. In this paper, we propose a deep cross-modal hashing method named hierarchical semantic structure preserving hashing (HSSPH), which directly exploits the label hierarchy information to learn discriminative hash codes. Specifically, HSSPH learns a set of class-wise hash codes for each layer. By augmenting class-wise codes with labels, it generates layer-wise prototype codes which reflect the semantic structure of each layer. In order to enhance the discriminative ability of hash codes, HSSPH supervises the hash codes learning with both labels and semantic structures to preserve the hierarchical semantics. Besides, efficient optimization algorithms are developed to directly learn the discrete hash codes for each instance and each class. Extensive experiments on two benchmark datasets show the superiority of HSSPH over several state-of-the-art methods. Di Wang 0011, Caiping Zhang, Quan Wang 0006, Yumin Tian, Lihuo He, Lin Zhao 0003 |
IEEE Trans. Multim. | 3 |
| 2022 | TelecomNet: Tag-Based Weakly-Supervised Modally Cooperative Hashing Network for Image RetrievalabstractWe are concerned with using user-tagged images to learn proper hashing functions for image retrieval. The benefits are two-fold: (1) we could obtain abundant training data for deep hashing models; (2) tagging data possesses richer semantic information which could help better characterize similarity relationships between images. However, tagging data suffers from noises, vagueness and incompleteness. Different from previous unsupervised or supervised hashing learning, we propose a novel weakly-supervised deep hashing framework which consists of two stages: weakly-supervised pre-training and supervised fine-tuning. The second stage is as usual. In the first stage, we propose two formulations Tag-basEd weakLy-supErvised Modally COoperative hashing Network (TelecomNet) and Generalized TelecomNet (GTelecomNet). Rather than performing supervision on tags, TelecomNet first learns an observed semantic embedding vector for each image from attached tags and then uses it to guide hashing learning. GTelecomNet introduces a novel semantic network to exploit more precise semantic information. By carefully designing the optimization problem, they can well leverage tagging information and image content for hashing learning. The framework is general and does not depend on specific deep hashing methods. Empirical results on real world datasets show that they significantly increase the performance of state-of-the-art deep hashing methods. Wei Zhao 0019, Ziyu Guan, Xunlian Wu, Wanqing Zhao, Qiguang Miao, Xiaofei He 0001, Quan Wang 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2022 | Pseudo-Label Guided Collective Matrix Factorization for Multiview ClusteringabstractMultiview clustering has aroused increasing attention in recent years since real-world data are always comprised of multiple features or views. Despite the existing clustering methods having achieved promising performance, there still remain some challenges to be solved: 1) most existing methods are unscalable to large-scale datasets due to the high computational burden of eigendecomposition or graph construction and 2) most methods learn latent representations and cluster structures separately. Such a two-step learning scheme neglects the correlation between the two learning stages and may obtain a suboptimal clustering result. To address these challenges, a pseudo-label guided collective matrix factorization (PLCMF) method that jointly learns latent representations and cluster structures is proposed in this article. The proposed PLCMF first performs clustering on each view separately to obtain pseudo-labels that reflect the intraview similarities of each view. Then, it adds a pseudo-label constraint on collective matrix factorization to learn unified latent representations, which preserve the intraview and interview similarities simultaneously. Finally, it intuitively incorporates latent representation learning and cluster structure learning into a joint framework to directly obtain clustering results. Besides, the weight of each view is learned adaptively according to data distribution in the joint framework. In particular, the joint learning problem can be solved with an efficient iterative updating method with linear complexity. Extensive experiments on six benchmark datasets indicate the superiority of the proposed method over state-of-the-art multiview clustering methods in both clustering accuracy and computational efficiency. Di Wang 0011, Songwei Han, Quan Wang 0006, Lihuo He, Yumin Tian, Xinbo Gao 0001 |
IEEE Trans. Cybern. | 3 |
| 2022 | Microservice Deployment in Edge Computing Based on Deep Q LearningabstractThe microservice deployment strategy is promising in reducing the overall service response time in the microservice-oriented edge computing platform. However, existing works ignore the effect of different interaction frequencies among microservices and the decrease in service execution performance caused by the increased node loads. In this article, we first model the invocation relationships among microservices as an undirected and weighted interaction graph to characterize the communication overhead. Then, we propose a multi-objective microservice deployment problem (MMDP) in edge computing. MMDP aims to minimize the communication overhead while achieving load balance between edge nodes. Without the requirement for domain experts, we propose Reward Sharing Deep Q Learning (RSDQL), a learning-based algorithm, to solve MMDP and obtain the optimal deployment strategy. In addition, to improve the scalability of the services, we propose an Elastic Scaling algorithm (ES) based on heuristics to deal with the dynamic pressure of requests. Finally, we conduct a series of experiments in Kubernetes to evaluate the performance of our approach. Experimental results indicate that, compared with interaction-aware strategy and Kubernetes default strategy, RSDQL has shorter response times, more balanced resource loads, and makes services scale elastically according to the request pressure. Wenkai Lv, Quan Wang 0006, Pengfei Yang 0001, Yunqing Ding, Bijie Yi, Chengmin Lin |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2021 | Work in Progress: Power-aware Scheduling Strategy for Multiple DAGs in the Heterogeneous CloudabstractHigh energy consumption has become a major problem of cloud platform. Most of the current task scheduling methods neglect the heterogeneity of cloud platform, which may consume more power consumption of heterogeneous cloud platform. In this paper, we propose a power-aware scheduling strategy (PASS) for multiple DAGs workflow in the heterogeneous cloud with the goal of minimizing the energy consumption. First, we predict the PM energy consumption considering VM status after scheduling tasks, and then we formulate the power-aware DAGs task scheduling as a NP-hard problem, which tries to minimize the energy consumption of heterogeneous cloud platform. Second, we propose a multiple DAGs workflow scheduling algorithm to solve the formulated NP-hard problem. We consider the combination of the coarse-grained sorting for DAGs and the fine-grained sorting for sub-tasks to obtain the optimal sorting of DAG workflows. We then assign tasks to the appropriate computing nodes considering the heterogeneity of the cloud platform to minimize its energy consumption. Third, experiments are conducted to evaluate PASS, and the experimental results verify its efficiency. Hui Zhao 0003, Shangshu Li, Quan Wang 0006, Jing Wang 0028 |
RTAS | 3 |
| 2021 | Zero-watermarking method for resisting rotation attacks in 3D models
Gang Liu 0006, Quan Wang 0006, Lianqin Wu, Rong Pan 0004, Bo Wan 0002, Yumin Tian |
Neurocomputing | 2 |
| 2021 | A binary harmony search algorithm as channel selection method for motor imagery-based BCIabstractChannel selection is a key topic in brain-computer interface (BCI). Task-irrelevant and redundant channels used in BCI may lead to low classification accuracy, high computational complexity, and inconvenience for application. By selecting optimal channels, the performance of BCI could enhance significantly. In this paper, a new binary harmony search (BHS) is proposed to select the optimal channel sets and optimize the system accuracy. The BHS is implemented on the training data sets to select the optimal channels and the test data sets are used to evaluate the classification performance on the selected channels. The sparse representation-based classification, linear discriminant analysis, and support vector machine are performed on the common spatial pattern (CSP) features for motor imagery (MI) classification. Two public EEG datasets are employed to validate the proposed BHS method. The paired t-test is conducted on the test classification performance between the BHS and traditional CSP with all channels. The results reveal that the proposed BHS method significantly improved classification accuracy as compared to the conventional CSP method (p < 0.05). This study proposed the BHS method to select the optimal channels in MI -based BCI. On the one hand, the results confirm the validity of the BHS algorithm as a channel selection method for motor imagery data. On the other hand, the BHS method with costing shorter computation time relatively yields a better average test accuracy than the steady-state genetic algorithms. The proposed method could significantly improve the practicability and convenience of the BCI system. Quan Wang 0006, Zan Yue, Yaping Huai, Jing Wang 0028 |
Neurocomputing | 2 |
| 2021 | Supervoxel-Based Region Growing Segmentation for Point Cloud DataabstractPoint cloud segmentation is a crucial fundamental step in 3D reconstruction, object recognition and scene understanding. This paper proposes a supervoxel-based point cloud segmentation algorithm in region growing principle to solve the issues of inaccurate boundaries and nonsmooth segments in the existing methods. To begin with, the input point cloud is voxelized and then pre-segmented into sparse supervoxels by flow constrained clustering, considering the spatial distance and local geometry between voxels. Afterwards, plane fitting is applied to the over-segmented supervoxels and seeds for region growing are selected with respect to the fitting residuals. Starting from pruned seed patches, adjacent supervoxels are merged in region growing style to form the final segments, according to the normalized similarity measure that integrates the smoothness and shape constraints of supervoxels. We determine the values of parameters via experimental tests, and the final results show that, by voxelizing and pre-segmenting the point clouds, the proposed algorithm is robust to noises and can obtain smooth segmentation regions with accurate boundaries in high efficiency. Nan Luo, Quan Wang 0006 |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2021 | Retrieving point cloud models of target objects in a scene from photographed images
Nan Luo, Quan Wang 0006, Bo Wan 0002 |
Multim. Tools Appl. | 3 |
| 2021 | Regulating synchronous oscillations of cerebellar granule cells by different types of inhibitionabstractSynchronous oscillations in neural populations are considered being controlled by inhibitory neurons. In the granular layer of the cerebellum, two major types of cells are excitatory granular cells (GCs) and inhibitory Golgi cells (GoCs). GC spatiotemporal dynamics, as the output of the granular layer, is highly regulated by GoCs. However, there are various types of inhibition implemented by GoCs. With inputs from mossy fibers, GCs and GoCs are reciprocally connected to exhibit different network motifs of synaptic connections. From the view of GCs, feedforward inhibition is expressed as the direct input from GoCs excited by mossy fibers, whereas feedback inhibition is from GoCs via GCs themselves. In addition, there are abundant gap junctions between GoCs showing another form of inhibition. It remains unclear how these diverse copies of inhibition regulate neural population oscillation changes. Leveraging a computational model of the granular layer network, we addressed this question to examine the emergence and modulation of network oscillation using different types of inhibition. We show that at the network level, feedback inhibition is crucial to generate neural oscillation. When short-term plasticity was equipped on GoC-GC synapses, oscillations were largely diminished. Robust oscillations can only appear with additional gap junctions. Moreover, there was a substantial level of cross-frequency coupling in oscillation dynamics. Such a coupling was adjusted and strengthened by GoCs through feedback inhibition. Taken together, our results suggest that the cooperation of distinct types of GoC inhibition plays an essential role in regulating synchronous oscillations of the GC population. With GCs as the sole output of the granular network, their oscillation dynamics could potentially enhance the computational capability of downstream neurons. Yuanhong Tang, Lingling An, Quan Wang 0006, Jian K. Liu |
PLoS Comput. Biol. | 3 |
| 2021 | Modulation of the dynamics of cerebellar Purkinje cells through the interaction of excitatory and inhibitory feedforward pathwaysabstractThe dynamics of cerebellar neuronal networks is controlled by the underlying building blocks of neurons and synapses between them. For which, the computation of Purkinje cells (PCs), the only output cells of the cerebellar cortex, is implemented through various types of neural pathways interactively routing excitation and inhibition converged to PCs. Such tuning of excitation and inhibition, coming from the gating of specific pathways as well as short-term plasticity (STP) of the synapses, plays a dominant role in controlling the PC dynamics in terms of firing rate and spike timing. PCs receive cascade feedforward inputs from two major neural pathways: the first one is the feedforward excitatory pathway from granule cells (GCs) to PCs; the second one is the feedforward inhibition pathway from GCs, via molecular layer interneurons (MLIs), to PCs. The GC-PC pathway, together with short-term dynamics of excitatory synapses, has been a focus over past decades, whereas recent experimental evidence shows that MLIs also greatly contribute to controlling PC activity. Therefore, it is expected that the diversity of excitation gated by STP of GC-PC synapses, modulated by strong inhibition from MLI-PC synapses, can promote the computation performed by PCs. However, it remains unclear how these two neural pathways are interacted to modulate PC dynamics. Here using a computational model of PC network installed with these two neural pathways, we addressed this question to investigate the change of PC firing dynamics at the level of single cell and network. We show that the nonlinear characteristics of excitatory STP dynamics can significantly modulate PC spiking dynamics mediated by inhibition. The changes in PC firing rate, firing phase, and temporal spike pattern, are strongly modulated by these two factors in different ways. MLIs mainly contribute to variable delays in the postsynaptic action potentials of PCs while modulated by excitation STP. Notably, the diversity of synchronization and pause response in the PC network is governed not only by the balance of excitation and inhibition, but also by the synaptic STP, depending on input burst patterns. Especially, the pause response shown in the PC network can only emerge with the interaction of both pathways. Together with other recent findings, our results show that the interaction of feedforward pathways of excitation and inhibition, incorporated with synaptic short-term dynamics, can dramatically regulate the PC activities that consequently change the network dynamics of the cerebellar circuit. Yuanhong Tang, Lingling An, Qingqi Pei, Quan Wang 0006, Jian K. Liu |
PLoS Comput. Biol. | 5 |
| 2021 | kNN-based feature learning network for semantic segmentation of point cloud data
Nan Luo, Yifeng Wang 0004, Yumin Tian, Quan Wang 0006, Chuan Jing |
Pattern Recognit. Lett. | 5 |
| 2021 | ABSAC: Attribute-Based Access Control Model Supporting Anonymous Access for Smart CitiesabstractSmart cities require new access control models for Internet of Things (IoT) devices that preserve user privacy while guaranteeing scalability and efficiency. Researchers believe that anonymous access can protect the private information even if the private information is not stored in authorization organization. Many attribute-based access control (ABAC) models that support anonymous access expose the attributes of the subject to the authorization organization during the authorization process, which allows the authorization organization to obtain the attributes of the subject and infer the identity of the subject. The ABAC with anonymous access proposed in this paper called ABSAC strengthens the identity-less of ABAC by combining homomorphic attribute-based signatures (HABSs) which does not send the subject attributes to the authorization organization, reducing the risk of subject identity re-identification. It is a secure anonymous access framework. Tests show that the performance of ABSAC implementation is similar to ABAC’s performance. Runnan Zhang, Gang Liu 0006, Shancang Li, Yongheng Wei, Quan Wang 0006 |
Secur. Commun. Networks | 5 |
| 2021 | Popularity-Based and Version-Aware Caching Scheme at Edge Servers for Multi-Version VoD SystemsabstractRecently, many video-on-demand (VoD) providers have begun storing multiple versions of the same video to offer multiple-quality video services with different bitrates to users, called multi-version VoD. To improve users' quality of experience (QoE), it is a good idea to cache videos at edge servers in multi-version VoD systems. However, determining which versions of which videos should be cached or replaced in an edge server is still a major challenge for a multi-version VoD system because of its limited cache storage. In this paper, we propose a popularity-based and version-aware caching scheme (PVCS) at edge servers for multi-version VoD systems. First, based on video popularity, we formulate cache placement as a knapsack problem under constraints such as the cache storage and transcoding computation of the edge server, which aims to maximize the cache hit ratio. Second, we use the transcoding relations among versions to calculate a version-aware caching profit when caching a certain version or multiple versions of a video. The version-aware caching profit is the basis for the subsequent cache replacement algorithm. Third, we propose two algorithms, the video cache placement (VCP) algorithm and the video cache replacement (VCrP) algorithm, to solve the cache placement and replacement problems respectively. VCP utilizes the Lagrangian relaxation algorithm to decide which video files should be cached initially, and VCrP decides which video files cached at the edge server will be replaced dynamically based on the version-aware profit. In this way, the PVCS can improve the cache hit ratio and decrease the average start-up delay. Our simulation results have shown that the PVCS outperforms the other schemes in terms of the cache hit ratio and the average start-up delay. Hui Zhao 0003, Quan Wang 0006, Jing Wang 0028, Bo Wan 0002, Zili Wu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Improved Bell-LaPadula Model With Break the Glass MechanismabstractThe Bell-LaPadula (BLP) model is a widely used access control model for the multilevel security system. The researchers proposed many modified BLP models to express privileges that cannot be expressed by the BLP model. However, these models are not compatible with the BLP model, leading to the transportation cost-prohibitive and difficult to be practically applied. In this article, an improved BLP model incorporated the break the glass (BTG) mechanism is proposed to overcome the limitations of the standard BLP and other modified BLP models. The improved model inherits some of the advantages of BTG, such as policy dynamic modification and fine-grained access control, which gives it wide availability. Additionally, in the implementation, BTG is used as an independent function attached to the original BLP; the proposed BLP model can be easily implemented in systems where BLP models have been implemented. The results of the analysis and simulations showed that the proposed BLP model improves the ability of expressing policy of BLP and achieves fine-grained access control without compromise in security. Compared with other modified BLP models, the proposed BLP model could express policy more effectively and is compatible with the original BLP model. Runnan Zhang, Gang Liu 0006, Hongzhaoning Kang, Quan Wang 0006, Yumin Tian |
IEEE Trans. Reliab. | 4 |
| 2020 | VM Performance Maximization and PM Load Balancing Virtual Machine Placement in CloudabstractVirtual machine placement (VMP) technology is widely used in cloud computing systems. The existing VMP methods mainly aimed at improving the cloud resource utilization, such as load balancing among physical machines (PMs), but they may result in virtual machine (VM) performance degradation because of the great resource contention among VMs running on top of the same PM. In contract to existing VMP algorithms, this paper proposes a virtual machine (VM) Performance maximization and physical machine (PM) Load balancing Virtual Machine Placement method (PLVMP) in cloud, which tries to maximize VM performance and balance PM workload from both users' and cloud providers' perspectives. First, we study the relationship between PM workload and VM performance to train a new and improved VM performance model, which can predict VM performance more accurately and offer help to the following VMP. Second, we take VM performance maximization and PM workload balancing into account to formulate the VMP as an optimization problem, which tries to maximize VM performance for users and make load balancing among PMs for cloud providers. Third, we propose a greedy-based algorithm to solve the VMP problem efficiently. We then evaluate PLVMP with other VMP methods on CloudSim platform and a real OpenStack platform. The results show PLVMP can maximize the VM performance significantly and make a good load balancing among PMs. Hui Zhao 0003, Quan Wang 0006, Jing Wang 0028, Bo Wan 0002, Shangshu Li |
CCGRID | 2 |
| 2020 | Online Collective Matrix Factorization Hashing for Large-Scale Cross-Media RetrievalabstractCross-modal hashing has been widely investigated recently for its efficiency in large-scale cross-media retrieval. However, most existing cross-modal hashing methods learn hash functions in a batch-based learning mode. Such mode is not suitable for large-scale data sets due to the large memory consumption and loses its efficiency when training streaming data. Online cross-modal hashing can deal with the above problems by learning hash model in an online learning process. However, existing online cross-modal hashing methods cannot update hash codes of old data by the newly learned model. In this paper, we propose Online Collective Matrix Factorization Hashing (OCMFH) based on collective matrix factorization hashing (CMFH), which can adaptively update hash codes of old data according to dynamic changes of hash model without accessing to old data. Specifically, it learns discriminative hash codes for streaming data by collective matrix factorization in an online optimization scheme. Unlike conventional CMFH which needs to load the entire data points into memory, the proposed OCMFH retrains hash functions only by newly arriving data points. Meanwhile, it generates hash codes of new data and updates hash codes of old data by the latest updated hash model. In such way, hash codes of new data and old data are well-matched. Furthermore, a zero mean strategy is developed to solve the mean-varying problem in the online hash learning process. Extensive experiments on three benchmark data sets demonstrate the effectiveness and efficiency of OCMFH on online cross-media retrieval. Di Wang 0011, Quan Wang 0006, Yaqiang An, Xinbo Gao 0001, Yumin Tian |
SIGIR | 2 |
| 2020 | Policy Evaluation and Dynamic Management Based on Matching Tree for XACMLabstractAs a widely recognized policy language of access control, the eXtensible Access Control Markup Language (XACML) is widely used with its fine-grained and easy-to-read. With the application of XACML, researchers find that the XACML based policy evaluation and policy management methods can no longer meet the current large-scale requests for efficient access and dynamic management requirements. To improve the performance of policy evaluation based on XACML, we propose a policy evaluation method based on the matching tree to search policy efficiently and avoid the extra consumption of invalid policy participation. Furthermore, we propose a policy dynamic management method based on the matching tree to reduce the scale of the policy to be disabled for management, by adding locks in the tree node and the information mapping table. Through theoretical derivation and the factors that may affect its evaluation performance, we verify the improvement of evaluation efficiency. The simulation also shows the improvement of the evaluation engine based on the matching tree compared with OuenAz. Hongzhaoning Kang, Gang Liu 0006, Quan Wang 0006, Runnan Zhang, Zichao Zhong, Yumin Tian |
TrustCom | 3 |
| 2020 | Joint and individual matrix factorization hashing for large-scale cross-modal retrieval
Di Wang 0011, Quan Wang 0006, Lihuo He, Xinbo Gao 0001, Yumin Tian |
Pattern Recognit. | 2 |
| 2020 | Accurate and Reliable Facial Expression Recognition Using Advanced Softmax Loss With Fixed WeightsabstractAn important challenge for facial expression recognition (FER) is that real-world training data are usually imbalanced. Although many deep learning approaches have been proposed to enhance the discriminative power of deep expression features and enable a good predictive effect, few works have focused on the multiclass imbalance problem. When supervised by softmax loss (SL), which is widely used in FER, the classifier is often biased against minority categories (i.e., smaller interclass angular distances). In this letter, we present advanced softmax loss (ASL) to mitigate the bias induced by data imbalance and hence increase accuracy and reliability. The proposed ASL essentially magnifies the interclass diversity in the angular space to enhance discriminative power in every category. The proposed loss can easily be implemented in any deep network. Extensive experiments on the FER2013 and real-world affective faces (RAF) databases demonstrate that ASL is significantly more accurate and reliable than many state-of-the-art approaches and that it can easily be plugged into other methods and improves their performance. Ping Jiang 0004, Gang Liu 0006, Quan Wang 0006, Jiang Wu 0004 |
IEEE Signal Process. Lett. | 3 |
| 2020 | Fast and Efficient Facial Expression Recognition Using a Gabor Convolutional NetworkabstractAutomatic facial expression recognition (FER) is a fundamental topic in computer vision. Many studies have indicated that facial emotion changes are strongly related to certain regions of interest (ROIs), such as the mouth, eyes, eyebrows, and nose; therefore, the features of these facial ROIs are very important for identifying expressions. Since Gabor filters are very efficient in extracting visual content, Gabor orientation filters (GoFs) modulated by Gabor kernels and traditional convolutional filters can capture such ROI information better than conventional convolutional filters. Consequently, this letter presents a light Gabor convolutional network (GCN) consisting of only four Gabor convolutional layers and two linear layers for FER tasks. Extensive experiments on the FER2013, FERPlus and Real-world Affective Faces (RAF) databases demonstrate that the proposed method achieves good recognition accuracy and requires very low computational costs. The source code can be found at https://github.com/general515/Facial_Expression_Recognition_Using _GCN. Ping Jiang 0004, Bo Wan 0002, Quan Wang 0006, Jian Wu 0001 |
IEEE Signal Process. Lett. | 3 |
| 2020 | Efficient Noise-Level Estimation Based on Principal Image TextureabstractBlind noise-level estimation (NLE) is a fundamental issue in digital image processing. This paper provides a noise-level estimator for additive white Gaussian noise and multiplicative Gaussian noise using principal texture patches (PTPs). The proposed algorithm first identifies the principal texture of the noisy image by using the principal component analysis, and then, it chooses PTPs to calculate the noise level. The major contributions of this paper toward addressing the challenges in the NLE literature include: 1) analyzing the nonlinear relationship between the smallest eigenvalue, true noise level, number of chosen patches, and block size; 2) extracting PTPs; and 3) computing noise level using PTPs. The abundant experimental results show that the proposed method works well for various natural images over a large range of noise levels and it performs well for multiplicative noise. Compared with some state-of-the-art approaches, our method has the best performance and a faster running speed. Ping Jiang 0004, Quan Wang 0006, Jiang Wu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | A Survey on Energy-Efficient Strategies in Static Wireless Sensor NetworksabstractA comprehensive analysis on the energy-efficient strategy in static Wireless Sensor Networks (WSNs) that are not equipped with any energy harvesting modules is conducted in this article. First, a novel generic mathematical definition of Energy Efficiency (EE) is proposed, which takes the acquisition rate of valid data, the total energy consumption, and the network lifetime of WSNs into consideration simultaneously. To the best of our knowledge, this is the first time that the EE of WSNs is mathematically defined. The energy consumption characteristics of each individual sensor node and the whole network are expounded at length. Accordingly, the concepts concerning EE, namely the Energy-Efficient Means, the Energy-Efficient Tier, and the Energy-Efficient Perspective, are proposed. Subsequently, the relevant energy-efficient strategies proposed from 2002 to 2019 are tracked and reviewed. Specifically, they respectively are classified into five categories: the Energy-Efficient Media Access Control protocol, the Mobile Node Assistance Scheme, the Energy-Efficient Clustering Scheme, the Energy-Efficient Routing Scheme, and the Compressive Sensing--based Scheme. A detailed elaboration on both of the basic principle and the evolution of them is made. Finally, further analysis on the categories is made and the related conclusion is drawn. To be specific, the interdependence among them, the relationships between each of them, and the Energy-Efficient Means, the Energy-Efficient Tier, and the Energy-Efficient Perspective are analyzed in detail. In addition, the specific applicable scenarios for each of them and the relevant statistical analysis are detailed. The proportion and the number of citations for each category are illustrated by the statistical chart. In addition, the existing opportunities and challenges facing WSNs in the context of the new computing paradigm and the feasible direction concerning EE in the future are pointed out. Deyu Lin, Quan Wang 0006, Weidong Min, Zhiqiang Zhang 0001 |
ACM Trans. Sens. Networks | 2 |
| 2020 | A PUF-based unified identity verification framework for secure IoT hardware via device authentication
Quan Wang 0006 |
World Wide Web | 2 |
| 2019 | A Multidimensional Reputation Evaluation Model for Mobile Crowd SensingabstractThe participant's reputation is vital to improve the quality of service for Mobile Crowd Sensing (MCS). A multidimensional reputation evaluation model was proposed in this paper to evaluate the participant's reputation more objectively. Different from the existing strategies, the service delay and the count of the successful as well as the failed transactions were additionally utilized to evaluate the participant's reputation. An algorithm based on Analytic Hierarchical Process (AHP) was presented to establish the reputation evaluation weight matrix. Besides, a fuzzy logic based mechanism was proposed to normalize the value of the four criteria and a dual-threshold mechanism was designed to achieve admission control more properly. Finally, extensive simulations were conducted and the simulation results confirmed the effectiveness of the reputation evaluation model. Deyu Lin, Quan Wang 0006, Pengfei Yang 0001, Zhiqiang Zhang 0001 |
IWCMC | 2 |
| 2019 | Work-in-Progress: Version-Aware Video Caching Strategy for Multi-version VoD SystemsabstractRecently, many video-on-demand (VoD) providers store multiple versions of the same videos to offer multiple-quality video services with different bitrates to users, called as multi-version VoD. To decrease the start-up delay for users, it is a good idea to cache videos at caching server that is in close proximity. However, how to decide which versions of which videos should be cached and replaced in caching server is still one major challenge for multi-version VoD systems because of limited caching storage. In this paper, we propose a version-aware video caching strategy for multi-version VoD systems, which aims to reduce start-up delay and improve cache hit ratio. First, we take into account the transcoding delay among versions and transmit delay from content server to caching server to calculate version-aware caching profit when caching a certain version or multiple versions of a video. It is the basis for the following caching replacement algorithm. Second, we propose version-aware video caching (VaVC) algorithm to decide which versions of which videos will be replaced based on the version-aware caching profit dynamically. In this way, VaVC can reduce start-up delay and improve the cache hit ratio. Our simulation results have shown that VaVC outperforms the others in both the start-up delay and the cache hit ratio. Hui Zhao 0003, Zili Wu, Quan Wang 0006, Jing Wang 0028, Weizhan Zhang |
RTSS | 3 |
| 2019 | Partially shared cache and adaptive replacement algorithm for NoC-based many-core systems
Pengfei Yang 0001, Quan Wang 0006, Hongwei Ye, Zhiqiang Zhang 0001 |
J. Syst. Archit. | 2 |
| 2019 | Semi-paired and semi-supervised multimodal hashing via cross-modality label propagation
Di Wang 0011, Bin Shang, Quan Wang 0006, Bo Wan 0002 |
Multim. Tools Appl. | 3 |
| 2019 | Queue-based and learning-based dynamic resources allocation for virtual streaming media server cluster of multi-version VoD system
Hui Zhao 0003, Jing Wang 0028, Quan Wang 0006 |
Multim. Tools Appl. | 3 |
| 2019 | Community Detection in Multi-Layer Networks Using Joint Nonnegative Matrix FactorizationabstractMany complex systems are composed of coupled networks through different layers, where each layer represents one of many possible types of interactions. A fundamental question is how to extract communities in multi-layer networks. The current algorithms either collapses multi-layer networks into a single-layer network or extends the algorithms for single-layer networks by using consensus clustering. However, these approaches have been criticized for ignoring the connection among various layers, thereby resulting in low accuracy. To attack this problem, a quantitative function (multi-layer modularity density) is proposed for community detection in multi-layer networks. Afterward, we prove that the trace optimization of multi-layer modularity density is equivalent to the objective functions of algorithms, such as kernel$K$-means, nonnegative matrix factorization (NMF), spectral clustering and multi-view clustering, for multi-layer networks, which serves as the theoretical foundation for designing algorithms for community detection. Furthermore, aSemi-SupervisedjointNonnegativeMatrixFactorization algorithm (S2-jNMF) is developed by simultaneously factorizing matrices that are associated with multi-layer networks. Unlike the traditional semi-supervised algorithms, the partial supervision is integrated into the objective of the S2-jNMF algorithm. Finally, through extensive experiments on both artificial and real world networks, we demonstrate that the proposed method outperforms the state-of-the-art approaches for community detection in multi-layer networks. Xiaoke Ma 0001, Di Dong, Quan Wang 0006 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2018 | An confidentiality and integrity scheme for the distributed shared memory of embedded multi-core system: work in progress
Pengfei Yang 0001, Quan Wang 0006, Xiaokun Huang, Xin Mi |
CASES | 2 |
| 2018 | Resource Allocation for Virtual Streaming Media Server Cluster in Cloud-based Multi-version VoDabstractWith the rapid development of mobile Internet and smart devices, VoD (video on demand) providers build media cloud to offer multi-bitrate video streaming services to users at a reduced cost, called as cloud-based multi-version VoD. In cloud-based multi-version VoD, we need to solve the problem of allocating appropriate resources for virtual streaming media server cluster with the aim of optimizing the user experience and reducing the service cost. To address this problem, a resource allocation for virtual streaming media server cluster in cloud-based multi-version VoD is proposed in this paper. We firstly analyze the user historical learning logs to mine the user behavior characteristics, including the average user request arrival rate, the video playing time distribution, and the video popularity distribution, etc. Then, based on the user behavior characteristics and the queueing theory, a resource allocation model for the virtual streaming media server cluster is introduced. It predicts the user arrival rate at first and then allocates appropriate resources dynamically to solve the resources allocation irrationality problem. Simulation results have proved the proposed method can allocate appropriate resources for virtual streaming media server cluster, which can ensure the user experience satisfaction and improve the resources utilization. Hui Zhao 0003, Jing Wang 0028, Quan Wang 0006, Nan Luo, Weizhan Zhang |
CSCWD | 4 |
| 2018 | Deep Multi-View Concept LearningabstractMulti-view data is common in real-world datasets, where different views describe distinct perspectives. To better summarize the consistent and complementary information in multi-view data, researchers have proposed various multi-view representation learning algorithms, typically based on factorization models. However, most previous methods were focused on shallow factorization models which cannot capture the complex hierarchical information. Although a deep multi-view factorization model has been proposed recently, it fails to explicitly discern consistent and complementary information in multi-view data and does not consider conceptual labels. In this work we present a semi-supervised deep multi-view factorization method, named Deep Multi-view Concept Learning (DMCL). DMCL performs nonnegative factorization of the data hierarchically, and tries to capture semantic structures and explicitly model consistent and complementary information in multi-view data at the highest abstraction level. We develop a block coordinate descent algorithm for DMCL. Experiments conducted on image and document datasets show that DMCL performs well and outperforms baseline methods. Ziyu Guan, Wei Zhao 0019, Yunfei Niu, Quan Wang 0006, Zhiheng Wang 0001 |
IJCAI | 5 |
| 2018 | Effective outlier matches pruning algorithm for rigid pairwise point cloud registration using distance disparity matrixabstractThis study focuses on fast and robust outlier matches removal strategy to improve the efficiency and precision of initial alignment and further the quality of pairwise registration. Starts from the point matches obtained via feature detecting and matching, the distance disparity matrix derived from Euclidean invariants of rigid transformation is introduced, based on which a fast and effective pruning method is proposed to eliminate the outlier correspondences, especially the sharp ones. Then, the remaining matches are sent into the enhanced least‐square backward method to estimate an initial transformation in lesser attempts. Since most of the outliers are rejected, presented backward method could provide a finer alignment to input point clouds in higher efficiency than existing methods, and the following refining procedure converges to a more precise registration consuming fewer iterations, which have been proved in designed experiments. The thresholds employed in the pipeline are all automatically determined according to the actual resolution of input point clouds. Users are just required to control the error precision through a scale factor, in which way the inaccuracy and inconvenience of manually threshold defining are avoided. Nan Luo, Quan Wang 0006 |
IET Comput. Vis. | 2 |
| 2018 | Robust and Flexible Discrete Hashing for Cross-Modal Similarity SearchabstractMultimodal hashing approaches have gained great success on large-scale cross-modal similarity search applications, due to their appealing computation and storage efficiency. However, it is still a challenge work to design binary codes to represent the original features with good performance in an unsupervised manner. We argue that there are some limitations that need to be further considered for unsupervised multimodal hashing: 1) most existing methods drop the discrete constraints to simplify the optimization, which will cause large quantization error; 2) many methods are sensitive to outliers and noises since they use ℓ2-norm in their objective functions which can amplify the errors; and 3) the weight of each modality, which greatly influences the retrieval performance, is manually or empirically determined and may not fully fit the specific training set. The above limitations may significantly degrade the retrieval accuracy of unsupervised multimodal hashing methods. To address these problems, in this paper, a novel hashing model is proposed to efficiently learn robust discrete binary codes, which is referred as Robust and Flexible Discrete Hashing (RFDH). In the proposed RFDH model, binary codes are directly learned based on discrete matrix decomposition, so that the large quantization error caused by relaxation is avoided. Moreover, the ℓ2,1-norm is used in the objective function to improve the robustness, such that the learned model is not sensitive to data outliers and noises. In addition, the weight of each modality is adaptively adjusted according to training data. Hence the important modality will get large weights during the hash learning procedure. Owing to above merits of RFDH, it can generate more effective hash codes. Besides, we introduce two kinds of hash function learning methods to project unseen instances into hash codes. Extensive experiments on several well-known large databases demonstrate superior performance of the proposed hash model over most state-of-the-art unsupervised multimodal hashing methods. Di Wang 0011, Quan Wang 0006, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Weakly-Supervised Deep Embedding for Product Review Sentiment AnalysisabstractProduct reviews are valuable for upcoming buyers in helping them make decisions. To this end, different opinion mining techniques have been proposed, where judging a review sentence's orientation (e.g., positive or negative) is one of their key challenges. Recently, deep learning has emerged as an effective means for solving sentiment classification problems. A neural network intrinsically learns a useful representation automatically without human efforts. However, the success of deep learning highly relies on the availability of large-scale training data. We propose a novel deep learning framework for product review sentiment classification which employs prevalently available ratings as weak supervision signals. The framework consists of two steps: (1) learning a high level representation (an embedding space) which captures the general sentiment distribution of sentences through rating information; and (2) adding a classification layer on top of the embedding layer and use labeled sentences for supervised fine-tuning. We explore two kinds of low level network structure for modeling review sentences, namely, convolutional feature extractors and long short-term memory. To evaluate the proposed framework, we construct a dataset containing 1.1M weakly labeled review sentences and 11,754 labeled review sentences from Amazon. Experimental results show the efficacy of the proposed framework and its superiority over baselines. Wei Zhao 0019, Ziyu Guan, Long Chen 0007, Xiaofei He 0001, Deng Cai 0001, Beidou Wang, Quan Wang 0006 |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2018 | Learning to Map Social Network Users by Unified Manifold Alignment on HypergraphabstractNowadays, a lot of people possess accounts on multiple online social networks, e.g., Facebook and Twitter. These networks are overlapped, but the correspondences between their users are not explicitly given. Mapping common users across these social networks will be beneficial for applications such as cross-network recommendation. In recent years, a lot of mapping algorithms have been proposed which exploited social and/or profile relations between users from different networks. However, there is still a lack of unified mapping framework which can well exploit high-order relational information in both social structures and profiles. In this paper, we propose a unified hypergraph learning framework named unified manifold alignment on hypergraph (UMAH) for this task. UMAH models social structures and user profile relations in a unified hypergraph where the relative weights of profile hyperedges are determined automatically. Given a set of training user correspondences, a common subspace is learned by preserving the hypergraph structure as well as the correspondence relations of labeled users. UMAH intrinsically performs semisupervised manifold alignment with profile information for calibration. For a target user in one network, UMAH ranks all the users in the other network by their probabilities of being the corresponding user (measured by similarity in the subspace). In experiments, we evaluate UMAH on three real world data sets and compare it to state-of-art baseline methods. Experimental results have demonstrated the effectiveness of UMAH in mapping users across networks. Wei Zhao 0019, Shulong Tan, Ziyu Guan, Boxuan Zhang 0002, Maoguo Gong, Zhengwen Cao, Quan Wang 0006 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2018 | Power-Aware and Performance-Guaranteed Virtual Machine Placement in the CloudabstractCloud service providers offer virtual machines (VMs) as services to users over Internet. As VMs are running on physical machines (PMs), PM power consumption needs to be considered. Meanwhile, VMs running on the same PM share physical resources, and there exists great resource contention, which results in VM performance degradation. Therefore, how to place VMs to reduce PM power consumption and guarantee VM performance is still one major challenge. However, existing VMPs did not study VM performance degradation, so they could not guarantee VM performance. To solve the high power consumption and VMs performance degradation problems, this paper explores the balance between saving PM power and guaranteeing VM performance, and proposes a power-aware and performance-guaranteed VMP (PPVMP). First, we investigate the relationship between power consumption and CPU utilization to build a non-linear power model, which is helpful for the following VMP. Second, we construct VM performance models to present the VM performance degradation trend. Third, based on these models, we formulate VMP as a bi-objective optimization problem, which tries to minimize PM power consumption and guarantee VM performance. We then propose an algorithm based on ant colony optimization to solve it. Finally, the results show the efficiency of our algorithm. Hui Zhao 0003, Jing Wang 0028, Quan Wang 0006, Weizhan Zhang |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2017 | A game theory based energy efficient clustering routing protocol for WSNs
Deyu Lin, Quan Wang 0006 |
Wirel. Networks | 2 |
| 2016 | Parallel design and implementation of Error Diffusion Algorithm and IP core for FPGA
Pengfei Yang 0001, Quan Wang 0006 |
Multim. Tools Appl. | 2 |