Kejie Wang

dblp:92/129 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ViMAEdit: Vision-Guided and Mask-Enhanced Adaptive Editing Algorithm for Prompt-Based Image Editing
Kejie Wang, Xuemeng Song, Meng Liu 0006, Jin Yuan 0002, Weili Guan
IEEE Trans. Circuits Syst. Video Technol.1
2026 Dual Alignment-Enhanced Fashion Vision-Language Pre-Training
abstract
Fashion vision-language pre-training (VLP) models have demonstrated remarkable capabilities in excelling at a wide range of fashion cross-modal tasks. However, current models still face three notable limitations (1) an inability to discern varying levels of consistency between textual descriptions and multi-view images, (2) a deficiency in explicit fine-grained alignment between images and text, and (3) a lack of specific supervision mechanisms for facilitating global joint embedding learning. To address these limitations, we propose a novel dual alignment-enhanced fashion VLP model. This model delves deeply into the rich resources of multi-view images and semantic attributes associated with each fashion item. Notably, we introduce two novel pre-training tasks: Multi-grained Adaptive Image-Text Alignment (MAITA) and Joint Embedding-oriented Alignment (JEA). MAITA focuses on optimizing the text/image encoder by orchestrating adaptive alignment between multi-view images and input text. This encompasses both coarse-grained and fine-grained alignment strategies to enrich semantic understanding, while JEA is devised to supervise the fine-grained semantic learning process of the global multimodal joint embedding. Experimental results spanning four diverse downstream tasks, including cross-modal retrieval, text-guided image retrieval, category recognition, and subcategory recognition, substantiate the significant performance superiority of our model over prior state-of-the-art fashion VLP models.
Weili Guan, Kejie Wang, Xuemeng Song, Kaihao Zhang, Xiaojun Chang, Shengping Zhang
ACM Trans. Multim. Comput. Commun. Appl.2
2026 TQ-SPUF: A Software PUF Design Based on HEVC Transform and Quantization Module for Device Security and Video Anti-Tampering
abstract
Addressing security threats faced by high efficiency video coding (HEVC) video devices in open deployment environments, this article proposes a HEVC Transform and Quantization (TQ)-based Software Physical Unclonable Function (TQ-SPUF). By overclocking the HEVC TQ module, this function induces clock violations on the critical path, thereby generating a stable and unique device fingerprint. Meanwhile, a two-stage postprocessing scheme combining majority voting and butterfly-xoroperations is employed to generate stable and uniform device fingerprints. Based on the proposed TQ-SPUF, a lightweight challenge-response authentication protocol was designed to achieve device identity binding. Furthermore, the PUF-gated session key is utilized to drive a dual-perturbation selective encryption scheme, which introduces both fixed- and random-position perturbations into the I-frame network abstraction layer units. In this way, semantic information exploitable was effectively disrupted, thereby preventing content recovery or tampering. Experimental results show that the proposed TQ-SPUF achieves 98.84% randomness and 49.15% uniqueness, passing part of the NIST test. Moreover, the encrypted videos exhibit an average peak signal-to-noise ratio (PSNR) of 10.4313dB and an average structural similarity index measure (SSIM) of 0.2935. This indicates that the encrypted content is visually incomprehensible, thereby achieving video anti-tampering protection.
Kejie Wang, Yuejun Zhang, Shuang Hu 0001, Ziyu Zhou 0001, Zhenkai Zhou, Huihong Zhang, Pengjun Wang
IEEE Trans. Very Large Scale Integr. Syst.1
2024 SFT-Net: A Network for Detecting Fatigue From EEG Signals by Combining 4D Feature Flow and Attention Mechanism
abstract
Fatigued driving is a leading cause of traffic accidents, and accurately predicting driver fatigue can significantly reduce their occurrence. However, modern fatigue detection models based on neural networks often face challenges such as poor interpretability and insufficient input feature dimensions. This article proposes a novel Spatial-Frequency-Temporal Network (SFT-Net) method for detecting driver fatigue using electroencephalogram (EEG) data. Our approach integrates EEG signals' spatial, frequency, and temporal information to improve recognition performance. We transform the differential entropy of five frequency bands of EEG signals into a 4D feature tensor to preserve these three types of information. An attention module is then used to recalibrate the spatial and frequency information of each input 4D feature tensor time slice. The output of this module is fed into a depthwise separable convolution (DSC) module, which extracts spatial and frequency features after attention fusion. Finally, long short-term memory (LSTM) is used to extract the temporal dependence of the sequence, and the final features are output through a linear layer. We validate the effectiveness of our model on the SEED-VIG dataset, and experimental results demonstrate that SFT-Net outperforms other popular models for EEG fatigue detection. Interpretability analysis supports the claim that our model has a certain level of interpretability. Our work addresses the challenge of detecting driver fatigue from EEG data and highlights the importance of integrating spatial, frequency, and temporal information.
Dongrui Gao, Kejie Wang, Manqing Wang, Jiliu Zhou, Yongqing Zhang 0001
IEEE J. Biomed. Health Informatics2
2024 Learning to Agree on Vision Attention for Visual Commonsense Reasoning
abstract
Visual Commonsense Reasoning (VCR) remains a significant yet challenging research problem in the realm of visual reasoning. A VCR model generally aims at answering a textual question regarding an image, followed by the rationale prediction for the preceding answering process. Though these two processes are sequential and intertwined, existing methods always consider them as two independent matching-based instances. They, therefore, ignore the pivotal relationship between the two processes, leading to sub-optimal model performance. This paper presents a novel visual attention alignment method to efficaciously handle these two processes in a unified framework. To achieve this, we first design a re-attention module for aggregating the vision attention map produced in each process. Thereafter, the resultant two sets of attention maps are carefully aligned to guide the two processes to make decisions based on the same image regions. We apply this method to both conventional attention and the recent Transformer models and carry out extensive experiments on the VCR benchmark dataset. The results demonstrate that with the attention alignment module, our method achieves a considerable improvement over the baseline methods, evidently revealing the feasibility of the coupling of the two processes as well as the effectiveness of the proposed method.
Kejie Wang, Fan Liu 0008, Liqiang Nie, Mohan Kankanhalli
IEEE Trans. Multim.3
2023 Do Vision-Language Transformers Exhibit Visual Commonsense? An Empirical Study of VCR
abstract
Visual Commonsense Reasoning (VCR) calls for explanatory reasoning behind question answering over visual scenes. To achieve this goal, a model is required to provide an acceptable rationale as the reason for the predicted answers. Progress on the benchmark dataset stems largely from the recent advancement of Vision-Language Transformers (VL Transformers). These models are first pre-trained on some generic large-scale vision-text datasets, and then the learned representations are transferred to the downstream VCR task. Despite their attractive performance, this paper posits that the VL Transformers do not exhibit visual commonsense, which is the key to VCR. In particular, our empirical results pinpoint several shortcomings of existing VL Transformers: small gains from pre-training, unexpected language bias, limited model architecture for the two inseparable sub-tasks, and neglect of the important object-tag correlation. With these findings, we tentatively suggest some future directions from the aspect of dataset, evaluation metric, and training tricks. We believe this work could make researchers revisit the intuition and goals of VCR, and thus help tackle the remaining challenges in visual reasoning.
Kejie Wang, Xiaolin Chen 0001, Liqiang Nie, Mohan Kankanhalli
ACM Multimedia3
2023 Egocentric Early Action Prediction via Multimodal Transformer-Based Dual Action Prediction
abstract
Egocentric early action prediction, which aims to recognize the on-going action in the video captured in the first-person view as early as possible before the action is fully executed, is a new yet challenging task due to the limited partial video input. Pioneer studies focused on solving this task with LSTMs as the backbone and simply compiling the observed video segment and unobserved video segment into a single vector, which hence suffer from two key limitations: lack the non-sequential relation modeling with the video snippet sequence and the correlation modeling between the observed and unobserved video segment. To address these two limitations, in this paper, we propose a novel multimodal TransfoRmer-based duAl aCtion prEdiction (mTRACE) model for the task of egocentric early action prediction, which consists of two key modules: the early (observed) segment action prediction module and the future (unobserved) segment action prediction module. Both modules take Transformer encoders as the backbone for encoding all the potential relations among the input video snippets, and involve several single-modal and multi-modal classifiers for comprehensive supervision. Different from previous work, each of the two modules outputs two multi-modal feature vectors: one for encoding the current input video segment, and the other one for predicting the missing video segment. For optimization, we design a two-stage training scheme, including the mutual enhancement stage and end-to-end aggregation stage. The former stage alternatively optimizes the two action prediction modules, where the correlation between the observed and unobserved video segment is modeled with a consistency regularizer, while the latter seamlessly aggregates the two modules to fully utilize the capacity of the two modules. Extensive experiments have demonstrated the superiority of our proposed model. We have released the codes and the corresponding parameters to benefit other researchers athttps://trace729.wixsite.com/trace.
Weili Guan, Xuemeng Song, Kejie Wang, Haokun Wen, Hongda Ni, Yaowei Wang 0001, Xiaojun Chang
IEEE Trans. Circuits Syst. Video Technol.3
2023 Joint Answering and Explanation for Visual Commonsense Reasoning
abstract
Visual Commonsense Reasoning (VCR), deemed as one challenging extension of Visual Question Answering (VQA), endeavors to pursue a higher-level visual comprehension. VCR includes two complementary processes: question answering over a given image and rationale inference for answering explanation. Over the years, a variety of VCR methods have pushed more advancements on the benchmark dataset. Despite significance of these methods, they often treat the two processes in a separate manner and hence decompose VCR into two irrelevant VQA instances. As a result, the pivotal connection between question answering and rationale inference is broken, rendering existing efforts less faithful to visual reasoning. To empirically study this issue, we perform some in-depth empirical explorations in terms of both language shortcuts and generalization capability. Based on our findings, we then propose a plug-and-play knowledge distillation enhanced framework to couple the question answering and rationale inference processes. The key contribution lies in the introduction of a new branch, which serves as a relay to bridge the two processes. Given that our framework is model-agnostic, we apply it to the existing popular baselines and validate its effectiveness on the benchmark dataset. As demonstrated in the experimental results, when equipped with our method, these baselines all achieve consistent and significant performance improvements, evidently verifying the viability of processes coupling.
Kejie Wang, Yinwei Wei, Liqiang Nie, Mohan Kankanhalli
IEEE Trans. Image Process.3