Wenjuan Gong

dblp:60/6056 · DBLP profile ↗
← Back
19ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0001-7805-3629ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Databases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1Computer networks · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 Retrieval-driven Reasoning for Deliberative Visual Classification
abstract
Vision-Language Models (VLMs) have demonstrated remarkable capabilities in visual classification tasks. Existing methods for enhancing VLMs on this task often rely heavily on direct category-to-image matching, which limits generalization and results in suboptimal performance. In addition, these methods provide no understanding of why a specific category is chosen. To address these limitations, we introduce a new deliberative visual classification task that decomposes the classification process into multiple deliberative steps and leverages Large Language Models (LLMs) to perform explicit reasoning before the final decision. Specifically, we propose a Retrieval-driven Reasoning model (RdR) with two components, i.e., retrieval database construction and deliberative category prediction. The first component leverages LLMs to extract category-relevant descriptors and constructs a retrieval database for effective image–descriptor matching. The second component facilitates multiple deliberative steps and performs explicit reasoning based on the retrieved descriptors to augment the category prediction. Extensive experiments on multiple datasets demonstrate that RdR consistently outperforms strong baselines, highlighting its robustness and generalization ability.
Jianye Xie, Lianyong Qi, Fan Wang 0020, Wenjuan Gong, Danxin Wang, Wan-Chun Dou, Yang Cao 0019, Shichao Pei, Xiaokang Zhou
AAAI5
2025 Facial Expression Generation from Text with FaceCLIP
Wenwen Fu, Wenjuan Gong, Chen-Yang Yu, Wei Wang 0115, Jordi Gonzàlez 0001
J. Comput. Sci. Technol.2
2025 Lite transformer with medium self attention for efficient traffic sign recognition
Junbi Xiao, Wenjuan Gong, Jianhang Liu
J. Vis. Commun. Image Represent.3
2025 Audio-visual scene recognition using attention-based graph convolutional model
Yikai Wu 0004, Wenjuan Gong, Jordi Gonzàlez 0001
Multim. Tools Appl.4
2024 FedSOKD-TFA: Federated Learning with Stage-Optimal Knowledge Distillation and Three-Factor Aggregation
Jianhao Liu, Wenjuan Gong, Tingbo Shi, Kechen Li, Jordi Gonzàlez 0001
ICPR (2)2
2024 MCLEMCD: multimodal collaborative learning encoder for enhanced music classification from dances
Wenjuan Gong, Qingshuang Yu, Wendong Huang, Peng Cheng 0008, Jordi Gonzàlez 0001
Multim. Syst.1
2024 Meta-MMFNet: Meta-learning-based Multi-model Fusion Network for Micro-expression Recognition
abstract
Despite its wide applications in criminal investigations and clinical communications with patients suffering from autism, automatic micro-expression recognition remains a challenging problem because of the lack of training data and imbalanced classes problems. In this study, we proposed a meta-learning-based multi-model fusion network (Meta-MMFNet) to solve the existing problems. The proposed method is based on the metric-based meta-learning pipeline, which is specifically designed for few-shot learning and is suitable for model-level fusion. The frame difference and optical flow features were fused, deep features were extracted from the fused feature, and finally in the meta-learning-based framework, weighted sum model fusion method was applied for micro-expression classification. Meta-MMFNet achieved better results than state-of-the-art methods on four datasets. The code is available at https://github.com/wenjgong/meta-fusion-based-method .
Wenjuan Gong, Yue Zhang 0087, Wei Wang 0115, Peng Cheng 0008, Jordi Gonzàlez 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2022 Deep learning-based microexpression recognition: a survey
Wenjuan Gong, Zhihong An, Noha M. Elfiky
Neural Comput. Appl.1
2020 Discriminative Correlation Filter for Long-Time Tracking
abstract
Abstract Object tracking is a very important step in building an intelligent video monitoring system that can protect people’s lives and property. In recent years, although visual tracking has made great progress in terms of speed and accuracy, there are still few real-time high-precision tracking algorithms. Although discriminative correlation filters have excellent performance in tracking speed, there are deficiencies in handling fast motion. This leads to the inability to achieve long-term stable tracking results. The long-time tracking with discriminative correlation filter (LT-DCF) was proposed to solve these deficiencies. We use larger size detection image blocks and smaller size filters to increase the proportion of real samples to solve the boundary effects of fast motion. And we combine the histogram of oriented gradient (HOG) feature detection and scale-invariant feature transform (SIFT) key point detection to solve the obstacles caused by scale variations. The detector with deep feature flow is then incorporated into the tracker to detect key frames to improve tracking accuracy. This method has achieved more than 75% of the distance accuracy and 70% of the overlapping success rate on the VOT2015 and VOT2016 datasets, and the stable tracking video length can reach 6895 frames.
Faming Gong, Hanbing Yue, Xiangbing Yuan, Wenjuan Gong, Tao Song 0001
Comput. J.4
2018 Service-aware adaptive link load balancing mechanism for Software-Defined Networking
Fengjun Shang, Lin Mao, Wenjuan Gong
Future Gener. Comput. Syst.3
2018 A Belief-Theoretical Approach to Example-Based Pose Estimation
abstract
In example-based human pose estimation, the configuration of an evolving object is sought given visual evidence, having to rely uniquely on a set of sample images. We assume here that, at each time instant of a training session, a number of feature measurements is extracted from the available images, while ground truth is provided in the form of the true object pose. In this scenario, a sensible approach consists in learning maps from features to poses, using the information provided by the training set. In particular, multivalued mappings linking feature values to set of training poses can be constructed. To this purpose we propose a belief modeling regression (BMR) approach in which a probability measure on any individual feature space maps to a convex set of probabilities on the set of training poses, in a form of a belief function. Given a test image, its feature measurements translate into a collection of belief functions on the set of training poses which, when combined, yield there an entire family of probability distributions. From the latter either a single central pose estimate or a set of extremal ones can be computed, together with a measure of how reliable the estimate is. Contrarily to other competing models, in BMR the sparsity of the training samples can be taken into account to model the level of uncertainty associated with these estimates. We illustrate BMR's performance in an application to human pose recovery, showing how it outperforms our implementation of both relevant vector machine and Gaussian process regression. Finally, we discuss motivation and advantages of the proposed approach with respect to its most direct competitors.
Wenjuan Gong, Fabio Cuzzolin
IEEE Trans. Fuzzy Syst.1
2017 Resource requests prediction in the cloud computing environment with a deep belief network
abstract
Summary Accurate resource requests prediction is essential to achieve optimal job scheduling and load balancing for cloud Computing. Existing prediction approaches fall short in providing satisfactory accuracy because of high variances of cloud metrics. We propose a deep belief network (DBN)‐based approach to predict cloud resource requests. We design a set of experiments to find the most influential factors for prediction accuracy and the best DBN parameter set to achieve optimal performance. The innovative points of the proposed approach is that it introduces analysis of variance and orthogonal experimental design techniques into the parameter learning of DBN. The proposed approach achieves high accuracy with mean square error of [10−6,10−5], approximately 72%reduction compared with the traditional autoregressive integrated moving average predictor, and has better prediction accuracy compared with the state‐of‐art fractal modeling approach. Copyright © 2016 John Wiley & Sons, Ltd.
Weishan Zhang, Pengcheng Duan, Laurence T. Yang, Feng Xia 0001, Qinghua Lu 0001, Wenjuan Gong, Su Yang 0001
Softw. Pract. Exp.7
2016 Distributed embedded deep learning based real-time video processing
abstract
There arises the needs for fast processing of continuous video data using embedded devices, for example the one needed for UAV aerial photography. In this paper, we proposed a distributed embedded platform built with NVIDIA Jetson TX1 using deep learning techniques for real time video processing, mainly for object detection. We design a Storm based distributed real-time computation platform and ran object detection algorithm based on convolutional neural networks. We have evaluated the performance of our platform by conducting real-time object detection on surveillance video. Compared with the high end GPU processing of NVIDIA TITAN X, our platform achieves the same processing speed but a much lower power consumption when doing the same work. At the same time, our platform had a good scalability and fault tolerance, which is suitable for intelligent mobile devices such as unmanned aerial vehicles or self-driving cars.
Weishan Zhang, Dehai Zhao, Liang Xu 0009, Wenjuan Gong, Jiehan Zhou
SMC5
2016 A Load-Aware Pluggable Cloud Framework for Real-Time Video Processing
abstract
A large number of video applications require real-time response. The high-speed video processing then requires a distributed and parallelized framework utilizing all possible computing resources, i.e., both Central Processing Unit (CPU) and Graphics Processing Unit (GPU) at their best. The CPU-GPU collaboration may cause resource imbalance where GPU-based jobs consume less computing resources while occupying more memory compared with CPU-based jobs. In this paper, we propose a load-aware pluggable cloud framework for real-time video processing where CPU-GPU switching based on workload status can be performed at runtime. Furthermore, we design aspect-oriented monitors to collect framework metrics and propose a distance coverage algorithm to detect performance degradation in order to make sure that the framework runs optimally to achieve good performance when a load-aware task switching is made. We have comprehensively evaluated the framework and the evaluation results show that the proposed framework has good performance, reusability, pluggability, and scalability.
Weishan Zhang, Pengcheng Duan, Wenjuan Gong, Qinghua Lu 0001, Su Yang 0001
IEEE Trans. Ind. Informatics3
2015 A video cloud platform combing online and offline cloud computing technologies
Weishan Zhang, Liang Xu 0009, Pengcheng Duan, Wenjuan Gong, Qinghua Lu 0001, Su Yang 0001
Pers. Ubiquitous Comput.4
2013 Belief modeling regression for pose estimation
Fabio Cuzzolin, Wenjuan Gong
FUSION2
2012 3D human pose estimation using 2D body part detectors
Adela Barbulescu, Wenjuan Gong, Jordi Gonzàlez 0001, Thomas B. Moeslund, F. Xavier Roca
ICPR2
2008 Continuous Collision Detection between Two 2DCurved-Edge Polygons under Rational Motions
Wenjuan Gong, Changhe Tu
GMP1
2004 Customer Behavior Pattern Discovering with Web Mining
Wenjuan Gong, Yoshihiro Kawamura
APWeb2