Jianhou Gan

dblp:00/10364 · DBLP profile ↗
← Back
26ranked-venue papers
0as first author
22since 2021 · last 2027
0000-0002-1287-857XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2027 Training-free multi-scale super-resolution with diffusion models
Aiping Zhang, Yuning Cui 0001, Jianhou Gan, Wenqi Ren
Expert Syst. Appl.3
2026 Knowledge graph question generation based on crucial semantic information
Mingtao Zhou, Juxiang Zhou, Jianhou Gan, Jun Wang 0101, Jiatian Mei
Data Knowl. Eng.3
2026 KPTUltra : Dual-Enhanced Knowledgeable Prompt Tuning for Few-Shot Text Classification in Low-Resource Scenarios
abstract
ABSTRACT Prompt tuning‐based few‐shot text classification aims to improve model performance by constructing high‐quality verbalizer. However, existing methods suffer from high subjective bias, insufficient semantic coverage, and uneven representation ability of label words, which limits the further improvements in classification performance. To address these challenges, we propose KPTUltra. The model synergistically integrates multiple pre‐trained models and contrastive learning through class‐sensitive ranking method (CSR) to construct a robust semantic embedding space. Additionally, a genetic algorithm is employed to optimise the mapping between label word and class, enhancing screening stability and semantic matching. Secondly, we introduce a genetic algorithm‐based adaptive label word weight optimization mechanism (GAAWO), which dynamically adjusts both the composition and the weight distribution of label words in the latent space. This enables fine‐grained control and effectively reduces the impact of low‐representative label words. Extensive experiments on multiple few‐shot text classification benchmarks demonstrate that KPTUltra outperforms state‐of‐the‐art baseline methods, achieving superior overall performance.
Wenlong Zha, Mingtao Zhou, Juxiang Zhou, Jianhou Gan, Di Wu 0068
Expert Syst. J. Knowl. Eng.4
2026 Structure-semantic fusion contrastive transformer for robust knowledge graph embedding
Jianhou Gan, Wenqi Ren, Jun Wang 0101
Expert Syst. Appl.2
2026 Dual-strategy retrieval-augmented generation for anomaly recovery in large language models
Jinxin Lv, Jianhou Gan, Wenqi Ren, Jun Wang 0101
Pattern Recognit.3
2026 Early Success Prediction for Online Learners: A Semisupervised and Self-Training Approach
abstract
The accurate and early prediction of online learners’ final success significantly aids administrators in providing timely assistance, thereby enhancing the retention rate for online educational platforms. However, a primary challenge lies in the fact that most learners’ behaviors remain unlabeled until the end of semester, hindering researchers from utilizing traditional supervised learning methods to forecast their final success in the initial weeks. To address these challenges, we introduce a framework called data augmentation-based semisupervised and self-training method for early success prediction (DSSP). Specifically, DSSP initially gathers labeled datasets with a limited number of instances from completed online courses and then employs a generative adversarial network (GAN)-based approach to augment an auxiliary dataset, which maintains the same mixed Gaussian distributions as the labeled set. Furthermore, we apply a self-training method to capture the latent behavior pattern from unlabeled learners’ dataset and then train a multilayer perceptron (MLP)-based semisupervised predictor via consistency loss till convergence, so that the model can predict the online learners’ final success when given their behavior at the first few weeks of the semester. The extensive experiments on the public dataset OULAD (https://analyse.kmi.open.ac.uk/opendataset) demonstrate the effectiveness of our proposed methods.
Jia Hao 0001, Jiatian Mei, Jianhou Gan
IEEE Trans. Comput. Soc. Syst.3
2026 CDIR: LoRA-Inspired Attention for Efficient Composite Degradation Image Restoration
abstract
Specialized image restoration methods have been extensively explored, each targeting a specific type of degradation. However, real-world images often suffer from composite degradations, prompting growing interest in unified restoration approaches. While recent unified models have shown promising results, many are hindered by high computational complexity, limiting their deployment in resource-constrained settings. Motivated by the parameter-efficient design of Low-Rank Adaptation (LoRA), we propose an efficient attention module specifically designed for composite degradation image restoration. The proposed method adopts a dual-branch architecture, where one branch processes features at full resolution, and the other operates with reduced spatial and channel dimensions to improve efficiency. To better adapt to diverse degradation patterns, the latter branch is further divided into two sub-branches, each incorporating dynamic operations guided by local and contextual priors. These context priors are iteratively updated within each module, drawing inspiration from feedback mechanisms in reinforcement learning, thereby enabling the model to effectively perceive and handle multiple degradation types within a unified structure. Additionally, we introduce a multi-scale feed-forward network to further enhance both performance and computational efficiency. Extensive experiments on two composite degradation benchmarks demonstrate that our proposed network, CDIR, achieves state-of-the-art performance with significantly reduced complexity and fast inference speed. In addition, CDIR shows strong adaptability to various task-specific image restoration scenarios, such as dehazing, desnowing, and deraining. It also performs robustly on domain-specific applications such as ultra-high-definition (UHD), remote sensing, and medical image restoration, highlighting its versatility and practical applicability.
Yuning Cui 0001, Wenqi Ren, Boxin Shi, Jianhou Gan, Alois C. Knoll
IEEE Trans. Image Process.4
2025 Weakly Semi-supervised Classroom Teacher Visual Tracking by Single-Point Annotations
Di Wu 0068, Jianhou Gan, Jiatian Mei, Jun Wang 0101, Juxiang Zhou
CGI (1)2
2025 Semi-Supervised State-Space Model with Dynamic Stacking Filter for Real-World Video Deraining
abstract
Significant progress has been made in video restoration under rainy conditions over the past decade, largely propelled by advancements in deep learning. Nevertheless, existing methods that depend on paired data struggle to generalize effectively to real-world scenarios, primarily due to the disparity between synthetic and authentic rain effects. To address these limitations, we propose a dual-branch spatiotemporal state-space model to enhance rain streak removal in video sequences. Specifically, we design spatial and temporal state-space model layers to extract spatial features and incorporate temporal dependencies across frames, respectively. To improve multi-frame feature fusion, we derive a dynamic stacking filter, which adaptively approximates statistical filters for superior pixel-wise feature refinement. Moreover, we develop a median stacking loss to enable semi-supervised learning by generating pseudo-clean patches based on the sparsity prior of rain. To further explore the capacity of deraining models in supporting other vision-based tasks in rainy environments, we introduce a novel real-world benchmark focused on object detection and tracking in rainy conditions. Our method is extensively evaluated across multiple benchmarks containing numerous synthetic and real-world rainy videos, consistently demonstrating its superiority in quantitative metrics, visual quality, efficiency, and its utility for downstream tasks.
Shangquan Sun, Wenqi Ren, Juxiang Zhou, Jianhou Gan, Xiaochun Cao
CVPR5
2025 Enhancing Handwritten Mathematical Expression Recognition with Structure and Counting Aware Network
abstract
The encoder-decoder architecture has been widely adopted by many handwritten mathematical expression recognition (HMER) models. However, these methods may struggle to accurately recognize expressions with complex structures or long sequences due to the lack of global information utilization. To address this issue, we propose a novel network based on the encoder-decoder framework, called the Structure and Counting Aware Network (SCAN), to enable global information awareness and achieve more precise recognition. Specifically, we introduce the Skeleton Shaping and Character Counting Module (SSCCM) to extract the skeleton structure of expressions and the frequency distribution of characters. These two types of information are integrated with the decoder output in the calibration module to refine the predictions. Experimental results demonstrate that SCAN significantly outperforms the existing state-of-the-art (SOTA) models in terms of expression recognition rate (ExpRate) on the CROHME 2014/2016/2019 and HME100K test sets. The code is publicly available at https://github.com/Kerston12138/SCAN.
Shiqi Mou, Juxiang Zhou, Jun Wang 0101, Jianhou Gan
ICME5
2025 FaceInsight: A Multimodal Large Language Model for Face Perception
abstract
Recent advances in multimodal large language models (MLLMs) have demonstrated strong capabilities in understanding general visual content. However, these general-domain MLLMs perform poorly in face perception tasks, often producing inaccurate or misleading responses to face-specific queries. To address this gap, we propose FaceInsight, a versatile face perception MLLM that provides fine-grained information. Our approach introduces visual textual alignment of facial knowledge to model both uncertain dependencies and deterministic relationships among facial information, mitigating the limitations of language-driven reasoning. Additionally, we incorporate face segmentation maps as an auxiliary perceptual modality, enriching visual input with localized structural cues to enhance semantic understanding. Comprehensive experiments show that FaceInsight consistently outperforms nine compared MLLMs under both training-free and fine-tuned settings.
Jingzhi Li 0002, Changjiang Luo, Ruoyu Chen 0001, Hua Zhang 0008, Wenqi Ren, Jianhou Gan, Xiaochun Cao
ACM Multimedia6
2025 Text semantic-guided adaptive feature aggregation for image-text retrieval
Yajie Gu, Jianhou Gan, Jiatian Mei, Chuanzhi Zhang
Multim. Syst.3
2024 Delay-guaranteed Mobile Augmented Reality Task Offloading in Edge-assisted Environment
abstract
With the introduction of Augmented Reality (AR) into mobile devices, it becomes a trend to develop mobile AR applications in various fields. However, the execution of AR task demands the extensive resources of computation, memory and storage, which makes it difficult for mobile terminals with constrained hardware resources to carry out AR applications within the limited delay. In response to this challenge, we propose a mobile AR offloading method under the edge-assisted environment. Firstly, we divide an AR task into consecutive subtasks, and then collect the features of hardware, software, configuration, and runtime environments from the edge servers to be offloaded. With the features, we construct an AR subtask Execution delay Prediction Bayesian Network (EPBN) to predict the execution delay of different subtasks on each edge platform. Based on the prediction, we model the task offloading as the NP-hard Traveling Salesman Problem (TSP), and then propose a PSO-GA based solution by adopting the heuristic algorithm of Particle Swarm Optimization (PSO) to encode the offloading strategy and using Genetic Algorithm (GA) for particle update. The extensive experiments prove that the average performances of EPBN outperform the others with 17.23%, 23.97%, and 20.67% on micro-P, micro-R, and micro-F1 respectively, and the PSO-GA ensures that the offloading latency is reduced by nearly 5% compared to the suboptimal algorithm.
Jia Hao 0001, Jianhou Gan
Ad Hoc Networks3
2024 Clarification question generation diversity and specificity enhancement based on question keyword prediction
Mingtao Zhou, Juxiang Zhou, Jianhou Gan, Wei Gao 0012
Appl. Intell.3
2024 Classroom teacher action recognition based on spatio-temporal dual-branch feature fusion
Di Wu 0068, Jun Wang 0101, Shaodong Zou, Juxiang Zhou, Jianhou Gan
Comput. Vis. Image Underst.6
2024 Rank-based multimodal immune algorithm for many-objective optimization problems
Jianhou Gan, Juxiang Zhou, Wei Gao 0012
Eng. Appl. Artif. Intell.2
2024 TracKGE: Transformer with Relation-pattern Adaptive Contrastive Learning for Knowledge Graph Embedding
Jun Wang 0101, Juxiang Zhou, Jianhou Gan
Knowl. Based Syst.6
2022 Fine-grained semantic ethnic costume high-resolution image colorization with conditional GAN
abstract
Grayscale image colorization, especially for ethnic costume images, is highly challenging due to its rich and complex color features. The existing image colorization methods usually take the costume image as a whole in practical applications that lead to the ignorance of the semantic information of different parts of the costume. It is known that each part's color distribution of the ethnic costume is different. So, the color mapping of other parts is also diverse, which is determined by distinctive ethnic characteristics. This study introduces fine-grained level semantic information and proposes a high-resolution image colorization model for ethnic costumes targeting enhancement, inspired by semantic-level colorization. The semantic information of different regions of ethnic costumes has a significant impact on the performance of the coloring task. Using Pix2PixHD as the backbone network, we create a new network architecture that maintains color distribution correspondence and spatial consistency of costume images using fine-grained semantic information. In our network, we take the splice result of fine-grained semantic for ethnic costume and grayscale image as the conditions and then feed them into the generative adversarial networks. We also discuss and analyze the influences of the grayscale channel and fine-grained semantics on discriminator. Extensive experiments demonstrate that our method performs well compared with other state-of-the-art automatic colorization methods.
Di Wu 0068, Jianhou Gan, Juxiang Zhou, Jun Wang 0101, Wei Gao 0012
Int. J. Intell. Syst.2
2022 Clothing attribute recognition via a holistic relation network
abstract
Clothing attribute prediction is a fundamental image classification task in the field of computer vision. Motivated by the human recognition system, we investigate the task relationship and spatial importance where people usually utilize these useful clues to assist in recognizing clothing attributes. In this paper, we propose a novel Holistic Relation Network (HRNet) for clothing attribute recognition, considering the fusion of multiple relations, including spatial and spatial relation via a spatial relation module, spatial and task relation via a task attention module, and task and task relation via a graph context reasoning module. Specifically, we first use the backbone network to extract features from the input image, two types of attention models followed will further learn the features, then the graph context reasoning module will be used to further enhance the features, and finally, a classifier exploited to classify the clothing attributes with the learned representation information. Without using manual image feature filtering methods, this paper aims to achieve clothing attribute recognition by deeply exploring the relationships among different clothing attribute recognition tasks. In this paper, we use double-branches of the attention model to model the relevance of spatial context information and learn more discriminative feature representations from multitask features for clothing attribute prediction. Derived from the prior knowledge learned from the two above-mentioned attention models, we further propose a graph-relation model constructing relationships among different clothing attribute tasks by integrating the spatial association relationships among multitask. The proposed HRNet only uses image-level annotation but it owns a good ability for obtaining distinguishing feature representations. We obtain state-of-the-art performance, which is demonstrated by extensive experiments on three mainstream benchmarks, for example, woman clothing data set, man clothing data set, and shop-domain clothing data set.
Di Wu 0068, Juxiang Zhou, Jianhou Gan, Wei Gao 0012, Hao Li 0188
Int. J. Intell. Syst.4
2022 Fine-Grained Image Classification Based on Cross-Attention Network
abstract
Due to the high similarity of fine-grained image subclasses, small inter-class changes and large intra-class changes are caused, which leads to the difficulty of fine-grained image classification task. However, existing convolutional neural networks have been unable to effectively solve this problem. Aiming at the above-mentioned fine-grained image classification problem, this paper proposes a multi-scale and multi-level ViT model. First, through data augmentation techniques, the accuracy of fine-grained image classification can be effectively improved. Secondly, the small-scale input and large-scale input of the model make the input image have more feature ex-pressions. The subsequent multi-layeredness effectively utilizes the results of the previous layer of ViT, so that the data of the previous layer can be more effectively used in the next layer of ViT. Finally, cross-attention allows the results of two scale inputs to be fused in a reasonable way. The proposed model is competitive with current mainstream state-of-the-art methods on multiple datasets.
Juxiang Zhou, Jianhou Gan, Sen Luo, Wei Gao 0012
Int. J. Semantic Web Inf. Syst.3
2021 Image retrieval based on aggregated deep features weighted by regional significance and channel sensitivity
Juxiang Zhou, Jianhou Gan, Wei Gao 0012, Antoni Liang
Inf. Sci.2
2021 Design and Research of Intelligent Question-Answering(Q&A) System Based on High School Course Knowledge Graph
Jianhou Gan, Ning Lei
Mob. Networks Appl.3
2019 Image retrieval based on effective feature extraction and diffusion process
Juxiang Zhou, Xiaodong Liu 0001, Wanquan Liu, Jianhou Gan
Multim. Tools Appl.4
2019 Learning Structural Representations via Dynamic Object Landmarks Discovery for Sketch Recognition and Retrieval
abstract
State-of-the-art methods on sketch classification and retrieval are based on deep convolutional neural network to learn representations. Although deep neural networks have the ability to model images with hierarchical representations by convolution kernels, they can not automatically extract the structural representations of object categories in a human-perceptible way. Furthermore, sketch images usually have large scale visual variations caused by the styles of drawing or viewpoints, which make it difficult to develop generalized representations using the fixed computational mode of convolutional kernel. In this paper, our aim is to address the problem of fixed computational mode in feature extraction process without extra supervision. We propose a novel architecture to dynamically discover the object landmarks and learn the discriminative structural representations. Our model is composed of two components: a representative landmark discovering module that localizes the key points on the object, and a category-aware representation learning module that develops the category-specific features. Specifically, we develop a structure-aware offset layer to dynamically localize the representative landmarks, which is optimized based on the category labels without extra supervision. After that, a diversity branch is introduced to extract the global discriminative features for each category. Finally, we employ a multi-task loss function to develop an end-to-end trainable architecture. At testing time, we fuse all the predictions with different number of landmarks to achieve the final results. Through extensive experiments, we compare our model with several state-of-the-art methods on two challenging datasets TU-Berlin and Sketchy for sketch classification and retrieval, and the experimental results demonstrate the effectiveness of our proposed model.
Hua Zhang 0008, Peng She, Yong Liu 0018, Jianhou Gan, Xiaochun Cao, Hassan Foroosh
IEEE Trans. Image Process.4
2018 Multi-ethnical Chinese facial characterization and analysis
Cunrui Wang, Qingling Zhang 0001, Xiaodong Duan, Jianhou Gan
Multim. Tools Appl.4
2015 Improved Bacterial Foraging Optimization Algorithm with Information Communication Mechanism for Nurse Scheduling
Ben Niu 0002, Jing Liu 0029, Jianhou Gan, Lingyun Yuan
ICIC (2)4