VLDB 2026 Research / reviewers in the wild / expert
Jianhou Gan
dblp:00/10364
· DBLP profile ↗
26ranked-venue papers
0as first author
22since 2021 · last 2027
0000-0002-1287-857XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Training-free multi-scale super-resolution with diffusion models
Aiping Zhang, Yuning Cui 0001, Jianhou Gan, Wenqi Ren |
Expert Syst. Appl. | 3 |
| 2026 | Knowledge graph question generation based on crucial semantic information
Mingtao Zhou, Juxiang Zhou, Jianhou Gan, Jun Wang 0101, Jiatian Mei |
Data Knowl. Eng. | 3 |
| 2026 | KPTUltra : Dual-Enhanced Knowledgeable Prompt Tuning for Few-Shot Text Classification in Low-Resource ScenariosabstractABSTRACT Prompt tuning‐based few‐shot text classification aims to improve model performance by constructing high‐quality verbalizer. However, existing methods suffer from high subjective bias, insufficient semantic coverage, and uneven representation ability of label words, which limits the further improvements in classification performance. To address these challenges, we propose KPTUltra. The model synergistically integrates multiple pre‐trained models and contrastive learning through class‐sensitive ranking method (CSR) to construct a robust semantic embedding space. Additionally, a genetic algorithm is employed to optimise the mapping between label word and class, enhancing screening stability and semantic matching. Secondly, we introduce a genetic algorithm‐based adaptive label word weight optimization mechanism (GAAWO), which dynamically adjusts both the composition and the weight distribution of label words in the latent space. This enables fine‐grained control and effectively reduces the impact of low‐representative label words. Extensive experiments on multiple few‐shot text classification benchmarks demonstrate that KPTUltra outperforms state‐of‐the‐art baseline methods, achieving superior overall performance. Wenlong Zha, Mingtao Zhou, Juxiang Zhou, Jianhou Gan, Di Wu 0068 |
Expert Syst. J. Knowl. Eng. | 4 |
| 2026 | Structure-semantic fusion contrastive transformer for robust knowledge graph embedding
Jianhou Gan, Wenqi Ren, Jun Wang 0101 |
Expert Syst. Appl. | 2 |
| 2026 | Dual-strategy retrieval-augmented generation for anomaly recovery in large language models
Jinxin Lv, Jianhou Gan, Wenqi Ren, Jun Wang 0101 |
Pattern Recognit. | 3 |
| 2026 | Early Success Prediction for Online Learners: A Semisupervised and Self-Training ApproachabstractThe accurate and early prediction of online learners’ final success significantly aids administrators in providing timely assistance, thereby enhancing the retention rate for online educational platforms. However, a primary challenge lies in the fact that most learners’ behaviors remain unlabeled until the end of semester, hindering researchers from utilizing traditional supervised learning methods to forecast their final success in the initial weeks. To address these challenges, we introduce a framework called data augmentation-based semisupervised and self-training method for early success prediction (DSSP). Specifically, DSSP initially gathers labeled datasets with a limited number of instances from completed online courses and then employs a generative adversarial network (GAN)-based approach to augment an auxiliary dataset, which maintains the same mixed Gaussian distributions as the labeled set. Furthermore, we apply a self-training method to capture the latent behavior pattern from unlabeled learners’ dataset and then train a multilayer perceptron (MLP)-based semisupervised predictor via consistency loss till convergence, so that the model can predict the online learners’ final success when given their behavior at the first few weeks of the semester. The extensive experiments on the public dataset OULAD (https://analyse.kmi.open.ac.uk/opendataset) demonstrate the effectiveness of our proposed methods. Jia Hao 0001, Jiatian Mei, Jianhou Gan |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2026 | CDIR: LoRA-Inspired Attention for Efficient Composite Degradation Image RestorationabstractSpecialized image restoration methods have been extensively explored, each targeting a specific type of degradation. However, real-world images often suffer from composite degradations, prompting growing interest in unified restoration approaches. While recent unified models have shown promising results, many are hindered by high computational complexity, limiting their deployment in resource-constrained settings. Motivated by the parameter-efficient design of Low-Rank Adaptation (LoRA), we propose an efficient attention module specifically designed for composite degradation image restoration. The proposed method adopts a dual-branch architecture, where one branch processes features at full resolution, and the other operates with reduced spatial and channel dimensions to improve efficiency. To better adapt to diverse degradation patterns, the latter branch is further divided into two sub-branches, each incorporating dynamic operations guided by local and contextual priors. These context priors are iteratively updated within each module, drawing inspiration from feedback mechanisms in reinforcement learning, thereby enabling the model to effectively perceive and handle multiple degradation types within a unified structure. Additionally, we introduce a multi-scale feed-forward network to further enhance both performance and computational efficiency. Extensive experiments on two composite degradation benchmarks demonstrate that our proposed network, CDIR, achieves state-of-the-art performance with significantly reduced complexity and fast inference speed. In addition, CDIR shows strong adaptability to various task-specific image restoration scenarios, such as dehazing, desnowing, and deraining. It also performs robustly on domain-specific applications such as ultra-high-definition (UHD), remote sensing, and medical image restoration, highlighting its versatility and practical applicability. Yuning Cui 0001, Wenqi Ren, Boxin Shi, Jianhou Gan, Alois C. Knoll |
IEEE Trans. Image Process. | 4 |
| 2025 | Weakly Semi-supervised Classroom Teacher Visual Tracking by Single-Point Annotations
Di Wu 0068, Jianhou Gan, Jiatian Mei, Jun Wang 0101, Juxiang Zhou |
CGI (1) | 2 |
| 2025 | Semi-Supervised State-Space Model with Dynamic Stacking Filter for Real-World Video DerainingabstractSignificant progress has been made in video restoration under rainy conditions over the past decade, largely propelled by advancements in deep learning. Nevertheless, existing methods that depend on paired data struggle to generalize effectively to real-world scenarios, primarily due to the disparity between synthetic and authentic rain effects. To address these limitations, we propose a dual-branch spatiotemporal state-space model to enhance rain streak removal in video sequences. Specifically, we design spatial and temporal state-space model layers to extract spatial features and incorporate temporal dependencies across frames, respectively. To improve multi-frame feature fusion, we derive a dynamic stacking filter, which adaptively approximates statistical filters for superior pixel-wise feature refinement. Moreover, we develop a median stacking loss to enable semi-supervised learning by generating pseudo-clean patches based on the sparsity prior of rain. To further explore the capacity of deraining models in supporting other vision-based tasks in rainy environments, we introduce a novel real-world benchmark focused on object detection and tracking in rainy conditions. Our method is extensively evaluated across multiple benchmarks containing numerous synthetic and real-world rainy videos, consistently demonstrating its superiority in quantitative metrics, visual quality, efficiency, and its utility for downstream tasks. Shangquan Sun, Wenqi Ren, Juxiang Zhou, Jianhou Gan, Xiaochun Cao |
CVPR | 5 |
| 2025 | Enhancing Handwritten Mathematical Expression Recognition with Structure and Counting Aware NetworkabstractThe encoder-decoder architecture has been widely adopted by many handwritten mathematical expression recognition (HMER) models. However, these methods may struggle to accurately recognize expressions with complex structures or long sequences due to the lack of global information utilization. To address this issue, we propose a novel network based on the encoder-decoder framework, called the Structure and Counting Aware Network (SCAN), to enable global information awareness and achieve more precise recognition. Specifically, we introduce the Skeleton Shaping and Character Counting Module (SSCCM) to extract the skeleton structure of expressions and the frequency distribution of characters. These two types of information are integrated with the decoder output in the calibration module to refine the predictions. Experimental results demonstrate that SCAN significantly outperforms the existing state-of-the-art (SOTA) models in terms of expression recognition rate (ExpRate) on the CROHME 2014/2016/2019 and HME100K test sets. The code is publicly available at https://github.com/Kerston12138/SCAN. Shiqi Mou, Juxiang Zhou, Jun Wang 0101, Jianhou Gan |
ICME | 5 |
| 2025 | FaceInsight: A Multimodal Large Language Model for Face PerceptionabstractRecent advances in multimodal large language models (MLLMs) have demonstrated strong capabilities in understanding general visual content. However, these general-domain MLLMs perform poorly in face perception tasks, often producing inaccurate or misleading responses to face-specific queries. To address this gap, we propose FaceInsight, a versatile face perception MLLM that provides fine-grained information. Our approach introduces visual textual alignment of facial knowledge to model both uncertain dependencies and deterministic relationships among facial information, mitigating the limitations of language-driven reasoning. Additionally, we incorporate face segmentation maps as an auxiliary perceptual modality, enriching visual input with localized structural cues to enhance semantic understanding. Comprehensive experiments show that FaceInsight consistently outperforms nine compared MLLMs under both training-free and fine-tuned settings. Jingzhi Li 0002, Changjiang Luo, Ruoyu Chen 0001, Hua Zhang 0008, Wenqi Ren, Jianhou Gan, Xiaochun Cao |
ACM Multimedia | 6 |
| 2025 | Text semantic-guided adaptive feature aggregation for image-text retrieval
Yajie Gu, Jianhou Gan, Jiatian Mei, Chuanzhi Zhang |
Multim. Syst. | 3 |
| 2024 | Delay-guaranteed Mobile Augmented Reality Task Offloading in Edge-assisted EnvironmentabstractWith the introduction of Augmented Reality (AR) into mobile devices, it becomes a trend to develop mobile AR applications in various fields. However, the execution of AR task demands the extensive resources of computation, memory and storage, which makes it difficult for mobile terminals with constrained hardware resources to carry out AR applications within the limited delay. In response to this challenge, we propose a mobile AR offloading method under the edge-assisted environment. Firstly, we divide an AR task into consecutive subtasks, and then collect the features of hardware, software, configuration, and runtime environments from the edge servers to be offloaded. With the features, we construct an AR subtask Execution delay Prediction Bayesian Network (EPBN) to predict the execution delay of different subtasks on each edge platform. Based on the prediction, we model the task offloading as the NP-hard Traveling Salesman Problem (TSP), and then propose a PSO-GA based solution by adopting the heuristic algorithm of Particle Swarm Optimization (PSO) to encode the offloading strategy and using Genetic Algorithm (GA) for particle update. The extensive experiments prove that the average performances of EPBN outperform the others with 17.23%, 23.97%, and 20.67% on micro-P, micro-R, and micro-F1 respectively, and the PSO-GA ensures that the offloading latency is reduced by nearly 5% compared to the suboptimal algorithm. Jia Hao 0001, Jianhou Gan |
Ad Hoc Networks | 3 |
| 2024 | Clarification question generation diversity and specificity enhancement based on question keyword prediction
Mingtao Zhou, Juxiang Zhou, Jianhou Gan, Wei Gao 0012 |
Appl. Intell. | 3 |
| 2024 | Classroom teacher action recognition based on spatio-temporal dual-branch feature fusion
Di Wu 0068, Jun Wang 0101, Shaodong Zou, Juxiang Zhou, Jianhou Gan |
Comput. Vis. Image Underst. | 6 |
| 2024 | Rank-based multimodal immune algorithm for many-objective optimization problems
Jianhou Gan, Juxiang Zhou, Wei Gao 0012 |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | TracKGE: Transformer with Relation-pattern Adaptive Contrastive Learning for Knowledge Graph Embedding
Jun Wang 0101, Juxiang Zhou, Jianhou Gan |
Knowl. Based Syst. | 6 |
| 2022 | Fine-grained semantic ethnic costume high-resolution image colorization with conditional GANabstractGrayscale image colorization, especially for ethnic costume images, is highly challenging due to its rich and complex color features. The existing image colorization methods usually take the costume image as a whole in practical applications that lead to the ignorance of the semantic information of different parts of the costume. It is known that each part's color distribution of the ethnic costume is different. So, the color mapping of other parts is also diverse, which is determined by distinctive ethnic characteristics. This study introduces fine-grained level semantic information and proposes a high-resolution image colorization model for ethnic costumes targeting enhancement, inspired by semantic-level colorization. The semantic information of different regions of ethnic costumes has a significant impact on the performance of the coloring task. Using Pix2PixHD as the backbone network, we create a new network architecture that maintains color distribution correspondence and spatial consistency of costume images using fine-grained semantic information. In our network, we take the splice result of fine-grained semantic for ethnic costume and grayscale image as the conditions and then feed them into the generative adversarial networks. We also discuss and analyze the influences of the grayscale channel and fine-grained semantics on discriminator. Extensive experiments demonstrate that our method performs well compared with other state-of-the-art automatic colorization methods. Di Wu 0068, Jianhou Gan, Juxiang Zhou, Jun Wang 0101, Wei Gao 0012 |
Int. J. Intell. Syst. | 2 |
| 2022 | Clothing attribute recognition via a holistic relation networkabstractClothing attribute prediction is a fundamental image classification task in the field of computer vision. Motivated by the human recognition system, we investigate the task relationship and spatial importance where people usually utilize these useful clues to assist in recognizing clothing attributes. In this paper, we propose a novel Holistic Relation Network (HRNet) for clothing attribute recognition, considering the fusion of multiple relations, including spatial and spatial relation via a spatial relation module, spatial and task relation via a task attention module, and task and task relation via a graph context reasoning module. Specifically, we first use the backbone network to extract features from the input image, two types of attention models followed will further learn the features, then the graph context reasoning module will be used to further enhance the features, and finally, a classifier exploited to classify the clothing attributes with the learned representation information. Without using manual image feature filtering methods, this paper aims to achieve clothing attribute recognition by deeply exploring the relationships among different clothing attribute recognition tasks. In this paper, we use double-branches of the attention model to model the relevance of spatial context information and learn more discriminative feature representations from multitask features for clothing attribute prediction. Derived from the prior knowledge learned from the two above-mentioned attention models, we further propose a graph-relation model constructing relationships among different clothing attribute tasks by integrating the spatial association relationships among multitask. The proposed HRNet only uses image-level annotation but it owns a good ability for obtaining distinguishing feature representations. We obtain state-of-the-art performance, which is demonstrated by extensive experiments on three mainstream benchmarks, for example, woman clothing data set, man clothing data set, and shop-domain clothing data set. Di Wu 0068, Juxiang Zhou, Jianhou Gan, Wei Gao 0012, Hao Li 0188 |
Int. J. Intell. Syst. | 4 |
| 2022 | Fine-Grained Image Classification Based on Cross-Attention NetworkabstractDue to the high similarity of fine-grained image subclasses, small inter-class changes and large intra-class changes are caused, which leads to the difficulty of fine-grained image classification task. However, existing convolutional neural networks have been unable to effectively solve this problem. Aiming at the above-mentioned fine-grained image classification problem, this paper proposes a multi-scale and multi-level ViT model. First, through data augmentation techniques, the accuracy of fine-grained image classification can be effectively improved. Secondly, the small-scale input and large-scale input of the model make the input image have more feature ex-pressions. The subsequent multi-layeredness effectively utilizes the results of the previous layer of ViT, so that the data of the previous layer can be more effectively used in the next layer of ViT. Finally, cross-attention allows the results of two scale inputs to be fused in a reasonable way. The proposed model is competitive with current mainstream state-of-the-art methods on multiple datasets. Juxiang Zhou, Jianhou Gan, Sen Luo, Wei Gao 0012 |
Int. J. Semantic Web Inf. Syst. | 3 |
| 2021 | Image retrieval based on aggregated deep features weighted by regional significance and channel sensitivity
Juxiang Zhou, Jianhou Gan, Wei Gao 0012, Antoni Liang |
Inf. Sci. | 2 |
| 2021 | Design and Research of Intelligent Question-Answering(Q&A) System Based on High School Course Knowledge Graph
Jianhou Gan, Ning Lei |
Mob. Networks Appl. | 3 |
| 2019 | Image retrieval based on effective feature extraction and diffusion process
Juxiang Zhou, Xiaodong Liu 0001, Wanquan Liu, Jianhou Gan |
Multim. Tools Appl. | 4 |
| 2019 | Learning Structural Representations via Dynamic Object Landmarks Discovery for Sketch Recognition and RetrievalabstractState-of-the-art methods on sketch classification and retrieval are based on deep convolutional neural network to learn representations. Although deep neural networks have the ability to model images with hierarchical representations by convolution kernels, they can not automatically extract the structural representations of object categories in a human-perceptible way. Furthermore, sketch images usually have large scale visual variations caused by the styles of drawing or viewpoints, which make it difficult to develop generalized representations using the fixed computational mode of convolutional kernel. In this paper, our aim is to address the problem of fixed computational mode in feature extraction process without extra supervision. We propose a novel architecture to dynamically discover the object landmarks and learn the discriminative structural representations. Our model is composed of two components: a representative landmark discovering module that localizes the key points on the object, and a category-aware representation learning module that develops the category-specific features. Specifically, we develop a structure-aware offset layer to dynamically localize the representative landmarks, which is optimized based on the category labels without extra supervision. After that, a diversity branch is introduced to extract the global discriminative features for each category. Finally, we employ a multi-task loss function to develop an end-to-end trainable architecture. At testing time, we fuse all the predictions with different number of landmarks to achieve the final results. Through extensive experiments, we compare our model with several state-of-the-art methods on two challenging datasets TU-Berlin and Sketchy for sketch classification and retrieval, and the experimental results demonstrate the effectiveness of our proposed model. Hua Zhang 0008, Peng She, Yong Liu 0018, Jianhou Gan, Xiaochun Cao, Hassan Foroosh |
IEEE Trans. Image Process. | 4 |
| 2018 | Multi-ethnical Chinese facial characterization and analysis
Cunrui Wang, Qingling Zhang 0001, Xiaodong Duan, Jianhou Gan |
Multim. Tools Appl. | 4 |
| 2015 | Improved Bacterial Foraging Optimization Algorithm with Information Communication Mechanism for Nurse Scheduling
Ben Niu 0002, Jing Liu 0029, Jianhou Gan, Lingyun Yuan |
ICIC (2) | 4 |