EDBT 2026 Demo / reviewers in the wild / expert
Ji Gan
dblp:227/7216
· DBLP profile ↗
29ranked-venue papers
10as first author
24since 2021 · last 2026
0000-0001-6041-588XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 6 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 11 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Transformer-Based 3-D Hand Pose Estimation via Bidirectional Multiscale Fusion and Learnable Anchor GuidanceabstractAccurate 3D hand pose estimation faces inherent challenges, including self-occlusions, joint similarities, and high degrees of freedom. Although most existing CNN-based or Transformer-based methods leverage global contexts, they often fail to capture fine-grained local details and robustly handle occluded joints. To address these limitations, we propose a novel Transformer framework for depth-based 3D hand pose estimation, which incorporates two key designs: the Bidirectional Multiscale Fusion and Learnable Anchor Guidance. Firstly, we propose bidirectional multiscale fusion that sequentially propa-gates features from different encoder levels in both top-down and bottom-up directions, followed by a final aggregation of all scale features. Such an intricate feature interaction eventually enables effective joint modeling of fine-grained local details (e.g., fingertip positions) and high-level semantic context information (e.g., palm orientation). Secondly, we introduce the learnable anchor query as prior guidance to dynamically guide the decoder to localize ambiguous joints better. The learnable anchors are derived from joint-specific attention maps under 3D ground-truth supervision and then are concatenated with static grid anchors to form hybrid anchors, which effectively enable more precise 3D hand pose estimation, especially for occlusions. To further demonstrate the practicality of our framework for IoT-oriented deployment, we conduct edge-device experiments to validate its deployment feasibility. Experiments on benchmark datasets (including NYU, ICVL, MSRA, and DexYCB) demonstrate the superiority of our method over previous state-of-the-art approaches. Ji Gan, Weiqiang Wang 0001, Feng Gao 0005, Jiaxu Leng, Haosheng Chen 0001, Xinbo Gao 0001 |
IEEE Internet Things J. | 2 |
| 2026 | PiercingEye: Dual-Space Video Violence Detection With Hyperbolic Vision-Language GuidanceabstractExisting weakly supervised video violence detection (VVD) methods primarily rely on Euclidean representation learning, which often struggles to distinguish visually similar yet semantically distinct events due to limited hierarchical modeling and insufficient ambiguous training samples. To address this challenge, we propose PiercingEye, a novel dual-space learning framework that synergizes Euclidean and hyperbolic geometries to enhance discriminative feature representation. Specifically, PiercingEye introduces a layer-sensitive hyperbolic aggregation strategy with hyperbolic Dirichlet energy constraints to progressively model event hierarchies, and a cross-space attention mechanism to facilitate complementary feature interactions between Euclidean and hyperbolic spaces. Furthermore, to mitigate the scarcity of ambiguous samples, we leverage large language models to generate logic-guided ambiguous event descriptions, enabling explicit supervision through a hyperbolic vision-language contrastive loss that prioritizes high-confusion samples via dynamic similarity-aware weighting. Extensive experiments on XD-Violence and UCF-Crime benchmarks demonstrate that PiercingEye achieves state-of-the-art performance, with particularly strong results on a newly curated ambiguous event subset, validating its superior capability in fine-grained violence detection. Jiaxu Leng, Zhanjie Wu, Mingpi Tan, Mengjingcheng Mo, Jiankang Zheng, Ji Gan, Xinbo Gao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | HandJoKe: Joint-Guided Keypoint Denoising Transformer for Depth-Based 3D Hand Pose EstimationabstractExisting depth-based 3D hand pose estimation methods typically estimate hand joints from either 2D depth images or 3D point clouds, whereas the approaches that fuse multimodal data remain underexplored. Furthermore, previous methods often struggle to learn geometric-facilitated features and precise joint correlations, especially for occluded hands, due to the lack of explicit prior guidance and insufficient cross-dimensional interaction. By taking advantage of multi-modal fusion, cross-dimensional interaction, and prior guidance, we propose a novel joint-guided keypoint denoising Transformer (named HandJoKe) to achieve more precise hand pose estimation, which can iteratively estimate hand poses based on keypoint features from both 2D depth images and 3D point clouds under explicit joint guidance within only several denoising steps. Rather than directly applying existing multi-modal fusion to perform redundant interactions among many background pixels and irrelevant points, HandJoKe focuses on modeling correlations and capturing dependencies among local informative hand regions (i.e., keypoints), thus attaining higher learning capability with lower computation redundancy. Moreover, a novel joint-guided denoising estimation strategy is introduced to adequately fuse cross-modal keypoint features under explicit joint guidance, achieving geometric-facilitated cross-modal keypoint interaction in both 2D and 3D spaces. The effectiveness of joint guidance can be further strengthened through iterative denoising, since it can subsequently update cross-modal keypoint features based on previous denoised hand poses and thus can help better locate confused joints, especially for occluded hands. Extensive experiments show that HandJoKe has achieved state-of-the-art performance on four public challenging benchmarks, including single-hand datasets NYU and ICVL, and hand-object datasets DexYCB and HO3D. Ji Gan, Jiaxu Leng, Weisheng Li 0001, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Transferring and Refining Visual-Semantic Priors via Graph-Enhanced CLIP for 3D Hand Pose Estimationabstract3D hand pose estimation is crucial for many human-computer interaction applications. However, existing deep neural networks (DNNs) for 3D hand pose estimation suffer from poor generalizability due to data scarcity and a lack of domain-specific knowledge. In contrast, humans remain far better than DNNs at learning; Humans require fewer samples for learning new concepts under the guidance of their prior knowledge. Inspired by this, we propose a graph-enhanced CLIP to deliver visual-semantic priors to DNNs, and provide refined domain-specific knowledge for better 3D hand pose estimation. Specifically, we first introduce a pre-trained CLIP to guide the hand estimation model in learning the semantic-aware visual features, and text-free contrastive learning is proposed to effectively transfer high-level visual-semantic priors from the pre-trained large multimodal models. Notably, our strategy is data-agnostic and avoids designing hand-crafted text prompts for various visual inputs. Second, we introduce novel graph Transformers to refine the domain-specific knowledge by fully exploiting the local adjacent relations of hand joints and capturing the global structure representations of hand poses. The introduced graph Transformers are supposed to further refine the generalized CLIP feature for the downstream task (i.e., hand pose estimation) with better performance. Experiments show that our proposed graph-enhanced CLIP achieves state-of-the-art performances on benchmark datasets, demonstrating its effectiveness for 3D hand pose estimation. The source code is available athttps://github.com/TLiu2832/TRVSP-GE-CLIP. Ji Gan, Jiaxu Leng, Xinbo Gao 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | Attention-Modulated Transformer for In-Air Handwriting Recognition
Ji Gan |
ICIC (15) | 3 |
| 2025 | Structure-Aware Handwritten Text Recognition via Graph-Enhanced Cross-Modal Mutual LearningabstractExisting handwriting recognition methods only focus on learning visual patterns by modeling low-level relationships of adjacent pixels, while overlooking the intrinsic geometric structures of characters. In this paper, we propose a novel graph-enhanced cross-modal mutual learning network GCM to fully process handwritten text images alongside their corresponding geometric graphs, which consists of one shared cross-modal encoder and two parallel inverse decoders. Specifically, the encoder simultaneously extracts visual and geometric information from the cross-modal inputs, and the decoders fuse the multi-modal features for prediction under the guidance of cross-modal fusion. Moreover, two parallel decoders sequentially aggregate cross-modal features in inverse orders (V→G and G→V) but are enhanced through mutual distillation at each time-step, which involves one-to-one knowledge transfer and fully leverages complementary cross-modal information from both directions. Notably, only one branch of GCM is activated in inference, thus avoiding the increase of the model parameters and computation costs for testing. Experiments show that our method outperforms previous state-of-the-art methods on public benchmarks such as IAM, RIMES, and ICDAR-2013 when no extra training data is utilized. Ji Gan, Yupeng Zhou, Jiaxu Leng, Xinbo Gao 0001 |
IJCAI | 1 |
| 2025 | A2Seek: Towards Reasoning-Centric Benchmark for Aerial Anomaly UnderstandingabstractWhile unmanned aerial vehicles (UAVs) offer wide-area, high-altitude coverage for anomaly detection, they face challenges such as dynamic viewpoints, scale variations, and complex scenes. Existing datasets and methods, mainly designed for fixed ground-level views, struggle to adapt to these conditions, leading to significant performance drops in drone-view scenarios.To bridge this gap, we introduce A2Seek (Aerial Anomaly Seek), a large-scale, reasoning-centric benchmark dataset for aerial anomaly understanding. This dataset covers various scenarios and environmental conditions, providing high-resolution real-world aerial videos with detailed annotations, including anomaly categories, frame-level timestamps, region-level bounding boxes, and natural language explanations for causal reasoning. Building on this dataset, we propose A2Seek-R1, a novel reasoning framework that generalizes R1-style strategies to aerial anomaly understanding, enabling a deeper understanding of “Where” anomalies occur and “Why” they happen in aerial frames.To this end, A2Seek-R1 first employs a graph-of-thought (GoT)-guided supervised fine-tuning approach to activate the model's latent reasoning capabilities on A2Seek. Then, we introduce Aerial Group Relative Policy Optimization (A-GRPO) to design rule-based reward functions tailored to aerial scenarios. Furthermore, we propose a novel “seeking” mechanism that simulates UAV flight behavior by directing the model's attention to informative regions.Extensive experiments demonstrate that A2Seek-R1 achieves up to a 22.04\% improvement in AP for prediction accuracy and a 13.9\% gain in mIoU for anomaly localization, exhibiting strong generalization across complex environments and out-of-distribution scenarios. Our dataset and code are released at https://2-mo.github.io/A2Seek/. Mengjingcheng Mo, Xinyang Tong, Mingpi Tan, Jiaxu Leng, Jiankang Zheng, Haosheng Chen 0001, Ji Gan, Weisheng Li 0001, Xinbo Gao 0001 |
NeurIPS | 8 |
| 2025 | Dual-Space Video Person Re-identification
Jiaxu Leng, Changjiang Kuang, Ji Gan, Haosheng Chen 0001, Xinbo Gao 0001 |
Int. J. Comput. Vis. | 4 |
| 2025 | Ranking-based adaptive query generation for DETRs in crowded pedestrian detection
Feng Gao 0005, Jiaxu Leng, Ji Gan, Xinbo Gao 0001 |
Neurocomputing | 3 |
| 2025 | RC-DETR: Improving DETRs in crowded pedestrian detection via rank-based contrastive learning
Feng Gao 0005, Jiaxu Leng, Ji Gan, Xinbo Gao 0001 |
Neural Networks | 3 |
| 2025 | GCapNet-FSD: A heterogeneous Graph Capsule Network for Few-Shot object Detection
Jiaxu Leng, Qianru Chen, Taiyue Chen, Feng Gao 0005, Ji Gan, Changjun Gu, Xinbo Gao 0001 |
Neural Networks | 5 |
| 2025 | Shape-centered representation learning for visible-infrared person re-identification
Jiaxu Leng, Ji Gan, Mengjingcheng Mo, Xinbo Gao 0001 |
Pattern Recognit. | 3 |
| 2025 | Dual-Space Normalizing Flow for Unsupervised Video Anomaly DetectionabstractConventional reconstruction-based video anomaly detection (VAD) methods implicitly model normality in latent spaces, which is limited by the generalization ability of latent features. Normalizing Flow (NF)-based methods have been introduced to address this issue, as they explicitly model the distribution of input data and achieve significant performance in VAD. However, existing NF-based methods are confined to Euclidean space, limiting their ability to model action hierarchies. While effective at capturing local joint dynamics and short-term temporal variations, they fail to encode kinematic dependencies and long-term pose evolution, ultimately struggling to discern ambiguous anomalies that deviate minimally from normal motion. In contrast, hyperbolic representation learning, with its ability to model hierarchical and complex relationships among actions, offers a promising solution to enhance the discriminative power between similar skeletal actions. Motivated by this, we propose a novel Dual-Space Normalizing Flow (DSNF) method. Specifically, we design a Dual-Space Parallel Graph Convolutional Network (DSPGCN) that synergistically integrates the strengths of both Euclidean and hyperbolic geometries to simultaneously capture local detail features of poses and intrinsic hierarchical relationships of actions. To enhance the model's focus on discriminative features, we design an Adaptive Weighted Approximation Mass (AWAM) loss that dynamically adjusts weights to impose stronger constraints on regions with low discriminability in the dual space, encouraging the model to focus more on key discriminative features in hyperbolic space that reflect complex relationships between actions. Extensive experiments on public datasets demonstrate the effectiveness and robustness of our method in various VAD scenarios. Jiaxu Leng, Mingpi Tan, Changjiang Kuang, Zhanjie Wu, Ji Gan, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 6 |
| 2025 | Difficulty-Guided Variant Degradation Learning for Blind Image Super-ResolutionabstractRecent blind super-resolution (BSR) methods are explored to handle unknown degradations and achieve impressive performance. However, the prevailing assumption in most BSR methods is the spatial invariance of degradation kernels across the entire image, which leads to significant performance declines when faced with spatially variant degradations caused by object motion or defocusing. Additionally, these methods do not account for the human visual system's tendency to focus differently on areas of varying perceptual difficulty, as they uniformly process each pixel during reconstruction. To cope with these issues, we propose a difficulty-guided variant degradation learning network for BSR, named difficulty-guided degradation learning (DDL)-BSR, which explores the relationship between reconstruction difficulty and degradation estimation. Accordingly, the proposed DDL-BSR consists of three customized networks: reconstruction difficulty prediction (RDP), space-variant degradation estimation (SDE), and degradation and difficulty-informed reconstruction (DDR). Specifically, RDP learns the reconstruction difficulty with the proposed reconstruction-distance supervision. Then, SDE is designed to estimate space-variant degradation kernels according to the difficulty map. Finally, both degradation kernels and reconstruction difficulty are fed into DDR, which takes into account such two prior knowledge information to guide super-resolution (SR). Experimental analysis on various synthetic datasets demonstrates that DDL-BSR invariably surpasses state-of-the-art (SOTA) methods, producing SR images with enhanced realism and texture quality. Code is available at https://github.com/JiaWang0704/DDL-BSR. Jiaxu Leng, Jia Wang 0036, Mengjingcheng Mo, Ji Gan, Wen Lu 0004, Xinbo Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Structure-Aware in-Air Handwritten Text Recognition with Graph-Guided Cross-Modality TranslatorabstractIn-air handwriting as a new human-computer interaction way plays an important role in many virtual/mixed-reality applications. Existing methods for in-air handwritten text recognition (IAHTR) typically directly process handwriting trajectories with deep neural networks. However, those methods all simply learn discriminative patterns by modelling low-level relationships between adjacent points of trajectories, while completely ignoring the inherent geometric structures of characters. Instead, we propose a novel Graph-guided Cross-modality Translator for IAHTR, which further explicitly exploits the geometric structures of characters for guiding the decoding of trajectories via graph-guided cross-modality attention mechanism without introducing extra annotation costs. Experiments on benchmarks IAHEW-UCAS2016 & IAM-OnDB show that our method has achieved state-of-the-art performance for handwritten text recognition. Yuyan Chen, Ji Gan, Jiaxu Leng, Yan Zhang 0108, Xinbo Gao 0001 |
ICASSP | 3 |
| 2024 | Dual Space Embedding Learning For Weakly Supervised Audio-Visual Violence DetectionabstractIn this paper, we propose Dual Space Embedding Learning (DSEL) for weakly supervised audio-visual violence detection, which excavates violence information deeply in both Euclidean and Hyperbolic spaces to distinguish violence from non-violence semantically and alleviate the asynchronous issue of violent cues in audio-visual patterns. Specifically, we first design a dual space visual feature interaction module (DSVFI) to deeply investigate the violence information in visual modality, which contains richer information compared to audio counterpart. Then, considering the modality asynchrony between the two modalities, we employ a late modality fusion method and design an asynchrony-aware audio-visual fusion module (AAF), in which visual features receive the violent prompt from the audio features after interacting among snippets and learning the violence information from each other. Experimental results show that our method achieves state-of-the-art performance on XD-Violence. Zhanjie Wu, Mengjingcheng Mo, Ji Gan, Jiaxu Leng, Xinbo Gao 0001 |
ICME | 4 |
| 2024 | Beyond Euclidean: Dual-Space Representation Learning for Weakly Supervised Video Violence DetectionabstractWhile numerous Video Violence Detection (VVD) methods have focused on representation learning in Euclidean space, they struggle to learn sufficiently discriminative features, leading to weaknesses in recognizing normal events that are visually similar to violent events (i.e., ambiguous violence). In contrast, hyperbolic representation learning, renowned for its ability to model hierarchical and complex relationships between events, has the potential to amplify the discrimination between visually similar events. Inspired by these, we develop a novel Dual-Space Representation Learning (DSRL) method for weakly supervised VVD to utilize the strength of both Euclidean and hyperbolic geometries, capturing the visual features of events while also exploring the intrinsic relations between events, thereby enhancing the discriminative capacity of the features. DSRL employs a novel information aggregation strategy to progressively learn event context in hyperbolic spaces, which selects aggregation nodes through layer-sensitive hyperbolic association degrees constrained by hyperbolic Dirichlet energy. Furthermore, DSRL attempts to break the cyber-balkanization of different spaces, utilizing cross-space attention to facilitate information interactions between Euclidean and hyperbolic space to capture better discriminative features for final violence detection. Comprehensive experiments demonstrate the effectiveness of our proposed DSRL. Jiaxu Leng, Zhanjie Wu, Mingpi Tan, Ji Gan, Haosheng Chen 0001, Xinbo Gao 0001 |
NeurIPS | 5 |
| 2024 | Blind Image Quality Index With Cross-Domain Interaction and Cross-Scale IntegrationabstractWith the assistance of Convolutional Neural Networks (CNNs), Image Quality Assessment (IQA) models have made great progress in evaluating both simulated distortion and authentic distortion. However, most of the existing IQA models only learn the features of distorted images, and thus do not make full use of the available feature representation of other domains. Furthermore, the common multi-scale fusion strategies are relatively simple, such as downsampling and concatenating, which further limits the prediction performance. To this end, we propose a novel blind image quality index with cross-domain interaction and cross-scale integration, which is designed based on the combination of CNN and Transformer. First, the hierarchical spatial-domain and gradient-domain representations are obtained through a typical CNN architecture. Then, based on the proposed gradient-query cross-attention, these two types of features are fully interacted in the Cross-Domain Interaction (CDI) module. To represent the distortion information more comprehensively, the Cross-Scale Integration (CSI) module is proposed to combine the information between different scales progressively. Finally, the quality score is obtained through a simple regression module. The experimental results on five public IQA databases of both simulated and authentic scenes show that the proposed model outperforms the compared state-of-the-art metrics. In addition, cross-database experiments show that the proposed model has strong generalization performance. Bo Hu 0008, Leida Li, Ji Gan, Weisheng Li 0001, Xinbo Gao 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Selecting Learnable Training Samples is All DETRs Need in Crowded Pedestrian DetectionabstractDEtection TRansformer (DETR) and its variants (DETRs) achieved impressive performance in general object detection. However, in crowded pedestrian detection, the performance of DETRs is still unsatisfactory due to the inappropriate sample selection method which results in more false positives. To settle the issue, we propose a simple but effective sample selection method for DETRs, Sample Selection for Crowded Pedestrians (SSCP), which consists of the constraint-guided label assignment scheme (CGLA) and the utilizability-aware focal loss (UAFL). Our core idea is to select learnable samples for DETRs and adaptively regulate the loss weights of samples based on their utilizability. Specifically, in CGLA, we proposed a new cost function to ensure that only learnable positive training samples are retained and the rest are negative training samples. Further, considering the utilizability of samples, we designed UAFL to adaptively assign different loss weights to learnable positive samples depending on their gradient ratio and IoU. Experimental results show that the proposed SSCP effectively improves the baselines without introducing any overhead in inference. Especially, Iter Deformable DETR is improved to 39.7(-2.0)% MR on Crowdhuman and 31.8(-0.4)% MR on Citypersons. Feng Gao 0005, Jiaxu Leng, Ji Gan, Xinbo Gao 0001 |
ACM Multimedia | 3 |
| 2023 | Reduced-reference image deblurring quality assessment based on multi-scale feature enhancement and aggregation
Bo Hu 0008, Shuaijian Wang, Xinbo Gao 0001, Leida Li, Ji Gan, Xixi Nie |
Neurocomputing | 5 |
| 2023 | Characters as graphs: Interpretable handwritten Chinese character recognition via Pyramid Graph Transformer
Ji Gan, Yuyan Chen, Bo Hu 0008, Jiaxu Leng, Weiqiang Wang 0001, Xinbo Gao 0001 |
Pattern Recognit. | 1 |
| 2023 | HiGAN+: Handwriting Imitation GAN with Disentangled RepresentationsabstractHumans remain far better than machines at learning, where humans require fewer examples to learn new concepts and can use those concepts in richer ways. Take handwriting as an example, after learning from very limited handwriting scripts, a person can easily imagine what the handwritten texts would like with other arbitrary textual contents (even for unseen words or texts). Moreover, humans can also hallucinate to imitate calligraphic styles from just a single reference handwriting sample (that even have never seen before). Humans can do such hallucinations, perhaps because they can learn to disentangle the textual contents and calligraphic styles from handwriting images. Inspired by this, we propose a novel handwriting imitation generative adversarial network (HiGAN+) for realistic handwritten text synthesis based on disentangled representations. The proposed HiGAN+ can achieve a precise one-shot handwriting style transfer by introducing the writer-specific auxiliary loss and contextual loss, and it also attains a good global & local consistency by refining local details of synthetic handwriting images. Extensive experiments, including human evaluations, on the benchmark dataset validate our superiority in terms of visual quality, scalability, compactness, and style transferability compared with the state-of-the-art GANs for handwritten text synthesis. Ji Gan, Weiqiang Wang 0001, Jiaxu Leng, Xinbo Gao 0001 |
ACM Trans. Graph. | 1 |
| 2022 | ICNet: Joint Alignment and Reconstruction via Iterative Collaboration for Video Super-ResolutionabstractMost previous frameworks either cost too much time or adopt some fixed modules resulting in alignment error in video super-resolution (VSR). In this paper, we propose a novel many-to-many VSR framework with Iterative Collaboration (ICNet), which employs the concurrent operation by iterative collaboration between alignment and reconstruction proving to be more efficient and effective than existing recurrent and sliding-window frameworks. With the proposed iterative collaboration, alignment can be conducted on super-resolved features from reconstruction while accurate alignment boosts reconstruction in return. In each iteration, the features of low-resolution video frames are first fed into the alignment and reconstruction subnetworks, which can generate temporal aligned features and spatial super-resolved features. Then, both outputs are fed into the proposed Tidy Two-stream Fusion (TTF) subnetwork that shares inter-frame temporal information and intra-frame spatial information without redundancy. Moreover, we design the Frequency Separation Reconstruction (FSR) subnetwork to not only model high-frequency and low-frequency information separately but also take benefit of each other for better reconstruction. Extensive experiments on benchmark datasets demonstrate that the proposed ICNet outperforms state-of-the-art VSR methods in terms of PSNR/SSIM values and visual quality, respectively. Jiaxu Leng, Jia Wang 0036, Xinbo Gao 0001, Bo Hu 0008, Ji Gan, Chenqiang Gao |
ACM Multimedia | 5 |
| 2021 | HiGAN: Handwriting Imitation Conditioned on Arbitrary-Length Texts and Disentangled StylesabstractGiven limited handwriting scripts, humans can easily visualize (or imagine) what the handwritten words/texts would look like with other arbitrary textual contents. Moreover, a person also is able to imitate the handwriting styles of provided reference samples. Humans can do such hallucinations, perhaps because they can learn to disentangle the calligraphic styles and textual contents from given handwriting scripts. However, computers cannot study to do such flexible handwriting imitation with existing techniques. In this paper, we propose a novel handwriting imitation generative adversarial network (HiGAN) to mimic such hallucinations. Specifically, HiGAN can generate variable-length handwritten words/texts conditioned on arbitrary textual contents, which are unconstrained to any predefined corpus or out-of-vocabulary words. Moreover, HiGAN can flexibly control the handwriting styles of synthetic images by disentangling calligraphic styles from the reference samples. Experiments on handwriting benchmarks validate our superiority in terms of visual quality and scalability when comparing to the state-of-the-art methods for handwritten word/text synthesis. The code and pre-trained models can be found at https://github.com/ganji15/HiGAN. Ji Gan, Weiqiang Wang 0001 |
AAAI | 1 |
| 2020 | In-air handwritten Chinese text recognition with temporal convolutional recurrent network
Ji Gan, Weiqiang Wang 0001, Ke Lu 0002 |
Pattern Recognit. | 1 |
| 2020 | Compressing the CNN architecture for in-air handwritten Chinese character recognition
Ji Gan, Weiqiang Wang 0001, Ke Lu 0002 |
Pattern Recognit. Lett. | 1 |
| 2019 | A new perspective: Recognizing online handwritten Chinese characters via 1-dimensional CNN
Ji Gan, Weiqiang Wang 0001, Ke Lu 0002 |
Inf. Sci. | 1 |
| 2019 | In-air handwritten English word recognition using attention recurrent translator
Ji Gan, Weiqiang Wang 0001 |
Neural Comput. Appl. | 1 |
| 2018 | A Unified CNN-RNN Approach for in-Air Handwritten English Word RecognitionabstractAs a new human-computer interaction application, in-air handwriting allows the user to write in the air in a natural way. In this paper, we propose a unified CNN-RNN approach for in-air handwritten English word recognition (IAHEWR), which integrates the advantages of both convolutional neural networks (CNNs) and recurrent neural networks (RNNs). Specifically, the proposed approach follows an encoder-decoder framework, where the encoder is a deep CNN for efficiently processing the input temporal-sequential features, and the decoder is a RNN for accurately generating the target character sequence. We evaluate the proposed approach on an in-air handwritten English word dataset IAHEW-UCAS2016, and the experimental results demonstrate that the proposed approach achieves the comparable recognition accuracy and much higher computation efficiency when compared with the state-of-the-art approach for IAHEWR. Ji Gan, Weiqiang Wang 0001, Ke Lu 0002 |
ICME | 1 |