Zhilong Ji

dblp:263/6772 · also Zhi-Long Ji · DBLP profile ↗
← Back
32ranked-venue papers
0as first author
32since 2021 · last 2026
0000-0002-8799-3409ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 23 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
YearPublicationVenuePosition
2026 VitaGlyph: Vitalizing Artistic Typography with Flexible Dual-branch Diffusion Models
abstract
Artistic typography is a technique that enables one to visualize the meaning of an input character in an imaginable and readable manner. With powerful text-to-image diffusion models, existing methods directly design the overall geometry and texture of input character, making it challenging to ensure both creativity and legibility. In this paper, we introduce a dual-branch, training-free method called VitaGlyph, enabling flexible artistic typography with controllable geometry changes while maintaining legibility. The key insight of VitaGlyph is to treat the input character as a scene composed of a Subject and its Surrounding, which are rendered with varying degrees of geometric transformation. To enhance the visual appeal and creativity of the generated artistic typography, the Subject flexibly expresses the essential concept of the input character, while the Surrounding enriches relevant background without altering the shape. Specifically, we implement VitaGlyph through a three-phase framework: (i) Knowledge Acquisition leverages large language models to design text descriptions for the Subject and Surrounding. (ii) Regional Interpretation detects the part that matches the subject description most closely and refines the structure using Semantic Typography. (iii) Attentional Compositional Generation separately renders the textures of the Subject and Surrounding and blends them in an attention-based manner. Experiments demonstrate that VitaGlyph not only achieves better artistry and legibility, but also manages to depict multiple customized concepts, facilitating more creative and pleasing artistic typography generation. Our code is available at https://github.com/Carlofkl/VitaGlyph.
Kailai Feng, Yabo Zhang, Haodong Yu, Zhilong Ji, Jinfeng Bai, Wangmeng Zuo
WACV4
2025 Explicit Relational Reasoning Network for Scene Text Detection
abstract
Connected component (CC) is a proper text shape representation that aligns with human reading intuition. However, CC-based text detection methods have recently faced a developmental bottleneck that their time-consuming post-processing is difficult to eliminate. To address this issue, we introduce an explicit relational reasoning network (ERRNet) to elegantly model the component relationships without post-processing. Concretely, we first represent each text instance as multiple ordered text components, and then treat these components as objects in sequential movement. In this way, scene text detection can be innovatively viewed as a tracking problem. From this perspective, we design an end-to-end tracking decoder to achieve a CC-based method dispensing with post-processing entirely. Additionally, we observe that there is an inconsistency between classification confidence and localization quality, so we propose a Polygon Monte-Carlo method to quickly and accurately evaluate the localization quality. Based on this, we introduce a position-supervised classification loss to guide the task-aligned learning of ERRNet. Experiments on challenging benchmarks demonstrate the effectiveness of our ERRNet. It consistently achieves state-of-the-art accuracy while holding highly competitive inference speed.
Zhineng Chen, Yongkun Du, Zhilong Ji, Kai Hu 0002, Jinfeng Bai, Xieping Gao 0001
AAAI4
2025 Enhancing Multimodal Continual Instruction Tuning with BranchLoRA
abstract
Multimodal Continual Instruction Tuning (MCIT) aims to finetune Multimodal Large Language Models (MLLMs) to continually align with human intent across sequential tasks. Existing approaches often rely on the Mixture-of-Experts (MoE) LoRA framework to preserve previous instruction alignments. However, these methods are prone to Catastrophic Forgetting (CF), as they aggregate all LoRA blocks via simple summation, which compromises performance over time. In this paper, we identify a critical parameter inefficiency in the MoELoRA framework within the MCIT context. Based on this insight, we propose BranchLoRA, an asymmetric framework to enhance both efficiency and performance. To mitigate CF, we introduce a flexible tuning-freezing mechanism within BranchLoRA, enabling branches to specialize in intra-task knowledge while fostering inter-task collaboration. Moreover, we incrementally incorporate task-specific routers to ensure an optimal branch distribution over time, rather than favoring the most recent task. To streamline inference, we introduce a task selector that automatically routes test inputs to the appropriate router without requiring task identity. Extensive experiments on the latest MCIT benchmark demonstrate that BranchLoRA significantly outperforms MoELoRA and maintains its superiority across various MLLM sizes.
Duzhen Zhang, Yong Ren 0006, Zhongzhi Li, Yahan Yu, Jiahua Dong 0001, Chenxing Li, Zhilong Ji, Jinfeng Bai
ACL (1)7
2025 CMMaTH: A Chinese Multi-modal Math Skill Evaluation Benchmark for Foundation Models
abstract
With the rapid advancements in multimodal large language models, evaluating their multimodal mathematical capabilities continues to receive wide attention. Although datasets such as MathVista have been introduced for evaluating mathematical capabilities in multimodal scenarios, there remains a lack of evaluation tools and datasets tailored for fine-grained assessment in Chinese K12 education. To systematically evaluate the ability of multimodal large models to solve Chinese multimodal mathematical problems, we propose a Chinese Multi-modal Math Skill Evaluation Benchmark (CMMaTH), containing 23,856 multimodal K12 math related questions, making it the largest Chinese multimodal mathematical problem benchmark to date. CMMaTH includes questions ranging from elementary to high school levels, offering greater diversity in problem types, solution goals, visual elements, detailed knowledge points, and standard solution annotations. To facilitate stable, fast, and cost-free model evaluation, we have developed an open-source tool called GradeGPT, which is integrated with the CMMaTH dataset. Our data and code are available at https://github.com/zzli2022/CMMaTH.
Zhongzhi Li, Mingliang Zhang 0005, Pei-Jie Wang, Jian Xu 0027, Rui-Song Zhang, Yin Fei, Zhilong Ji, Jinfeng Bai, Zhenru Pan
COLING7
2025 Integrating Visual Interpretation and Linguistic Reasoning for Geometric Problem Solving
Zixian Guo, Ming Liu 0018, Qilong Wang 0001, Zhilong Ji, Jinfeng Bai, Lei Zhang 0006, Wangmeng Zuo
ICCV4
2025 SolidGeo: Measuring Multimodal Spatial Math Reasoning in Solid Geometry
abstract
Geometry is a fundamental branch of mathematics and plays a crucial role in evaluating the reasoning capabilities of multimodal large language models (MLLMs). However, existing multimodal mathematics benchmarks mainly focus on plane geometry and largely ignore solid geometry, which requires spatial reasoning and is more challenging than plane geometry. To address this critical gap, we introduce SolidGeo, the first large-scale benchmark specifically designed to evaluate the performance of MLLMs on mathematical reasoning tasks in solid geometry. SolidGeo consists of 3,113 real-world K–12 and competition-level problems, each paired with visual context and annotated with difficulty levels and fine-grained solid geometry categories. Our benchmark covers a wide range of 3D reasoning subjects such as projection, unfolding, spatial measurement, and spatial vector, offering a rigorous testbed for assessing solid geometry. Through extensive experiments, we observe that MLLMs encounter substantial challenges in solid geometry math tasks, with a considerable performance gap relative to human capabilities on SolidGeo. Moreover, we analyze the performance, inference effiency and error patterns of various models, offering insights into the solid geometric mathematical reasoning capabilities of MLLMs. We hope SolidGeo serves as a catalyst for advancing MLLMs toward deeper geometric reasoning and spatial intelligence. The dataset is released at https://huggingface.co/datasets/HarryYancy/SolidGeo/
Zhongzhi Li, Dekang Ran, Zhilong Ji, Jinfeng Bai, Cheng-Lin Liu 0001
NeurIPS7
2025 Personalized Image Generation with Deep Generative Models: A Decade Survey
abstract
Recent advances in generative models have significantly facilitated the development of personalized content creation. Given a small set of images containing a user-specific concept, personalized image generation allows the user to create images that incorporate that concept while adhering to provided text descriptions. The technologies used for personalization have evolved alongside the development of generative models, with their distinct and interrelated components. In this survey, we present a comprehensive review of generalized personalized image generation across various generative models, including traditional GANs, contemporary text-to-image diffusion models, and emerging multi-modal autoregressive (AR) models. We first define a unified framework that standardizes the personalization process across different generative models, encompassing three key components: inversion spaces, inversion methods, and personalization schemes. This unified framework offers a structured approach to dissecting and comparing personalization techniques across different generative architectures. Building upon our framework, we provide an in-depth analysis of personalization techniques within each generative model, highlighting their unique contributions and innovations. Through comparative analysis, we elucidate the current landscape of personalized image generation, identifying commonalities and distinguishing features of existing methods. Finally, we discuss open challenges in the field and propose potential directions for future research. We keep a bibliography of related works at https://github.com/csyxwei/Awesome-Personalized-Image-Generation.
Yuxiang Wei 0001, Yiheng Zheng, Yabo Zhang, Ming Liu 0018, Zhilong Ji, Lei Zhang 0006, Wangmeng Zuo
Comput. Vis. Media5
2025 Real face foundation representation learning for generalized deepfake detection
Liang Shi 0002, Jie Zhang 0071, Zhilong Ji, Jinfeng Bai, Shiguang Shan
Pattern Recognit.3
2024 Decoupled Textual Embeddings for Customized Image Generation
abstract
Customized text-to-image generation, which aims to learn user-specified concepts with a few images, has drawn significant attention recently. However, existing methods usually suffer from overfitting issues and entangle the subject-unrelated information (e.g., background and pose) with the learned concept, limiting the potential to compose concept into new scenes. To address these issues, we propose the DETEX, a novel approach that learns the disentangled concept embedding for flexible customized text-to-image generation. Unlike conventional methods that learn a single concept embedding from the given images, our DETEX represents each image using multiple word embeddings during training, i.e., a learnable image-shared subject embedding and several image-specific subject-unrelated embeddings. To decouple irrelevant attributes (i.e., background and pose) from the subject embedding, we further present several attribute mappers that encode each image as several image-specific subject-unrelated embeddings. To encourage these unrelated embeddings to capture the irrelevant information, we incorporate them with corresponding attribute words and propose a joint training strategy to facilitate the disentanglement. During inference, we only use the subject embedding for image generation, while selectively using image-specific embeddings to retain image-specified attributes. Extensive experiments demonstrate that the subject embedding obtained by our method can faithfully represent the target concept, while showing superior editability compared to the state-of-the-art methods. Our code will be available at https://github.com/PrototypeNx/DETEX.
Yufei Cai, Yuxiang Wei 0001, Zhilong Ji, Jinfeng Bai, Hu Han 0001, Wangmeng Zuo
AAAI3
2024 LRANet: Towards Accurate and Efficient Scene Text Detection with Low-Rank Approximation Network
abstract
Recently, regression-based methods, which predict parameterized text shapes for text localization, have gained popularity in scene text detection. However, the existing parameterized text shape methods still have limitations in modeling arbitrary-shaped texts due to ignoring the utilization of text-specific shape information. Moreover, the time consumption of the entire pipeline has been largely overlooked, leading to a suboptimal overall inference speed. To address these issues, we first propose a novel parameterized text shape method based on low-rank approximation. Unlike other shape representation methods that employ data-irrelevant parameterization, our approach utilizes singular value decomposition and reconstructs the text shape using a few eigenvectors learned from labeled text contours. By exploring the shape correlation among different text contours, our method achieves consistency, compactness, simplicity, and robustness in shape representation. Next, we propose a dual assignment scheme for speed acceleration. It adopts a sparse assignment branch to accelerate the inference speed, and meanwhile, provides ample supervised signals for training through a dense assignment branch. Building upon these designs, we implement an accurate and efficient arbitrary-shaped text detector named LRANet. Extensive experiments are conducted on several challenging benchmarks, demonstrating the superior accuracy and efficiency of LRANet compared to state-of-the-art methods. Code is available at: https://github.com/ychensu/LRANet.git
Zhineng Chen, Zhiwen Shao, Yuning Du, Zhilong Ji, Jinfeng Bai, Yong Zhou 0003, Yu-Gang Jiang 0001
AAAI5
2024 CK12: A Rounded K12 Knowledge Graph Based Benchmark for Chinese Holistic Cognition Evaluation
abstract
New NLP benchmarks are urgently needed to align with the rapid development of large language models (LLMs). We present a meticulously designed evaluation benchmark that leverages the knowledge graph. This evaluation comprises 584 level-1 knowledge points and 1,989 level-2 knowledge points, thereby encompassing a comprehensive spectrum of the K12 education domain knowledge. The primary objective is to comprehensively assess the high-level comprehension aptitude and reasoning capabilities of LLMs operating within the Chinese context. Our evaluation incorporates five distinct question types with 39,452 questions. We test the current mainstream LLMs by three distinct modes. Firstly, four prompt evaluation modes were employed to assess the fundamental capacity. Additionally, for choice questions, a result-oriented evaluation approach was designed through data augmentation to assess the model's proficiency in advanced knowledge and reasoning. Moreover, a subset with reasoning process is derived, and the process-oriented testing method is used to test the model's interpretability and higher-order reasoning capacity. We further show models' capability in our knowledge points, and anticipate the evaluation can assist in the assessment of the strengths and deficiencies of LLMs on knowledge points, thus fostering their development within the Chinese context. Our Dataset will be publicly available in https://github.com/tal-tech/chinese-k12-evaluation.
Weihao You, Zhilong Ji, Jinfeng Bai
AAAI4
2024 HPNet: Dynamic Trajectory Forecasting with Historical Prediction Attention
abstract
Predicting the trajectories of road agents is essential for autonomous driving systems. The recent mainstream methods follow a static paradigm, which predicts the future trajectory by using a fixed duration of historical frames. These methods make the predictions independently even at adjacent time steps, which leads to potential instability and temporal inconsistency. As successive time steps have largely overlapping historical frames, their forecasting should have intrinsic correlation, such as overlapping predicted trajectories should be consistent, or be different but share the same motion goal depending on the road situation. Motivated by this, in this work, we introduce HPNet, a novel dynamic trajectory forecasting method. Aiming for stable and accurate trajectory forecasting, our method leverages not only historical frames including maps and agent states, but also historical predictions. Specifically, we newly design a Historical Prediction Attention module to automatically encode the dynamic relationship between successive predictions. Besides, it also extends the attention range beyond the currently visible window benefitting from the use of historical predictions. The proposed Historical Prediction Attention together with the Agent Attention and Mode Attention is further formulated as the Triple Factorized Attention module, serving as the core design of HPNet. Experiments on the Argoverse and INTERACTION datasets show that HP-Net achieves state-of-the-art performance, and generates accurate and stable future trajectories. Our code are available at https://github.com/XiaolongTang23/HPNet.
Meina Kan, Shiguang Shan, Zhilong Ji, Jinfeng Bai, Xilin Chen 0001
CVPR4
2024 MasterWeaver: Taming Editability and Face Identity for Personalized Text-to-Image Generation
Yuxiang Wei 0001, Zhilong Ji, Jinfeng Bai, Lei Zhang 0006, Wangmeng Zuo
ECCV (51)2
2024 MuMath-Code: Combining Tool-Use Large Language Models with Multi-perspective Data Augmentation for Mathematical Reasoning
abstract
The tool-use Large Language Models (LLMs) that integrate with external Python interpreters have significantly enhanced mathematical reasoning capabilities for open-source LLMs, while tool-free methods chose another track: augmenting math reasoning data.However, a great method to integrate the above two research paths and combine their advantages remains to be explored.In this work, we firstly include new math questions via multi-perspective data augmenting methods and then synthesize code-nested solutions to them.The open LLMs (e.g., Llama-2) are finetuned on the augmented dataset to get the resulting models, MuMath-Code (µ-Math-Code).During the inference phase, our MuMath-Code generates code and interacts with the external python interpreter to get the execution results.Therefore, MuMath-Code leverages the advantages of both the external tool and data augmentation.To fully leverage the advantages of our augmented data, we propose a two-stage training strategy: In Stage-1, we finetune Llama-2 on pure CoT data to get an intermediate model, which then is trained on the code-nested data in Stage-2 to get the resulting MuMath-Code.Our MuMath-Code-7B achieves 83.8% on GSM8K and 52.4% on MATH, while MuMath-Code-70B model achieves new state-of-the-art performance among open methods-achieving 90.7% on GSM8K and 55.1% on MATH.Extensive experiments validate the combination of tool use and data augmentation, as well as our two-stage training strategy.We release the proposed dataset along with the associated code for public use: https://github.com/ youweihao-tal/MuMath-Code.
Weihao You, Zhilong Ji, Guoqiang Zhong 0001, Jinfeng Bai
EMNLP3
2024 DPA-2D: Depth Propagation and Alignment with 2D Observations Guidance for Human Mesh Recovery
abstract
In the current state of 3D human pose and shape estimation, the two-stage methods involving 3D joints estimation and inverse kinematics (IK) to obtain more accurate human pose, have gained significant traction. However, the effectiveness of the final optimized depth is greatly influenced by the initial 3D joints estimation. The 3D joints errors primarily caused by depth estimation. Our findings indicate that 2D observations (2D joint locations and Confidence) offer valuable supplementary data. Depth estimation errors are greater for low-confidence joints, but high-confidence joints yield more accuracy. In light of this, we propose a Biased Depth Propagation (BDP) module that leverages the confidence of 2D joints to enhance depth estimation. This module enables us to improve the depth estimation of low-confidence joints by propagating the depth information from high-confidence joints. Furthermore, small variations in the SMPL parameters can significantly impact the reconstructed meshes causing misalignment with image evidence. Therefore, we propose the Refined Joint Aligned (RJA) module to fine-tune the SMPL parameters by utilizing enhanced joint ROI features captured by the estimated 2D joint locations. Finally, we have validated the performance of our method on popular datasets such as Human3.6m and 3DPW for human mesh recovery task, and achieved SOTA results, demonstrate the effectiveness of our Depth Propagation and Alignment with 2D Observations Guidance(DPA-2D).
Weihao You, Zhilong Ji, Jinfeng Bai
FG3
2024 Collaborative Domain Alignment for Multi-source Domain Adaptation
Meina Kan, Zhilong Ji, Jinfeng Bai, Shiguang Shan, Xilin Chen 0001
ICPR (27)3
2023 ReCoT: Regularized Co-Training for Facial Action Unit Recognition with Noisy Labels
Hu Han 0001, Shiguang Shan, Zhilong Ji, Jinfeng Bai, Xilin Chen 0001
BMVC4
2023 Inferring and Leveraging Parts from Object Shape for Improving Semantic Image Synthesis
abstract
Despite the progress in semantic image synthesis, it remains a challenging problem to generate photo-realistic parts from input semantic map. Integrating part segmentation map can undoubtedly benefit image synthesis, but is bothersome and inconvenient to be provided by users. To improve part synthesis, this paper presents to infer Parts from Object ShapE (iPOSE) and leverage it for improving semantic image synthesis. However, albeit several part segmentation datasets are available, part annotations are still not provided for many object categories in semantic image synthesis. To circumvent it, we resort to few-shot regime to learn a PartNet for predicting the object part map with the guidance of pre-defined support part maps. PartNet can be readily generalized to handle a new object category when a small number (e.g., 3) of support part maps for this category are provided. Furthermore, part semantic modulation is presented to incorporate both inferred part map and semantic map for image synthesis. Experiments show that our iPOSE not only generates objects with rich part details, but also enables to control the image synthesis flexibly. And our iPOSE performs favorably against the state-of-the-art methods in terms of quantitative and qualitative evaluation. Our code will be publicly available at https://github.com/csyxwei/iPOSE.
Yuxiang Wei 0001, Zhilong Ji, Xiaohe Wu, Jinfeng Bai, Lei Zhang 0006, Wangmeng Zuo
CVPR2
2023 Texts as Images in Prompt Tuning for Multi-Label Image Recognition
abstract
Prompt tuning has been employed as an efficient way to adapt large vision-language pre-trained models (e.g. CLIP) to various downstream tasks in data-limited or label-limited settings. Nonetheless, visual data (e.g., images) is by default prerequisite for learning prompts in existing methods. In this work, we advocate that the effectiveness of image-text contrastive learning in aligning the two modalities (for training CLIP) further makes it feasible to treat texts as images for prompt tuning and introduce TaI prompting. In contrast to the visual data, text descriptions are easy to collect, and their class labels can be directly derived. Particularly, we apply TaI prompting to multi-label image recognition, where sentences in the wild serve as alternatives to images for prompt tuning. Moreover, with TaI, double-grained prompt tuning (TaI-DPT) is further presented to extract both coarse-grained and fine-grained embeddings for enhancing the multi-label recognition performance. Experimental results show that our proposed TaI-DPT outperforms zero-shot CLIP by a large margin on multiple benchmarks, e.g., MS-COCO, VOC2007, and NUS-WIDE, while it can be combined with existing methods of prompting from images to improve recognition performance further. The code is released at https://github.com/guozix/TaI-DPT.
Zixian Guo, Bowen Dong 0001, Zhilong Ji, Jinfeng Bai, Yiwen Guo, Wangmeng Zuo
CVPR3
2023 ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image Generation
abstract
In addition to the unprecedented ability in imaginary creation, large text-to-image models are expected to take customized concepts in image generation. Existing works generally learn such concepts in an optimization-based manner, yet bringing excessive computation or memory burden. In this paper, we instead propose a learning-based encoder, which consists of a global and a local mapping networks for fast and accurate customized text-to-image generation. In specific, the global mapping network projects the hierarchical features of a given image into multiple "new" words in the textual word embedding space, i.e., one primary word for well-editable concept and other auxiliary words to exclude irrelevant disturbances (e.g., background). In the meantime, a local mapping network injects the encoded patch features into cross attention layers to provide omitted details, without sacrificing the editability of primary concepts. We compare our method with existing optimization-based approaches on a variety of user-defined concepts, and demonstrate that our method enables high-fidelity inversion and more robust editability with a significantly faster encoding process. Our code is publicly available at https://github.com/csyxwei/ELITE.
Yuxiang Wei 0001, Yabo Zhang, Zhilong Ji, Jinfeng Bai, Lei Zhang 0006, Wangmeng Zuo
ICCV3
2023 Semantic Graph Representation Learning for Handwritten Mathematical Expression Recognition
Zhilong Ji, Jinfeng Bai, Xiang Bai
ICDAR (1)3
2023 ViSA: Visual and Semantic Alignment for Robust Scene Text Recognition
Zhenru Pan, Zhilong Ji, Xiao Liu 0040, Jinfeng Bai, Cheng-Lin Liu 0001
ICDAR (2)2
2023 Decoupling Visual-Semantic Features Learning with Dual Masked Autoencoder for Self-Supervised Scene Text Recognition
Zhilong Ji, Jinfeng Bai
ICDAR (2)2
2023 CCLAP: Controllable Chinese Landscape Painting Generation Via Latent Diffusion Model
abstract
With the development of deep generative models, recent years have seen great success of Chinese landscape painting generation. However, few works focus on controllable Chinese landscape painting generation due to the lack of data and limited modeling capabilities. In this work, we propose a controllable Chinese landscape painting generation method named CCLAP, which can generate painting with specific content and style based on Latent Diffusion Model. Specifically, it consists of two cascaded modules, i.e., content generator and style aggregator. The content generator module guarantees the content of generated paintings specific to the input text. While the style aggregator module is to generate paintings of a style corresponding to a reference image. Moreover, a new dataset of Chinese landscape paintings named CLAP is collected for comprehensive evaluation. Both the qualitative and quantitative results demonstrate that our method achieves state-of-the-art performance, especially in artfully-composed and artistic conception. Codes are available at https://github.com/Robin-WZQ/CCLAP.
Jie Zhang 0071, Zhilong Ji, Jinfeng Bai, Shiguang Shan
ICME3
2023 Only Classification Head Is Sufficient for Medical Image Segmentation
Hongbin Wei, Zhiwei Hu, Zhilong Ji, Hongpeng Jia, Lihe Zhang, Huchuan Lu
PRCV (13)4
2022 Syntax-Aware Network for Handwritten Mathematical Expression Recognition
abstract
Handwritten mathematical expression recognition (HMER) is a challenging task that has many potential applications. Recent methods for HMER have achieved outstanding performance with an encoder-decoder architecture. However, these methods adhere to the paradigm that the prediction is made “from one character to another”, which inevitably yields prediction errors due to the complicated structures of mathematical expressions or crabbed handwritings. In this paper, we propose a simple and efficient method for HMER, which is the first to incorporate syntax information into an encoder-decoder network. Specifically, we present a set of grammar rules for converting the LaTeX markup sequence of each expression into a parsing tree; then, we model the markup sequence prediction as a tree traverse process with a deep neural network. In this way, the proposed method can effectively describe the syntax context of expressions, alleviating the structure prediction errors of HMER. Experiments on three benchmark datasets demonstrate that our method achieves better recognition performance than prior arts. To further validate the effectiveness of our method, we create a large-scale dataset consisting of 100k handwritten mathematical expression images acquired from ten thousand writers. The source code, new dataset††https://ai.100tal.com/dataset, and pre-trained models of this work will be publicly available.
Xiao Liu 0040, Wondimu Dikubab, Zhilong Ji, Zhongqin Wu, Xiang Bai
CVPR5
2022 When Counting Meets HMER: Counting-Aware Network for Handwritten Mathematical Expression Recognition
Bohan Li 0010, Dingkang Liang, Xiao Liu 0040, Zhilong Ji, Jinfeng Bai, Wenyu Liu 0001, Xiang Bai
ECCV (28)5
2022 A Vision Transformer Based Scene Text Recognizer with Multi-grained Encoding and Decoding
Zhilong Ji, Jinfeng Bai
ICFHR2
2022 Towards Diverse and Faithful One-shot Adaption of Generative Adversarial Networks
abstract
One-shot generative domain adaption aims to transfer a pre-trained generator on one domain to a new domain using one reference image only. However, it remains very challenging for the adapted generator (i) to generate diverse images inherited from the pre-trained generator while (ii) faithfully acquiring the domain-specific attributes and styles of the reference image. In this paper, we present a novel one-shot generative domain adaption method, i.e., DiFa, for diverse generation and faithful adaptation. For global-level adaptation, we leverage the difference between the CLIP embedding of the reference image and the mean embedding of source images to constrain the target generator. For local-level adaptation, we introduce an attentive style loss which aligns each intermediate token of an adapted image with its corresponding token of the reference image. To facilitate diverse generation, selective cross-domain consistency is introduced to select and retain domain-sharing attributes in the editing latent $\mathcal{W}+$ space to inherit the diversity of the pre-trained generator. Extensive experiments show that our method outperforms the state-of-the-arts both quantitatively and qualitatively, especially for the cases of large domain gap. Moreover, our DiFa can easily be extended to zero-shot generative domain adaption with appealing results.
Yabo Zhang, Mingshuai Yao, Yuxiang Wei 0001, Zhilong Ji, Jinfeng Bai, Wangmeng Zuo
NeurIPS4
2021 Local Global Relational Network for Facial Action Units Recognition
abstract
Many existing facial action units (AUs) recognition approaches often enhance the AU representation by combining local features from multiple independent branches, each corresponding to a different AU. However, such multi-branch combination-based methods usually neglect potential mutual assistance and exclusion relationship between AU branches or simply employ a pre-defined and fixed knowledge-graph as a prior. In addition, extracting features from pre-defined AU regions of regular shapes limits the representation ability. In this paper, we propose a novel Local Global Relational Network (LGRNet) for facial AU recognition. LGRNet mainly consists of two novel structures, i.e., a skip-BiLSTM module which models the latent mutual assistance and exclusion relationship among local AU features from multiple branches to enhance the feature robustness, and a feature fusion&refining module which explores the complementarity between local AUs and the whole face in order to refine the local AU features to improve the discriminability. Experiments on the BP4D and DISFA AU datasets show that the proposed approach outperforms the state-of-the-art methods by a large margin.
Xuri Ge, Hu Han 0001, Joemon M. Jose, Zhilong Ji, Zhongqin Wu, Xiao Liu 0040
FG5
2021 Orthogonal Jacobian Regularization for Unsupervised Disentanglement in Image Generation
abstract
Unsupervised disentanglement learning is a crucial issue for understanding and exploiting deep generative models. Recently, SeFa tries to find latent disentangled directions by performing SVD on the first projection of a pretrained GAN. However, it is only applied to the first layer and works in a post-processing way. Hessian Penalty minimizes the off-diagonal entries of the output’s Hessian matrix to facilitate disentanglement, and can be applied to multi-layers. However, it constrains each entry of output independently, making it not sufficient in disentangling the latent directions (e.g., shape, size, rotation, etc.) of spatially correlated variations. In this paper, we propose a simple Orthogonal Jacobian Regularization (OroJaR) to encourage deep generative model to learn disentangled representations. It simply encourages the variation of output caused by perturbations on different latent dimensions to be orthogonal, and the Jacobian with respect to the input is calculated to represent this variation. We show that our OroJaR also encourages the output’s Hessian matrix to be diagonal in an indirect manner. In contrast to the Hessian Penalty, our OroJaR constrains the output in a holistic way, making it very effective in disentangling latent dimensions corresponding to spatially correlated variations. Quantitative and qualitative experimental results show that our method is effective in disentangled and controllable image generation, and performs favorably against the state-of-the-art methods. Our code is available at https://github.com/csyxwei/OroJaR.
Yuxiang Wei 0001, Yupeng Shi, Xiao Liu 0040, Zhilong Ji, Zhongqin Wu, Wangmeng Zuo
ICCV4
2021 Structured Multi-modal Feature Embedding and Alignment for Image-Sentence Retrieval
abstract
The current state-of-the-art image-sentence retrieval methods implicitly align the visual-textual fragments, like regions in images and words in sentences, and adopt attention modules to highlight the relevance of cross-modal semantic correspondences. However, the retrieval performance remains unsatisfactory due to a lack of consistent representation in both semantics and structural spaces. In this work, we propose to address the above issue from two aspects: (i) constructing intrinsic structure (along with relations) among the fragments of respective modalities, e.g., "dog → play → ball" in semantic structure for an image, and (ii) seeking explicit inter-modal structural and semantic correspondence between the visual and textual modalities.
Xuri Ge, Fuhai Chen, Joemon M. Jose, Zhilong Ji, Zhongqin Wu, Xiao Liu 0040
ACM Multimedia4