Jinfeng Bai

dblp:120/7270 · also Jin-Feng Bai · DBLP profile ↗
← Back
43ranked-venue papers
3as first author
38since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 1 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 2 first-author · 25 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021
YearPublicationVenuePosition
2026 VitaGlyph: Vitalizing Artistic Typography with Flexible Dual-branch Diffusion Models
abstract
Artistic typography is a technique that enables one to visualize the meaning of an input character in an imaginable and readable manner. With powerful text-to-image diffusion models, existing methods directly design the overall geometry and texture of input character, making it challenging to ensure both creativity and legibility. In this paper, we introduce a dual-branch, training-free method called VitaGlyph, enabling flexible artistic typography with controllable geometry changes while maintaining legibility. The key insight of VitaGlyph is to treat the input character as a scene composed of a Subject and its Surrounding, which are rendered with varying degrees of geometric transformation. To enhance the visual appeal and creativity of the generated artistic typography, the Subject flexibly expresses the essential concept of the input character, while the Surrounding enriches relevant background without altering the shape. Specifically, we implement VitaGlyph through a three-phase framework: (i) Knowledge Acquisition leverages large language models to design text descriptions for the Subject and Surrounding. (ii) Regional Interpretation detects the part that matches the subject description most closely and refines the structure using Semantic Typography. (iii) Attentional Compositional Generation separately renders the textures of the Subject and Surrounding and blends them in an attention-based manner. Experiments demonstrate that VitaGlyph not only achieves better artistry and legibility, but also manages to depict multiple customized concepts, facilitating more creative and pleasing artistic typography generation. Our code is available at https://github.com/Carlofkl/VitaGlyph.
Kailai Feng, Yabo Zhang, Haodong Yu, Zhilong Ji, Jinfeng Bai, Wangmeng Zuo
WACV5
2025 Explicit Relational Reasoning Network for Scene Text Detection
abstract
Connected component (CC) is a proper text shape representation that aligns with human reading intuition. However, CC-based text detection methods have recently faced a developmental bottleneck that their time-consuming post-processing is difficult to eliminate. To address this issue, we introduce an explicit relational reasoning network (ERRNet) to elegantly model the component relationships without post-processing. Concretely, we first represent each text instance as multiple ordered text components, and then treat these components as objects in sequential movement. In this way, scene text detection can be innovatively viewed as a tracking problem. From this perspective, we design an end-to-end tracking decoder to achieve a CC-based method dispensing with post-processing entirely. Additionally, we observe that there is an inconsistency between classification confidence and localization quality, so we propose a Polygon Monte-Carlo method to quickly and accurately evaluate the localization quality. Based on this, we introduce a position-supervised classification loss to guide the task-aligned learning of ERRNet. Experiments on challenging benchmarks demonstrate the effectiveness of our ERRNet. It consistently achieves state-of-the-art accuracy while holding highly competitive inference speed.
Zhineng Chen, Yongkun Du, Zhilong Ji, Kai Hu 0002, Jinfeng Bai, Xieping Gao 0001
AAAI6
2025 Enhancing Multimodal Continual Instruction Tuning with BranchLoRA
abstract
Multimodal Continual Instruction Tuning (MCIT) aims to finetune Multimodal Large Language Models (MLLMs) to continually align with human intent across sequential tasks. Existing approaches often rely on the Mixture-of-Experts (MoE) LoRA framework to preserve previous instruction alignments. However, these methods are prone to Catastrophic Forgetting (CF), as they aggregate all LoRA blocks via simple summation, which compromises performance over time. In this paper, we identify a critical parameter inefficiency in the MoELoRA framework within the MCIT context. Based on this insight, we propose BranchLoRA, an asymmetric framework to enhance both efficiency and performance. To mitigate CF, we introduce a flexible tuning-freezing mechanism within BranchLoRA, enabling branches to specialize in intra-task knowledge while fostering inter-task collaboration. Moreover, we incrementally incorporate task-specific routers to ensure an optimal branch distribution over time, rather than favoring the most recent task. To streamline inference, we introduce a task selector that automatically routes test inputs to the appropriate router without requiring task identity. Extensive experiments on the latest MCIT benchmark demonstrate that BranchLoRA significantly outperforms MoELoRA and maintains its superiority across various MLLM sizes.
Duzhen Zhang, Yong Ren 0006, Zhongzhi Li, Yahan Yu, Jiahua Dong 0001, Chenxing Li, Zhilong Ji, Jinfeng Bai
ACL (1)8
2025 CMMaTH: A Chinese Multi-modal Math Skill Evaluation Benchmark for Foundation Models
abstract
With the rapid advancements in multimodal large language models, evaluating their multimodal mathematical capabilities continues to receive wide attention. Although datasets such as MathVista have been introduced for evaluating mathematical capabilities in multimodal scenarios, there remains a lack of evaluation tools and datasets tailored for fine-grained assessment in Chinese K12 education. To systematically evaluate the ability of multimodal large models to solve Chinese multimodal mathematical problems, we propose a Chinese Multi-modal Math Skill Evaluation Benchmark (CMMaTH), containing 23,856 multimodal K12 math related questions, making it the largest Chinese multimodal mathematical problem benchmark to date. CMMaTH includes questions ranging from elementary to high school levels, offering greater diversity in problem types, solution goals, visual elements, detailed knowledge points, and standard solution annotations. To facilitate stable, fast, and cost-free model evaluation, we have developed an open-source tool called GradeGPT, which is integrated with the CMMaTH dataset. Our data and code are available at https://github.com/zzli2022/CMMaTH.
Zhongzhi Li, Mingliang Zhang 0005, Pei-Jie Wang, Jian Xu 0027, Rui-Song Zhang, Yin Fei, Zhilong Ji, Jinfeng Bai, Zhenru Pan
COLING8
2025 Integrating Visual Interpretation and Linguistic Reasoning for Geometric Problem Solving
Zixian Guo, Ming Liu 0018, Qilong Wang 0001, Zhilong Ji, Jinfeng Bai, Lei Zhang 0006, Wangmeng Zuo
ICCV5
2025 SolidGeo: Measuring Multimodal Spatial Math Reasoning in Solid Geometry
abstract
Geometry is a fundamental branch of mathematics and plays a crucial role in evaluating the reasoning capabilities of multimodal large language models (MLLMs). However, existing multimodal mathematics benchmarks mainly focus on plane geometry and largely ignore solid geometry, which requires spatial reasoning and is more challenging than plane geometry. To address this critical gap, we introduce SolidGeo, the first large-scale benchmark specifically designed to evaluate the performance of MLLMs on mathematical reasoning tasks in solid geometry. SolidGeo consists of 3,113 real-world K–12 and competition-level problems, each paired with visual context and annotated with difficulty levels and fine-grained solid geometry categories. Our benchmark covers a wide range of 3D reasoning subjects such as projection, unfolding, spatial measurement, and spatial vector, offering a rigorous testbed for assessing solid geometry. Through extensive experiments, we observe that MLLMs encounter substantial challenges in solid geometry math tasks, with a considerable performance gap relative to human capabilities on SolidGeo. Moreover, we analyze the performance, inference effiency and error patterns of various models, offering insights into the solid geometric mathematical reasoning capabilities of MLLMs. We hope SolidGeo serves as a catalyst for advancing MLLMs toward deeper geometric reasoning and spatial intelligence. The dataset is released at https://huggingface.co/datasets/HarryYancy/SolidGeo/
Zhongzhi Li, Dekang Ran, Zhilong Ji, Jinfeng Bai, Cheng-Lin Liu 0001
NeurIPS8
2025 Real face foundation representation learning for generalized deepfake detection
Liang Shi 0002, Jie Zhang 0071, Zhilong Ji, Jinfeng Bai, Shiguang Shan
Pattern Recognit.4
2024 Decoupled Textual Embeddings for Customized Image Generation
abstract
Customized text-to-image generation, which aims to learn user-specified concepts with a few images, has drawn significant attention recently. However, existing methods usually suffer from overfitting issues and entangle the subject-unrelated information (e.g., background and pose) with the learned concept, limiting the potential to compose concept into new scenes. To address these issues, we propose the DETEX, a novel approach that learns the disentangled concept embedding for flexible customized text-to-image generation. Unlike conventional methods that learn a single concept embedding from the given images, our DETEX represents each image using multiple word embeddings during training, i.e., a learnable image-shared subject embedding and several image-specific subject-unrelated embeddings. To decouple irrelevant attributes (i.e., background and pose) from the subject embedding, we further present several attribute mappers that encode each image as several image-specific subject-unrelated embeddings. To encourage these unrelated embeddings to capture the irrelevant information, we incorporate them with corresponding attribute words and propose a joint training strategy to facilitate the disentanglement. During inference, we only use the subject embedding for image generation, while selectively using image-specific embeddings to retain image-specified attributes. Extensive experiments demonstrate that the subject embedding obtained by our method can faithfully represent the target concept, while showing superior editability compared to the state-of-the-art methods. Our code will be available at https://github.com/PrototypeNx/DETEX.
Yufei Cai, Yuxiang Wei 0001, Zhilong Ji, Jinfeng Bai, Hu Han 0001, Wangmeng Zuo
AAAI4
2024 Leveraging Local Variance for Pseudo-Label Selection in Semi-supervised Learning
abstract
Semi-supervised learning algorithms that use pseudo-labeling have become increasingly popular for improving model performance by utilizing both labeled and unlabeled data. In this paper, we offer a fresh perspective on the selection of pseudo-labels, inspired by theoretical insights. We suggest that pseudo-labels with a high degree of local variance are more prone to inaccuracies. Based on this premise, we introduce the Local Variance Match (LVM) method, which aims to optimize the selection of pseudo-labels in semi-supervised learning (SSL) tasks. Our methodology is validated through a series of experiments on widely-used image classification datasets, such as CIFAR-10, CIFAR-100, and SVHN, spanning various labeled data quantity scenarios. The empirical findings show that the LVM method substantially outpaces current SSL techniques, achieving state-of-the-art results in many of these scenarios. For instance, we observed an error rate of 5.41% on CIFAR-10 with a single label for each class, 35.87% on CIFAR-100 when using four labels per class, and 1.94% on SVHN with four labels for each class. Notably, the standout error rate of 5.41% is less than 1% shy of the performance in a fully-supervised learning environment. In experiments on ImageNet with 100k labeled data, the LVM also reached state-of-the-art outcomes. Additionally, the efficacy of the LVM method is further validated by its stellar performance in speech recognition experiments.
Zeping Min, Jinfeng Bai, Chengfei Li
AAAI2
2024 LRANet: Towards Accurate and Efficient Scene Text Detection with Low-Rank Approximation Network
abstract
Recently, regression-based methods, which predict parameterized text shapes for text localization, have gained popularity in scene text detection. However, the existing parameterized text shape methods still have limitations in modeling arbitrary-shaped texts due to ignoring the utilization of text-specific shape information. Moreover, the time consumption of the entire pipeline has been largely overlooked, leading to a suboptimal overall inference speed. To address these issues, we first propose a novel parameterized text shape method based on low-rank approximation. Unlike other shape representation methods that employ data-irrelevant parameterization, our approach utilizes singular value decomposition and reconstructs the text shape using a few eigenvectors learned from labeled text contours. By exploring the shape correlation among different text contours, our method achieves consistency, compactness, simplicity, and robustness in shape representation. Next, we propose a dual assignment scheme for speed acceleration. It adopts a sparse assignment branch to accelerate the inference speed, and meanwhile, provides ample supervised signals for training through a dense assignment branch. Building upon these designs, we implement an accurate and efficient arbitrary-shaped text detector named LRANet. Extensive experiments are conducted on several challenging benchmarks, demonstrating the superior accuracy and efficiency of LRANet compared to state-of-the-art methods. Code is available at: https://github.com/ychensu/LRANet.git
Zhineng Chen, Zhiwen Shao, Yuning Du, Zhilong Ji, Jinfeng Bai, Yong Zhou 0003, Yu-Gang Jiang 0001
AAAI6
2024 CK12: A Rounded K12 Knowledge Graph Based Benchmark for Chinese Holistic Cognition Evaluation
abstract
New NLP benchmarks are urgently needed to align with the rapid development of large language models (LLMs). We present a meticulously designed evaluation benchmark that leverages the knowledge graph. This evaluation comprises 584 level-1 knowledge points and 1,989 level-2 knowledge points, thereby encompassing a comprehensive spectrum of the K12 education domain knowledge. The primary objective is to comprehensively assess the high-level comprehension aptitude and reasoning capabilities of LLMs operating within the Chinese context. Our evaluation incorporates five distinct question types with 39,452 questions. We test the current mainstream LLMs by three distinct modes. Firstly, four prompt evaluation modes were employed to assess the fundamental capacity. Additionally, for choice questions, a result-oriented evaluation approach was designed through data augmentation to assess the model's proficiency in advanced knowledge and reasoning. Moreover, a subset with reasoning process is derived, and the process-oriented testing method is used to test the model's interpretability and higher-order reasoning capacity. We further show models' capability in our knowledge points, and anticipate the evaluation can assist in the assessment of the strengths and deficiencies of LLMs on knowledge points, thus fostering their development within the Chinese context. Our Dataset will be publicly available in https://github.com/tal-tech/chinese-k12-evaluation.
Weihao You, Zhilong Ji, Jinfeng Bai
AAAI5
2024 HPNet: Dynamic Trajectory Forecasting with Historical Prediction Attention
abstract
Predicting the trajectories of road agents is essential for autonomous driving systems. The recent mainstream methods follow a static paradigm, which predicts the future trajectory by using a fixed duration of historical frames. These methods make the predictions independently even at adjacent time steps, which leads to potential instability and temporal inconsistency. As successive time steps have largely overlapping historical frames, their forecasting should have intrinsic correlation, such as overlapping predicted trajectories should be consistent, or be different but share the same motion goal depending on the road situation. Motivated by this, in this work, we introduce HPNet, a novel dynamic trajectory forecasting method. Aiming for stable and accurate trajectory forecasting, our method leverages not only historical frames including maps and agent states, but also historical predictions. Specifically, we newly design a Historical Prediction Attention module to automatically encode the dynamic relationship between successive predictions. Besides, it also extends the attention range beyond the currently visible window benefitting from the use of historical predictions. The proposed Historical Prediction Attention together with the Agent Attention and Mode Attention is further formulated as the Triple Factorized Attention module, serving as the core design of HPNet. Experiments on the Argoverse and INTERACTION datasets show that HP-Net achieves state-of-the-art performance, and generates accurate and stable future trajectories. Our code are available at https://github.com/XiaolongTang23/HPNet.
Meina Kan, Shiguang Shan, Zhilong Ji, Jinfeng Bai, Xilin Chen 0001
CVPR5
2024 MasterWeaver: Taming Editability and Face Identity for Personalized Text-to-Image Generation
Yuxiang Wei 0001, Zhilong Ji, Jinfeng Bai, Lei Zhang 0006, Wangmeng Zuo
ECCV (51)3
2024 MuMath-Code: Combining Tool-Use Large Language Models with Multi-perspective Data Augmentation for Mathematical Reasoning
abstract
The tool-use Large Language Models (LLMs) that integrate with external Python interpreters have significantly enhanced mathematical reasoning capabilities for open-source LLMs, while tool-free methods chose another track: augmenting math reasoning data.However, a great method to integrate the above two research paths and combine their advantages remains to be explored.In this work, we firstly include new math questions via multi-perspective data augmenting methods and then synthesize code-nested solutions to them.The open LLMs (e.g., Llama-2) are finetuned on the augmented dataset to get the resulting models, MuMath-Code (µ-Math-Code).During the inference phase, our MuMath-Code generates code and interacts with the external python interpreter to get the execution results.Therefore, MuMath-Code leverages the advantages of both the external tool and data augmentation.To fully leverage the advantages of our augmented data, we propose a two-stage training strategy: In Stage-1, we finetune Llama-2 on pure CoT data to get an intermediate model, which then is trained on the code-nested data in Stage-2 to get the resulting MuMath-Code.Our MuMath-Code-7B achieves 83.8% on GSM8K and 52.4% on MATH, while MuMath-Code-70B model achieves new state-of-the-art performance among open methods-achieving 90.7% on GSM8K and 55.1% on MATH.Extensive experiments validate the combination of tool use and data augmentation, as well as our two-stage training strategy.We release the proposed dataset along with the associated code for public use: https://github.com/ youweihao-tal/MuMath-Code.
Weihao You, Zhilong Ji, Guoqiang Zhong 0001, Jinfeng Bai
EMNLP5
2024 DPA-2D: Depth Propagation and Alignment with 2D Observations Guidance for Human Mesh Recovery
abstract
In the current state of 3D human pose and shape estimation, the two-stage methods involving 3D joints estimation and inverse kinematics (IK) to obtain more accurate human pose, have gained significant traction. However, the effectiveness of the final optimized depth is greatly influenced by the initial 3D joints estimation. The 3D joints errors primarily caused by depth estimation. Our findings indicate that 2D observations (2D joint locations and Confidence) offer valuable supplementary data. Depth estimation errors are greater for low-confidence joints, but high-confidence joints yield more accuracy. In light of this, we propose a Biased Depth Propagation (BDP) module that leverages the confidence of 2D joints to enhance depth estimation. This module enables us to improve the depth estimation of low-confidence joints by propagating the depth information from high-confidence joints. Furthermore, small variations in the SMPL parameters can significantly impact the reconstructed meshes causing misalignment with image evidence. Therefore, we propose the Refined Joint Aligned (RJA) module to fine-tune the SMPL parameters by utilizing enhanced joint ROI features captured by the estimated 2D joint locations. Finally, we have validated the performance of our method on popular datasets such as Human3.6m and 3DPW for human mesh recovery task, and achieved SOTA results, demonstrate the effectiveness of our Depth Propagation and Alignment with 2D Observations Guidance(DPA-2D).
Weihao You, Zhilong Ji, Jinfeng Bai
FG4
2024 Collaborative Domain Alignment for Multi-source Domain Adaptation
Meina Kan, Zhilong Ji, Jinfeng Bai, Shiguang Shan, Xilin Chen 0001
ICPR (27)4
2023 ReCoT: Regularized Co-Training for Facial Action Unit Recognition with Noisy Labels
Hu Han 0001, Shiguang Shan, Zhilong Ji, Jinfeng Bai, Xilin Chen 0001
BMVC5
2023 Inferring and Leveraging Parts from Object Shape for Improving Semantic Image Synthesis
abstract
Despite the progress in semantic image synthesis, it remains a challenging problem to generate photo-realistic parts from input semantic map. Integrating part segmentation map can undoubtedly benefit image synthesis, but is bothersome and inconvenient to be provided by users. To improve part synthesis, this paper presents to infer Parts from Object ShapE (iPOSE) and leverage it for improving semantic image synthesis. However, albeit several part segmentation datasets are available, part annotations are still not provided for many object categories in semantic image synthesis. To circumvent it, we resort to few-shot regime to learn a PartNet for predicting the object part map with the guidance of pre-defined support part maps. PartNet can be readily generalized to handle a new object category when a small number (e.g., 3) of support part maps for this category are provided. Furthermore, part semantic modulation is presented to incorporate both inferred part map and semantic map for image synthesis. Experiments show that our iPOSE not only generates objects with rich part details, but also enables to control the image synthesis flexibly. And our iPOSE performs favorably against the state-of-the-art methods in terms of quantitative and qualitative evaluation. Our code will be publicly available at https://github.com/csyxwei/iPOSE.
Yuxiang Wei 0001, Zhilong Ji, Xiaohe Wu, Jinfeng Bai, Lei Zhang 0006, Wangmeng Zuo
CVPR4
2023 Texts as Images in Prompt Tuning for Multi-Label Image Recognition
abstract
Prompt tuning has been employed as an efficient way to adapt large vision-language pre-trained models (e.g. CLIP) to various downstream tasks in data-limited or label-limited settings. Nonetheless, visual data (e.g., images) is by default prerequisite for learning prompts in existing methods. In this work, we advocate that the effectiveness of image-text contrastive learning in aligning the two modalities (for training CLIP) further makes it feasible to treat texts as images for prompt tuning and introduce TaI prompting. In contrast to the visual data, text descriptions are easy to collect, and their class labels can be directly derived. Particularly, we apply TaI prompting to multi-label image recognition, where sentences in the wild serve as alternatives to images for prompt tuning. Moreover, with TaI, double-grained prompt tuning (TaI-DPT) is further presented to extract both coarse-grained and fine-grained embeddings for enhancing the multi-label recognition performance. Experimental results show that our proposed TaI-DPT outperforms zero-shot CLIP by a large margin on multiple benchmarks, e.g., MS-COCO, VOC2007, and NUS-WIDE, while it can be combined with existing methods of prompting from images to improve recognition performance further. The code is released at https://github.com/guozix/TaI-DPT.
Zixian Guo, Bowen Dong 0001, Zhilong Ji, Jinfeng Bai, Yiwen Guo, Wangmeng Zuo
CVPR4
2023 Unveiling the Implicit Toxicity in Large Language Models
abstract
The open-endedness of large language models (LLMs) combined with their impressive capabilities may lead to new safety issues when being exploited for malicious use.While recent studies primarily focus on probing toxic outputs that can be easily detected with existing toxicity classifiers, we show that LLMs can generate diverse implicit toxic outputs that are exceptionally difficult to detect via simply zero-shot prompting.Moreover, we propose a reinforcement learning (RL) based attacking method to further induce the implicit toxicity in LLMs.Specifically, we optimize the language model with a reward that prefers implicit toxic outputs to explicit toxic and non-toxic ones.Experiments on five widely-adopted toxicity classifiers demonstrate that the attack success rate can be significantly improved through RL fine-tuning.For instance, the RL-finetuned LLaMA-13B model achieves an attack success rate of 90.04% on BAD and 62.85% on Davinci003.Our findings suggest that LLMs pose a significant threat in generating undetectable implicit toxic outputs.We further show that fine-tuning toxicity classifiers on the annotated examples from our attacking method can effectively enhance their ability to detect LLM-generated implicit toxic language.The code is publicly available at https://github. com/thu-coai/Implicit-Toxicity.
Jiaxin Wen, Pei Ke, Hao Sun 0012, Zhexin Zhang, Chengfei Li, Jinfeng Bai, Minlie Huang
EMNLP6
2023 DSPGAN: A Gan-Based Universal Vocoder for High-Fidelity TTS by Time-Frequency Domain Supervision from DSP
abstract
Recent development of neural vocoders based on the generative adversarial neural network (GAN) has shown obvious advantages of generating raw waveform conditioned on mel-spectrogram with fast inference speed and lightweight networks. Whereas, it is still challenging to train a universal neural vocoder that can synthesize high-fidelity speech from various scenarios with unseen speakers, languages, and speaking styles. In this paper, we propose DSP- GAN, a GAN-based universal vocoder for high-fidelity speech synthesis by applying the time-frequency domain supervision from digital signal processing (DSP). To eliminate the mismatch problem caused by the ground-truth spectrograms in the training phase and the predicted spectrograms in the inference phase, we leverage the mel-spectrogram extracted from the waveform generated by a DSP module, rather than the predicted mel-spectrogram from the Text-to-Speech (TTS) acoustic model, as the time-frequency domain supervision to the GAN-based vocoder. We also utilize sine excitation as the time-domain supervision to improve the harmonic modeling and eliminate various artifacts of the GAN-based vocoder. Experiments show that DSPGAN significantly outperforms the compared approaches and it can generate high-fidelity speech for various TTS models trained using diverse data.1
Yongmao Zhang, Jian Cong, Hanzhao Li, Lei Xie 0001, Jinfeng Bai
ICASSP8
2023 A Synthetic Corpus Generation Method for Neural Vocoder Training
abstract
Nowadays, neural vocoders are preferred for their ability to synthesize high-fidelity audio. However, training a neural vocoder requires a massive corpus of high-quality real audio, and the audio recording process is often labor-intensive. In this work, we propose a synthetic corpus generation method for neural vocoder training, which can easily generate synthetic audio with an unlimited number at nearly no cost. We explicitly model the prior characteristics of audio from multiple target domains simultaneously (e.g., speeches, singing voices, and instrumental pieces) to equip the generated audio data with these characteristics. And we show that our synthetic corpus allows the neural vocoder to achieve competitive results without any real audio in the training process. To validate the effectiveness of our proposed method, we performed empirical experiments on both speech and music utterances in subjective and objective metrics. The experimental results show that the neural vocoder trained with the synthetic corpus produced by our method can generalize to multiple target scenarios and has excellent singing voice (MOS: 4.20) and instrumental piece (MOS: 4.00) synthesis results.
Zilin Wang 0002, Jun Chen 0024, Sipan Li, Jinfeng Bai, Zhiyong Wu 0001, Helen M. Meng
ICASSP5
2023 ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image Generation
abstract
In addition to the unprecedented ability in imaginary creation, large text-to-image models are expected to take customized concepts in image generation. Existing works generally learn such concepts in an optimization-based manner, yet bringing excessive computation or memory burden. In this paper, we instead propose a learning-based encoder, which consists of a global and a local mapping networks for fast and accurate customized text-to-image generation. In specific, the global mapping network projects the hierarchical features of a given image into multiple "new" words in the textual word embedding space, i.e., one primary word for well-editable concept and other auxiliary words to exclude irrelevant disturbances (e.g., background). In the meantime, a local mapping network injects the encoded patch features into cross attention layers to provide omitted details, without sacrificing the editability of primary concepts. We compare our method with existing optimization-based approaches on a variety of user-defined concepts, and demonstrate that our method enables high-fidelity inversion and more robust editability with a significantly faster encoding process. Our code is publicly available at https://github.com/csyxwei/ELITE.
Yuxiang Wei 0001, Yabo Zhang, Zhilong Ji, Jinfeng Bai, Lei Zhang 0006, Wangmeng Zuo
ICCV4
2023 ICDAR 2023 Competition on Recognition of Multi-line Handwritten Mathematical Expressions
Chenyang Gao, Shiyu Yao, Jinfeng Bai, Xiang Bai, Cheng-Lin Liu 0001
ICDAR (2)4
2023 Semantic Graph Representation Learning for Handwritten Mathematical Expression Recognition
Zhilong Ji, Jinfeng Bai, Xiang Bai
ICDAR (1)4
2023 ViSA: Visual and Semantic Alignment for Robust Scene Text Recognition
Zhenru Pan, Zhilong Ji, Xiao Liu 0040, Jinfeng Bai, Cheng-Lin Liu 0001
ICDAR (2)4
2023 Decoupling Visual-Semantic Features Learning with Dual Masked Autoencoder for Self-Supervised Scene Text Recognition
Zhilong Ji, Jinfeng Bai
ICDAR (2)4
2023 CCLAP: Controllable Chinese Landscape Painting Generation Via Latent Diffusion Model
abstract
With the development of deep generative models, recent years have seen great success of Chinese landscape painting generation. However, few works focus on controllable Chinese landscape painting generation due to the lack of data and limited modeling capabilities. In this work, we propose a controllable Chinese landscape painting generation method named CCLAP, which can generate painting with specific content and style based on Latent Diffusion Model. Specifically, it consists of two cascaded modules, i.e., content generator and style aggregator. The content generator module guarantees the content of generated paintings specific to the input text. While the style aggregator module is to generate paintings of a style corresponding to a reference image. Moreover, a new dataset of Chinese landscape paintings named CLAP is collected for comprehensive evaluation. Both the qualitative and quantitative results demonstrate that our method achieves state-of-the-art performance, especially in artfully-composed and artistic conception. Codes are available at https://github.com/Robin-WZQ/CCLAP.
Jie Zhang 0071, Zhilong Ji, Jinfeng Bai, Shiguang Shan
ICME4
2023 TPS++: Attention-Enhanced Thin-Plate Spline for Scene Text Recognition
abstract
Text irregularities pose significant challenges to scene text recognizers. Thin-Plate Spline (TPS)-based rectification is widely regarded as an effective means to deal with them. Currently, the calculation of TPS transformation parameters purely depends on the quality of regressed text borders. It ignores the text content and often leads to unsatisfactory rectified results for severely distorted text. In this work, we introduce TPS++, an attention-enhanced TPS transformation that incorporates the attention mechanism to text rectification for the first time. TPS++ formulates the parameter calculation as a joint process of foreground control point regression and content-based attention score estimation, which is computed by a dedicated designed gated-attention block. TPS++ builds a more flexible content-aware rectifier, generating a natural text correction that is easier to read by the subsequent recognizer. Moreover, TPS++ shares the feature backbone with the recognizer in part and implements the rectification at feature-level rather than image-level, incurring only a small overhead in terms of parameters and inference time. Experiments on public benchmarks show that TPS++ consistently improves the recognition and achieves state-of-the-art accuracy. Meanwhile, it generalizes well on different backbones and recognizers. Code is at https://github.com/simplify23/TPS_PP.
Tianlun Zheng, Zhineng Chen, Jinfeng Bai, Hongtao Xie 0001, Yu-Gang Jiang 0001
IJCAI3
2023 DistilXLSR: A Light Weight Cross-Lingual Speech Representation Model
Haoyu Wang 0014, Siyuan Wang 0002, Weiqiang Zhang 0001, Jinfeng Bai
INTERSPEECH4
2023 Dual Contrastive Prediction for Incomplete Multi-View Representation Learning
abstract
In this article, we propose a unified framework to solve the following two challenging problems in incomplete multi-view representation learning: i) how to learn a consistent representation unifying different views, and ii) how to recover the missing views. To address the challenges, we provide an information theoretical framework under which the consistency learning and data recovery are treated as a whole. With the theoretical framework, we propose a novel objective function which jointly solves the aforementioned two problems and achieves a provable sufficient and minimal representation. In detail, the consistency learning is performed by maximizing the mutual information of different views through contrastive learning, and the missing views are recovered by minimizing the conditional entropy through dual prediction. To the best of our knowledge, this is one of the first works to theoretically unify the cross-view consistency learning and data recovery for representation learning. Extensive experimental results show that the proposed method remarkably outperforms 20 competitive multi-view learning methods on six datasets in terms of clustering, classification, and human action recognition. The code could be accessed from https://pengxi.me.
Yijie Lin 0001, Yuanbiao Gou, Xiaotian Liu, Jinfeng Bai, Jiancheng Lv 0001, Xi Peng 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Robust Multi-View Clustering With Incomplete Information
abstract
The success of existing multi-view clustering methods heavily relies on the assumption of view consistency and instance completeness, referred to as the complete information. However, these two assumptions would be inevitably violated in data collection and transmission, thus leading to the so-called Partially View-unaligned Problem (PVP) and Partially Sample-missing Problem (PSP). To overcome such incomplete information challenges, we propose a novel method, termed robuSt mUlti-view clusteRing with incomplEte information (SURE), which solves PVP and PSP under a unified framework. In brief, SURE is a novel contrastive learning paradigm which uses the available pairs as positives and randomly chooses some cross-view samples as negatives. To reduce the influence of the false negatives caused by random sampling, SURE is with a noise-robust contrastive loss that theoretically and empirically mitigates or even eliminates the influence of the false negatives. To the best of our knowledge, this could be the first successful attempt that simultaneously handles PVP and PSP using a unified solution. In addition, this could be one of the first studies on the noisy correspondence problem (i.e., the false negatives) which is a novel paradigm of noisy labels. Extensive experiments demonstrate the effectiveness and efficiency of SURE comparing with 10 state-of-the-art approaches on the multi-view clustering task.
Mouxing Yang, Yunfan Li 0003, Peng Hu 0002, Jinfeng Bai, Jiancheng Lv 0001, Xi Peng 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 When Counting Meets HMER: Counting-Aware Network for Handwritten Mathematical Expression Recognition
Bohan Li 0010, Dingkang Liang, Xiao Liu 0040, Zhilong Ji, Jinfeng Bai, Wenyu Liu 0001, Xiang Bai
ECCV (28)6
2022 Time-Domain Audio-Visual Speech Separation on Low Quality Videos
abstract
Incorporating visual information is a promising approach to improve the performance of speech separation. Many related works have been conducted and provide inspiring results. However, low quality videos appear commonly in real scenarios, which may significantly degrade the performance of normal audio-visual speech separation system. In this paper, we propose a new structure to fuse the audio and visual features, which uses the audio feature to select relevant visual features by utilizing the attention mechanism. A Conv-TasNet based model is combined with the proposed attention-based multi-modal fusion, trained with proper data augmentation and evaluated with 3 categories of low quality videos. The experimental results show that our system outperforms the baseline which simply concatenates the audio and visual features when training with normal or low quality data, and is robust to low quality video inputs at inference time.
Chenda Li, Jinfeng Bai, Zhongqin Wu, Yanmin Qian
ICASSP3
2022 A Vision Transformer Based Scene Text Recognizer with Multi-grained Encoding and Decoding
Zhilong Ji, Jinfeng Bai
ICFHR4
2022 TALCS: An open-source Mandarin-English code-switching corpus and a speech recognition baseline
abstract
This paper introduces a new corpus of Mandarin-English code-switching speech recognition-TALCS corpus, suitable for training and evaluating code-switching speech recognition systems.TALCS corpus is derived from real online one-to-one English teaching scenes in TAL education group, which contains roughly 587 hours of speech sampled at 16 kHz.To our best knowledge, TALCS corpus is the largest well labeled Mandarin-English code-switching open source automatic speech recognition (ASR) dataset in the world.In this paper, we will introduce the recording procedure in detail, including audio capturing devices and corpus environments.And the TALCS corpus is freely available for download under the permissive license 1 .Using TALCS corpus, we conduct ASR experiments in two popular speech recognition toolkits to make a baseline system, including ESPnet and Wenet.The Mixture Error Rate (MER) performance in the two speech recognition toolkits is compared in TALCS corpus.The experimental results implies that the quality of audio recordings and transcriptions are promising and the baseline system is workable.
Chengfei Li, Shuhao Deng, Yaoping Wang, Guangjing Wang 0004, Yaguang Gong, Changbin Chen, Jinfeng Bai
INTERSPEECH7
2022 Towards Diverse and Faithful One-shot Adaption of Generative Adversarial Networks
abstract
One-shot generative domain adaption aims to transfer a pre-trained generator on one domain to a new domain using one reference image only. However, it remains very challenging for the adapted generator (i) to generate diverse images inherited from the pre-trained generator while (ii) faithfully acquiring the domain-specific attributes and styles of the reference image. In this paper, we present a novel one-shot generative domain adaption method, i.e., DiFa, for diverse generation and faithful adaptation. For global-level adaptation, we leverage the difference between the CLIP embedding of the reference image and the mean embedding of source images to constrain the target generator. For local-level adaptation, we introduce an attentive style loss which aligns each intermediate token of an adapted image with its corresponding token of the reference image. To facilitate diverse generation, selective cross-domain consistency is introduced to select and retain domain-sharing attributes in the editing latent $\mathcal{W}+$ space to inherit the diversity of the pre-trained generator. Extensive experiments show that our method outperforms the state-of-the-arts both quantitatively and qualitatively, especially for the cases of large domain gap. Moreover, our DiFa can easily be extended to zero-shot generative domain adaption with appealing results.
Yabo Zhang, Mingshuai Yao, Yuxiang Wei 0001, Zhilong Ji, Jinfeng Bai, Wangmeng Zuo
NeurIPS5
2022 Unsupervised Neural Rendering for Image Hazing
abstract
Image hazing aims to render a hazy image from a given clean one, which could be applied to a variety of practical applications such as gaming, filming, photographic filtering, and image dehazing. To generate plausible haze, we study two less-touched but challenging problems in hazy image rendering, namely, i) how to estimate the transmission map from a single image without auxiliary information, and ii) how to adaptively learn the airlight from exemplars, i.e., unpaired real hazy images. To this end, we propose a neural rendering method for image hazing, dubbed as HazeGEN. To be specific, HazeGEN is a knowledge-driven neural network which estimates the transmission map by leveraging a new prior, i.e., there exists the structure similarity (e.g., contour and luminance) between the transmission map and the input clean image. To adaptively learn the airlight, we build a neural module based on another new prior, i.e., the rendered hazy image and the exemplar are similar in the airlight distribution. To the best of our knowledge, this could be the first attempt to deeply render hazy images in an unsupervised fashion. Compared with existing haze generation methods, HazeGEN renders the hazy images in an unsupervised, learnable, and controllable manner, thus avoiding the labor-intensive efforts in paired data collection and the domain-shift issue in haze generation. Extensive experiments show the promising performance of our method comparing with some baselines in both qualitative and quantitative comparisons. The code is available at https://github.com/XLearning-SCU.
Boyun Li, Yijie Lin 0001, Jinfeng Bai, Peng Hu 0002, Jiancheng Lv 0001, Xi Peng 0001
IEEE Trans. Image Process.3
2014 Chinese Image Character Recognition Using DNN and Machine Simulated Training Samples
Jinfeng Bai, Zhineng Chen, Bailan Feng, Bo Xu 0002
ICANN1
2014 Chinese Image Text Recognition on grayscale pixels
abstract
This paper presents a novel scheme for Chinese text recognition in images and videos. It's different from traditional paradigms that binarize text images, fed the binarized text to an OCR engine and get the recognized results. The proposed scheme, named grayscale based Chinese Image Text Recognition (gCITR), implements the recognition directly on grayscale pixels via the following steps: image text over-segmentation, building recognition graph, Chinese character recognition and beam search determination. The advantages of gCITR lie in: (1) it does not heavily rely on the performance of binarization, which is not robust in practical and thus severely affects the performance of OCR, (2) grayscale image retains more information of the text thus facilitates the recognition. Experimental results on text from 13 TV news videos demonstrate the effectiveness of the proposed gCITR, from which significant performance gains are observed.
Jinfeng Bai, Zhineng Chen, Bailan Feng, Bo Xu 0002
ICASSP1
2014 Image character recognition using deep convolutional neural network learned from different languages
abstract
This paper proposes a shared-hidden-layer deep convolutional neural network (SHL-CNN) for image character recognition. In SHL-CNN, the hidden layers are made common across characters from different languages, performing a universal feature extraction process that aims at learning common character traits existed in different languages such as strokes, while the final softmax layer is made language dependent, trained based on characters from the destination language only. This paper is the first attempt to introduce the SHL-CNN framework to image character recognition. Under the SHL-CNN framework, we discuss several issues including architecture of the network, training of the network, from which a suitable SHL-CNN model for image character recognition is empirically learned. The effectiveness of the learned SHL-CNN is verified on both English and Chinese image character recognition tasks, showing that the SHL-CNN can reduce recognition errors by 16–30% relatively compared with models trained by characters of only one language using conventional CNN, and by 35.7% relatively compared with state-of-the-art methods. In addition, the shared hidden layers learned are also useful for unseen image character recognition tasks.
Jinfeng Bai, Zhineng Chen, Bailan Feng, Bo Xu 0002
ICIP1
2014 CeleLabel: an interactive system for annotating celebrities in web videos
abstract
Manual annotation of celebrities in Web videos is an essential task in many people-related Web services. The task, however, poses a significant challenge even to skillful annotators, mainly due to the large quantity of unfamiliar and greatly varied celebrities, and the lack of a customized system for it. This work develops CeleLabel, an interactive system for manually annotating celebrities in the Web video domain. The peculiarity of CeleLabel is to exploit and display multiple types of information that could assist the annotation, including video content, context surrounding and within a video, celebrity images on the Web, and human factors. Using the system, annotators can interactively switch between two views, i.e., merging similar faces and labeling faces with names, to approach the annotation. User studies show that the CeleLabel leads to a much better labeling efficiency and satisfaction.
Zhineng Chen, Jinfeng Bai, Chong-Wah Ngo, Bailan Feng, Bo Xu 0002
ACM Multimedia2
2012 Multi-modal information fusion for news story segmentation in broadcast video
abstract
With the fast development of high-speed network and digital video recording technologies, broadcast video has been playing a more and more important role in our daily life. In this paper, we propose a novel news story segmentation scheme which can segment broadcast video into story units with multi-modal information fusion (MMIF) strategy. Compared with traditional methods, the proposed scheme extracts a wealth of semantic-level features including anchor person, topic caption, face, silence, acoustic change, audio keywords and textual content. Parallel to this, we make use of a multi-modal information fusion strategy for news story boundary characterization by joining these visual, audio and textual cues. Encouraging experimental results on News Vision dataset demonstrate the effectiveness of the proposed scheme.
Bailan Feng, Peng Ding 0003, Jiansong Chen, Jinfeng Bai, Bo Xu 0002
ICASSP4