VLDB 2026 Research / reviewers in the wild / expert
Ziyong Feng
dblp:120/4362
· DBLP profile ↗
31ranked-venue papers
3as first author
20since 2021 · last 2026
0009-0007-8689-8366ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 2 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 1 first-author · 15 since 2021Databases, data management, data science and information retrieval · 3Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding LearningabstractUniversal multimodal embedding models are essential in various tasks. Existing approaches typically use in-batch mining to identify hard negatives by measuring the similarity of query-candidate pairs. However, these methods often struggle to capture subtle semantic differences among candidates and lack diversity in negative samples. Moreover, the embeddings exhibit limited discriminative ability in distinguishing false and hard negatives. In this paper, we leverage the advanced understanding capabilities of MLLMs to enhance representation learning, and present a novel Universal Multimodal Embedding(UniME-V2) model. Our approach first constructs a potential hard negative set through global retrieval. We then introduce the MLLM-as-a-Judge mechanism, which utilizes MLLMs to assess the semantic alignment of query-candidate pairs and generate soft semantic matching scores. These scores serve as a foundation for hard negative mining, mitigating the impact of false negatives and enabling the identification of diverse, high-quality hard negatives. Furthermore, the semantic matching scores are used as soft labels to mitigate the rigid one-to-one mapping constraint. By aligning the similarity matrix with the soft semantic matching score matrix, the model learns semantic distinctions among candidates, significantly enhancing its discriminative capacity. To further improve performance, we propose UniME-V2, a reranking model trained on our mined hard negatives through a joint pairwise and listwise optimization approach. We conduct comprehensive experiments on the MMEB benchmark and multiple retrieval tasks, demonstrating that our method achieves state-of-the-art performance across all tasks. Tiancheng Gu, Kaicheng Yang 0002, Kaichen Zhang, Xiang An, Ziyong Feng, Tom Weidong Cai, Jiankang Deng, Lidong Bing |
AAAI | 5 |
| 2026 | ViCToR: Improving Visual Comprehension via Token Reconstruction for Pretraining LMMsabstractLarge Multimodal Models (LMMs) often face a modality representation gap during pretraining: while language embeddings remain stable, visual representations are highly sensitive to contextual noise (e.g., background clutter). To address this issue, we introduce a visual comprehension stage, which we call ViCToR (Visual Comprehension via Token Reconstruction), a novel pretraining framework for LMMs. ViCToR employs a learnable visual token pool and utilizes the Hungarian matching algorithm to select semantically relevant tokens from this pool for visual token replacement. Furthermore, by integrating a visual token reconstruction loss with dense semantic supervision, ViCToR can learn tokens which retain high visual detail, thereby enhancing the large language model's (LLM's) understanding of visual information. After pretraining on 3 million publicly accessible images and captions, ViCToR achieves state-of-the-art results, improving over LLaVA-NeXT-8B by 10.4%, 3.2%, and 7.2% on the MMStar, SEEDI, and RealWorldQA benchmarks, respectively. Yin Xie, Kaicheng Yang 0002, Peirou Liang, Xiang An, Yongle Zhao, Ziyong Feng, Roy Miles, Ismail Elezi, Jiankang Deng |
AAAI | 7 |
| 2025 | CLIP-CID: Efficient CLIP Distillation via Cluster-Instance DiscriminationabstractContrastive Language-Image Pre-training (CLIP) has achieved excellent performance over a wide range of tasks. However, the effectiveness of CLIP heavily relies on a substantial corpus of pre-training data, resulting in notable consumption of computational resources. Although knowledge distillation has been widely applied in single modality models, how to efficiently expand knowledge distillation to vision-language foundation models with extensive data remains relatively unexplored. In this paper, we introduce CLIP-CID, a novel distillation mechanism that effectively transfers knowledge from a large vision-language foundation model to a smaller model. We initially propose a simple but efficient image semantic balance method to reduce transfer learning bias and improve distillation efficiency. This method filters out 43.7% of image-text pairs from the LAION400M while maintaining superior performance. After that, we leverage cluster-instance discrimination to facilitate knowledge transfer from the teacher model to the student model, thereby empowering the student model to acquire a holistic semantic comprehension of the pre-training data. Experimental results demonstrate that CLIP-CID achieves state-of-the-art performance on various downstream tasks including linear probe and zero-shot classification. Kaicheng Yang 0002, Tiancheng Gu, Xiang An, Haiqiang Jiang, Xiangzi Dai, Ziyong Feng, Tom Weidong Cai, Jiankang Deng |
AAAI | 6 |
| 2025 | Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person RetrievalabstractAlthough Contrastive Language-Image Pretraining (CLIP) exhibits strong performance across diverse vision tasks, its application to person representation learning faces two critical challenges: (i) the scarcity of large-scale annotated vision-language data focused on person-centric images, and (ii) the inherent limitations of global contrastive learning, which struggles to maintain discriminative local features crucial for fine-grained matching while remaining vulnerable to noisy text tokens.This work advances CLIP for person representation learning through synergistic improvements in data curation and model architecture.First, we develop a noise-resistant data construction pipeline that leverages the in-context learning capabilities of MLLMs to automatically filter and caption web-sourced images.This yields WebPerson, a large-scale dataset of 5M high-quality person-centric image-text pairs.Second, we introduce the GA-DMS (Gradient-Attention Guided Dual-Masking Synergetic) framework, which improves cross-modal alignment by adaptively masking noisy textual tokens based on the gradient-attention similarity score.Additionally, we incorporate masked token prediction objectives that compel the model to predict informative text tokens, enhancing fine-grained semantic representation learning.Extensive experiments show that GA-DMS achieves state-of-the-art performance across multiple benchmarks. Tianlu Zheng, Xiang An, Ziyong Feng, Kaicheng Yang 0002, Qichuan Ding |
EMNLP | 4 |
| 2025 | UniVerse: Unleashing the Scene Prior of Video Diffusion Models for Robust Radiance Field ReconstructionabstractThis paper tackles the challenge of robust reconstruction, i.e., the task of reconstructing a 3D scene from a set of inconsistent multi-view images. Some recent works have attempted to simultaneously remove image inconsistencies and perform reconstruction by integrating image degradation modeling into neural 3D scene representations. However, these methods rely heavily on dense observations for robustly optimizing model parameters. To address this issue, we propose to decouple robust reconstruction into two subtasks: restoration and reconstruction, which naturally simplifies the optimization process. To this end, we introduce UniVerse, a unified framework for robust reconstruction based on a video diffusion model. Specifically, UniVerse first converts inconsistent images into initial videos, then uses a specially designed video diffusion model to restore them into consistent images, and finally reconstructs the 3D scenes from these restored images. Compared with case-by-case per-view degradation modeling, the diffusion model learns a general scene prior from large-scale data, making it applicable to diverse image inconsistencies. Extensive experiments on both synthetic and real-world datasets demonstrate the strong generalization capability and superior performance of our method in robust reconstruction. Moreover, UniVerse can control the style of the reconstructed 3D scene. Project page: https://jin-cao-tma.github.io/UniVerse.github.io/ Hongrui Wu, Ziyong Feng, Hujun Bao, Xiaowei Zhou 0001, Sida Peng |
ICCV | 3 |
| 2025 | HUST: High-Fidelity Unbiased Skin Tone Estimation via Texture Quantization
Zimin Ran, Xingyu Ren, Xiang An, Kaicheng Yang 0002, Ziyong Feng, Jing Yang 0038, Rolandos Alexandros Potamias, Linchao Zhu, Jiankang Deng |
ICCV | 5 |
| 2025 | MotionStreamer: Streaming Motion Generation via Diffusion-Based Autoregressive Model in Causal Latent Space
Lixing Xiao, Shunlin Lu, Huaijin Pi, Liang Pan, Yueer Zhou, Ziyong Feng, Xiaowei Zhou 0001, Sida Peng, Jingbo Wang 0003 |
ICCV | 7 |
| 2025 | Region-based Cluster Discrimination for Visual Representation LearningabstractLearning visual representations is foundational for a broad spectrum of downstream tasks. Although recent vision-language contrastive models, such as CLIP and SigLIP, have achieved impressive zero-shot performance via large-scale vision-language alignment, their reliance on global representations constrains their effectiveness for dense prediction tasks, such as grounding, OCR, and segmentation. To address this gap, we introduce Region-Aware Cluster Discrimination (RICE), a novel method that enhances region-level visual and OCR capabilities. We first construct a billion-scale candidate region dataset and propose a Region Transformer layer to extract rich regional semantics. We further design a unified region cluster discrimination loss that jointly supports object and OCR learning within a single classification framework, enabling efficient and scalable distributed training on large-scale data. Extensive experiments show that RICE consistently outperforms previous methods on tasks, including segmentation, dense detection, and visual perception for Multimodal Large Language Models (MLLMs). The pre-trained models have been released at https://github.com/deepglint/MVT. Yin Xie, Kaicheng Yang 0002, Xiang An, Yongle Zhao, Weimo Deng, Zimin Ran, Ziyong Feng, Roy Miles, Ismail Elezi, Jiankang Deng |
ICCV | 9 |
| 2025 | Dual-Level Open-Vocabulary 3D Scene Representation for Instance-Aware Robot NavigationabstractAdvanced scene understanding is crucial for robots to navigate robustly in complex 3D environments. Recent works utilize large Vision-Language Models (VLMs) to embed semantic information into reconstructed maps, thereby creating open-vocabulary scene representations for instance-aware robot navigation. However, existing methods primarily generate point-wise feature vectors for maps, which inadequately capture the intricate scene contents necessary for navigation tasks, including holistic and relational object information. To address this limitation, we propose a novel Dual-Level Open-Vocabulary 3D (DLOV-3D) scene representation framework to improve robot navigation performance. Our framework integrates both pixel-level and image-level features into spatial scene representations, facilitating a more comprehensive understanding of the scene. By incorporating an adaptive revalidation mechanism, DLOV-3D achieves precise instance-aware navigation based on free-form queries that describe object properties such as color, shape, and relational references. Notably, when combined with Large Language Models (LLMs), DLOV-3D supports long-sequence multi-instance robot navigation guided by natural language instructions. Extensive experimental results demonstrate that DLOV-3D achieves new state-of-the-art performance in instance-aware robot navigation. Tianlu Zheng, Kaicheng Yang 0002, Yilong Dou, Ziyong Feng, Qichuan Ding |
IROS | 4 |
| 2025 | Breaking the Modality Barrier: Universal Embedding Learning with Multimodal LLMsabstractThe Contrastive Language-Image Pre-training (CLIP) framework has become a widely used approach for multimodal representation learning, particularly in image-text retrieval and clustering. However, its efficacy is constrained by three key limitations: (1) text token truncation, (2) isolated image-text encoding, and (3) deficient compositionality due to bag-of-words behavior. While recent Multimodal Large Language Models (MLLMs) have demonstrated significant advances in generalized vision-language understanding, their potential for learning transferable multimodal representations remains underexplored. In this work, we present UniME (Universal Multimodal Embedding), a novel two-stage framework that leverages MLLMs to learn discriminative representations for diverse downstream tasks. In the first stage, we perform textual discriminative knowledge distillation from a powerful LLM-based teacher model to enhance the embedding capability of the MLLM's language component. In the second stage, we introduce hard negative enhanced instruction tuning to further advance discriminative representation learning. Specifically, we initially mitigate false negative contamination and then sample multiple hard negatives per instance within each batch, forcing the model to focus on challenging samples. This approach not only improves discriminative power but also enhances instruction-following ability in downstream tasks. We conduct extensive experiments on the MMEB benchmark and multiple retrieval tasks, including short & long caption retrieval and compositional retrieval. Results demonstrate that UniME achieves consistent performance improvement across all tasks, exhibiting superior discriminative and compositional capabilities. The code will be released in https://garygutc.github.io/UniME. Tiancheng Gu, Kaicheng Yang 0002, Ziyong Feng, Yanzhao Zhang, Dingkun Long, Yingda Chen, Tom Weidong Cai, Jiankang Deng |
ACM Multimedia | 3 |
| 2025 | RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation ParadigmabstractAfter pre-training on extensive image-text pairs, Contrastive Language-Image Pre-training (CLIP) demonstrates promising performance on a wide variety of benchmarks. However, a substantial volume of multimodal interleaved documents remains underutilized for contrastive vision-language representation learning. To fully leverage these unpaired documents, we initially establish a Real-World Data Extraction pipeline to extract high-quality images and texts. Then we design a hierarchical retrieval method to efficiently associate each image with multiple semantically relevant realistic texts. To further enhance fine-grained visual information, we propose an image semantic augmented generation module for synthetic text production. Furthermore, we employ a semantic balance sampling strategy to improve dataset diversity, enabling better learning of long-tail concepts. Based on these innovations, we construct RealSyn, a dataset combining realistic and synthetic texts, available in three scales: 15M, 30M, and 100M. We compare our dataset with other widely used datasets of equivalent scale for CLIP training. Models pre-trained on RealSyn consistently achieve state-of-the-art performance across various downstream tasks, including linear probe, zero-shot transfer, zero-shot robustness, and zero-shot retrieval. Furthermore, extensive experiments confirm that RealSyn significantly enhances contrastive vision-language representation learning and demonstrates robust scalability. The code will be released in https://garygutc.github.io/RealSyn. Tiancheng Gu, Kaicheng Yang 0002, Chaoyi Zhang, Yin Xie, Xiang An, Ziyong Feng, Dongnan Liu, Tom Weidong Cai, Jiankang Deng |
ACM Multimedia | 6 |
| 2025 | Decoupled Global-Local Alignment for Improving Compositional UnderstandingabstractContrastive Language-Image Pre-training (CLIP) has achieved success on multiple downstream tasks by aligning image and text modalities. However, the nature of global contrastive learning limits CLIP's ability to comprehend compositional concepts, such as relations and attributes. Although recent studies employ global hard negative samples to improve compositional understanding, these methods significantly compromise the model's inherent general capabilities by forcibly distancing textual negative samples from images in the embedding space. To overcome this limitation, we introduce a Decoupled Global-Local Alignment (DeGLA) framework that improves compositional understanding while substantially mitigating losses in general capabilities. To optimize the retention of the model's inherent capabilities, we incorporate a self-distillation mechanism within the global alignment process, aligning the learnable image-text encoder with a frozen teacher model derived from an exponential moving average. Under the constraint of self-distillation, it effectively mitigates the catastrophic forgetting of pretrained knowledge during fine-tuning. To improve compositional understanding, we first leverage the in-context learning capability of Large Language Models (LLMs) to construct about 2M high-quality negative captions across five types. Subsequently, we propose the Image-Grounded Contrast (IGC) loss and Text-Grounded Contrast (TGC) loss to enhance vision-language compositionally. Experimental results across both general and compositional reasoning tasks validate the effectiveness of the DeGLA framework. Our code is released at https://github.com/xiaoxing2001/DeGLA. Xiaoxing Hu, Kaicheng Yang 0002, Ziyong Feng, Yupei Wang |
ACM Multimedia | 5 |
| 2025 | Grounding Deliberate Reasoning in Multimodal Large Language Models
Yuxuan Liu 0011, Dehu Li, Xiang An, Weimo Deng, Ziyong Feng, Yongle Zhao, Yin Xie |
MMM (2) | 6 |
| 2025 | UniViT: Unifying Image and Video Understanding in One Vision EncoderabstractDespite the impressive progress of recent pretraining methods on multimodal tasks, existing methods are inherently biased towards either spatial modeling (e.g., CLIP) or temporal modeling (e.g., V-JEPA), limiting their joint capture of spatial details and temporal dynamics. To this end, we propose UniViT, a cluster-driven unified self-supervised learning framework that effectively captures the structured semantics of both image spatial content and video temporal dynamics through event-level and object-level clustering and discrimination. Specifically, we leverage offline clustering to generate semantic clusters across both modalities. For videos, multi-granularity event-level clustering progressively expands from single-event to structured multi-event segments, capturing coarse-to-fine temporal semantics; for images, object-level clustering captures fine-grained spatial semantics. However, while global clustering provides semantically consistent clusters, it lacks modeling of structured semantic relations (e.g., temporal event structures). To address this, we introduce a contrastive objective that leverages these semantic clusters as pseudo-label supervision to explicitly enforce structural constraints, including temporal event relations and spatial object co-occurrences, capturing structured semantics beyond categories. Meanwhile, UniViT jointly embeds structured object-level and event-level semantics into a unified representation space. Furthermore, UniViT introduces two key components: (i) Unified Rotary Position Embedding integrates relative positional embedding with frequency-aware dimension allocation to support position-invariant semantic learning and enhance the stability of structured semantics in the discrimination stage; and (ii) Variable Spatiotemporal Streams adapt to inputs of varying frame lengths, addressing the rigidity of conventional fixed-input approaches. Extensive experiments across varying model scales demonstrate that UniViT achieves state-of-the-art performance on linear probing, attentive probing, question answering, and spatial understanding tasks. Xiang An, Yin Xie, Kaicheng Yang 0002, Zimin Ran, Muhammad Imran Razzak, Ziyong Feng, Behzad Bozorgtabar, Jiankang Deng, ZongYuan Ge |
NeurIPS | 11 |
| 2025 | ORID: Organ-Regional Information Driven Framework for Radiology Report GenerationabstractThe objective of Radiology Report Generation (RRG) is to automatically generate coherent textual analyses of diseases based on radiological images, thereby alleviating the workload of radiologists. Current AI-based methods for RRG primarily focus on modifications to the encoder-decoder model architecture. To advance these approaches, this paper introduces an Organ-Regional Information Driven (ORID) framework which can effectively integrate multi-modal information and reduce the influence of noise from unrelated organs. Specifically, based on the LLaVA-Med, we first construct an RRG-related instruction dataset to improve organ-regional diagnosis description ability and get the LLaVA-Med-RRG. After that, we propose an organ-based cross-modal fusion module to effectively combine the information from the organ-regional diagnosis description and radiology image. To further reduce the influence of noise from unrelated organs on the radiology report generation, we introduce an organ importance coefficient analysis module, which leverages Graph Neural Network (GNN) to examine the interconnections of the cross-modal information of each organ region. Extensive experiments and comparisons with state-of-the-art methods across various evaluation metrics demonstrate the superior performance of our proposed method. Tiancheng Gu, Kaicheng Yang 0002, Xiang An, Ziyong Feng, Dongnan Liu, Tom Weidong Cai |
WACV | 4 |
| 2024 | Multi-label Cluster Discrimination for Visual Representation Learning
Xiang An, Kaicheng Yang 0002, Xiangzi Dai, Ziyong Feng, Jiankang Deng |
ECCV (27) | 4 |
| 2024 | RWKV-CLIP: A Robust Vision-Language Representation LearnerabstractContrastive Language-Image Pre-training (CLIP) has significantly improved performance in various vision-language tasks by expanding the dataset with image-text pairs obtained from the web.This paper further explores CLIP from the perspectives of data and model architecture.To mitigate the impact of the noise data and enhance the quality of large-scale image-text data crawled from the internet, we introduce a diverse description generation framework that can leverage Large Language Models (LLMs) to combine and refine information from web-based image-text pairs, synthetic captions, and detection tags.Additionally, we propose RWKV-CLIP, the first RWKV-driven vision-language representation learning model that combines the effective parallel training of transformers with the efficient inference of RNNs.Extensive experiments across different model scales and pre-training datasets demonstrate that RWKV-CLIP is a robust vision-language representation learner and it achieves state-of-the-art performance across multiple downstream tasks, including linear probing, zero-shot classification, and zero-shot image-text retrieval.To facilitate future research, the code and pre-trained models are released at https: //github.com/deepglint/RWKV-CLIP. Tiancheng Gu, Kaicheng Yang 0002, Xiang An, Ziyong Feng, Dongnan Liu, Tom Weidong Cai, Jiankang Deng |
EMNLP | 4 |
| 2023 | ALIP: Adaptive Language-Image Pre-training with Synthetic CaptionabstractContrastive Language-Image Pre-training (CLIP) has significantly boosted the performance of various vision-language tasks by scaling up the dataset with image-text pairs collected from the web. However, the presence of intrinsic noise and unmatched image-text pairs in web data can potentially affect the performance of representation learning. To address this issue, we first utilize the OFA model to generate synthetic captions that focus on the image content. The generated captions contain complementary information that is beneficial for pre-training. Then, we propose an Adaptive Language-Image Pre-training (ALIP), a bi-path model that integrates supervision from both raw text and synthetic caption. As the core components of ALIP, the Language Consistency Gate (LCG) and Description Consistency Gate (DCG) dynamically adjust the weights of samples and image-text/caption pairs during the training process. Meanwhile, the adaptive contrastive loss can effectively reduce the impact of noise data and enhances the efficiency of pre-training data. We validate ALIP with experiments on different scales of models and pre-training datasets. Experiments results show that ALIP achieves state-of-the-art performance on multiple downstream tasks including zero-shot image-text retrieval and linear probe. To facilitate future research, the code and pre-trained models are released at https://github.com/deepglint/ALIP. Kaicheng Yang 0002, Jiankang Deng, Xiang An, Ziyong Feng, Jia Guo 0003, Jing Yang 0038, Tongliang Liu |
ICCV | 5 |
| 2023 | Unicom: Universal and Compact Representation Learning for Image Retrieval
Xiang An, Jiankang Deng, Kaicheng Yang 0002, Jaiwei Li, Ziyong Feng, Jia Guo 0003, Jing Yang 0038, Tongliang Liu |
ICLR | 5 |
| 2022 | Killing Two Birds with One Stone: Efficient and Robust Training of Face Recognition CNNs by Partial FCabstractLearning discriminative deep feature embeddings by using million-scale in-the-wild datasets and margin-based softmax loss is the current state-of-the-art approach for face recognition. However, the memory and computing cost of the Fully Connected (FC) layer linearly scales up to the number of identities in the training set. Besides, the largescale training data inevitably suffers from inter-class conflict and long-tailed distribution. In this paper, we propose a sparsely updating variant of the FC layer, named Partial FC (PFC). In each iteration, positive class centers and a random subset of negative class centers are selected to compute the margin-based softmax loss. All class centers are still maintained throughout the whole training process, but only a subset is selected and updated in each iteration. Therefore, the computing requirement, the probability of inter-class conflict, and the frequency of passive update on tail class centers, are dramatically reduced. Extensive experiments across different training data and backbones (e.g. CNN and ViT) confirm the effectiveness, robustness and efficiency of the proposed PFC. The source code is available at https://github.com/deepinsight/insightface/tree/master/recognition. Xiang An, Jiankang Deng, Jia Guo 0003, Ziyong Feng, Xuhan Zhu, Jing Yang 0038, Tongliang Liu |
CVPR | 4 |
| 2017 | Facial attractiveness prediction using psychologically inspired convolutional neural network (PI-CNN)abstractThis paper proposes a psychologically inspired convolutional neural network (PI-CNN) to achieve automatic facial beauty prediction. Different from the previous methods, the PI-CNN is a hierarchical model that facilitates both the facial beauty representation learning and predictor training. Inspired by the recent psychological studies, significant appearance features of facial detail, lighting and color were used to optimize the PI-CNN facial beauty predictor using a new cascaded fine-tuning method. Experiments indicate that the cascaded fine-tuned PI-CNN predictor is robust to facial appearance variances, and obtains the highest correlation of 0.87 in the SCUT-FBP benchmark database, which is superior to the related hand-designed feature and related deep learning methods. Jie Xu 0041, Lingyu Liang, Ziyong Feng, Duorui Xie, Huiyun Mao |
ICASSP | 4 |
| 2017 | Identifying Machine-Printed and Handwritten Texts Using DropRegion and Deep Convolutional NetworkabstractIn this paper, we propose a deep convolutional neural network to identify machine-printed and handwritten texts. We also propose a novel data augmentation technique called DropRegion to make up for the lack of available data and enhance the generalization of the model. DropRegion increases data diversity by randomly dropping one of the stroke-containing regions in each raw input text-line image. Two parameters are introduced to make DropRegion adjustable for different data. For distinguishing texts of mixture of five languages including English, Chinese, Japanese, Korean and Russian, we have successfully achieved a very promising accuracy of 99.07% after DropRegion is applied, which is a significantly better performance compared to traditional method (97.91%) and our deep convolutional network baseline (98.75%). Zhaoyang Yang, Ziyong Feng, Jun Sun 0004, Weiying Zhou |
ICDAR | 3 |
| 2017 | Robust shared feature learning for script and handwritten/machine-printed identification
Ziyong Feng, Zhaoyang Yang, Shuangping Huang, Jun Sun 0004 |
Pattern Recognit. Lett. | 1 |
| 2016 | Learning deep neural network using max-margin minimum classification errorabstractDeep neural networks (DNNs) have recently achieved state-of-the-art performance on various tasks, such as image classification, handwriting recognition, text spotting, and speech recognition. Most of these DNNs use softmax regression and cross-entropy loss to calculate the loss function for optimization. However, the loss function is merely expected to raise the output of the true class and reduce others without distinction. In this paper, we propose a new max-margin minimum classification error (M3CE) training method, which is inspired by the traditional minimum classification error, but is more appropriate for training DNNs. The proposed M3CE aims not only to increase the posteriori of the true class but also to decrease the output of the most confused class, which can cover any shortage of the cross-entropy loss. We evaluate the M3CE on two popular datasets, MNIST and CIFAR-10. Experimental results show that the M3CE complements cross-entropy efficiently and achieves better performance. Ziyong Feng, Zenghui Sun |
ICASSP | 1 |
| 2016 | Convolutional Multi-directional Recurrent Network for Offline Handwritten Text RecognitionabstractIn this paper, we propose a new network architecture called Convolutional Multi-directional Recurrent Network (CDRN) for offline handwritten text recognition. The conventional recurrent neural network model obtains the local context from limited directions, whereas we build up the multi-directional long short-term memory (MDirLSTM) module to abstract contextual information in various directions. Moreover, we develop a shortcut connection strategy in our proposed architecture for faster yet better convergence. In cooperation with the aforementioned methods, the proposed architecture also benefits from the following properties: (1) it obtains informative features of the input directly without involving hand-crafted features and segmentation, and (2) it is an end-to-end trainable model whose components are trained conjointly. We evaluate the performance of the proposed method on two databases: IAM words and IRONOFF. Our experimental results demonstrate a significant increase in recognition performance using MDirLSTM and shortcut connections, which suggests the effectiveness of these two proposed methods. Zenghui Sun, Zecheng Xie, Ziyong Feng, Shuye Zhang |
ICFHR | 4 |
| 2016 | Handwritten/Printed Receipt Classification Using Attention-Based Convolutional Neural NetworkabstractThis paper presents an approach for the classification of handwritten and printed receipts based on a convolutional neural network (CNN). One of the main challenges related to such classification is the diversity of the background interference in the receipt images. To overcome this problem, we propose a new technique named "attention-based CNN" (ABCNN), inspired by the concept of "attention" in visual neuroscience. This approach helps us to focus on the receipt in an image without bounding box annotation. Our experimental results showed that the proposed ABCNN (i) significantly improves the classification accuracy compared to normal CNN (from 95% to 98.25%), and (ii) enables the network to process images directly without object detection, and (iii) it is faster to train and test the network. Ziyong Feng, Shuye Zhang |
ICFHR | 4 |
| 2016 | Fully convolutional recurrent network for handwritten Chinese text recognitionabstractThis paper proposes an end-to-end framework, namely fully convolutional recurrent network (FCRN) for handwritten Chinese text recognition (HCTR). Unlike traditional methods that rely heavily on segmentation, our FCRN is trained with online text data directly and learns to associate the pen-tip trajectory with a sequence of characters. FCRN consists of four parts: a path-signature layer to extract signature features from the input pen-tip trajectory, a fully convolutional network to learn informative representation, a sequence modeling layer to make per-frame predictions on the input sequence and a transcription layer to translate the predictions into a label sequence. We also present a refined beam search method that efficiently integrates the language model to decode the FCRN and significantly improve the recognition results. We evaluate the performance of the proposed method on the test sets from the databases CASIA-OLHWDB and ICDAR 2013 Chinese handwriting recognition competition, and both achieve state-of-the-art performance with correct rates of 96.40% and 95.00%, respectively. Zecheng Xie, Zenghui Sun, Ziyong Feng, Shuye Zhang |
ICPR | 4 |
| 2016 | DropSample: A new training method to enhance deep convolutional neural networks for large-scale unconstrained handwritten Chinese character recognition
Dacheng Tao, Zecheng Xie, Ziyong Feng |
Pattern Recognit. | 5 |
| 2015 | Improved deep convolutional neural network for online handwritten Chinese character recognition using domain-specific knowledgeabstractDeep convolutional neural networks (DCNNs) have achieved great success in various computer vision and pattern recognition applications, including those for handwritten Chinese character recognition (HCCR). However, most current DCNN-based HCCR approaches treat the handwritten sample simply as an image bitmap, ignoring some vital domain-specific information that may be useful but that cannot be learnt by traditional networks. In this paper, we propose an enhancement of the DCNN approach to online HCCR by incorporating a variety of domain-specific knowledge, including deformation, non-linear normalization, imaginary strokes, path signature, and 8-directional features. Our contribution is twofold. First, these domain-specific technologies are investigated and integrated with a DCNN to form a composite network to achieve improved performance. Second, the resulting DCNNs with diversity in their domain knowledge are combined using a hybrid serial-parallel (HSP) strategy. Consequently, we achieve a promising accuracy of 97.20% and 96.87% on CASIA-OLHWDB1.0 and CASIA-OLHWDB1.1, respectively, outperforming the best results previously reported in the literature. Zecheng Xie, Ziyong Feng |
ICDAR | 4 |
| 2015 | Multi-font printed Chinese character recognition using multi-pooling convolutional neural networkabstractAlthough previous studies have achieved effective printed Chinese character recognition (PCCR) in the case a single font or a few different fonts, large scale multi-font PCCR remains a major challenge owing to the wide variety in the shape, layout, and grey-level distribution of single Chinese characters across different font styles. This paper applies multi-pooling and data augmentation with non-linear transformation to a convolutional neural network (CNN) for multi-font PCCR. We propose a multi-pooling layer on top of the final convolutional layer; this approach is found to be robust to spatial layout variations and deformations in multi-font printed Chinese characters. Experimental results show that multi-pooling significantly improves CNN performance. In addition, we adopt a distorted sample generation technique by applying non-linear warping functions along an original font image, which distorts the local density of image-based Chinese character strokes. We find that CNN performance is further boosted by the distorted samples technique. An input character image is transformed into four distorted images and the CNN learns the original image as well as the distorted samples to classify 3755 classes (level-1 set of GB2312-80) of printed Chinese characters in 280 widely varying fonts and 120 manually selected fonts. Outstanding recognition rates of 94.38% and 99.74% are achieved in the former and latter cases, respectively, which indicates the effectiveness of the proposed methods. Zhuoyao Zhong, Ziyong Feng |
ICDAR | 3 |
| 2015 | DLANet: A manifold-learning-based discriminative feature learning network for scene classification
Ziyong Feng, Dapeng Tao, Shuangping Huang |
Neurocomputing | 1 |