EDBT 2026 Demo / reviewers in the wild / expert
Zhineng Chen
dblp:96/7998
· DBLP profile ↗
95ranked-venue papers
11as first author
51since 2021 · last 2026
0000-0003-1543-6889ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 65 · 8 first-author · 38 since 2021Artificial intelligence and machine learning · 44 · 1 first-author · 34 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 1 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Complex Mathematical Expression Recognition: Benchmark, Large-Scale Dataset and Strong BaselineabstractMathematical Expression Recognition (MER) has made significant progress in recognizing simple expressions, but the robust recognition of complex mathematical expressions with many tokens and multiple lines remains a formidable challenge. In this paper, we first introduce CMER-Bench, a carefully constructed benchmark that categorizes expressions into three difficulty levels: easy, moderate, and complex. Leveraging CMER-Bench, we conduct a comprehensive evaluation of existing MER models and general-purpose multimodal large language models (MLLMs). The results reveal that while current methods perform well on easy and moderate expressions, their performance degrades significantly when handling complex mathematical expressions, mainly because existing public training datasets are primarily composed of simple samples. In response, we propose MER-17M and CMER-3M that are large-scale datasets emphasizing the recognition of complex mathematical expressions. The datasets provide rich and diverse samples to support the development of accurate and robust complex MER models. Furthermore, to address the challenges posed by the complicated spatial layout of complex expressions, we introduce a novel expression tokenizer, and a new representation called Structured Mathematical Language, which explicitly models the hierarchical and spatial structure of expressions beyond LaTeX format. Based on these, we propose a specialized model named CMERNet, built upon an encoder-decoder architecture and trained on CMER-3M. Experimental results show that CMERNet, with only 125 million parameters, significantly outperforms existing MER models and MLLMs on CMER-Bench. Weikang Bai, Yongkun Du, Yazhen Xie, Zhineng Chen |
AAAI | 5 |
| 2026 | MDiff4STR: Mask Diffusion Model for Scene Text RecognitionabstractMask Diffusion Models (MDMs) have recently emerged as a promising alternative to auto-regressive models (ARMs) for vision-language tasks, owing to their flexible balance of efficiency and accuracy. In this paper, for the first time, we introduce MDMs into the Scene Text Recognition (STR) task. We show that vanilla MDM lags behind ARMs in terms of accuracy, although it improves recognition efficiency. To bridge this gap, we propose MDiff4STR, a Mask Diffusion model enhanced with two key improvement strategies tailored for STR. Specifically, we identify two key challenges in applying MDMs to STR: noising gap between training and inference, and overconfident predictions during inference. Both significantly hinder the performance of MDMs. To mitigate the first issue, we develop six noising strategies that better align training with inference behavior. For the second, we propose a token-replacement noise mechanism that provides a non-mask noise type, encouraging the model to reconsider and revise overly confident but incorrect predictions. We conduct extensive evaluations of MDiff4STR on both standard and challenging STR benchmarks, covering diverse scenarios including irregular, artistic, occluded, and Chinese text, as well as whether the use of pretraining. Across these settings, MDiff4STR consistently outperforms popular STR models, surpassing state-of-the-art ARMs in accuracy, while maintaining fast inference with only three denoising steps. Code: https://github.com/Topdu/OpenOCR. Yongkun Du, Miaomiao Zhao, Songlin Fan, Zhineng Chen, Caiyan Jia, Yu-Gang Jiang 0001 |
AAAI | 4 |
| 2026 | SSR-SAM: Retrieval-Style Segment Anything Model for Semi-Supervised Ultra-High-Resolution Image SegmentationabstractAccurate segmentation of ultra-high-resolution (UHR) images, which often exceed tens of millions of pixels, is critically important in domains such as remote sensing and biomedical imaging. However, acquiring pixel-level annotations for such high-resolution images is prohibitively expensive and labor-intensive. While semi-supervised semantic segmentation can significantly reduce the annotation burden, its extension to UHR images holds great potential for addressing the unique challenges posed by sparse supervision. To this end, we propose SSR-SAM, a retrieval-style semi-supervised segmentation framework tailored for UHR images. Leveraging the promptable paradigm of the Segment Anything Model (SAM), SSR-SAM treats locally annotated regions as prompts to retrieve semantically consistent pixels across the entire image. Building upon this retrieval-style segmentation paradigm, we further introduce prompt-level perturbation, a novel trail to deploy consistency regularization for semi-supervised segmentation. It encourages the model to learn consistency across predictions guided by diverse visual-semantic prompts, thereby enhancing generalization on unlabeled data. We evaluate SSR-SAM on three UHR datasets: Inria Aerial, BCSS, and URUR. Experimental results show that SSR-SAM achieves clear performance gains over the labeled-only supervision, with average mIoU improvements of 4.9%, 4.15%, and 2.5%, respectively. Additionally, SSR-SAM possesses zero-shot segmentation capability, exhibiting potential for general retrieval-style segmentation tasks. Zhineng Chen, Kai Hu 0002, Xieping Gao 0001 |
AAAI | 3 |
| 2026 | ICPR 2026 Competition on Low-Resolution License Plate Recognition
Rayson Laroca, Valfride Nascimento, Donggun Kim 0004, Sanghyeok Chung, Subin Bae, Uihwan Seo, Seungsang Oh, Chi M. Phung, Minh G. Vo, Xingsong Ye, Yongkun Du, Zhineng Chen, Sunhee Heo, Hyangwoo Lee, Kihyun Na, Khanh V. Vu Nguyen, Sang T. Pham, Duc N. N. Phung, Trong P. Le, Vy N. Vo Tran, David Menotti |
ICPR (16) | 13 |
| 2026 | CSFMIL: Whole slide image classification with two-stage cross-scale fusion
Zhineng Chen, Feng-Jung Chen, Kai Hu 0002, Xieping Gao 0001 |
Neurocomputing | 3 |
| 2026 | LRANet++: Low-Rank Approximation Network for Accurate and Efficient Text SpottingabstractEnd-to-end text spotting aims to jointly optimize text detection and recognition within a unified framework. Despite significant progress, designing an accurate and efficient end-to-end text spotter for arbitrary-shaped text remains challenging. We identify the primary bottleneck as the lack of a reliable and efficient text detection method. To address this, we propose a novel parameterized text shape representation based on low-rank approximation for precise detection and a triple assignment detection head for fast inference. Specifically, unlike current data-irrelevant shape representation methods, we exploit shape correlations among labeled text boundaries to construct a robust low-rank subspace. By minimizing an $\ell _{1}$ℓ1-norm objective, we extract orthogonal vectors that capture the intrinsic text shape from noisy annotations, enabling precise reconstruction via the linear combination of only a few basis vectors. Next, the triple assignment scheme decouples training complexity from inference speed. It utilizes a deep sparse branch to guide an ultra-lightweight inference branch, while a dense branch provides rich parallel supervision. Building upon these advancements, we integrate the enhanced detection module with a lightweight recognition branch to form an end-to-end text spotting framework, termed LRANet++, capable of accurately and efficiently spotting arbitrary-shaped text. Extensive experiments on challenging benchmarks demonstrate the superiority of LRANet++ compared to state-of-the-art methods. Zhineng Chen, Yongkun Du, Zuxuan Wu, Hongtao Xie 0001, Yu-Gang Jiang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | TextDCTv2: Scene Text Detection With Patch-Based Discrete Cosine Transform RepresentationabstractRegression-based scene text detection methods, which mainly locate text by regressing parameterized curves, have become popular recently. However, parameterized curves typically require a pre-defined starting coordinate and direction for the contour curves to ensure consistent learning, which limits their flexibility in representing arbitrary-shaped text. Methods based on the parameterized mask would be a promising alternative. Text instance mask naturally has consistency and is less noisy than contour vertices representing the parameterized curves. However, this branch of methods still remains less studied. In this work, we propose a novel scene text detection framework, named TextDCTv2, which adopts the discrete cosine transform (DCT) to represent each text instance mask as a compact vector. Instead of simply applying DCT to the entire text instance mask, we split the mask into multiple patches and perform DCT on each patch individually. This allows us to focus on regions that are difficult to detect and then refine them without affecting other patches. Specifically, each patch is first transformed into a frequency domain matrix using DCT. Then, we introduce a direct current identifier to detect foreground, background and challenging patches based on direct current coefficients, i.e., the zero-frequency components of the DCT matrix that reflect the overall brightness of a patch. Next, for challenging patches, we employ a lightweight refinement regressor to further correct their DCT vectors and generate precise text boundaries. Through these designs, we achieve an efficient and precise text mask representation. Extensive experiments prove that our method is superior to most of existing methods on various datasets. Specifically, TextDCTv2 achieves an F-measure of 88.7 at 20.3 frames per second (FPS) on the Total-Text dataset, and an Fmeasure of 87.8 at 30.3 FPS on the CTW1500 dataset. Zhaolong Pan, Lina Tan, Peng Gao 0012, Zhineng Chen |
IEEE Trans. Multim. | 6 |
| 2025 | Out of Length Text Recognition with Sub-String MatchingabstractScene Text Recognition (STR) methods have demonstrated robust performance in word-level text recognition. However, in real applications the text image is sometimes long due to detected with multiple horizontal words. It triggers the requirement to build long text recognition models from readily available short (i.e., word-level) text datasets, which has been less studied previously. In this paper, we term this task Out of Length (OOL) text recognition. We establish the first Long Text Benchmark (LTB) to facilitate the assessment of different methods in long text recognition. Meanwhile, we propose a novel method called OOL Text Recognition with sub-String Matching (SMTR). SMTR comprises two cross-attention-based modules: one encodes a sub-string containing multiple characters into next and previous queries, and the other employs the queries to attend to the image features, matching the sub-string and simultaneously recognizing its next and previous character. SMTR can recognize text of arbitrary length by iterating the process above. To avoid being trapped in recognizing highly similar sub-strings, we introduce a regularization training to compel SMTR to effectively discover subtle differences between similar sub-strings for precise matching. In addition, we propose an inference augmentation strategy to alleviate confusion caused by identical sub-strings in the same text and improve the overall recognition efficiency. Extensive experimental results reveal that SMTR, even when trained exclusively on short text, outperforms existing methods in public short text benchmarks and exhibits a clear advantage on LTB. Yongkun Du, Zhineng Chen, Caiyan Jia, Xieping Gao 0001, Yu-Gang Jiang 0001 |
AAAI | 2 |
| 2025 | Distilling Knowledge from Heterogeneous Architectures for Semantic SegmentationabstractCurrent knowledge distillation (KD) methods for semantic segmentation focus on guiding the student to imitate the teacher's knowledge within homogeneous architectures. However, these methods overlook the diverse knowledge contained in architectures with different inductive biases, which is crucial for enabling the student to acquire a more precise and comprehensive understanding of the data during distillation. To this end, we propose for the first time a generic knowledge distillation method for semantic segmentation from a heterogeneous perspective, named HeteroAKD. Due to the substantial disparities between heterogeneous architectures, such as CNN and Transformer, directly transferring cross-architecture knowledge presents significant challenges. To eliminate the influence of architecture-specific information, the intermediate features of both the teacher and student are skillfully projected into an aligned logits space. Furthermore, to utilize diverse knowledge from heterogeneous architectures and deliver customized knowledge required by the student, a teacher-student knowledge mixing mechanism (KMM) and a teacher-student knowledge evaluation mechanism (KEM) are introduced. These mechanisms are performed by assessing the reliability and its discrepancy between heterogeneous teacher-student knowledge. Extensive experiments conducted on three main-stream benchmarks using various teacher-student pairs demonstrate that our HeteroAKD framework outperforms state-of-the-art KD methods in facilitating distillation between heterogeneous architectures. Yanglin Huang, Kai Hu 0002, Yuan Zhang 0022, Zhineng Chen, Xieping Gao 0001 |
AAAI | 4 |
| 2025 | Explicit Relational Reasoning Network for Scene Text DetectionabstractConnected component (CC) is a proper text shape representation that aligns with human reading intuition. However, CC-based text detection methods have recently faced a developmental bottleneck that their time-consuming post-processing is difficult to eliminate. To address this issue, we introduce an explicit relational reasoning network (ERRNet) to elegantly model the component relationships without post-processing. Concretely, we first represent each text instance as multiple ordered text components, and then treat these components as objects in sequential movement. In this way, scene text detection can be innovatively viewed as a tracking problem. From this perspective, we design an end-to-end tracking decoder to achieve a CC-based method dispensing with post-processing entirely. Additionally, we observe that there is an inconsistency between classification confidence and localization quality, so we propose a Polygon Monte-Carlo method to quickly and accurately evaluate the localization quality. Based on this, we introduce a position-supervised classification loss to guide the task-aligned learning of ERRNet. Experiments on challenging benchmarks demonstrate the effectiveness of our ERRNet. It consistently achieves state-of-the-art accuracy while holding highly competitive inference speed. Zhineng Chen, Yongkun Du, Zhilong Ji, Kai Hu 0002, Jinfeng Bai, Xieping Gao 0001 |
AAAI | 2 |
| 2025 | SynTab-LLaVA: Enhancing Multimodal Table Understanding with Decoupled SynthesisabstractDue to the limited scale of multimodal table understanding (MTU) data, model performance is constrained. A straightforward approach is to use multimodal large language models to obtain more samples, but this may cause hallucinations, generate incorrect sample pairs, and cost significantly. To address the above issues, we design a simple yet effective synthesis framework that consists of two independent steps: table image rendering and table question and answer (Q&A) pairs generation. We use table codes (HTML, LaTeX, Markdown) to synthesize images and generate Q&A pairs with large language model (LLM). This approach leverages LLMs high concurrency and low cost to boost annotation efficiency and reduce expenses. By inputting code instead of images, LLMs can directly access the content and structure of the table, reducing hallucinations in table understanding and improving the accuracy of generated Q&A pairs. Finally, we synthesize a large-scale MTU dataset, SynTab, containing 636K images and 1.8M samples costing within $200 in US dollars. We further introduce a generalist tabular multimodal model, SynTab-LLaVA. This model not only effectively extracts local textual content within the table but also enables global modeling of relationships between cells. SynTab-LLaVA achieves SOTA performance on 21 out of 24 in-domain and out-of-domain benchmarks, demonstrating the effectiveness and generalization of our method. The Code is available at SynTab-LLaVA. Bangbang Zhou, Zuan Gao, Boqiang Zhang, Zhineng Chen, Hongtao Xie 0001 |
CVPR | 6 |
| 2025 | Improving Irregular Text Recognition with Adaptive Feature CompressionabstractScene text recognition models typically do not handle text irregularities well, especially for connectionist temporal classification (CTC)-based ones. In CTC models, visual features must be compressed into a one-dimensional sequence to fit the CTC decoding. Current solutions adopt simple average pooling or feature shrinking for this compression, which is a bottleneck restricting their recognition capabilities. To tackle this, we introduce an adaptive feature compression block to compress the features adaptively. It leverages the attention mechanism to selectively preserve features related to text foreground and discard those belonging to text background. As a result, the compression adaptively retains features the mostly important to recognition, and CTC models could better deal with text irregularities when equipped with this block. Correspondingly, we design a novel text recognition model termed AFCTR by appending this adaptive feature compression block to existing CTC models. Experimental results on typical English benchmarks show that AFCTR outperforms existing popular models in terms of accuracy under multiple evaluation protocols. Moreover, AFCTR also preserves the efficiency advantage of CTC models. Zhineng Chen |
ICASSP | 2 |
| 2025 | SVTRv2: CTC Beats Encoder-Decoder Models in Scene Text RecognitionabstractConnectionist temporal classification (CTC)-based scene text recognition (STR) methods, e.g., SVTR, are widely employed in OCR applications, mainly due to their simple architecture, which only contains a visual model and a CTC-aligned linear classifier, and therefore fast inference. However, they generally exhibit worse accuracy than encoder-decoder-based methods (EDTRs) due to struggling with text irregularity and linguistic missing. To address these challenges, we propose SVTRv2, a CTC model endowed with the ability to handle text irregularities and model linguistic context. First, a multi-size resizing strategy is proposed to resize text instances to appropriate predefined sizes, effectively avoiding severe text distortion. Meanwhile, we introduce a feature rearrangement module to ensure that visual features accommodate the requirement of CTC, thus alleviating the alignment puzzle. Second, we propose a semantic guidance module. It integrates linguistic context into the visual features, allowing CTC model to leverage language information for accuracy improvement. This module can be omitted at the inference stage and would not increase the time cost. We extensively evaluate SVTRv2 in both standard and recent challenging benchmarks, where SVTRv2 is fairly compared to popular STR models across multiple scenarios, including different types of text irregularity, languages, long text, and whether employing pretraining. SVTRv2 surpasses most EDTRs across the scenarios in terms of accuracy and inference speed. Code: https://github.com/Topdu/OpenOCR. Yongkun Du, Zhineng Chen, Hongtao Xie 0001, Caiyan Jia, Yu-Gang Jiang 0001 |
ICCV | 2 |
| 2025 | Igd: Instructional Graphic Design With Multimodal Layer Generatio
Yadong Qu, Hongtao Xie 0001, Yongdong Zhang 0001, Shancheng Fang, Yuxin Wang 0002, Zhineng Chen |
ICCV | 7 |
| 2025 | TextSSR: Diffusion-Based Data Synthesis for Scene Text RecognitionabstractScene text recognition (STR) suffers from challenges of either less realistic synthetic training data or the difficulty of collecting sufficient high-quality real-world data, limiting the effectiveness of trained models. Meanwhile, despite producing holistically appealing text images, diffusion-based visual text generation methods struggle to synthesize accurate and realistic instance-level text at scale. To tackle this, we introduce TextSSR: a novel pipeline for Synthesizing Scene Text Recognition training data. TextSSR targets three key synthesizing characteristics: accuracy, realism, and scalability. It achieves accuracy through a proposed region-centric text generation with position-glyph enhancement, ensuring proper character placement. It maintains realism by guiding style and appearance generation using contextual hints from surrounding text or background. This character-aware diffusion architecture enjoys precise character-level control and semantic coherence preservation, without relying on natural language prompts. Therefore, TextSSR supports large-scale generation through combinatorial text permutations. Based on these, we present TextSSR-F, a dataset of 3.55 million quality-screened text instances. Extensive experiments show that STR models trained on TextSSR-F outperform those trained on existing synthetic datasets by clear margins on common benchmarks, and further improvements are observed when mixed with real-world training data. Code is available at https://github.com/YesianRohn/TextSSR. Xingsong Ye, Yongkun Du, Yunbo Tao, Zhineng Chen |
ICCV | 4 |
| 2025 | Stochasticity-aware No-Reference Point Cloud Quality AssessmentabstractThe evolution of point cloud processing algorithms necessitates an accurate assessment for their quality. Previous works consistently regard point cloud quality assessment (PCQA) as a MOS regression problem and devise a deterministic mapping, ignoring the stochasticity in generating MOS from subjective tests. This work presents the first probabilistic architecture for no-reference PCQA, motivated by the labeling process of existing datasets. The proposed method can model the quality judging stochasticity of subjects through a tailored conditional variational autoencoder (CVAE) and produces multiple intermediate quality ratings. These intermediate ratings simulate the judgments from different subjects and are then integrated into an accurate quality prediction, mimicking the generation process of a ground truth MOS. Specifically, our method incorporates a Prior Module, a Posterior Module, and a Quality Rating Generator, where the former two modules are introduced to model the judging stochasticity in subjective tests, while the latter is developed to generate diverse quality ratings. Extensive experiments indicate that our approach outperforms previous cutting-edge methods by a large margin and exhibits gratifying crossdataset robustness. Codes are available at https://git.openi.org.cn/OpenPointCloud/nrpcqa. Songlin Fan, Wei Gao 0003, Zhineng Chen, Ge Li 0002, Qicheng Wang |
IJCAI | 3 |
| 2025 | Proactive Deepfake Detection via Self-Verifiable Semantic WatermarkingabstractMalicious Deepfakes pose serious security risks by producing highly realistic forged faces. While numerous countermeasures have been developed to train binary Deepfake classifiers, their limited generalization capacity restricts practical deployment. To proactively defend against Deepfakes, we propose SVS-WM, a Self-Verifiable Semantic Watermarking strategy. The core idea behind SVS-WM is to embed pairs of correlated watermarks within facial semantics, leveraging the inherent fragility of these features, i.e., any semantic modification will disrupt the watermark correlation, thereby enabling robust Deepfake detection. SVS-WM employs a facial semantic disentanglement and reconstruction network, allowing semi-fragile watermarks to be embedded concurrently across multiple semantic levels, including identity and multi-levels of attributes. Specifically, pairs of pseudo-random noise watermarks are adaptively injected into facial attribute and identity features. During propagation stage, the protected image may encounter identity or facial attributes manipulations, we then detect Deepfakes by verifying the correlation result between the decoded attribute watermark and the extracted identity vector. This unique cross-verification mechanism enables authentication without requiring original reference watermark, thereby realizing blind Deepfake detection. Extensive experiments validate the effectiveness of our approach, achieving an average detection accuracy of 98.19% across diverse Deepfake manipulations. Peiqi Jiang, Bohan Lei, Lingyun Yu 0002, Zhineng Chen, Hongtao Xie 0001, Yongdong Zhang 0001 |
ACM Multimedia | 5 |
| 2025 | Confusion-Driven Self-Supervised Progressively Weighted Ensemble Learning for Non-Exemplar Class Incremental LearningabstractNon-exemplar class incremental learning (NECIL) aims to continuously assimilate new knowledge while retaining previously acquired knowledge in scenarios where prior examples are unavailable. A prevalent strategy within NECIL mitigates knowledge forgetting by freezing the feature extractor after training on the initial task. However, this freezing mechanism does not provide explicit training to differentiate between new and old classes, resulting in overlapping feature representations. To address this challenge, we propose a **C**onfusion-driven se**L**f-supervised pr**O**gressi**V**ely weighted **E**nsemble lea**R**ning (*CLOVER*) framework for NECIL. Firstly, we introduce a confusion-driven self-supervised learning approach that enhances representation extraction by guiding the model to distinguish between highly confusable classes, thereby reducing class representation overlap. Secondly, we develop a progressively weighted ensemble learning method that gradually adjusts weights to integrate diverse knowledge more effectively, further minimizing representation overlap. Finally, extensive experiments demonstrate that our proposed method achieves state-of-the-art results on the CIFAR100, TinyImageNet, and ImageNet-Subset NECIL benchmarks. Kai Hu 0002, Yuan Zhang 0022, Zhineng Chen, Xieping Gao 0001 |
NeurIPS | 4 |
| 2025 | Context Perception Parallel Decoder for Scene Text RecognitionabstractScene text recognition (STR) methods have struggled to attain high accuracy and fast inference speed. Auto-Regressive (AR)-based models implement the recognition in a character-by-character manner, showing superiority in accuracy but with slow inference speed. Alternatively, Parallel Decoding (PD)-based models infer all characters in a single decoding pass, offering faster inference speed but generally worse accuracy. To realize the dual goals of "AR-level accuracy and PD-level speed", we propose a Context Perception Parallel Decoder (CPPD) to perceive the related context and predict the character sequence in a PD pass. CPPD devises a character counting module to infer the occurrence count of each character, and a character ordering module to deduce the content-free reading order and positions. Meanwhile, the character prediction task associates the positions with characters. They together build a comprehensive recognition context, which benefits the decoder to focus accurately on characters with the attention mechanism, thereby improving the recognition accuracy. We construct a series of CPPD models and also plug the proposed modules into existing STR decoders. Experiments on both English and Chinese benchmarks demonstrate that the CPPD models achieve highly competitive accuracy while running much faster than existing leading models. Moreover, the plugged models achieve significant accuracy improvements. Yongkun Du, Zhineng Chen, Caiyan Jia, Xiaoting Yin, Chenxia Li, Yuning Du, Yu-Gang Jiang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Instruction-Guided Scene Text RecognitionabstractMulti-modal models have shown appealing performance in visual recognition tasks, as free-form text-guided training evokes the ability to understand fine-grained visual content. However, current models cannot be trivially applied to scene text recognition (STR) due to the compositional difference between natural and text images. We propose a novel instruction-guided scene text recognition (IGTR) paradigm that formulates STR as an instruction learning problem and understands text images by predicting character attributes, e.g., character frequency, position, etc. IGTR first devises instruction triplets, providing rich and diverse descriptions of character attributes. To effectively learn these attributes through question-answering, IGTR develops a lightweight instruction encoder, a cross-modal feature fusion module and a multi-task answer head, which guides nuanced text image understanding. Furthermore, IGTR realizes different recognition pipelines simply by using different instructions, enabling a character-understanding-based text reasoning paradigm that differs from current methods considerably. Experiments on English and Chinese benchmarks show that IGTR outperforms existing models by significant margins, while maintaining a small model size and fast inference speed. Moreover, by adjusting the sampling of instructions, IGTR offers an elegant way to tackle the recognition of rarely appearing and morphologically similar characters, which were previous challenges. Yongkun Du, Zhineng Chen, Caiyan Jia, Yu-Gang Jiang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Multi-Perspective Pseudo-Label Generation and Confidence-Weighted Training for Semi-Supervised Semantic SegmentationabstractSelf-training has been shown to achieve remarkable gains in semi-supervised semantic segmentation by creating pseudo-labels using unlabeled data. This approach, however, suffers from the quality of the generated pseudo-labels, and generating higher quality pseudo-labels is the main challenge that needs to be addressed. In this paper, we propose a novel method for semi-supervised semantic segmentation based on Multi-perspective pseudo-label Generation and Confidence-weighted Training (MGCT). First, we present a multi-perspective pseudo-label generation strategy that considers both global and local semantic perspectives. This strategy prioritizes pixels in all images by the global and local predictions, and subsequently generates pseudo-labels for different pixels in stages according to the ranking results. Our pseudo-label generation method shows superior suitability for semi-supervised semantic segmentation compared to other approaches. Second, we propose a confidence-weighted training method to alleviate performance degradation caused by unstable pixels. Our training method assigns confident weights to unstable pixels, which reduces the interference of unstable pixels during training and facilitates the efficient training of the model. Finally, we validate our approach on the PASCAL VOC 2012 and Cityscapes datasets, and the results indicate that we achieve new state-of-the-art performance on both datasets in all settings. Kai Hu 0002, Zhineng Chen, Yuan Zhang 0022, Xieping Gao 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | LRANet: Towards Accurate and Efficient Scene Text Detection with Low-Rank Approximation NetworkabstractRecently, regression-based methods, which predict parameterized text shapes for text localization, have gained popularity in scene text detection. However, the existing parameterized text shape methods still have limitations in modeling arbitrary-shaped texts due to ignoring the utilization of text-specific shape information. Moreover, the time consumption of the entire pipeline has been largely overlooked, leading to a suboptimal overall inference speed. To address these issues, we first propose a novel parameterized text shape method based on low-rank approximation. Unlike other shape representation methods that employ data-irrelevant parameterization, our approach utilizes singular value decomposition and reconstructs the text shape using a few eigenvectors learned from labeled text contours. By exploring the shape correlation among different text contours, our method achieves consistency, compactness, simplicity, and robustness in shape representation. Next, we propose a dual assignment scheme for speed acceleration. It adopts a sparse assignment branch to accelerate the inference speed, and meanwhile, provides ample supervised signals for training through a dense assignment branch. Building upon these designs, we implement an accurate and efficient arbitrary-shaped text detector named LRANet. Extensive experiments are conducted on several challenging benchmarks, demonstrating the superior accuracy and efficiency of LRANet compared to state-of-the-art methods. Code is available at: https://github.com/ychensu/LRANet.git Zhineng Chen, Zhiwen Shao, Yuning Du, Zhilong Ji, Jinfeng Bai, Yong Zhou 0003, Yu-Gang Jiang 0001 |
AAAI | 2 |
| 2024 | Learning to Rank Patches for Unbiased Image Redundancy ReductionabstractImages suffer from heavy spatial redundancy because pixels in neighboring regions are spatially correlated. Existing approaches strive to overcome this limitation by reducing less meaningful image regions. However, current leading methods rely on supervisory signals. They may compel models to preserve content that aligns with labeled categories and discard content belonging to unlabeled categories. This categorical inductive bias makes these methods less effective in real-world scenarios. To address this issue, we propose a self-supervised framework for image redundancy reduction called Learning to Rank Patches (LTRP). We observe that image reconstruction of masked image modeling models is sensitive to the removal of visible patches when the masking ratio is high (e.g., 90%). Building upon it, we implement LTRP via two steps: inferring the semantic density score of each patch by quantifying variation between reconstructions with and without this patch, and learning to rank the patches with the pseudo score. The entire process is self-supervised, thus getting out of the dilemma of categorical inductive bias. We design extensive experiments on different datasets and tasks. The results demonstrate that LTRP outperforms both supervised and other self-supervised methods due to the fair assessment of image content. Code is available at https://github.com/irsLu/1trp. Zhineng Chen, Peng Zhou 0009, Zuxuan Wu, Xieping Gao 0001, Yu-Gang Jiang 0001 |
CVPR | 2 |
| 2024 | Improving Text-Guided Object Inpainting with Semantic Pre-inpainting
Jingwen Chen 0001, Yingwei Pan, Yehao Li, Ting Yao 0003, Zhineng Chen, Tao Mei 0001 |
ECCV (46) | 6 |
| 2024 | DreamMesh: Jointly Manipulating and Texturing Triangle Meshes for Text-to-3D Generation
Haibo Yang 0002, Yang Chen 0048, Yingwei Pan, Ting Yao 0003, Zhineng Chen, Zuxuan Wu, Yu-Gang Jiang 0001, Tao Mei 0001 |
ECCV (59) | 5 |
| 2024 | Dual Contrastive Learning Guided Pathological Image Re-StainingabstractPathological virtual re-staining is a valuable research topic in AI-aided diagnosis, as it reduces the need for costly and time-consuming physical staining. However, existing methods still suffer from the insufficient ability to preserve tissue microstructure and cellular details, making the generated images less convincing. In this paper, we propose a CycleGAN-based dual contrastive learning re-staining method called DCLRStain. DCLRStain establishes dual contrastive learning between the source and re-stained image domains, conducting negative sampling within each image pair from both domains. It guides the model’s attention to finer content such as cellular details. Meanwhile, DCLRStain introduces a structural similarity-based loss term that further forces the tissue microstructure to be consistent between the source and re-stained images. Experimental results demonstrate that DCLRStain yields competitive quantitative scores compared to state-of-the-art models and maintains superior qualitative performance. Moreover, DCLRStain achieves higher accuracy in the downstream classification task. Yuexiao Liang, Zhineng Chen, Caiyan Jia, Xiongjun Ye, Xieping Gao 0001 |
ICASSP | 2 |
| 2024 | Hi3D: Pursuing High-Resolution Image-to-3D Generation with Video Diffusion ModelsabstractDespite having tremendous progress in image-to-3D generation, existing methods still struggle to produce multi-view consistent images with high-resolution textures in detail, especially in the paradigm of 2D diffusion that lacks 3D awareness. In this work, we present High-resolution Image-to-3D model (Hi3D), a new video diffusion based paradigm that redefines a single image to multi-view images as 3D-aware sequential image generation (i.e., orbital video generation). This methodology delves into the underlying temporal consistency knowledge in video diffusion model that generalizes well to geometry consistency across multiple views in 3D generation. Technically, Hi3D first empowers the pre-trained video diffusion model with 3D-aware prior (camera pose condition), yielding multi-view images with low-resolution texture details. A 3D-aware video-to-video refiner is learnt to further scale up the multi-view images with high-resolution texture details. Such high-resolution multi-view images are further augmented with novel views through 3D Gaussian Splatting, which are finally leveraged to obtain high-fidelity meshes via 3D reconstruction. Extensive experiments on both novel view synthesis and single view reconstruction demonstrate that our Hi3D manages to produce superior multi-view consistency images with highly-detailed textures. Source code and data are available at https://github.com/yanghb22-fdu/Hi3D-Official. Haibo Yang 0002, Yang Chen 0048, Yingwei Pan, Ting Yao 0003, Zhineng Chen, Chong-Wah Ngo, Tao Mei 0001 |
ACM Multimedia | 5 |
| 2024 | AdvQDet: Detecting Query-Based Adversarial Attacks with Adversarial Contrastive Prompt Tuning
Xin Wang 0119, Kai Chen 0027, Xingjun Ma, Zhineng Chen, Jingjing Chen 0001, Yu-Gang Jiang 0001 |
ACM Multimedia | 4 |
| 2024 | FreeEnhance: Tuning-Free Image Enhancement via Content-Consistent Noising-and-Denoising ProcessabstractThe emergence of text-to-image generation models has led to the recognition that image enhancement, performed as post-processing, would significantly improve the visual quality of the generated images. Exploring diffusion models to enhance the generated images nevertheless is not trivial and necessitates to delicately enrich plentiful details while preserving the visual appearance of key content in the original image. In this paper, we propose a novel framework, namely FreeEnhance, for content-consistent image enhancement using the off-the-shelf image diffusion models. Technically, FreeEnhance is a two-stage process that firstly adds random noise to the input image and then capitalizes on a pre-trained image diffusion model (i.e., Latent Diffusion Models) to denoise and enhance the image details. In the noising stage, FreeEnhance is devised to add lighter noise to the region with higher frequency to preserve the high-frequent patterns (e.g., edge, corner) in the original image. In the denoising stage, we present three target properties as constraints to regularize the predicted noise, enhancing images with high acutance and high visual quality. Extensive experiments conducted on the HPDv2 dataset demonstrate that our FreeEnhance outperforms the state-of-the-art image enhancement models in terms of quantitative metrics and human preference. More remarkably, FreeEnhance also shows higher human preference compared to the commercial image enhancement solution of Magnific AI. Zhaofan Qiu, Ting Yao 0003, Zhineng Chen, Yu-Gang Jiang 0001, Tao Mei 0001 |
ACM Multimedia | 5 |
| 2024 | Decoder Pre-Training with only Text for Scene Text RecognitionabstractScene text recognition (STR) pre-training methods have achieved remarkable progress, primarily relying on synthetic datasets. However, the domain gap between synthetic and real images poses a challenge in acquiring feature representations that align well with images on real scenes, thereby limiting the performance of these methods. We note that vision-language models like CLIP, pre-trained on extensive real image-text pairs, effectively align images and text in a unified embedding space, suggesting the potential to derive the representations of real images from text alone. Building upon this premise, we introduce a novel method named Decoder Pre-training with only text for STR (DPTR). DPTR treats text embeddings produced by the CLIP text encoder as pseudo visual embeddings and uses them to pre-train the decoder. An Offline Randomized Perturbation (ORP) strategy is introduced. It enriches the diversity of text embeddings by incorporating natural image embeddings extracted from the CLIP image encoder, effectively directing the decoder to acquire the potential representations of real images. In addition, we introduce a Feature Merge Unit (FMU) that guides the extracted visual embeddings focusing on the character foreground within the text image, thereby enabling the pre-trained decoder to work more efficiently and accurately. Extensive experiments across various STR decoders and language recognition tasks underscore the broad applicability and remarkable performance of DPTR, providing a novel insight for STR pre-training. Code is available at https://github.com/Topdu/OpenOCR. Yongkun Du, Zhineng Chen, Yu-Gang Jiang 0001 |
ACM Multimedia | 3 |
| 2024 | CDistNet: Perceiving Multi-domain Character Distance for Robust Text Recognition
Tianlun Zheng, Zhineng Chen, Shancheng Fang, Hongtao Xie 0001, Yu-Gang Jiang 0001 |
Int. J. Comput. Vis. | 2 |
| 2023 | Resolving Task Confusion in Dynamic Expansion Architectures for Class Incremental LearningabstractThe dynamic expansion architecture is becoming popular in class incremental learning, mainly due to its advantages in alleviating catastrophic forgetting. However, task confu- sion is not well assessed within this framework, e.g., the discrepancy between classes of different tasks is not well learned (i.e., inter-task confusion, ITC), and certain prior- ity is still given to the latest class batch (i.e., old-new con- fusion, ONC). We empirically validate the side effects of the two types of confusion. Meanwhile, a novel solution called Task Correlated Incremental Learning (TCIL) is pro- posed to encourage discriminative and fair feature utilization across tasks. TCIL performs a multi-level knowledge distil- lation to propagate knowledge learned from old tasks to the new one. It establishes information flow paths at both fea- ture and logit levels, enabling the learning to be aware of old classes. Besides, attention mechanism and classifier re- scoring are applied to generate more fair classification scores. We conduct extensive experiments on CIFAR100 and Ima- geNet100 datasets. The results demonstrate that TCIL con- sistently achieves state-of-the-art accuracy. It mitigates both ITC and ONC, while showing advantages in battle with catas- trophic forgetting even no rehearsal memory is reserved. Source code: https://github.com/YellowPancake/TCIL. Bingchen Huang, Zhineng Chen, Peng Zhou 0009, Zuxuan Wu |
AAAI | 2 |
| 2023 | Self-distillation Augmented Masked Autoencoders for Histopathological Image UnderstandingabstractSelf-supervised learning (SSL) has drawn increasing attention in histopathological image analysis in recent years. Compared to contrastive learning which is troubled with the false negative problem, i.e., semantically similar images are selected as negative samples, masked autoencoders (MAE) build SSL from a generative paradigm which is probably a more appropriate pretraining. In this paper, we introduce MAE to histopathological image understanding, and moreover, verify the effect of visible patches in this task. Specifically, a novel SD-MAE model is proposed to enable a self-distillation augmented MAE. Besides the reconstruction loss on masked image patches, SD-MAE further imposes the self-distillation loss on visible patches to enhance the representational capacity of encoder located in the shallow layers. It generates a more effective feature pre-training and benefits downstream applications. We apply SD-MAE to histopathological image classification, cell segmentation and cell detection. Experiments demonstrate that SD-MAE shows highly competitive performance compared with other SSL methods in these tasks. Code is available at https://github.com/irsLu/SD-MAE/ Zhineng Chen, Shengtian Zhou, Kai Hu 0002, Xieping Gao 0001 |
BIBM | 2 |
| 2023 | Prototypical Residual Networks for Anomaly Detection and LocalizationabstractAnomaly detection and localization are widely used in industrial manufacturing for its efficiency and effectiveness. Anomalies are rare and hard to collect and supervised models easily over-fit to these seen anomalies with a handful of abnormal samples, producing unsatisfactory performance. On the other hand, anomalies are typically subtle, hard to discern, and of various appearance, making it difficult to detect anomalies and let alone locate anomalous regions. To address these issues, we propose a framework called Prototypical Residual Network (PRN), which learns feature residuals of varying scales and sizes between anomalous and normal patterns to accurately reconstruct the segmentation maps of anomalous regions. PRN mainly consists of two parts: multi-scale prototypes that explicitly represent the residual features of anomalies to normal patterns; a multisize self-attention mechanism that enables variable-sized anomalous feature learning. Besides, we present a variety of anomaly generation strategies that consider both seen and unseen appearance variance to enlarge and diversify anomalies. Extensive experiments on the challenging and widely used MVTec AD benchmark show that PRN outperforms current state-of-the-art unsupervised and supervised methods. We further report SOTA results on three additional datasets to demonstrate the effectiveness and generalizability of PRN. Hui Zhang 0090, Zuxuan Wu, Zheng Wang 0059, Zhineng Chen, Yu-Gang Jiang 0001 |
CVPR | 4 |
| 2023 | Bi-directional Feature Fusion Generative Adversarial Network for Ultra-high Resolution Pathological Image Virtual Re-stainingabstractThe cost of pathological examination makes virtual restaining of pathological images meaningful. However, due to the ultra-high resolution of pathological images, traditional virtual restaining methods have to divide a WSI image into patches for model training and inference. Such a limitation leads to the lack of global information, resulting in observable differences in color, brightness and contrast when the re-stained patches are merged to generate an image of larger size. We summarize this issue as the square effect. Some existing methods try to solve this issue through overlapping between patches or simple postprocessing. But the former one is not that effective, while the latter one requires carefully tuning. In order to eliminate the square effect, we design a bidirectional feature fusion generative adversarial network (BFF-GAN) with a global branch and a local branch. It learns the inter-patch connections through the fusion of global and local features plus patch-wise attention. We perform experiments on both the private dataset RCC and the public dataset ANHIR. The results show that our model achieves competitive performance and is able to generate extremely real images that are deceptive even for experienced pathologists, which means it is of great clinical significance. Zhineng Chen, Gongwei Wang, Xiongjun Ye, Yu-Gang Jiang 0001 |
CVPR | 2 |
| 2023 | Multi-Object Localization and Irrelevant-Semantic Separation for Nuclei Segmentation in Histopathology ImagesabstractAutomated segmentation of nuclei in histopathology images is critical for cancer diagnosis and prognosis. Due to the high variability of nuclei morphology, numerous nuclei overlapping, and the wide existence of nuclei clusters, this task still remains challenging. In this paper, we propose an effective nuclei segmentation method for histopathology images based on a novel neural network for multi-object localization and irrelevant-semantic separation (MI-Net), which includes a multi-object localization module (MOLM), a deep boundary awareness module (DBAM), and an irrelevant semantic separation module (ISSM). Specifically, the MOLM is used to calibrate features to capture more comprehensive information about the location of nuclei. It alleviates the problem of cell adhesion in histopathology images. The DBAM is intended to extract boundary information and solve the blurred boundary problem effectively. The ISSM is used to separate foreground features from background features and effectively address the problem of complex backgrounds in histopathology images. It also solves the low contrast problem between the target objects and the background. We evaluate MI-Net on the famous MoNuSeg dataset and compare it with ten state-of-the-art methods. MI-Net surpasses the best-performed method by margins from 1.12% to 2.29% over standard evaluation metrics, showing its effectiveness on nuclei segmentation. Ya Tang, Xiongjun Ye, Xuanya Li, Zhineng Chen |
ICASSP | 4 |
| 2023 | MRN: Multiplexed Routing Network for Incremental Multilingual Text RecognitionabstractMultilingual text recognition (MLTR) systems typically focus on a fixed set of languages, which makes it difficult to handle newly added languages or adapt to ever-changing data distribution. In this paper, we propose the Incremental MLTR (IMLTR) task in the context of incremental learning (IL), where different languages are introduced in batches. IMLTR is particularly challenging due to rehearsal-imbalance, which refers to the uneven distribution of sample characters in the rehearsal set, used to retain a small amount of old data as past memories. To address this issue, we propose a Multiplexed Routing Network (MRN). MRN trains a recognizer for each language that is currently seen. Subsequently, a language domain predictor is learned based on the rehearsal set to weigh the recognizers. Since the recognizers are derived from the original data, MRN effectively reduces the reliance on older data and better fights against catastrophic forgetting, the core issue in IL. We extensively evaluate MRN on MLT17 and MLT19 datasets. It outperforms existing general-purpose IL methods by large margins, with average accuracy improvements ranging from 10.3% to 35.8% under different settings. Code is available at https://github.com/simplify23/MRN. Tianlun Zheng, Zhineng Chen, Bingchen Huang, Wei Zhang 0031, Yu-Gang Jiang 0001 |
ICCV | 2 |
| 2023 | TPS++: Attention-Enhanced Thin-Plate Spline for Scene Text RecognitionabstractText irregularities pose significant challenges to scene text recognizers. Thin-Plate Spline (TPS)-based rectification is widely regarded as an effective means to deal with them. Currently, the calculation of TPS transformation parameters purely depends on the quality of regressed text borders. It ignores the text content and often leads to unsatisfactory rectified results for severely distorted text. In this work, we introduce TPS++, an attention-enhanced TPS transformation that incorporates the attention mechanism to text rectification for the first time. TPS++ formulates the parameter calculation as a joint process of foreground control point regression and content-based attention score estimation, which is computed by a dedicated designed gated-attention block. TPS++ builds a more flexible content-aware rectifier, generating a natural text correction that is easier to read by the subsequent recognizer. Moreover, TPS++ shares the feature backbone with the recognizer in part and implements the rectification at feature-level rather than image-level, incurring only a small overhead in terms of parameters and inference time. Experiments on public benchmarks show that TPS++ consistently improves the recognition and achieves state-of-the-art accuracy. Meanwhile, it generalizes well on different backbones and recognizers. Code is at https://github.com/simplify23/TPS_PP. Tianlun Zheng, Zhineng Chen, Jinfeng Bai, Hongtao Xie 0001, Yu-Gang Jiang 0001 |
IJCAI | 2 |
| 2023 | 3DStyle-Diffusion: Pursuing Fine-grained Text-driven 3D Stylization with 2D Diffusion Modelsabstract3D content creation via text-driven stylization has played a fundamental challenge to multimedia and graphics community. Recent advances of cross-modal foundation models (e.g., CLIP) have made this problem feasible. Those approaches commonly leverage CLIP to align the holistic semantics of stylized mesh with the given text prompt. Nevertheless, it is not trivial to enable more controllable stylization of fine-grained details in 3D meshes solely based on such semantic-level cross-modal supervision. In this work, we propose a new 3DStyle-Diffusion model that triggers fine-grained stylization of 3D meshes with additional controllable appearance and geometric guidance from 2D Diffusion models. Technically, 3DStyle-Diffusion first parameterizes the texture of 3D mesh into reflectance properties and scene lighting using implicit MLP networks. Meanwhile, an accurate depth map of each sampled view is achieved conditioned on 3D mesh. Then, 3DStyle-Diffusion leverages a pre-trained controllable 2D Diffusion model to guide the learning of rendered images, encouraging the synthesized image of each view semantically aligned with text prompt and geometrically consistent with depth map. This way elegantly integrates both image rendering via implicit MLP networks and diffusion process of image synthesis in an end-to-end fashion, enabling a high-quality fine-grained stylization of 3D meshes. We also build a new dataset derived from Objaverse and the evaluation protocol for this task. Through both qualitative and quantitative experiments, we validate the capability of our 3DStyle-Diffusion. Source code and data are available at https://github.com/yanghb22-fdu/3DStyle-Diffusion-Official. Haibo Yang 0002, Yang Chen 0048, Yingwei Pan, Ting Yao 0003, Zhineng Chen, Tao Mei 0001 |
ACM Multimedia | 5 |
| 2022 | Genre-Conditioned Long-Term 3D Dance Generation Driven by MusicabstractDancing to music is an artistic behavior of humans, however, letting machines generate dances from music is still challenging. Most existing works have been made progress in tackling the problem of motion prediction conditioned by music, yet they rarely consider the importance of the musical genre. In this paper, we focus on generating long-term 3D dance from music with a specific genre. Specifically, we construct a pure transformer-based architecture to correlate motion features and music features. To utilize the genre information, we propose to embed the genre categories into the transformer decoder so that it can guide every frame. Moreover, different from previous inference schemes, we introduce the motion queries to output the dance sequence in parallel that significantly improves the efficiency. Extensive experiments on AIST++[1] dataset show that our model outperforms state-of-the-art methods with a much faster inference speed. Yuhang Huang 0006, Junjie Zhang 0002, Qian Bao, Dan Zeng 0001, Zhineng Chen, Wu Liu 0005 |
ICASSP | 6 |
| 2022 | SVTR: Scene Text Recognition with a Single Visual ModelabstractDominant scene text recognition models commonly contain two building blocks, a visual model for feature extraction and a sequence model for text transcription. This hybrid architecture, although accurate, is complex and less efficient. In this study, we propose a Single Visual model for Scene Text recognition within the patch-wise image tokenization framework, which dispenses with the sequential modeling entirely. The method, termed SVTR, firstly decomposes an image text into small patches named character components. Afterward, hierarchical stages are recurrently carried out by component-level mixing, merging and/or combining. Global and local mixing blocks are devised to perceive the inter-character and intra-character patterns, leading to a multi-grained character component perception. Thus, characters are recognized by a simple linear prediction. Experimental results on both English and Chinese scene text recognition tasks demonstrate the effectiveness of SVTR. SVTR-L (Large) achieves highly competitive accuracy in English and outperforms existing methods by a large margin in Chinese, while running faster. In addition, SVTR-T (Tiny) is an effective and much smaller model, which shows appealing speed at inference. The code is publicly available at https://github.com/PaddlePaddle/PaddleOCR. Yongkun Du, Zhineng Chen, Caiyan Jia, Xiaoting Yin, Tianlun Zheng, Chenxia Li, Yuning Du, Yu-Gang Jiang 0001 |
IJCAI | 2 |
| 2022 | AS-Net: Attention Synergy Network for skin lesion segmentation
Kai Hu 0002, Dapeng Xiong, Zhineng Chen |
Expert Syst. Appl. | 5 |
| 2022 | Bridge-Net: Context-involved U-net with patch-based loss weight mapping for retinal blood vessel segmentation
Yuan Zhang 0022, Zhineng Chen, Kai Hu 0002, Xuanya Li, Xieping Gao 0001 |
Expert Syst. Appl. | 3 |
| 2021 | A Hybrid Feature Enhancement Method for Gl And Segmentation In Histopathology ImagesabstractAccurate and automatic gland segmentation can help pathologists diagnose the malignancy of colorectal cancers. However, it remains a challenging task because of the large morphological differences between the glands and the presence of sticky glands. In this paper, a hybrid feature enhancement network (HFE-Net) for glandular segmentation is proposed, which includes a multi-scale local feature extraction block (MSLFEB) and a global feature enhancement block (GFEB). Specifically, the MSLFEB is used to extract multiscale features through different sizes of the receptive field to reduce the loss of the local information and effectively alleviate glandular adhesion. The GFEB is used to transfer the underlying features to the decoder by considering the global semantic information. Furthermore, we design a focal and variance (FV) loss function to alleviate the class imbalance and constraint the pixels within the same instance. Finally, we evaluate the proposed method on the 2015 MICCAI GlaS challenge dataset and the CRAG colorectal adenocarcinoma dataset. The results show that our HFE-Net can achieve competitive results with fewer computing resources when compared with the state-of-the-art gland segmentation methods. Xiangjiang Wu, Xuanya Li, Kai Hu 0002, Zhineng Chen, Xieping Gao 0001 |
ICASSP | 4 |
| 2021 | Bag of Tricks for Building an Accurate and Slim Object Detector for Embedded ApplicationsabstractObject detection is an essential computer vision task that possesses extensive application prospects in on-road applications. Copious novel methods have been proposed in this branch recently. However, the majority of them have high computational cost, making them intractable to be deployed on embedded devices. In this paper, taking YOLOv5s, the smallest model in the YOLOv5 family, as the baseline, we explore a bag of tricks that improve the detection performance for a specified on-road application, under the premise of ensuring that it does not increase the computational cost of YOLOv5s. Specifically, we introduce relevantly external data to deal with the problems of sample imbalance. Meanwhile, knowledge distillation is employed to transfer knowledge from a cumbersome model to a compact model, where a united distillation scheme is developed to enhance the effectiveness. In addition, a pseudo-label based training strategy is utilized to further learn from the biggest YOLOv5 model. We have applied the above tricks to the Embedded Deep Learning Object Detection Model Compression Competition for Traffic in Asian Countries held in conjunction with ICMR 2021. The experiments have shown that all the tricks are useful. Their combination have built an accurate and slim detection model. It is highly competitive and has been ranked 2nd place in the competition. We believe the tricks are also meaningful for building other application-oriented object detectors. Yongkun Du, Zhineng Chen, Caiyan Jia, Xuanya Li, Yu-Gang Jiang 0001 |
ICMR | 2 |
| 2021 | Res2-Unet: An Enhanced Network for Generalized Nuclear Segmentation in Pathological Images
Xuanya Li, Zhineng Chen, Changgen Peng |
MMM (2) | 3 |
| 2021 | ERV-Net: An efficient 3D residual neural network for brain tumor segmentation
Xuanya Li, Kai Hu 0002, Yuan Zhang 0022, Zhineng Chen, Xieping Gao 0001 |
Expert Syst. Appl. | 5 |
| 2021 | Deep supervised learning using self-adaptive auxiliary loss for COVID-19 diagnosis from imbalanced CT images
Kai Hu 0002, Zhineng Chen, Xuanya Li, Yuan Zhang 0022, Xieping Gao 0001 |
Neurocomputing | 5 |
| 2021 | Object-difference drived graph convolutional networks for visual question answering
Zhendong Mao 0001, Zhineng Chen, Bin Wang 0004 |
Multim. Tools Appl. | 3 |
| 2021 | Adaptive multi-branch correlation filters for robust visual tracking
Lei Huang 0010, Zhiqiang Wei 0002, Jie Nie, Zhineng Chen |
Neural Comput. Appl. | 5 |
| 2021 | A hierarchical and multi-view registration of serial histopathological images
Zhineng Chen, Kai Hu 0002, Shaoping Ling, Xieping Gao 0001 |
Pattern Recognit. Lett. | 1 |
| 2020 | Nuclei Segmentation in Histopathology Images Using Rotation Equivariant and Multi-level Feature Aggregation Neural NetworkabstractThe histopathological analysis is the gold standard for assessing the presence and many complex diseases, like tumors. As one of the essential part of tumors, the shape, staining, and tissue distribution of the nuclei plays an important role in tumor diagnosis. However, due to nuclei congestion and possible occlusion, nuclei segmentation remains challenging. In this paper, we propose an automatic and effective nuclei segmentation method in histopathology images based on rotation equivariant and multi-level feature aggregation neural network (REMFANet). First, considering the inherent rotation equivariant of digital pathological images, we introduce group equivariant convolutions to improve the performance of the automatic segmentation of pathological images. Second, to eliminate the semantic gap between shallow features and deep features in encoder-decoder structural models, we propose a multi-level feature aggregation strategy based on U-Net 3+. Specifically, (1) we design a new decoder module to restore pixel-level predictions more accurately; (2) we propose an improved long-skip connection mode to provide richer semantic information in the decoder; (3) we also construct a semantic enhancement block to enhance the robustness of lowlevel semantic information. Finally, we evaluate our REMFA-Net on the MoNuSeg dataset and compare the results with seven state-of-the-art methods. Experimental results demonstrate the superiority of the proposed method over other models for the nuclei segmentation in histopathology images. Xuanya Li, Kai Hu 0002, Zhineng Chen, Xieping Gao 0001 |
BIBM | 4 |
| 2020 | HMOE-Net: Hybrid Multi-scale Object Equalization Network for Intracerebral Hemorrhage Segmentation in CT ImagesabstractIn this paper, we propose a novel Hybrid Multi-scale Object Equalization Network (HMOE-Net) to segment intracerebral hemorrhage (ICH) regions. In particular, we design a shallow feature extraction network (SFENet) and a deep feature extraction network (DFENet) to solve the problem of equalization learning of hybrid multi-scale object features. The multi-level feature extraction (MLFE) blocks are presented in DFENet to explore multi-level semantic features more effectively. Furthermore, we adopt a progressive feature extraction strategy combining SFENet and DFENet to further consider the differences of various ICH regions and achieve the equalization feature learning of multi-scale objects. To verify the effectiveness of HMOE-Net, we collect a clinical ICH dataset with a total of 500 CT cases from three hospitals for the evaluation. The experimental results show that HMOE-Net is superior to six state-of-the-art methods and achieves accurate segmentation for multi-scale ICH regions. Xizhi He, Kai Chen 0027, Kai Hu 0002, Zhineng Chen, Xuanya Li, Xieping Gao 0001 |
BIBM | 4 |
| 2020 | EffiDiag: an Efficient Framework for Breast Cancer Diagnosis in Multi-Gigapixel Whole Slide ImagesabstractBreast cancer diagnosis in multi-gigapixel whole slide images (WSIs) is an important task that highly relevant to cancer grading and prognosis. In recent years, many computer-aided diagnosis methods were proposed and achieved promising performance. However, they mostly suffer from heavy computational burden that becomes a significant barrier to clinical practice. Efficient solutions are urgently demanded but still less studied. In this paper, we propose a novel framework named EffiDiag for a fast and lightweight breast cancer diagnosis. To this end, a loss-modified U-net is developed at first to enable a fast suspected cancer Region Of Interest (ROI) localization. Therefore the subsequent patch-based classification, which commonly executes at the finest magnification hundreds of thousands times per WSI for cancer identification, could be carried out on these ROIs only rather than the whole WSI for speedup. Meanwhile, a super-efficient convolutional neural network (CNN) is devised to optimize the classification speed and resource consumption per classification. Experiments on the Camelyonl6 benchmark demonstrate, by integrating the two contributions into a well-established approach, 47x inference acceleration is obtained with limited accuracy drop, yet with much less resource consumption even compared to popular lightweight networks. Junda Ren, Zhineng Chen, Kai Hu 0002, Fen Xiao, Xuanya Li, Xieping Gao 0001 |
BIBM | 3 |
| 2020 | Signet Ring Cell Detection with Classification Reinforcement Detection Network
Caiyan Jia, Zhineng Chen, Xieping Gao 0001 |
ISBRA | 3 |
| 2020 | Automatic segmentation of intracerebral hemorrhage in CT images using encoder-decoder convolutional neural network
Kai Hu 0002, Kai Chen 0027, Xizhi He, Yuan Zhang 0022, Zhineng Chen, Xuanya Li, Xieping Gao 0001 |
Inf. Process. Manag. | 5 |
| 2020 | Context propagation embedding network for weakly supervised semantic segmentation
Yajun Xu, Zhendong Mao 0001, Zhineng Chen |
Multim. Tools Appl. | 3 |
| 2020 | Graph-based neural networks for explainable image privacy inference
Guang Yang 0031, Juan Cao 0001, Zhineng Chen, Junbo Guo, Jintao Li 0001 |
Pattern Recognit. | 3 |
| 2020 | Bidirectional Attention-Recognition Model for Fine-Grained Object ClassificationabstractFine-grained object classification (FGOC) is a challenging research topic in multimedia computing with machine learning, which faces two pivotal conundrums: focusing attention on the discriminate part regions, and then processing recognition with the part-based features. Existing approaches generally adopt a unidirectional two-step structure, that first locate the discriminate parts and then recognize the part-based features. However, they neglect the truth that part localization and feature recognition can be reinforced in a bidirectional process. In this paper, we propose a novel bidirectional attention-recognition model (BARM) to actualize the bidirectional reinforcement for FGOC. The proposed BARM consists of one attention agent for discriminate part regions proposing and one recognition agent for feature extraction and recognition. Meanwhile, a feedback flow is creatively established to optimize the attention agent directly by recognition agent. Therefore, in BARM the attention agent and the recognition agent can reinforce each other in a bidirectional way and the overall framework can be trained end-to-end without neither object nor parts annotations. Moreover, a novel Multiple Random Erasing data augmentation is proposed, and it exhibits impressive pertinency and superiority for FGOC. Conducted on several extensive FGOC benchmarks, BARM outperforms the present state-of-the-art methods in classification accuracy. Furthermore, BARM exhibits a clear interpretability and keeps consistent with the human perception in visualization experiments. Chuanbin Liu 0001, Hongtao Xie 0001, Zhengjun Zha, Lingyun Yu 0002, Zhineng Chen, Yongdong Zhang 0001 |
IEEE Trans. Multim. | 5 |
| 2019 | NRTR: A No-Recurrence Sequence-to-Sequence Model for Scene Text RecognitionabstractScene text recognition has attracted a great many researches due to its importance to various applications. Existing methods mainly adopt recurrence or convolution based networks. Though have obtained good performance, these methods still suffer from two limitations: slow training speed due to the internal recurrence of RNNs, and high complexity due to stacked convolutional layers for long-term feature extraction. This paper, for the first time, proposes a no-recurrence sequence-to-sequence text recognizer, named NRTR, that dispenses with recurrences and convolutions entirely. NRTR follows the encoder-decoder paradigm, where the encoder uses stacked self-attention to extract image features, and the decoder applies stacked self-attention to recognize texts based on encoder output. NRTR relies solely on self-attention mechanism thus could be trained with more parallelization and less complexity. Considering scene image has large variation in text and background, we further design a modality-transform block to effectively transform 2D input images to 1D sequences, combined with the encoder to extract more discriminative features. NRTR achieves state-of-the-art or highly competitive performance on both regular and irregular benchmarks, while requires only a small fraction of training time compared to the best model from the literature (at least 8 times faster). Fenfen Sheng, Zhineng Chen, Bo Xu 0002 |
ICDAR | 2 |
| 2019 | A Single-Shot Oriented Scene Text Detector with Learnable AnchorsabstractCurrent regression based text detectors mainly use fixed anchors, where scales and positions can not be changed during network training. As scene texts tend to have large variation in orientations, aspect ratios and sizes, fixed anchors are insufficient to cover all varieties. This paper proposes a novel text detector with learnable anchors, named LATD. LATD contains two prediction branches. One aims to refine scales and locations of anchors according to the characteristics of scene texts. The other one receives refined anchors as defaults and regresses their offsets to text regions. These two branches are optimized jointly without sacrifices much speed. Meanwhile, we explore the class-imbalance issue between texts and backgrounds, and replace softmax loss with focal loss. Extensive experiments on both oriented and horizontal benchmarks demonstrate the effectiveness of LATD with new state-of-the-art performance. By visualizing qualitative results, as expected, LATD provides more accurate locations and lower rate of missed detections. Fenfen Sheng, Zhineng Chen, Tao Mei 0001, Bo Xu 0002 |
ICME | 2 |
| 2019 | A Novel End-to-End Multiple Tagging Model for Knowledge ExtractionabstractIt is an emerging research topic in NLP to joint extraction of knowledge including entities and relations from unstructured text and representing them as meaningful triplets. Despite significant progresses made by recent deep neural network based solutions, these methods still confront the overlapping issue that different relational triplets may have overlapped entities in a sentence, and it is troublesome to address this issue by current solutions. In this paper, we propose a novel multiple tagging model to address the overlapping issue and extract knowledge from unstructured text. Specifically, we devise a multiple tagging scheme that transforms the problem of joint entity and relation extraction into a multiple sequence tagging problem. By using GRU as the building block for encoding-decoding, the proposed model is capable of handling the triplet overlapping problem because the decoder layer allows one entity to take part in more than one triplet. The whole network is end-to-end trianable and outputs all triplets in a sentence directly. Experimental results on the NYT and KBP benchmarks demonstrate that the proposed model siginificantly improves the recall of triplet, and consequently, achieving the new state-of-the-art in the task of triplet extraction on both datasets. Yunhua Song, Hongyun Bao, Zhineng Chen |
IJCNN | 3 |
| 2019 | ACE-Net: Biomedical Image Segmentation with Augmented Contracting and Expansive Paths
Yanhao Zhu, Zhineng Chen, Hongtao Xie 0001, Wenming Guo, Yongdong Zhang 0001 |
MICCAI (1) | 2 |
| 2019 | W2VV++: Fully Deep Learning for Ad-hoc Video SearchabstractAd-hoc video search (AVS) is an important yet challenging problem in multimedia retrieval. Different from previous concept-based methods, we propose a fully deep learning method for query representation learning. The proposed method requires no explicit concept modeling, matching and selection. The backbone of our method is the proposed W2VV++ model, a super version of Word2VisualVec (W2VV) previously developed for visual-to-text matching. W2VV++ is obtained by tweaking W2VV with a better sentence encoding strategy and an improved triplet ranking loss. With these simple yet important changes, W2VV++ brings in a substantial improvement. As our participation in the TRECVID 2018 AVS task and retrospective experiments on the TRECVID 2016 and 2017 data show, our best single model, with an overall inferred average precision (infAP) of 0.157, outperforms the state-of-the-art. The performance can be further boosted by model ensemble using late average fusion, reaching a higher infAP of 0.163. With W2VV++, we establish a new baseline for ad-hoc video search. Xirong Li 0001, Chaoxi Xu, Gang Yang 0001, Zhineng Chen, Jianfeng Dong |
ACM Multimedia | 4 |
| 2019 | psDirector: An Automatic Director for Watching View Generation from Panoramic Soccer Video
Caiyan Jia, Zhineng Chen, Xiaoyan Gu 0001, Hongyun Bao |
MMM (2) | 3 |
| 2019 | Name-face association with web facial image supervision
Zhineng Chen, Wei Zhang 0031, Hongtao Xie 0001, Xiaoyan Gu 0001 |
Multim. Syst. | 1 |
| 2019 | Block compressed sampling of image signals by saliency based adaptive partitioning
Siwang Zhou, Zhineng Chen, Qian Zhong |
Multim. Tools Appl. | 2 |
| 2019 | Automated pulmonary nodule detection in CT images using deep convolutional neural networks
Hongtao Xie 0001, Dongbao Yang, Nannan Sun, Zhineng Chen, Yongdong Zhang 0001 |
Pattern Recognit. | 4 |
| 2019 | Pyrboxes: An efficient multi-scale scene text detector with feature pyramids
Fenfen Sheng, Zhineng Chen, Wei Zhang 0031, Bo Xu 0002 |
Pattern Recognit. Lett. | 2 |
| 2019 | Double-Bit Quantization and Index Hashing for Nearest Neighbor SearchabstractAs binary code is storage efficient and fast to compute, it has become a trend to compact real-valued data to binary codes for the nearest neighbors (NN) search in a large-scale database. However, the use of binary code for the NN search leads to low retrieval accuracy. To increase the discriminability of the binary codes of existing hash functions, in this paper, we propose a framework of double-bit quantization and index hashing for an effective NN search. The main contributions of our framework are: first, a novel double-bit quantization (DBQ) is designed to assign more bits to each dimension for higher retrieval accuracy; second, a double-bit index hashing (DBIH) is presented to efficiently index binary codes generated by DBQ; and third, a weighted distance measurement for DBQ binary codes is put forward to re-rank the search results from DBIH. The empirical results on three benchmark databases demonstrate the superiority of our framework over existing approaches in terms of both retrieval accuracy and query efficiency. Specifically, we observe an absolute improvement on precision of 10%-25% in most cases and the query speed increases over 30 times compared to traditional binary embedding methods and linear scan, respectively. Hongtao Xie 0001, Zhendong Mao 0001, Yongdong Zhang 0001, Chenggang Yan 0001, Zhineng Chen |
IEEE Trans. Multim. | 6 |
| 2019 | Structure-Aware Deep Learning for Product Image ClassificationabstractAutomatic product image classification is a task of crucial importance with respect to the management of online retailers. Motivated by recent advancements of deep Convolutional Neural Networks (CNN) on image classification, in this work we revisit the problem in the context of product images with the existence of a predefined categorical hierarchy and attributes, aiming to leverage the hierarchy and attributes to improve classification accuracy. With these structure-aware clues, we argue that more advanced deep models could be developed beyond the flat one-versus-all classification performed by conventional CNNs. To this end, novel efforts of this work include a salient-sensitive CNN that gazes into the product foreground by inserting a dedicated spatial attention module; a multiclass regression-based refinement that is expected to predict more accurately by merging prediction scores from multiple preceding CNNs, each corresponding to a distinct classifier in the hierarchy; and a multitask deep learning architecture that effectively explores correlations among categories and attributes for categorical label prediction. Experimental results on nearly 1 million real-world product images basically validate the effectiveness of the proposed efforts individually and jointly, from which performance gains are observed. Zhineng Chen, Shanshan Ai, Caiyan Jia |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2018 | FFDet: a Fully Convolutional Network for Coral Reef Fish Detection by Layer FusionabstractUnderwater coral reef fish detection is topic receiving increasingly attention due to its importance in various applications like fish biodiversity monitoring, marine resource managements, etc. However, compared with studies on generic object detection, existing methods on this task are not mature so far where advanced deep models and technologies are seldom considered. This paper presents FFDet, a fully convolutional network for coral reef fish detection by layer fusion. FFDet consists of a single shot multibox detector (SSD)-based backbone, but different with SSD, it devises a novel feature fusion module to aggregate adjacent prediction layers for enhanced feature representation. Thus, instead of using the prediction layers one-by-one, the enhanced features each combining information from multiple layers, are leveraged to detect fishes at different scales. We argue that the proposed module is capable of encoding both strong semantics and detail context information. Experimental results on SeaCLEF dataset show that FFDet not only outperforms SSD in performance by sacrificing only a little efficiency, but also better than another two popular end-to-end deep models in both detection performance and speed, especially on detecting large-sized fishes. Cuncun Shi, Caiyan Jia, Zhineng Chen |
VCIP | 3 |
| 2018 | Distant supervision for relation extraction with hierarchical selective attention
Peng Zhou 0009, Jiaming Xu 0001, Zhenyu Qi 0003, Hongyun Bao, Zhineng Chen, Bo Xu 0002 |
Neural Networks | 5 |
| 2017 | Binarized Mode Seeking for Scalable Visual Pattern DiscoveryabstractThis paper studies visual pattern discovery in large-scale image collections via binarized mode seeking, where images can only be represented as binary codes for efficient storage and computation. We address this problem from the perspective of binary space mode seeking. First, a binary mean shift (bMS) is proposed to discover frequent patterns via mode seeking directly in binary space. The binomial-based kernel and binary constraint are introduced for binarized analysis. Second, we further extend bMS to a more general form, namely contrastive binary mean shift (cbMS), which maximizes the contrastive density in binary space, for finding informative patterns that are both frequent and discriminative for the dataset. With the binarized algorithm and optimization, our methods demonstrate significant computation (50×) and storage (32×) improvement compared to standard techniques operating in Euclidean space, while the performance does not largely degenerate. Furthermore, cbMS discovers more informative patterns by suppressing low discriminative modes. We evaluate our methods on both annotated ILSVRC (1M images) and un-annotated blind Flickr (10M images) datasets with million scale images, which demonstrates both the scalability and effectiveness of our algorithms for discovering frequent and informative patterns in large scale collection. Wei Zhang 0031, Xiaochun Cao, Rui Wang 0032, Yuanfang Guo, Zhineng Chen |
CVPR | 5 |
| 2017 | End-to-End Chinese Image Text Recognition with Attention Model
Fenfen Sheng, Chuanlei Zhai, Zhineng Chen, Bo Xu 0002 |
ICONIP (3) | 3 |
| 2017 | Large-Scale Product Classification via Spatial Attention Based CNN Learning and Multi-class Regression
Shanshan Ai, Caiyan Jia, Zhineng Chen |
MMM (1) | 3 |
| 2017 | Micro-Expression Recognition by Aggregating Local Spatio-Temporal Patterns
Bailan Feng, Zhineng Chen, Xiangsheng Huang |
MMM (1) | 3 |
| 2017 | Detecting Uyghur text in complex background images with convolutional neural network
Shancheng Fang, Hongtao Xie 0001, Zhineng Chen, Shiai Zhu, Xiaoyan Gu 0001, Xingyu Gao 0001 |
Multim. Tools Appl. | 3 |
| 2017 | Robust and parallel Uyghur text localization in complex background images
Yun Song, Hongtao Xie 0001, Zhineng Chen, Xingyu Gao 0001 |
Mach. Vis. Appl. | 4 |
| 2017 | Who Are Your "Real" Friends: Analyzing and Distinguishing Between Offline and Online Friendships From Social Multimedia DataabstractThe Internet has extended the physical boundary of people's social circles to manage an inordinate number of online friends. It is recognized that only a fraction of these online friends are also known with each other in offline circumstances, i.e., the offline friends. An important type of offline friend, onsite offline friend, is defined and addressed in this paper. We explores the possibility of utilizing users' online photo sharing-related behaviors and network topologies to analyze and distinguish between online and onsite offline friendships. Different from traditional social science studies which rely on survey-based data, we employ users' tagged people on the shared Instagram photos as the ground-truth for onsite offline friends. This enables a large-scale and objective analysis and experimental evaluation, which compares between different factors and identifies the features that are key to onsite offline friend identification. Dongyuan Lu, Jitao Sang 0001, Zhineng Chen, Min Xu 0001, Tao Mei 0001 |
IEEE Trans. Multim. | 3 |
| 2015 | Improving Automatic Name-Face Association using Celebrity Images on the WebabstractThis paper investigates the task of automatically associating faces appearing in images (or videos) with their names. Our novelty lies in the use of celebrity Web images to facilitate the task. Specifically, we first propose a method named Image Matching (IM), which uses the faces in images returned from name queries over an image search engine as the gallery set of the names, and a probe face is classified as one of the names, or none of them, according to their matching scores and compatibility characterized by a proposed Assigning-Thresholding (AT) pipeline. Noting IM could provide guidance for association for the well-established Graph-based Association (GA), we further propose two methods that jointly utilize the two kinds of complementary cues. They are: the early fusion of IM and GA (EF-IMGA) that takes the IM score as an additional information source to help the association in GA, and the late fusion of IM and GA (LF-IMGA) that combines the scores from both IM and GA obtained individually to make the association. Evaluations on datasets of captioned news images and Web videos both show the proposed methods, especially the two fused ones, provide significant improvements over GA. Zhineng Chen, Bailan Feng, Chong-Wah Ngo, Caiyan Jia, Xiangsheng Huang |
ICMR | 1 |
| 2014 | Chinese Image Character Recognition Using DNN and Machine Simulated Training Samples
Jinfeng Bai, Zhineng Chen, Bailan Feng, Bo Xu 0002 |
ICANN | 2 |
| 2014 | Chinese Image Text Recognition on grayscale pixelsabstractThis paper presents a novel scheme for Chinese text recognition in images and videos. It's different from traditional paradigms that binarize text images, fed the binarized text to an OCR engine and get the recognized results. The proposed scheme, named grayscale based Chinese Image Text Recognition (gCITR), implements the recognition directly on grayscale pixels via the following steps: image text over-segmentation, building recognition graph, Chinese character recognition and beam search determination. The advantages of gCITR lie in: (1) it does not heavily rely on the performance of binarization, which is not robust in practical and thus severely affects the performance of OCR, (2) grayscale image retains more information of the text thus facilitates the recognition. Experimental results on text from 13 TV news videos demonstrate the effectiveness of the proposed gCITR, from which significant performance gains are observed. Jinfeng Bai, Zhineng Chen, Bailan Feng, Bo Xu 0002 |
ICASSP | 2 |
| 2014 | Image character recognition using deep convolutional neural network learned from different languagesabstractThis paper proposes a shared-hidden-layer deep convolutional neural network (SHL-CNN) for image character recognition. In SHL-CNN, the hidden layers are made common across characters from different languages, performing a universal feature extraction process that aims at learning common character traits existed in different languages such as strokes, while the final softmax layer is made language dependent, trained based on characters from the destination language only. This paper is the first attempt to introduce the SHL-CNN framework to image character recognition. Under the SHL-CNN framework, we discuss several issues including architecture of the network, training of the network, from which a suitable SHL-CNN model for image character recognition is empirically learned. The effectiveness of the learned SHL-CNN is verified on both English and Chinese image character recognition tasks, showing that the SHL-CNN can reduce recognition errors by 16–30% relatively compared with models trained by characters of only one language using conventional CNN, and by 35.7% relatively compared with state-of-the-art methods. In addition, the shared hidden layers learned are also useful for unseen image character recognition tasks. Jinfeng Bai, Zhineng Chen, Bailan Feng, Bo Xu 0002 |
ICIP | 2 |
| 2014 | CeleLabel: an interactive system for annotating celebrities in web videosabstractManual annotation of celebrities in Web videos is an essential task in many people-related Web services. The task, however, poses a significant challenge even to skillful annotators, mainly due to the large quantity of unfamiliar and greatly varied celebrities, and the lack of a customized system for it. This work develops CeleLabel, an interactive system for manually annotating celebrities in the Web video domain. The peculiarity of CeleLabel is to exploit and display multiple types of information that could assist the annotation, including video content, context surrounding and within a video, celebrity images on the Web, and human factors. Using the system, annotators can interactively switch between two views, i.e., merging similar faces and labeling faces with names, to approach the annotation. User studies show that the CeleLabel leads to a much better labeling efficiency and satisfaction. Zhineng Chen, Jinfeng Bai, Chong-Wah Ngo, Bailan Feng, Bo Xu 0002 |
ACM Multimedia | 1 |
| 2014 | Video to Article Hyperlinking by Multiple Tag Property Exploration
Zhineng Chen, Bailan Feng, Hongtao Xie 0001, Rong Zheng 0005, Bo Xu 0002 |
MMM (1) | 1 |
| 2014 | Name-Face Association in Web Videos: A Large-Scale Dataset, Baselines, and Open Issues
Zhineng Chen, Chong-Wah Ngo, Wei Zhang 0031, Juan Cao 0001, Yu-Gang Jiang 0001 |
J. Comput. Sci. Technol. | 1 |
| 2014 | Multiple style exploration for story unit segmentation of broadcast news video
Bailan Feng, Zhineng Chen, Rong Zheng 0005, Bo Xu 0002 |
Multim. Syst. | 2 |
| 2013 | A general Framework of video segmentation to logical unit based on conditional random fieldsabstractSegmenting video into logical units like scenes in movies and topic units in News videos is an essential prerequisite for a wide range of video related applications. In this paper, a novel approach for logical unit segmentation based on conditional random fields (CRFs) is presented. In comparison with previous approaches that handle scenes and topic units separately, the proposed approach deals with them in a general framework. Specifically, four types of shots are defined and represented by four middle-level features, i.e., shot difference, scene transition, shot theme and audio type. Then, the problem of logical unit segmentation is novelly formulated as a problem of identifying the type of shot based on the extracted features, by leveraging the CRFs model. The proposed framework effectively integrate visual, audio and contextual features, and it is able to produce ideal result for both scene and topic unit segmentation. The effectiveness of the proposed approach is verified on seven mainstream types of videos, from which average F-measures of 88% and 86% on scenes and topic units are reported respectively, illustrating that the proposed method can accurately segment logical units in different genres of videos. Bailan Feng, Zhineng Chen, Bo Xu 0002 |
ICMR | 3 |
| 2012 | Community as a connector: associating faces with celebrity names in web videosabstractAssociating celebrity faces appearing in videos with their names is of increasingly importance with the popularity of both celebrity videos and related queries. However, the problem is not yet seriously studied in Web video domain. This paper proposes a Community connected Celebrity Name-Face Association approach (C-CNFA), where the community is regarded as an intermediate connector to facilitate the association. Specifically, with the names and faces extracted from Web videos, C-CNFA decomposes the association task into a three-step framework: community discovering, community matching and celebrity face tagging. To achieve the goal of efficient name-face association under this umbrella, algorithms such as the constrained density-based clustering and exemplar based voting are developed by leveraging different pieces of visual and contextual cues. The evaluation on 0.4 million faces and 144 celebrities shows the effectiveness of the proposed C-CNFA approach. Moreover, using the obtained associations, encouraging results are reported in celebrity video ranking. Zhineng Chen, Chong-Wah Ngo, Juan Cao 0001, Wei Zhang 0031 |
ACM Multimedia | 1 |
| 2011 | Web video retagging
Zhineng Chen, Juan Cao 0001, Tian Xia 0002, Yicheng Song, Yongdong Zhang 0001, Jintao Li 0001 |
Multim. Tools Appl. | 1 |
| 2010 | Web video categorization based on Wikipedia categories and content-duplicated open resourcesabstractThis paper presents a novel approach for web video categorization by leveraging Wikipedia categories (WikiCs) and open resources describing the same content as the video, i.e., content-duplicated open resources (CDORs). Note that current approaches only collect CDORs within one or a few media forms and ignore CDORs of other forms. We explore all these resources by utilizing WikiCs and commercial search engines. Given a web video, its discriminative Wikipedia concepts are first identified and classified. Then a textual query is constructed and from which CDORs are collected. Based on these CDORs, we propose to categorize web videos in the space spanned by WikiCs rather than that spanned by raw tags. Experimental results demonstrate the effectiveness of both the proposed CDOR collection method and the WikiC voting categorization algorithm. In addition, the categorization model built based on both WikiCs and CDORs achieves better performance compared with the models built based on only one of them as well as state-of-the-art approach. Zhineng Chen, Juan Cao 0001, Yicheng Song, Yongdong Zhang 0001, Jintao Li 0001 |
ACM Multimedia | 1 |
| 2010 | Tag transformerabstractHuman annotations (titles and tags) of web videos facilitate most web video applications. However, the raw tags are noisy, sparse and structureless, which limit the effectiveness of tags. In this paper, we propose a tag transformer schema to solve these problems. We first eliminate those imprecise and meaningless tags with Wikipedia, and then transform the remaining tags to the Wikipedia category set to gather a precise, complete and structural description of the tags. Our experimental results on web video categorization demonstrate the superiority of the transformed space. We also apply tag transformer into the first study of using Wikipedia category system to structurally recommend the related videos. The online user study of the demo system suggests that our method could bring fantastic experience to the web users. Yicheng Song, Juan Cao 0001, Zhineng Chen, Yongdong Zhang 0001, Jintao Li 0001 |
ACM Multimedia | 3 |
| 2010 | Multi-modal query expansion for web video searchabstractQuery expansion is an effective method to improve the usability of multimedia search. Most existing multimedia search engines are able to automatically expand a list of textual query terms based on text search techniques, which can be called textual query expansion (TQE). However, the annotations (title and tag) around web videos are generally noisier for text-only query expansion and search matching. In this paper, we propose a novel multi-modal query expansion (MMQE) framework for web video search to solve the issue. Compared with traditional methods, MMQE provides a more intuitive query suggestion by transforming tex-tual query to visual presentation based on visual clustering. Paral-lel to this, MMQE can enhance the process of search matching with strong pertinence of intent-specific query by joining textual, visual and social cues from both metadata and content of videos. Experimental results on real web videos from YouTube demon-strate the effectiveness of the proposed method. Bailan Feng, Juan Cao 0001, Zhineng Chen, Yongdong Zhang 0001, Shouxun Lin |
SIGIR | 3 |
| 2010 | Context-oriented web video tag recommendationabstractTag recommendation is a common way to enrich the textual annotation of multimedia contents. However, state-of-the-art recommendation methods are built upon the pair-wised tag relevance, which hardly capture the context of the web video, i.e., when who are doing what at where. In this paper we propose the context-oriented tag recommendation (CtextR) approach, which expands tags for web videos under the context-consistent constraint. Given a web video, CtextR first collects the multi-form WWW resources describing the same event with the video, which produce an informative and consistent context; and then, the tag recommendation is conducted based on the obtained context. Experiments on an 80,031 web video collection show CtextR recommends various relevant tags to web videos. Moreover, the enriched tags improve the performance of web video categorization. Zhineng Chen, Juan Cao 0001, Yicheng Song, Junbo Guo, Yongdong Zhang 0001, Jintao Li 0001 |
WWW | 1 |