Yingbin Zheng

dblp:04/6699 · DBLP profile ↗
← Back
48ranked-venue papers
6as first author
18since 2021 · last 2025
0000-0002-5590-9292ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 31 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 23 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author
YearPublicationVenuePosition
2025 An Exemplar-based Framework for Chinese Text Recognition
abstract
This paper introduces a novel exemplar-based framework for reading Chinese texts in natural scene or document images. We present the Deep Exemplar-based Chinese Text Recognizer, which is structured to first identify candidate characters as exemplars from each text-line, and subsequently recognize them by retrieving analogous exemplars from a database. With text-line level annotations, we design the exemplar discovery network to simultaneously recognize texts and capture individual character positions in a weak-supervision manner. The exemplar retrieval module is then crafted to identify the most similar exemplar and propagate the corresponding character label. This enables us to effectively rectify the misrecognized characters and boost the performance of scene text recognition. Experiments on four scenarios of Chinese texts demonstrate the effectiveness of our proposed framework.
Zhao Zhou, Xiangcheng Du, Yingbin Zheng, Xingjiao Wu, Cheng Jin 0001
AAAI3
2025 Achieving Ensemble-Like Performance in a Single Model: A Feature Diversification Framework for Image-Text Matching
abstract
Model ensembling is a widely used technique that enhances performance in image-text matching tasks by combining multiple models, each trained with different initializations. However, the inefficiencies associated with training several models and generating outputs from them constrain their practical applicability. In this paper, we argue that while the parameters of two randomly initialized models can differ significantly, their feature distributions can be similar at certain stages. By employing a proposed technique called cross-modal realignment, we demonstrate that features derived from differently initialized models maintain similarity at the feature extraction stage and can be effectively transformed by fine-tuning a small number of parameters. These findings provide an efficient way to achieve ensemble-like performance within a single model. Specifically, we propose a Feature Diversification Framework (FDF) that emulates the outputs of multiple model initializations to generate diverse features from a common shared feature. Firstly, we introduce feature conversion methods to transform shared features into a set of distinct features. Next, a realignment training strategy is presented to optimize negative pairs for realigning these transformed features, thereby enhancing their diversification to resemble the outputs of different models. Additionally, we propose a reweighting module that assigns weights to these features, enabling a weighted fusion approach for robust feature representation. Extensive experiments on the Flickr30K and MS-COCO datasets demonstrate the effectiveness and generalizability of our framework.
Zhao Zhou, Yingbin Zheng, Xiangcheng Du, Cheng Jin 0001
AAAI4
2025 Expanding the Scope of Negatives: Boosting Image-Text Matching with Negatives Distribution Guided Learning
abstract
Image-text matching is a crucial task that bridges visual and linguistic modalities. Recent research typically formulates it into the problem of maximizing the margin with the truly hardest negatives to enhance the learning efficiency and avoid the poor local optima. We argue that such formulation can lead to a serious limitation, i.e., under this formulation, conventional trainers would confine their horizon within the hardest negative examples, while other negative examples offer a range of semantic differences not present in the hardest negatives. In this paper, we propose an efficient negative distribution guided training framework for image-text matching to unlock the substantial promotion space left by the above limitation. Rather than simply incorporating additional negative examples into the training objective, which could diminish both the leading role of the hardest negatives in training and the effect of a large margin learning in producing a robust matching model, our central idea is to supply the objective with distributional information on the entire set of negative examples. To be precise, we first construct the sample similarity matrix based on several pretrained models to extract the distributional information of the entire negative sample dataset. Then we encode it into a margin regularization module to smooth the similarities differences of all negatives. This enhancement facilitates the capture of fine-grained semantic differences and guides the main learning process by maximizing the margin with hard negative examples. Furthermore, we propose a hardest negative rectification module to address the instability in hardest negative selection based on predicted similarity and to correct erroneous hardest negatives. We evaluate our method in combination with several state-of-the-art image-text matching methods, and our quantitative and qualitative experiments demonstrate its significant generalizability and effectiveness.
Zhao Zhou, Xiangcheng Du, Yingbin Zheng, Cheng Jin 0001
AAAI4
2025 Unleashing the Semantic Adaptability of Controlled Diffusion Model for Image Colorization
abstract
Recent data-driven image colorization methods have leveraged pre-trained Text-to-Image (T2I) diffusion models as generative prior, while still suffering from unsatisfactory and inaccurate semantic-level color control. To address these issues, we propose a Semantic Adaptation method (SeAda) that enhances the prior while considering the semantic discrepancy between color and grayscale image pairs. The SeAda employs a semantic adapter to produce refined semantic embeddings and a controlled T2I diffusion model to create reasonably colored images. Specifically, the semantic adapter transfers the embedding from grayscale to color domain, while the diffusion model utilizes the refined embedding and prior knowledge to achieve realistic and diverse results. We also design a three-staged training strategy to improve semantic comprehension and prior integration for further performance improvement. Extensive experiments on public datasets demonstrate that our method outperforms existing state-of-the-art techniques, yielding superior performance in image colorization.
Xiangcheng Du, Zhao Zhou, Yingbin Zheng, Xingjiao Wu, Peizhu Gong, Cheng Jin 0001
IJCAI4
2024 Fine-Grained Scene Image Classification with Modality-Agnostic Adapter
abstract
When dealing with the task of fine-grained scene image classification, most previous works lay much emphasis on global visual features when doing multi-modal feature fusion. In other words, models are deliberately designed based on prior intuitions about the importance of different modalities. In this paper, we present a new multi-modal feature fusion approach named MAA (Modality-Agnostic Adapter), trying to make the model learn the importance of different modalities in different cases adaptively, without giving a prior setting in the model architecture. More specifically, we eliminate the modal differences in distribution and then use a modality-agnostic Transformer encoder for a semantic-level feature fusion. Our experiments demonstrate that MAA achieves state-of-the-art results on benchmarks by applying the same modalities with previous methods. Besides, it is worth mentioning that new modalities can be easily added when using MAA and further boost the performance.
Zhao Zhou, Xiangcheng Du, Xingjiao Wu, Yingbin Zheng, Cheng Jin 0001
ICME5
2024 Minutes to Seconds: Speeded-up DDPM-based Image Inpainting with Coarse-to-Fine Sampling
abstract
For image inpainting, the existing Denoising Diffusion Probabilistic Model (DDPM) based method i.e. RePaint can produce high-quality images for any inpainting form. It utilizes a pre-trained DDPM as a prior and generates inpainting results by conditioning on the reverse diffusion process, namely denoising process. However, this process is significantly time-consuming. In this paper, we propose an efficient DDPM-based image inpainting method which includes three speed-up strategies. First, we utilize a pre-trained Light-Weight Diffusion Model (LWDM) to reduce the number of parameters. Second, we introduce a skip-step sampling scheme of Denoising Diffusion Implicit Models (DDIM) for the denoising process. Finally, we propose Coarse-to-Fine Sampling (CFS), which speeds up inference by reducing image resolution in the coarse stage and decreasing denoising timesteps in the refinement stage. We conduct extensive experiments on both faces and general-purpose image inpainting tasks, and our method achieves competitive performance with approximately 60 times speedup.
Xiangcheng Du, LeoWu TomyEnrique, Yingbin Zheng, Cheng Jin 0001
ICME5
2024 MultiColor: Image Colorization by Learning from Multiple Color Spaces
Xiangcheng Du, Zhao Zhou, Xingjiao Wu, Yingbin Zheng, Cheng Jin 0001
ACM Multimedia6
2024 Cross-domain document layout analysis using document style guide
Xingjiao Wu, Luwei Xiao, Xiangcheng Du, Yingbin Zheng, Xin Li 0110, Tianlong Ma, Cheng Jin 0001, Liang He 0001
Expert Syst. Appl.4
2023 DDT: Dual-branch Deformable Transformer for Image Denoising
abstract
Transformer is beneficial for image denoising tasks since it can model long-range dependencies to overcome the limitations presented by inductive convolutional biases. However, directly applying the transformer structure to remove noise is challenging because its complexity grows quadratically with the spatial resolution. In this paper, we propose an efficient Dual-branch Deformable Transformer (DDT) denoising network which captures both local and global interactions in parallel. We divide features with a fixed patch size and a fixed number of patches in local and global branches, respectively. In addition, we apply deformable attention operation in both branches, which helps the network focus on more important regions and further reduces computational complexity. We conduct extensive experiments on real-world and synthetic denoising tasks, and the proposed DDT achieves state-of-the-art performance with significantly fewer computational costs.
Kangliang Liu, Xiangcheng Du, Yingbin Zheng, Xingjiao Wu, Cheng Jin 0001
ICME4
2023 Modeling Stroke Mask for End-to-End Text Erasing
abstract
Scene text erasing aims to wipe text regions in scene images with reasonable background. Most previous approaches employ scene text detectors to assist localization of the text regions. However, detected text boxes contain both text strokes and background clutters, and directly in-painting on the whole boxes may remain text artifacts and make regions unnatural. In this paper, we present an end-to-end network that focuses on modeling text stroke masks that provide more accurate locations to compute erased images. The network consists of two stages, i.e., a basic network with stroke generation and a refinement network with stroke awareness. The basic network predicts the text stroke masks and initial erasing results simultaneously. The refinement network receives the masks as supervision to generate natural erased results. Experiments on both synthetic and real-world scene images demonstrate the effectiveness of our framework in producing high quality erasing results.
Xiangcheng Du, Zhao Zhou, Yingbin Zheng, Tianlong Ma, Xingjiao Wu, Cheng Jin 0001
WACV3
2023 ETR: An Efficient Transformer for Re-ranking in Visual Place Recognition
abstract
Visual place recognition is to estimate the geographical location of a given image, which is usually addressed by recognizing its similar reference images from a database. The reference images are usually retrieved via similarity search using global descriptor, and the local descriptors are used to re-rank the initial retrieved candidates. The local descriptors re-ranking can significantly improve the accuracy of global retrieval but comes at a high computational cost. To achieve a good trade-off between accuracy and efficiency, we propose an Efficient Transformer for Re-ranking (ETR), utilizing both global and local descriptors to re-rank the top candidates in a single shot. In contrast to traditional re-ranking methods, we leverage self-attention to capture relationships between local descriptors in a single image and cross-attention to explore the similarity of the image pairs. We show that the proposed model can be regarded as a general re-ranking algorithm for significantly boosting the performance of other global-only retrieval methods. Extensive experimental results show that our method outperforms state-of-the-arts and is orders of magnitude faster in terms of computational efficiency.
Heming Jing, Yingbin Zheng, Yuan Wu 0004, Cheng Jin 0001
WACV4
2023 Progressive scene text erasing with self-supervision
Xiangcheng Du, Zhao Zhou, Yingbin Zheng, Xingjiao Wu, Tianlong Ma, Cheng Jin 0001
Comput. Vis. Image Underst.3
2023 Reading Scene Text with Aggregated Temporal Convolutional Encoder
abstract
Reading scene text in the natural image is of fundamental importance in many real-world problems. Text recognition has a profound effect on information processing by enabling automated extraction and interpretation. Recent scene text recognition methods employ the encoder-decoder framework, which constructs the encoder by obtaining the visual representations based on the last layer of the backbone network and then feeding them into a sequence model. In this article, we propose a novel encoder structure that performs the feature extractor and the sequence modeling within a unified framework. The introduced Aggregated Temporal Convolutional Encoder (ATCE) first incorporates the temporal convolutional layers to consider the long-term temporal relationship in the encoder stage. The aggregation of these temporal convolution modules is designed to utilize visual features from different levels, by augmenting the standard architecture with deeper aggregation to better fuse information across modules. We also study the impact of different attention modules in convolutional blocks for learning accurate text representations. We conduct comparisons on several scene text recognition benchmarks for both Chinese and English; the experiments demonstrate the complementary ability with different decoder variants and the effectiveness of our proposed approach.
Tianlong Ma, Xiangcheng Du, Xingjiao Wu, Zhao Zhou, Yingbin Zheng, Cheng Jin 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.5
2022 Edge-aware deep image deblurring
Zhichao Fu, Yingbin Zheng, Tianlong Ma, Hao Ye 0005, Jing Yang 0023, Liang He 0001
Neurocomputing2
2021 A Coarse-to-fine Approach for Fast Super-Resolution with Flexible Magnification
abstract
We perform fast single image super-resolution with flexible magnification for natural images. A novel coarse-to-fine super-resolution framework is developed for the magnification that is factorized into a maximum integer component and the quotient. Specifically, our framework is embedded with a light-weight upscale network for super-resolution with the integer scale factor, followed by the fine-grained network to guide interpolation on feature maps as well as to generate the super-resolved image. Compared with the previous flexible magnification super-resolution approaches, the proposed framework achieves a tradeoff between computational complexity and performance. We conduct experiments using the coarse-to-fine framework on the standard benchmarks and demonstrate its superiority in terms of effectiveness and efficiency over previous approaches.
Zhichao Fu, Tianlong Ma, Yingbin Zheng, Hao Ye 0005, Liang He 0001
MMAsia4
2021 A Coarse-to-Fine Framework for Resource Efficient Video Recognition
Zuxuan Wu, Hengduo Li, Yingbin Zheng, Caiming Xiong, Yu-Gang Jiang 0001, Larry Davis 0001
Int. J. Comput. Vis.3
2021 Parallel pathway dense neural network with weighted fusion structure for brain tumor segmentation
Fangyan Ye, Yingbin Zheng, Hao Ye 0005, Xiaohao Han, Jun Wang 0006, Jian Pu
Neurocomputing2
2021 Document image layout analysis via explicit edge embedding network
Xingjiao Wu, Yingbin Zheng, Tianlong Ma, Hao Ye 0005, Liang He 0001
Inf. Sci.2
2020 Detecting Curve Text with Local Segmentation Network and Curve Connection
abstract
Curve text or arbitrary shape text is very common in real-world scenarios. In this paper, we propose a novel framework with the local segmentation network (LSN) followed by the curve connection to detect text in horizontal, oriented and curved forms. The LSN is composed of two elements, i.e., proposal generation to get the horizontal rectangle proposals with high overlap with text and text segmentation to find the arbitrary shape text region within proposals. The curve connection is then designed to connect the local mask to the detection results. We conduct experiments using the proposed framework on two real-world curve text detection datasets and demonstrate the effectiveness over previous approaches.
Zhao Zhou, Hao Ye 0005, Luhui Chen, Yingbin Zheng
IEEE BigData4
2020 Scene Text Recognition with Temporal Convolutional Encoder
abstract
Texts from scene images typically consist of several characters and exhibit a characteristic sequence structure. Existing methods capture the structure with the sequence-to-sequence models by an encoder to have the visual representations and then a decoder to translate the features into the label sequence. In this paper, we study text recognition framework by considering the long-term temporal dependencies in the encoder stage. We demonstrate that the proposed Temporal Convolutional Encoder with increased sequential extents improves the accuracy of text recognition. We also study the impact of different attention modules in convolutional blocks for learning accurate text representations. We conduct comparisons on seven datasets and the experiments demonstrate the effectiveness of our proposed approach.
Xiangcheng Du, Tianlong Ma, Yingbin Zheng, Hao Ye 0005, Xingjiao Wu, Liang He 0001
ICASSP3
2020 Feature channel enhancement for crowd counting
abstract
Crowd counting, i.e. count the number of people in a crowded visual space, is emerging as an essential research problem with public security. A key in the design of the crowd counting system is to create a stable and accurate robust model, which requires to process on the feature channels of the counting network. In this study, the authors present a featured channel enhancement (FCE) block for crowd counting. First, they use a feature extraction unit to obtain the information of each channel and encodes the information of each channel. Then use a non‐linear variation unit to deal with the encoded channel information, finally, normalise the data and affixed to each channel separately. With the use of the FCE, the positive characteristic channel can be enhanced and weak or negative channel information can be suppressed. The authors successfully incorporate the FCE with two compact networks on the standard benchmarks and prove that the proposed FCE achieves promising results.
Xingjiao Wu, Shuchen Kong, Yingbin Zheng, Hao Ye 0005, Jing Yang 0023, Liang He 0001
IET Image Process.3
2020 Fast video crowd counting with a Temporal Aware Network
Xingjiao Wu, Baohan Xu, Yingbin Zheng, Hao Ye 0005, Jing Yang 0023, Liang He 0001
Neurocomputing3
2020 Counting crowds with varying densities via adaptive scenario discovery framework
Xingjiao Wu, Yingbin Zheng, Hao Ye 0005, Wenxin Hu, Tianlong Ma, Jing Yang 0023, Liang He 0001
Neurocomputing2
2019 Scalable Document Image Information Extraction with Application to Domain-Specific Analysis
abstract
Document images are ubiquitous, but existing methods mainly focus on the text reading but not information understanding. In this paper, we propose a novel document image information extraction framework with application to domain-specific analysis. Key gains of our system result from the modularized implementation of the document analysis modules needed for different document analysis problems. Further, we provide an efficient text recognition approach that makes a trade-off between performance and running speed for document images and a novel information extraction method with both visual and semantic information. Our framework is scalable and customizable, and only a few annotations of the keyword-content mapping is needed towards domain-specific document analysis.
Yingbin Zheng, Shuchen Kong, Wanshan Zhu, Hao Ye 0005
IEEE BigData1
2019 Aggregating Rich Deep Semantic Features for Fine-Grained Place Classification
Tingyu Wei, Wenxin Hu, Xingjiao Wu, Yingbin Zheng, Hao Ye 0005, Jing Yang 0023, Liang He 0001
ICANN (3)4
2019 Adaptive Scenario Discovery for Crowd Counting
abstract
Crowd counting, i.e., estimation number of the pedestrian in crowd images, is emerging as an important research problem with the public security applications. A key component for the crowd counting systems is the construction of counting models which are robust to various scenarios under facts such as camera perspective and physical barriers. In this paper, we present an adaptive scenario discovery framework for crowd counting. The system is structured with two parallel pathways that are trained with different sizes of the receptive field to represent different scales and crowd densities. After ensuring that these components are present in the proper geometric configuration, a third branch is designed to adaptively recalibrate the pathway-wise responses by discovering and modeling the dynamic scenarios implicitly. Our system is able to represent highly variable crowd images and achieves state-of-the-art results in two challenging benchmarks.
Xingjiao Wu, Yingbin Zheng, Hao Ye 0005, Wenxin Hu, Jing Yang 0023, Liang He 0001
ICASSP2
2019 Cascaded Detail-Preserving Networks for Super-Resolution of Document Images
abstract
The accuracy of OCR is usually affected by the quality of the input document image and different kinds of marred document images hamper the OCR results. Among these scenarios, the low-resolution image is a common and challenging case. In this paper, we propose the cascaded networks for document image super-resolution. Our model is composed by the Detail-Preserving Networks with small magnification. The loss function with perceptual terms is designed to simultaneously preserve the original patterns and enhance the edge of the characters. These networks are trained with the same architecture and different parameters and then assembled into a pipeline model with a larger magnification. The low-resolution images can upscale gradually by passing through each Detail-Preserving Network until the final high-resolution images. Through extensive experiments on two scanning document image datasets, we demonstrate that the proposed approach outperforms recent state-of-the-art image super-resolution methods, and combining it with standard OCR system lead to signification improvements on the recognition results.
Zhichao Fu, Yingbin Zheng, Hao Ye 0005, Wenxin Hu, Jing Yang 0023, Liang He 0001
ICDAR3
2019 Video Emotion Recognition with Concept Selection
abstract
Understanding video content is a challenging problem in many applications, especially for emotion analysis. Diverse and complicated video contents are the major obstacles in video emotion understanding. In this paper, we propose a modality fusion framework to combine concept and content features from action, scene and object models. We conduct concept selection to investigate the relations between high-level concept features and emotions. The discriminative concepts play important roles in emotion recognition. Fusing different modality of content features further improve the performance. The extensive experiments show the state-of-the-art results on two challenging video emotion benchmarks.
Baohan Xu, Yingbin Zheng, Hao Ye 0005, Caili Wu, Gufei Sun
ICME2
2019 False positive rate control for positive unlabeled learning
Shuchen Kong, Weiwei Shen, Yingbin Zheng, Jian Pu, Jun Wang 0006
Neurocomputing3
2019 Dense Dilated Network for Video Action Recognition
abstract
The ability to recognize actions throughout a video is essential for surveillance, self-driving, and many other applications. Although many researchers have investigated deep neural networks to get a better result in video action recognition, these networks usually require a large number of well-labeled data to train. In this paper, we introduce a dense dilated network to collect action information from snippet-level to global-level. The dilated dense network is composed of the blocks with densely connected dilated convolutions layers. Our proposed framework is capable of fusing outputs from each layer to learn high-level representations, and these representations are robust even with only a few training snippets. We study different spatial and temporal modality fusing configurations and introduce a novel temporal guided fusion upon the dense dilated network which can further boost the performance. We conduct extensive experiments on two popular video action datasets: UCF101 and HMDB51. The experiments demonstrate the effectiveness of our proposed framework.
Baohan Xu, Hao Ye 0005, Yingbin Zheng, Tianyu Luwang, Yu-Gang Jiang 0001
IEEE Trans. Image Process.3
2018 Precise Temporal Action Localization by Evolving Temporal Proposals
abstract
Locating actions in long untrimmed videos has been a challenging problem in video content analysis. The performances of existing action localization approaches remain unsatisfactory in precisely determining the beginning and the end of an action. Imitating the human perception procedure with observations and refinements, we propose a novel three-phase action localization framework. Our framework is embedded with an Actionness Network to generate initial proposals through frame-wise similarity grouping, and then a Refinement Network to conduct boundary adjustment on these proposals. Finally, the refined proposals are sent to a Localization Network for further fine-grained location regression. The whole process can be deemed as multi-stage refinement using a novel non-local pyramid feature under various temporal granularities. We evaluate our framework on THUMOS14 benchmark and obtain a significant improvement over the state-of-the-arts approaches. Specifically, the performance gain is remarkable under precise localization with high IoU thresholds. Our proposed framework achieves [email protected]=0.5 of 34.2%.
Haonan Qiu, Yingbin Zheng, Hao Ye 0005, Yao Lu 0028, Feng Wang 0036, Liang He 0001
ICMR2
2018 Dense Dilated Network for Few Shot Action Recognition
abstract
Recently, video action recognition has been widely studied. Training deep neural networks requires a large amount of well-labeled videos. On the other hand, videos in the same class share high-level semantic similarity. In this paper, we introduce a novel neural network architecture to simultaneously capture local and long-term spatial temporal information. The dilated dense network is proposed with the blocks being composed of densely-connected dilated convolutions layers. The proposed framework is capable of fusing each layer's outputs to learn high-level representations, and the representations are robust even with only few training snippets. The aggregations of dilated dense blocks are also explored. We conduct extensive experiments on UCF101 and demonstrate the effectiveness of our proposed method, especially with few training examples.
Baohan Xu, Hao Ye 0005, Yingbin Zheng, Tianyu Luwang, Yu-Gang Jiang 0001
ICMR3
2018 Learning part-based mid-level representation for visual recognition
Baodi Yuan, Jian Tu, Yingbin Zheng, Yu-Gang Jiang 0001
Neurocomputing4
2018 Learning Multiviewpoint Context-Aware Representation for RGB-D Scene Classification
abstract
Effective visual representation plays an important role in the scene classification systems. While many existing methods are focused on the generic descriptors extracted from the RGB color channels, we argue the importance of depth context, since scenes are composed with spatial variability and depth is an essential component in understanding the geometry. In this letter, we present a novel depth representation for RGB-D scene classification based on a specific designed convolutional neural network (CNN). Contrast to previous deep models that transfer from pretrained RGB CNN models, we harness model by using the multiviewpoint depth image augmentation to overcome the data scarcity problem. The proposed CNN framework contains the dilated convolutions to expand the receptive field and a subsequent spatial pooling to aggregate multiscale contextual information. The combination of contextual design and multiviewpoint depth images are important toward a more compact representation, compared to directly using original depth images or off-the-shelf networks. Through extensive experiments on SUN RGB-D dataset, we demonstrate that the representation outperforms recent state of the arts, and combining it with standard CNN-based RGB features can lead to further improvements.
Yingbin Zheng, Hao Ye 0005, Li Wang 0033, Jian Pu
IEEE Signal Process. Lett.1
2018 Arbitrary-Oriented Scene Text Detection via Rotation Proposals
abstract
This paper introduces a novel rotation-based framework for arbitrary-oriented text detection in natural scene images. We present theRotation Region Proposal Networks, which are designed to generate inclined proposals with text orientation angle information. The angle information is then adapted for bounding box regression to make the proposals more accurately fit into the text region in terms of the orientation. TheRotation Region-of-Interestpooling layer is proposed to project arbitrary-oriented proposals to a feature map for a text region classifier. The whole framework is built upon a region-proposal-based architecture, which ensures the computational efficiency of the arbitrary-oriented text detection compared with previous text detection systems. We conduct experiments using the rotation-based framework on three real-world scene text detection datasets and demonstrate its superiority in terms of effectiveness and efficiency over previous approaches.
Jianqi Ma, Weiyuan Shao, Hao Ye 0005, Li Wang 0033, Hong Wang 0014, Yingbin Zheng, Xiangyang Xue 0001
IEEE Trans. Multim.6
2017 UA-DETRAC 2017: Report of AVSS2017 & IWT4S Challenge on Advanced Traffic Monitoring
abstract
The rapid advances of transportation infrastructure have led to a dramatic increase in the demand for smart systems capable of monitoring traffic and street safety. Fundamental to these applications are a community-based evaluation platform and benchmark for object detection and multi-object tracking. To this end, we organize the AVSS2017 Challenge on Advanced Traffic Monitoring, in conjunction with the International Workshop on Traffic and Street Surveillance for Safety and Security (IWT4S), to evaluate the state-of-the-art object detection and multi-object tracking algorithms in the relevance of traffic surveillance. Submitted algorithms are evaluated using the large-scale UA-DETRAC benchmark and evaluation protocol. The benchmark, the evaluation toolkit and the algorithm performance are publicly available from the website http://detrac-db.rit.albany.edu.
Siwei Lyu, Ming-Ching Chang, Dawei Du, Longyin Wen, Honggang Qi, Yuezun Li, Yi Wei 0006, Lipeng Ke, Tao Hu 0011, Marco Del Coco, Pierluigi Carcagnì, Dmitriy Anisimov, Erik Bochinski, Fabio Galasso, Filiz Bunyak, Hao Ye 0005, Hong Wang 0014, Kannappan Palaniappan, Koray Ozcan, Li Wang 0033, Liang Wang 0001, Martin Lauer, Nattachai Watcharapinchai, Nenghui Song, Noor Al-Shakarji, Sikandar Amin, Sitapa Watcharapinchai, Tatiana Khanova, Thomas Sikora, Tino Kutschbach, Volker Eiselein, Wei Tian 0001, Xiangyang Xue 0001, Xiaoyi Yu, Yao Lu 0028, Yingbin Zheng, Yongzhen Huang, Yuqi Zhang 0001
AVSS38
2017 Boosting Alzheimer diagnosis accuracy with the help of incomplete privileged information
abstract
Early and accurate diagnosis of Alzheimer's disease is beneficial to both preserve daily functioning and test possible new treatments. However, current diagnosis depends on dozens of factors, including the family member, past medical problems, tests of memory, blood and urine tests, brain scans and even cerebrospinal fluid specimens. Among them, regular features (e.g., blood and urine tests, brain scans) are simple and accurate. Privileged features (e.g., tests of memory, cerebrospinal fluid specimens) are inconvenient and uncomfortable to acquire, and thus incomplete in most cases. In this work, we propose a two-stage learning framework to predict merely using the regular feature with the aid of incomplete privileged data in the training time. In particular, we first complement the missing data of privileged features by exploring the relationship with regular features and labels. The recovered privileged features, regular features, as well as labels, are combined as a new set of privileged features. Then privileged learning via matching logits is applied to boost the diagnosis accuracy. Our experiments and comparison studies with competing techniques on both synthetic data and real benchmarks have corroborated the effectiveness and superiority of the proposed framework for biomedical applications.
Jian Pu, Jun Wang 0006, Yingbin Zheng, Hao Ye 0005, Weiwei Shen, Hongyuan Zha
BIBM3
2017 Evolving boxes for fast vehicle detection
abstract
We perform fast vehicle detection from traffic surveillance cameras. A novel deep learning framework, namely Evolving Boxes, is developed that proposes and refines the object boxes under different feature representations. Specifically, our framework is embedded with a light-weight proposal network to generate initial anchor boxes as well as to early discard unlikely regions; a fine-turning network produces detailed features for these candidate boxes. We show intriguingly that by applying different feature fusion techniques, the initial boxes can be refined for both localization and recognition. We evaluate our network on the recent DETRAC benchmark and obtain a significant improvement over the state-of-the-art Faster RCNN by 9.5% mAP. Further, our network achieves 9–13 FPS detection speed on a moderate commercial GPU.
Li Wang 0033, Yao Lu 0028, Hong Wang 0014, Yingbin Zheng, Hao Ye 0005, Xiangyang Xue 0001
ICME4
2016 Face Recognition via Active Annotation and Learning
abstract
In this paper, we introduce an active annotation and learning framework for the face recognition task. Starting with an initial label deficient face image training set, we iteratively train a deep neural network and use this model to choose the examples for further manual annotation. We follow the active learning strategy and derive the Value of Information criterion to actively select candidate annotation images. During these iterations, the deep neural network is incrementally updated. Experimental results conducted on LFW benchmark and MS-Celeb-1M challenge demonstrate the effectiveness of our proposed framework.
Hao Ye 0005, Weiyuan Shao, Hong Wang 0014, Jianqi Ma, Li Wang 0033, Yingbin Zheng, Xiangyang Xue 0001
ACM Multimedia6
2013 Understanding and Predicting Interestingness of Videos
abstract
The amount of videos available on the Web is growing explosively. While some videos are very interesting and receive high rating from viewers, many of them are less interesting or even boring. This paper conducts a pilot study on the understanding of human perception of video interestingness, and demonstrates a simple computational method to identify more interesting videos. To this end we first construct two datasets of Flickr and YouTube videos respectively. Human judgements of interestingness are collected and used as the ground-truth for training computational models. We evaluate several off-the-shelf visual and audio features that are potentially useful for predicting interestingness on both datasets. Results indicate that audio and visual features are equally important and the combination of both modalities shows very promising results.
Yu-Gang Jiang 0001, Rui Feng 0001, Xiangyang Xue 0001, Yingbin Zheng, Hanfang Yang
AAAI5
2013 Sparse Reconstruction for Weakly Supervised Semantic Segmentation
Ke Zhang 0028, Wei Zhang 0016, Yingbin Zheng, Xiangyang Xue 0001
IJCAI3
2012 Learning Hybrid Part Filters for Scene Recognition
Yingbin Zheng, Yu-Gang Jiang 0001, Xiangyang Xue 0001
ECCV (5)1
2012 A fast video event recognition system and its application to video search
abstract
Techniques for recognizing complex events in diverse Internet videos are important in many applications. State-of-the-art video event recognition approaches normally involve modules that demand extensive computation, which prevents their application to large scale problems. In this demonstration, we present a fast video event recognition system, which requires just a few seconds to process a general YouTube video with a few minutes of duration. The development of this system is grounded on several important findings from a large set of empirical studies, where we systematically evaluated many technical options for each critical module of a present-day video event recognition framework. Pooling the insights gained from this study leads to a speeded-up event recognition system that is 220-times faster than a decent baseline while still has a high degree of recognition accuracy. We also demonstrate the technical feasibility of using event recognition results as the sole clue for video search, where the similarity of videos is determined based on the consistency of the event recognition confidence scores. We showcase this capability using an Internet video dataset containing about 10 thousands of YouTube videos. Very promising results were observed.
Yu-Gang Jiang 0001, Qi Dai 0001, Yingbin Zheng, Xiangyang Xue 0001
ACM Multimedia3
2012 A simplified multi-class support vector machine with reduced dual optimization
Xisheng He, Zhe Wang 0002, Cheng Jin 0001, Yingbin Zheng, Xiangyang Xue 0001
Pattern Recognit. Lett.4
2011 Refining local descriptors by embedding semantic information for visual categorization
abstract
Local descriptor extraction and vector quantization are the important components of widely-used Bag-of-Features (BoF) model for visual categorization. This paper proposes a simple and efficient approach to refine the local descriptors for vector quantization by embedding semantic information. The original local descriptors are integrated by a sequence of category-independent and category-dependent basis. Particularly, the category-dependent basis is learned by minimizing the joint loss minimization over local descriptors from different categories with a shared regularization penalty, which can be formulated as a linear programming problem. The transferred descriptors are further quantized and aggregated to the visual vocabulary. Experiments are performed on PASCAL VOC 2007 benchmark and the quantitative comparisons with several state-of-the-art approaches demonstrate the effectiveness of our proposed approach.
Yingbin Zheng, Renzhong Wei, Hong Lu 0001, Xiangyang Xue 0001
ACM Multimedia1
2010 How context helps: A discriminative codeword selection method for object detection
abstract
We first propose in this paper to localize objects in images based on the models learned from the weakly labeled images. This task is termed as region of interest (ROI) detection. Local features such as SIFT or HOG are extracted and the discriminative words from clustered codewords based on SIFT and HOG are selected to model the objects. Then how to find the discriminative words to model the object is important. Existing ROI detection methods consider the information from the foreground objects by selecting the words appearing more in the images belonging to one specific image class. Considering the information from background/context is also helpful for object detection and classification, we propose to select the discriminative words which appear more in the foreground/object and less in the background/context. Second, another task is to give the class label (object in this setting) for a given image and also give the position of the object appearing in the image. This task is termed as objection detection. A normal way for this task after ROI is to extract features from the detected regions and not from the whole image. Since the discriminative words extracted during ROI detection has good discriminative ability, we propose to use these words for object detection. Experimental results on PASCAL VOC 2006 dataset and a larger dataset containing 29 classes demonstrate the effectiveness of the proposed method.
Renzhong Wei, Hong Lu 0001, Yingbin Zheng, Lei Cen, Cheng Jin 0001, Xiangyang Xue 0001, Weiguo Wu
ICIP3
2010 Semantic video indexing by fusing explicit and implicit context spaces
abstract
This paper addresses the problem of context-based concept fusion (CBCF) for concept detection and semantic video indexing. We introduce a novel framework based on constructing context spaces of concepts, such that the contextual correlations are used to improve the performance of concept detectors. Different from traditional CBCF approach, we present two kinds of such context spaces: explicit context space for modeling the correlation of pairwise concepts, and implicit context space for representing latent themes trained from a set of concepts. The final concept detection scores are then directly fused from explicit and implicit context spaces. Experiments are presented on TRECVid 2006 benchmark and the comparisons with several state-of-the-art approaches demonstrate the effectiveness of proposed framework.
Yingbin Zheng, Renzhong Wei, Hong Lu 0001, Xiangyang Xue 0001
ACM Multimedia1
2009 Incorporating Spatial Correlogram into Bag-of-Features Model for Scene Categorization
Yingbin Zheng, Hong Lu 0001, Cheng Jin 0001, Xiangyang Xue 0001
ACCV (1)1