Dongming Zhang 0004

dblp:96/5692-4 · DBLP profile ↗
← Back
56ranked-venue papers
4as first author
13since 2021 · last 2025
0000-0002-1237-7177ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 49 · 3 first-author · 12 since 2021Artificial intelligence and machine learning · 12 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2025 Face Forgery Detection Attack via Hybrid Perturbation with Dual Mask
abstract
In this paper, we attempt to attack deep forgery detection systems rigorously to uncover their shortcomings. Unlike existing methods that solely attack fake regions to increase their authenticity, the proposed identifies decisive regions to decision and creates a dual mask to carry out different attacks. The dual mask comprises two components: a positive mask, which identifies real regions for targeted attacks, and a negative mask, pinpointing fake regions for disruption. Perturbations applied to the positive mask aim to bolster confidence in real decisions, while those on the negative mask aim to diminish confidence in fake decisions, thereby inducing the detection system to make erroneous judgments. Additionally, to enhance the visual quality of the manipulated samples, we introduce a hybrid constraint optimization, which carefully balances the advantages and disadvantages of spatial perturbations and frequency-domain adversarial samples, facilitating the learning of perturbations that maintain high attack efficacy while preserving visual fidelity. These two modules are integrated into a novel GAN-based attacking network, HPDM, to generate adversarial samples and achieve improvements in attack transferability and imperceptibility. We evaluated our method through experiments on FFHQ and FF++. The results demonstrate that our approach consistently outperforms existing methods in terms of average attack efficacy.
Zixiang Wu, Dongming Zhang 0004, Zijie Yang
IJCNN2
2025 RealText: Realistic Text Image Generation based on Glyph and Scene Aware Inpainting
abstract
Text-to-image generation models can create diverse, high-quality images, but they frequently encounter challenges in accurately rendering text within those images due to the insufficient representation of desired text. In this study, we introduce RealText, a method for generating scene text images that excels in producing precise and realistic scene text images in any language. We disentangle scene text images generation into three stages: background and glyph image generation, text deformation, and whole image generation. Initially, we utilize prompts to guide the creation of well-organized background images. By identifying optimal text placements on these backgrounds, we render the glyph images of target text using user-specified font, effectively eliminating incorrect characters. In the next stage, we propose scene sensing to perceive text carrier surfaces and viewpoints through 3D scene reconstruction using depth and normal map to apply text deformation, thereby enhancing the realism of generated images. The final stage involves generating complete image with the aid of background and glyph guidance. Thanks to glyph disentangling, scene sensing, and text inpainting, we can exert more precise control over scene text image generation process. We have developed a unified framework which supports major generation models. Extensive experiments illustrate the exceptional performance of our method in generating images with multilingual text. The codes will soon be available at https://github.com/cccvl/RealText.
Zihou Liu, Dongming Zhang 0004, Jing Zhang 0023, Yongdong Zhang 0001
ACM Multimedia2
2025 Cross-Modal Contrastive Masked AutoEncoder for Compressed Video Pre-Training
abstract
In this paper, we propose a novel Transformer based approach, namely Cross-modal Contrastive Masked AutoEncoder (C2MAE), to Self-Supervised Learning (SSL) on compressed videos. A unified Transformer encoder is employed to discover relationships of visual tokens from RGBs, motion vectors and residuals. A hybrid SSL framework is proposed, which combines the complementary advantages of Masked Image Modeling (MIM) and Contrastive Learning (CL) pretext tasks, for powerful representation learning. The MIM branch extends VideoMAE by a new Fine-Grained Motion-aware Masking (FGMM) strategy and a modified Multi-modal Reconstruction (MR) task, where FGMM computes motion saliency maps as motion priors to guide the masks so that it well fits for the data properties in the compressed domain and the MR task highlights the reconstruction of raw videos by joint representations from corresponding compressed videos in addition to that in each single modality. The CL branch introduces the Contrastive Cross-modal Learning (CCL) module, and the features from a compressed video clip and the ones from its raw video counterpart are compared instead of widely used augmented data. Due to these designs, C2MAE significantly enhances interactions across modalities to compensate the sparsity of I-frames and the coarse and noisy nature of P-frames, thus delivering much stronger pre-trained models. Extensive experiments are conducted on the UCF-101, HMDB-51 and Kinetics-400 benchmarks with state-of-the-art results reported, demonstrating its effectiveness.
Jiaxin Chen 0002, Guohao Li 0010, Dongming Zhang 0004, Xiuguo Bao, Di Huang 0001
IEEE Trans. Image Process.4
2025 Masked Text Pre-Training for Scene Text Detection
Hongtao Xie 0001, Keran Wang, Bangbang Zhou, Yuxin Wang 0002, Weigang Qi, Yadong Qu, Zuan Gao, Dongming Zhang 0004
IEEE Trans. Multim.9
2024 Reference-Aware Adaptive Network for Image-Text Matching
abstract
Image-text matching aims to bridge vision and language areas, which is a crucial task in multi-modal intelligence. The core idea is to learn features of each modality and aggregate learned features as holistic representations to measure image-text relevance. Most existing methods involve cross-modal interaction during feature learning by modeling fine-grained relationships between two modalities for better results. However, these methods may obtain wrong attention scores when directly computing similarities between regions and words. Besides, current methods mainly rely on simple pooling operations for feature aggregation, which introduces interference from redundant information, resulting in inaccurate matching results. To alleviate these issues, we propose a novel reference-aware adaptive network for image-text matching by jointly using a reference attention module for feature learning and an adaptive aggregation module for feature aggregation. The proposed model enjoys several merits. First, the designed reference attention module effectively reduces wrong attention scores by introducing a set of references during cross-modal interaction. Second, the proposed adaptive aggregation module highlights useful information adaptively while suppressing redundant information during aggregation. Extensive experiments on two standard benchmarks demonstrate that our method performs favorably against state-of-the-art methods.
Guoxin Xiong, Tianzhu Zhang 0001, Dongming Zhang 0004, Yongdong Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 Bi-Source Reconstruction-Based Classification Network for Face Forgery Video Detection
abstract
Current methods for detecting deep fakes concentrate on specific patterns of forgery like noise characteristics, local textures, or frequency statistics. These approaches assume training and test sets exhibit similar data distributions, which bring severe performance drops and further limit broader applications when migrating unseen domains. Existing works show that reconstruction learning is effective in capturing unseen forgery clues. However, 2D reconstruction is insufficient and can not handle non-frontal face reconstruction, while 3D reconstruction provides more critical details of facial structure and finds accurate forgery regions. In this paper, we propose a bi-source reconstruction based classification network (BRCNet) to incorporate 2D and 3D reconstruction as the supervisions and learn the optimal feature representation. In detail, we employ an encoder-decoder architecture to facilitate reconstruction learning, enhancing the learned representations to detect forgery patterns that are unknown. To further capture forgery evidence across multiple scales, instead of using encoder features from the reconstruction network only, we build a feature improvement network to combine feature details from encoder and decoder features in a multi-scale fashion. In addition, we use the reconstruction difference to supervise the feature aggregation, which enables detecting the subtle and trivial discrepancies between fake and real video frames. Extensive experiments are conducted to validate the performance of our proposed method on several deep fake benchmarks. The results demonstrate the efficacy of our approach, offering promising results and showcasing its potential for practical applications. The source code is available at https://github.com/cccvl/BRCNet.
Dongming Zhang 0004, Chenqin Fu, Dingyu Lu, Yongdong Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2024 Semantic-Enhanced Proxy-Guided Hashing for Long-Tailed Image Retrieval
abstract
Hashing has been studied extensively for large-scale image retrieval due to its efficient computation and storage. Deep hashing methods typically train models with category-balanced data and suffer from a serious performance deterioration when dealing with long-tailed training samples. Recently, several long-tailed hashing methods focus on this newly emerging field for practical purpose. However, existing methods still face challenges that fixed category centers with limited semantic information cannot effectively improve the discriminative ability of tail-category hash codes. To tackle the issue, we propose a novel method called Semantic-enhanced Proxy-guided Hashing in this paper. We leverage two sets of learnable category proxies in the feature space and the Hamming space respectively, which can describe category semantics by getting updated continuously along with the whole model via back-propagation. Based on this, we introduce the Mahalanobis distance metric to characterize relationships accurately and enhance the semantic representation of both proxies and samples concurrently, improving the hash learning process. Moreover, we capture the multilateral correlations between proxies and samples in the feature space and extend a hypergraph neural network to transfer semantic knowledge from proxies to samples in the Hamming space. Extensive experiments show that our method achieves the state-of-the-art performance and surpasses existing methods by 1.47%–7.56% MAP on long-tailed benchmarks, demonstrating the superiority of learnable category proxies and the effectiveness of our proposed learning algorithm for long-tailed hashing.
Hongtao Xie 0001, Lei Zhang 0119, Pandeng Li, Dongming Zhang 0004, Yongdong Zhang 0001
IEEE Trans. Multim.5
2023 Unilaterally Aggregated Contrastive Learning with Hierarchical Augmentation for Anomaly Detection
abstract
Anomaly detection (AD), aiming to find samples that deviate from the training distribution, is essential in safety-critical applications. Though recent self-supervised learning based attempts achieve promising results by creating virtual outliers, their training objectives are less faithful to AD which requires a concentrated inlier distribution as well as a dispersive outlier distribution. In this paper, we propose Unilaterally Aggregated Contrastive Learning with Hierarchical Augmentation (UniCon-HA), taking into account both the requirements above. Specifically, we explicitly encourage the concentration of inliers and the dispersion of virtual outliers via supervised and unsupervised contrastive losses, respectively. Considering that standard contrastive data augmentation for generating positive views may induce outliers, we additionally introduce a soft mechanism to re-weight each augmented inlier according to its deviation from the inlier distribution, to ensure a purified concentration. Moreover, to prompt a higher concentration, inspired by curriculum learning, we adopt an easy-to-hard hierarchical augmentation strategy and perform contrastive aggregation at different depths of the network based on the strengths of data augmentation. Our method is evaluated under three AD settings including unlabeled one-class, unlabeled multi-class, and labeled multi-class, demonstrating its consistent superiority over other competitors.
Guodong Wang 0006, Yunhong Wang 0001, Jie Qin 0004, Dongming Zhang 0004, Xiuguo Bao, Di Huang 0001
ICCV4
2023 Dual Dynamic Proxy Hashing Network for Long-tailed Image Retrieval
abstract
Deep hashing has been extensively explored for image retrieval due to fast computation and efficient storage. Since conventional deep hashing methods are not suitable for the common scenario in real life that data exhibits a long-tailed distribution, several long-tailed hashing methods have been proposed recently. However, existing long-tail hashing methods seek to utilize fixed class centroids and cannot fully develop the discriminative ability of hash codes for tail-class samples. Specifically, fixed class centroids cannot characterize authentic semantics of tail classes or provide effective semantic information for hash codes learning under the long-tailed setting. To this end, we propose a novel Dual Dynamic Proxy Hashing Network (DDPHN) with two sets of learnable dynamic proxies, i.e. hash proxies and feature proxies, to improve the discrimination of hash codes for tail-class samples. Compared with fixed class centroids, learnable proxies can be optimized constantly via the proxy learning loss and depict accurate class semantics despite the scarcity of tail-class samples. Apart from low-dimensional binary hash proxies, we introduce high-dimensional continuous feature proxies that can describe semantic relationships more precisely, contributing to hash codes learning as well. To further leverage semantic information carried by proxies, we build a hypergraph by exploring neighborhood relationships in the feature space and then introduce a hypergraph neural network to transfer knowledge from proxies to samples in the Hamming space. Extensive experiments show the superiority of our learnable dynamic proxies and demonstrate that our method outperforms numerous deep hashing models and recent state-of-the-art long-tailed hashing methods.
Hongtao Xie 0001, Lei Zhang 0119, Pandeng Li, Dongming Zhang 0004, Yongdong Zhang 0001
ACM Multimedia5
2023 Masked Text Modeling: A Self-Supervised Pre-training Method for Scene Text Detection
abstract
Scene text detection has made great progress recently with the wide use of pre-training. Nonetheless, existing scene text detection methods still suffer from two problems: 1) Limited annotated real data reduces the feature robustness. 2) Detectors perform poorly on text lacking of visual information. In this paper, we explore the potential of the CLIP model, and propose a novel self-supervised Masked Text Modeling (MTM) pre-training method for scene text detection, which can be trained with unlabeled data and improve the linguistic reasoning ability for text occlusion. Different from previous randomly pixel-level masking methods, MTM performs a targeted text-aware masking process under an unsupervised manner. Specifically, MTM consists of text perception and masked text modeling. In the text perception step, benefiting from the text-friendliness of CLIP, a Text Perception Module is proposed to attend to text area by computing the similarity between the text and image tokens from CLIP model. In the masked text modeling step, a Text-aware Masking Strategy is designed to mask the text area, and the Masked Text Modeling Module is used to reconstruct the masked texts. MTM obtains the ability to reason the linguistic information of masked texts with the reconstruction. This robust feature extraction learned by MTM ensures a more discriminative representation for the text lacking of visual information. Moreover, a new text dataset named OcclusionText is proposed to evaluate the robustness for text occlusion of detection methods. Extensive experiments on public benchmarks demonstrate that our MTM can boost the performance of existing text detectors.
Keran Wang, Hongtao Xie 0001, Yuxin Wang 0002, Dongming Zhang 0004, Yadong Qu, Zuan Gao, Yongdong Zhang 0001
ACM Multimedia4
2023 S$^{2}$-Net:Semantic and Saliency Attention Network for Person Re-Identification
abstract
Person re-identification is still a challenging task when moving objects or another person occludes the probe person. Mainstream methods based on even partitioning apply an off-the-shelf human semantic parsing to highlight the non-collusion part. In this paper, we apply an attention branch to learn the human semantic partition to avoid misalignment introduced by even partitioning. In detail, we propose a semantic attention branch to learn 5 human semantic maps. We also note that some accessories or belongings, such as a hat, bag, may provide more informative clues to improve the person Re-ID. Human semantic parsing, however, usually treats non-human parts as distractions and discards them. To fetch the missing clues, we design a branch to capture the salient non-human parts. Finally, we merge the semantic and saliency attention to build an end-to-end network, named as S$^{2}$-Net. Specifically, to further improve Re-ID, we develop a trade-off weighting scheme between semantic and saliency attention and set the right weight with the actual scene. The extensive experiments show that S$^{2}$-Net gets the competitive performance. S$^{2}$-Net achieves 87.4% mAP on Market1501 and obtains 79.3%/56.1% rank-1/mAP on MSMT17 without semantic supervision. The source codes are available athttps://github.com/upgirlnana/S2Net.
Xuena Ren, Dongming Zhang 0004, Xiuguo Bao, Yongdong Zhang 0001
IEEE Trans. Multim.2
2022 Video Anomaly Detection by Solving Decoupled Spatio-Temporal Jigsaw Puzzles
Guodong Wang 0006, Yunhong Wang 0001, Jie Qin 0004, Dongming Zhang 0004, Xiuguo Bao, Di Huang 0001
ECCV (10)4
2022 Representation Learning for Compressed Video Action Recognition via Attentive Cross-modal Interaction with Motion Enhancement
abstract
Compressed video action recognition has recently drawn growing attention, since it remarkably reduces the storage and computational cost via replacing raw videos by sparsely sampled RGB frames and compressed motion cues (e.g., motion vectors and residuals). However, this task severely suffers from the coarse and noisy dynamics and the insufficient fusion of the heterogeneous RGB and motion modalities. To address the two issues above, this paper proposes a novel framework, namely Attentive Cross-modal Interaction Network with Motion Enhancement (MEACI-Net). It follows the two-stream architecture, i.e. one for the RGB modality and the other for the motion modality. Particularly, the motion stream employs a multi-scale block embedded with a denoising module to enhance representation learning. The interaction between the two streams is then strengthened by introducing the Selective Motion Complement (SMC) and Cross-Modality Augment (CMA) modules, where SMC complements the RGB modality with spatio-temporally attentive local motion features and CMA further combines the two modalities with selective feature augmentation. Extensive experiments on the UCF-101, HMDB-51 and Kinetics-400 benchmarks demonstrate the effectiveness and efficiency of MEACI-Net.
Jiaxin Chen 0002, Dongming Zhang 0004, Xiuguo Bao, Di Huang 0001
IJCAI3
2020 Towards Practical Compressed Video Action Recognition: A Temporal Enhanced Multi-Stream Network
abstract
Current compressed video action recognition methods are mainly based on complete data. However, in a real transmission scenario, the compressed video packets are usually disorderly received and even lost due to network jitters or congestion. To recognize actions in early phases with limited packets, e.g. for quickly forecasting possible potential risks, in this paper, we propose a Temporal Enhanced Multi-Stream Network (TEMSN) towards practical compressed video action recognition. First, we make use of three modalities in the compressed domain as complementary cues and build a multi-stream network to capture rich information from compressed video packets. Second, we design a temporal enhanced module based on an Encoder-Decoder structure, which is applied to each stream to infer missing packets, generating more accurate action dynamics. Thanks to the multiple modalities and their temporal enhancement, our approach better models actions with partial available compressed video packets. Experiments on the HMDB-51 and UCF-101 datasets validate its effectiveness and efficiency.
Longteng Kong, Dongming Zhang 0004, Xiuguo Bao, Di Huang 0001, Yunhong Wang 0001
ICPR3
2019 APE-GAN: Adversarial Perturbation Elimination with GAN
abstract
Although Deep Neural Networks could achieve state-of-the-art performance while recongnizing images, they often suffer a tremendous defeat from adversarial examples-inputs generated by utilizing imperceptible but intentional perturbations to samples from the datasets. So far, very few methods have provided a significant defense to adversarial examples. In this paper, an effective framework based Generative Adversarial Nets(GAN) is proposed to defense against the adversarial examples. The essense of the model is to eliminate the adversarial perturbations being highly aligned with the weight vectors of nueral models. Extensive experiments on benchmark datasets MNIST, CIFAR10 and ImageNet indicate that our framework is able to defense against adversarial examples effectively.
Guoqing Jin, Shiwei Shen, Dongming Zhang 0004, Yongdong Zhang 0001
ICASSP3
2018 Semantic Preserving Hash Coding Through VAE-GAN
abstract
This paper proposes a novel framework for fast image retrieval. The proposed framework combines variational autoencoder with generative adversarial network to generate content preserving images for learning-based hashing. By accepting real image and systhesized image in a pairwise form, a semantic perserving binary mapping model is learned using pairwise ranking loss under an adversarial generative process. Extensive experiments on several benchmark datasets demonstrate that the proposed method shows substantial improvement over the state-of-the-art hashing methods.
Guoqing Jin, Dongming Zhang 0004, Junbo Guo, Yike Ma, Yongdong Zhang 0001
ICIP2
2018 Eigenobject-wise saliency detection based on manifold ranking
Guoqing Jin, Dongming Zhang 0004, Yongdong Zhang 0001
Neurocomputing2
2018 Region similarity arrangement for large-scale image retrieval
Dongming Zhang 0004, Jingya Tang, Guoqing Jin, Yongdong Zhang 0001, Qi Tian 0001
Neurocomputing1
2017 Task-Driven Dynamic Fusion: Reducing Ambiguity in Video Description
abstract
Integrating complementary features from multiple channels is expected to solve the description ambiguity problem in video captioning, whereas inappropriate fusion strategies often harm rather than help the performance. Existing static fusion methods in video captioning such as concatenation and summation cannot attend to appropriate feature channels, thus fail to adaptively support the recognition of various kinds of visual entities such as actions and objects. This paper contributes to: 1)The first in-depth study of the weakness inherent in data-driven static fusion methods for video captioning. 2) The establishment of a task-driven dynamic fusion (TDDF) method. It can adaptively choose different fusion patterns according to model status. 3) The improvement of video captioning. Extensive experiments conducted on two well-known benchmarks demonstrate that our dynamic fusion method outperforms the state-of-the-art results on MSVD with METEOR scores 0.333, and achieves superior METEOR scores 0.278 on MSR-VTT-10K. Compared to single features, the relative improvement derived from our fusion method are 10.0% and 5.7% respectively on two datasets.
Xishan Zhang, Ke Gao 0012, Yongdong Zhang 0001, Dongming Zhang 0004, Jintao Li 0001, Qi Tian 0001
CVPR4
2017 Deep saliency map estimation of hand-crafted features
abstract
Saliency detection that utilizes deep convolutional neural networks to obtain high level features from original images has achieved considerable progress during the past years. However, few methods consider learning saliency cues from hand-crafted features. In this paper, we demonstrate that deep learning can produce good enough saliency detection results using only hand-crafted features. We propose a novel multi-context deep learning saliency detection algorithm, where only hand-crafted features are taken into account and modeled in a unified deep learning framework. Extensive experiments on benchmark datasets indicate significant and consistent improvements over the representative deep learning framework based saliency detection methods.
Guoqing Jin, Shiwei Shen, Dongming Zhang 0004, Wenjing Duan, Yongdong Zhang 0001
ICIP3
2017 Large-scale person re-identification as retrieval
abstract
This paper targets to bring together the research efforts on two fields that are growing actively in the past few years: multicamera person Re-Identification (ReID) and large-scale image retrieval. We demonstrate that the essentials of image retrieval and person ReID are the same, i.e., measuring the similarity between images. However, person ReID requires more discriminative and robust features to identify the subtle differences of different persons and overcome the large variance among images of the same person. Specifically, we propose a coarse-to-fine (C2F) framework and a Convolutional Neural Network structure named as Conv-Net to tackle the large-scale person ReID as an image retrieval task. Given a query person image, the C2F firstly employ Conv-Net to extract a compact descriptor and perform the coarse-level search. A robust descriptor conveying more spatial cues is hence extracted to perform the fine-level search. Extensive experimental results show that the proposed method outperforms existing methods on two public datasets. Further, the evaluation on a large-scale Person-520K dataset demonstrates that our work is significantly more efficient than existing works, e.g., only needs 180ms to identify a query person from 520K images.
Hantao Yao, Shiliang Zhang, Dongming Zhang 0004, Yongdong Zhang 0001, Jintao Li 0001, Yu Wang 0089, Qi Tian 0001
ICME3
2017 DSP: Discriminative Spatial Part modeling for Fine-Grained Visual Categorization
Hantao Yao, Dongming Zhang 0004, Jintao Li 0001, Jianshe Zhou, Shiliang Zhang, Yongdong Zhang 0001
Image Vis. Comput.2
2017 Trip Outfits Advisor: Location-Oriented Clothing Recommendation
abstract
When packing for a journey, have you ever asked “what clothes should I take with me?” Wearing appropriate and aesthetically pleasing clothing when traveling is a concern for many of us. Our data observation of photos from several popular travel websites reveals that people's choice of clothing items and their color combinations have strong correlations with the weather, the season, and the main type of attraction at the destination. This leads to an interesting and novel problem: can the correlation between clothing and locations be automatically learned from social photos and leveraged for location-oriented clothing recommendations? In this paper, we systematically study this problem and propose a hybrid multilabel convolutional neural network combined with the support vector machine (mCNN-SVM) approach to capture the intrinsic and complex correlations between clothing attributes and location attributes. Specifically, we adapt the CNN architecture to multilabel learning and fine-tune it using each fine-grained clothing item. Then, the recognized items are fed to the SVM to learn the correlations. Experiments on three fashion datasets and a benchmark journey outfit dataset show that our proposed approach outperforms several baselines by over 10.52-16.38% in terms of the mAP for clothing item recognition and outperforms several alternative methods by over 9.59-29.41% in terms of the mAP when ranking clothing by appropriateness for travel destinations. Finally, an interesting case study demonstrates the effectiveness of our method by answering what items to wear, how to match them, and how to dress in an aesthetically pleasing manner for a journey.
Xishan Zhang, Jia Jia 0001, Ke Gao 0012, Yongdong Zhang 0001, Dongming Zhang 0004, Jintao Li 0001, Qi Tian 0001
IEEE Trans. Multim.5
2016 Region similarity arrangement for image retrieval
abstract
We propose a promising method of geometric verification to improve the precision of Bag-of-Words (BoW) model in image retrieval. Most previous methods focus on the positions of interest points or the absolute differences of regions' scales and angles. In contrast, our method, named Region Similarity Arrangement (RSA), exploits the spatial arrangement of interest regions. For each image, RSA constructs a Region Property Space, regarding each region's (scale, angle) pair as a point in a polar coordinate system, and encodes the arrangement of these points into the BoW vector. From experimental results on Holidays, Oxford5K and Paris, RSA could get comparable results with sate-of-the-art methods. In addition, RSA increases no extra memory and negligible computational consumption compared with the baseline BoW approach.
Jingya Tang, Dongming Zhang 0004, Yongdong Zhang 0001, Qi Tian 0001
ICME2
2015 Binary feature from intensity quantization and weakly spatial contextual coding for image search
Dongye Zhuang, Dongming Zhang 0004, Jintao Li 0001, Qi Tian 0001
Inf. Sci.2
2014 Hamming embedding with fragile bits for image search
abstract
Recently, several binary descriptors are proposed, which represent interest points in image using binary codes. In these binary feature schemes, two descriptors are considered as a match, if the Hamming distance between them is below a threshold. Applying Hamming distance to measure the similarity between binary descriptors can extremely promote the computational efficiency. However, our experimental results presents that there exists a large number of bits in the binary feature vector cannot maintain the robustness while image conditions change. Rather than ignore the impacts of those unstable bits, we take into account the difference of robustness among the feature bits and propose a novel similarity measurement, which called the Fragile Bit Ratio (FBR). FBR is used in binary feature matching to measure how two features differ. High FBRs are associated with genuine matches between two binary features and low FBRs are associated with impostor ones. Based on this metric, we propose a new binary feature matching scheme to fuse the Hamming distance and Fragile Bit Ratio. In our approach, we match the descriptors using the Hamming distance threshold roughly, and then filtered by the Fragile Bits Ratio to refine the candidate set. In experiments, using Fragile Bits Radio can effectively remove the false matches and highly improve the accuracy of image search. Furthermore, our method can easily be integrated into the other well-established binary features schemes.
Dongye Zhuang, Dongming Zhang 0004, Jintao Li 0001, Ke Lu 0002, Qi Tian 0001
ICIP2
2014 Representative local features mining for large-scale near-duplicates retrieval
abstract
Local features have been widely used in many computer vision related researches, such as near-duplicate image and video retrieval. However, the storage and query cost of local features become prohibitive on large-scale database. In this paper, we propose a representative local features mining method to generate a compact but more effective feature subset. First, we do an unsupervised annotation for all similar images(or frames in video) in the database. Second, we compute a comprehensive score for every local feature. The score function combines the robustness and discrimination. Finally, we sort all the local features in an image by their scores and the low-score local features can be removed. The selected local features are robust and discriminative, which can guarantee the better retrieval quality than using full of the original feature set. By our method, the number of local features can be significantly reduced and a large amount of storage and computational cost can be saved. The experimental results show that we can use 30% of the features to get a better query performance than that of full feature set.
Xiaoguang Gu, Yongdong Zhang 0001, Dongming Zhang 0004, Jintao Li 0001
ICME3
2014 Salient region detection : Integrate both global and local cues
abstract
Visual saliency detection provides an alternative methodology to semantic image understanding in many applications such as region-based image retrieval and adaptive compression of images. In this paper, we propose an approach which utilizes both global and local cues to extract saliency information. Our method can achieve better performance than existing saliency detection methods in terms of precision and recall rates. The main contributions are threefold: 1) a new model which can better describe the color perception of human beings is proposed. Based on this model, a global color contrast cue is also presented. 2) as supplements, two other global cues and one local cues are also presented to capture as much saliency information as we can. 3) a CRF model is used to integrate these cues and generate the final saliency map. Experimental results indicate that our proposed approach is effective and practicable.
Tiancai Ye, Dongming Zhang 0004, Ke Gao 0012, Guoqing Jin, Yongdong Zhang 0001, Qingsheng Yuan
ICME2
2014 Real-Time Scene Text Detection Based on Stroke Model
abstract
In this paper we bring forth a novel stroke-based method which is simple and effective to detect texts in natural scenes. We first introduce a general mathematical model to describe character strokes from the perspective of the scale space along with difference of Gaussian filters. Then we detail a text line aggregation approach utilizing the inherent text layout. Afterwards, we set up the whole scheme with three main steps, i.e. stroke extraction, text line aggregation and verification. Finally, experiments show the advantage of our method. As strokes are considered to be the fundamental component of characters, compared to edge- or other connected-component-based methods, our method is much more reasonable.
Yi Liu 0148, Dongming Zhang 0004, Yongdong Zhang 0001, Shouxun Lin
ICPR2
2014 Monte Carlo Sampling based Salient Region Detection
abstract
In this paper, a simple and effective method is proposed for salient region detection. Based on the observation that salient regions tend to be compact, connected and surrounded, our original idea is to exploit these three kinds of prior knowledge. However, concepts of spatial structure (such as connectivity and surroundedness) only have definite meanings in binary images. Thus, a Monte Carlo Sampling based Saliency model is proposed. Our model has two main advantages over other methods. Firstly, the result of each sampling process is a binary map which can greatly simplify the combination with prior knowledge of spatial structure. Secondly, our method is naturally parallelized because every sampling process is independent with each other, which makes our method very efficient. Experimental results on two datasets show that, compared with eleven state-of-the-art methods, our approach has a competitive performance and also runs very fast.
Tiancai Ye, Dongming Zhang 0004, Guoqing Jin, Ke Gao 0012, Xiaoguang Gu, Yongdong Zhang 0001
ICMR2
2014 Encoder combined video moving object detection
Lingling Tong, Dongming Zhang 0004, Yongdong Zhang 0001
Neurocomputing3
2013 A video copy detection algorithm combining local feature's robustness and global feature's speed
abstract
This paper presents a novel algorithm for fast and robust video copy detection. The idea is to use local features to estimate the copy transformation parameters first and then use the estimated parameters to guide the global-feature-based matching at a later stage. It is based on the fact that the copy transformations generally remain unchanged in a continuous video clip even in the whole video. Local-feature-based matching can find the candidates which are difficult to be detected only using global features. Furthermore, the matched local feature points can provide enough information to estimate the copy transformations. After the copy transformations are estimated, the subsequent detection can be accelerated by doing global-feature-based matching. The experimental results show that the proposed algorithm can get the same good robustness as the local-feature-based method but the faster detection speed.
Xiaoguang Gu, Dongming Zhang 0004, Yongdong Zhang 0001, Jintao Li 0001, Lei Zhang 0119
ICASSP2
2013 Encoder/Decoder for Privacy Protection Video with Privacy Region Detection and Scrambling
Dongming Zhang 0004, Jintao Li 0001
MMM (2)2
2013 Stripe Model: An Efficient Method to Detect Multi-form Stripe Structures
Dongming Zhang 0004, Junbo Guo, Shouxun Lin
MMM (1)2
2013 Learning Affine Robust Binary Codes Based on Locality Preserving Hash
Wei Zhang 0043, Ke Gao 0012, Dongming Zhang 0004, Jintao Li 0001
MMM (1)3
2013 Distribution-Aware Locality Sensitive Hashing
Lei Zhang 0119, Yongdong Zhang 0001, Dongming Zhang 0004, Qi Tian 0001
MMM (2)3
2013 A Novel Binary Feature from Intensity Difference Quantization between Random Sample of Points
Dongye Zhuang, Dongming Zhang 0004, Jintao Li 0001, Qi Tian 0001
MMM (2)2
2013 Accurate off-line query expansion for large-scale mobile visual search
Ke Gao 0012, Yongdong Zhang 0001, Dongming Zhang 0004, Shouxun Lin
Signal Process.3
2013 An improved method of locality sensitive hashing for indexing large-scale and high-dimensional features
Xiaoguang Gu, Yongdong Zhang 0001, Lei Zhang 0119, Dongming Zhang 0004, Jintao Li 0001
Signal Process.4
2012 A method for detecting salient regions using integrated features
abstract
We develop a novel algorithm for detecting salient regions. By analyzing the advantages and disadvantages of the existing methods, five principles for designing salient region detection algorithms are summarized. Based on these principles, we propose a novel method that generates saliency map with highlighted salient regions by utilizing two different features, namely visual saliency value and spatial weight. The visual saliency value is determined based on local contrast differences and low-level feature frequencies. The spatial weight is computed by analyzing the size and location of salient regions. Experimental results show that the proposed algorithm outperforms 7 state-of-the-art methods on the public image set.
Zhendong Mao 0001, Yongdong Zhang 0001, Ke Gao 0012, Dongming Zhang 0004
ACM Multimedia4
2012 Geometric context-preserving progressive transmission in mobile visual search
abstract
Progressive transmission is very effective to reduce retrieval latency in mobile visual search. However, the acceleration effects of existing progressive transmission strategies are often limited because of the neglect of geometric information in the query image. This paper proposes an effective and efficient geometric context-preserving progressive transmission method, which is suitable for mobile visual search. Here a query image is divided into blocks and local features in the same block are used as query units rather than a single feature. Since clustered features with geometric information are more discriminative, only a few of them could support correct matching with high precision. Thus our method significantly decreases the number of features needed for transmission, and dramatically reduces the retrieval latency. Experiments on Stanford dataset for mobile visual search show that, with comparable precision, we uses 43% less retrieval time than existing progressive transmission method. Moreover, we establish and release a large-scale image dataset called MVSBench which is more difficult and suitable for mobile visual search. It contains 75500 images and considers many variations like view change, blur, scale, illumination and rotation. MVSBench is another major contribution of this paper, and our method also outperforms other strategies on this dataset.
Junhai Xia, Ke Gao 0012, Dongming Zhang 0004, Zhendong Mao 0001
ACM Multimedia3
2012 Query Range Sensitive Probability Guided Multi-probe Locality Sensitive Hashing
abstract
Locality Sensitive Hashing (LSH) is proposed to construct indexes for high-dimensional approximate similarity search. Multi-Probe LSH (MPLSH) is a variation of LSH which can reduce the number of hash tables. Based on the idea of MPLSH, this paper proposes a novel probability model and a query-adaptive algorithm to generate the optimal multi-probe sequence for range queries. Our probability model takes the query range into account to generate the probe sequence which is optimal for range queries. Furthermore, our algorithm does not use a fixed number of probe steps but a query-adaptive threshold to control the search quality. We do the experiments on an open dataset to evaluate our method. The experimental results show that our method can probe fewer points than MPLSH for getting the same recall. As a result, our method can get an average acceleration of 10% compared to MPLSH.
Xiaoguang Gu, Lei Zhang 0119, Dongming Zhang 0004, Yongdong Zhang 0001, Jintao Li 0001, Ning Bao
SNPD3
2012 Improved total variation minimization method for compressive sensing by intra-prediction
Jie Xu 0031, Jianwei Ma 0006, Dongming Zhang 0004, Yongdong Zhang 0001, Shouxun Lin
Signal Process.3
2011 Compressive sensing based video scrambling for privacy protection
abstract
Surveillance video privacy protection has drawn significant attention recently. In this paper, we describe a privacy protected video surveillance system which utilizes the emerging compressive sensing (CS) theory. Privacy regions are scrambled through block based CS sampling on quantized coefficients during compression. Security is ensured by key controlled chaotic sequence which is used to construct CS measurement matrix. To prevent drift error caused by scrambling, a coding restricted scheme is exploited. Experimental results show that the proposed system effectively protects privacy with the scene intelligible. Compared with the existing ones, this system has high security and dramatic coding efficiency improvement.
Lingling Tong, Yongdong Zhang 0001, Jintao Li 0001, Dongming Zhang 0004
VCIP5
2011 A pivot-based filtering algorithm for enhancing query performance of LSH
abstract
In recent years, Locality Sensitive Hashing (LSH) (and its variant Euclidean LSH) has become a popular index structure for large-scale and high-dimensional similarity search problem. In this paper, we analyze a phenomenon we called "Non-Uniform" that degrades the query performance of LSH and propose a pivot-based algorithm to improve the query performance. We also provide a method to get optimal pivot for even larger improvement. Experiments show that our algorithm significantly improves the query performance of LSH.
Lei Zhang 0119, Xiaoguang Gu, Yongdong Zhang 0001, Dongming Zhang 0004, Jintao Li 0001
VCIP4
2011 Robust Spatial Matching for Object Retrieval and Its Parallel Implementation on GPU
abstract
Spatial matching for object retrieval is often time-consuming and susceptible to viewpoint changes. To address this problem, we propose a novel spatial matching method that is robust to viewpoint changes and implement it on modern graphics processing unit (GPU) in parallel for real-time applications. Unlike previous spatial matching methods used in object retrieval, in which the affine transformation estimation is based on the gravity vector assumption, our method abandons this strong assumption by matching the affine covariant neighbors (ACNs) of corresponding local regions and estimating affine transformation from each single pair of corresponding local regions. Taking into account real-time applications, we implement the method on modern GPU in parallel to speed up the process. Computations are distributed evenly to threads with load balancing, and device memory accesses are optimized with bitmap-based parallel scan. Experimental results demonstrate that our method is more robust and more efficient than previous methods especially when the viewpoints are changed, and the parallel implementation on GPU obtains ten times speedup.
Wenying Wang, Dongming Zhang 0004, Yongdong Zhang 0001, Jintao Li 0001, Xiaoguang Gu
IEEE Trans. Multim.2
2010 Fast and robust spatial matching for object retrieval
abstract
Spatial matching for visual words based object retrieval often involves generating affine transformation hypotheses and then choosing the best hypothesis to measure the spatial consistency. In existing methods, generating an affine transformation hypothesis either requires three correspondences or assumes images are taken in restricted range of viewpoints in using a single correspondence. In this paper, we propose a novel spatial matching method, in which the transformation hypothesis can be estimated from only a single correspondence without the assumption of the viewpoints from which the images are taken. Firstly, affine covariant neighborhoods(ACNs) of features are used to eliminate possible false matches. Secondly, we decompose the affine transformation into three sub-transforms and conquer each sub-transform by exploiting the shape information and the ACNs of a single pair of corresponding features. Experiment results demonstrate that this method improves the average retrieval precision evidently with less computation in comparison with the previous methods.
Wenying Wang, Dongming Zhang 0004, Yongdong Zhang 0001, Jintao Li 0001
ICASSP2
2010 Parallel spatial matching for object retrieval implemented on GPU
abstract
Spatial matching for object retrieval is often time-consuming and susceptible to viewpoint changes. To address this problem, we propose a novel spatial matching method and implement it on modern GPU in parallel. Unlike previous spatial matching methods, in which the affine transformation estimation is based on the gravity vector assumption, our method abandons this strong assumption by matching the ACNs (affine covariant neighbors) of corresponding local regions and estimating affine transformation from a single pair of corresponding local regions. To speed up the process, we implement the method on modern GPU in parallel. Computations are distributed evenly to threads with load balancing, and the memory accesses are optimized and bitmap based parallel scan is exploited. Experimental results demonstrate that our method is more robust and more efficient than previous methods especially when the viewpoints are changed, and the parallel implementation on GPU obtains ten times speedup.
Wenying Wang, Dongming Zhang 0004, Yongdong Zhang 0001, Jintao Li 0001, Xiaoguang Gu
ICME2
2010 Trajectory-based visualization of web video topics
abstract
While there have been research efforts in organizing large scale web videos into clusters or topics, efficient browsing of web video topics remains a challenging problem not yet addressed. The related issues include how to efficiently browse and track the evolution of topics and eventually locate the videos of interest. In this demo paper, we introduce a novel interface for visualizing video topics as evolution trajectories. The trajectory visualization is capable of highlighting milestone events and depicting the topical hotness over time. The interface also allows multi-level browsing from topics to events and to videos, resulting in search exploration could be more efficiently conducted to locate videos of interest. In addition, recommendation of topics accordingly to three-hots content-hot, evolution-hot and potential-hot, can be easily supported by our system. A user study on three months. YouTube videos using our interface demonstrates the efficiency of our system in browsing web videos.
Juan Cao 0001, Chong-Wah Ngo, Yongdong Zhang 0001, Dongming Zhang 0004
ACM Multimedia4
2010 Compressive video sensing based on user attention model
abstract
We propose a compressive video sensing scheme based on user attention model (UAM) for real video sequences acquisition. In this work, for every group of consecutive video frames, we set the first frame as reference frame and build a UAM with visual rhythm analysis (VRA) to automatically determine region-of-interest (ROI) for non-reference frames. The determined ROI usually has significant movement and attracts more attention. Each frame of the video sequence is divided into non-overlapping blocks of 16 × 16 pixel size. Compressive video sampling is conducted in a block-by-block manner on each frame through a single operator and in a whole region manner on the ROIs through a different operator. Our video reconstruction algorithm involves alternating direction l1- norm minimization algorithm (ADM) for the frame difference of non-ROI blocks and minimum total-variance (TV) method for the ROIs. Experimental results showed that our method could significantly enhance the quality of reconstructed video and reduce the errors accumulated during the reconstruction.
Jie Xu 0031, Jianwei Ma 0006, Dongming Zhang 0004, Yongdong Zhang 0001, Shouxun Lin
PCS3
2009 Logo detection based on spatial-spectral saliency and partial spatial context
abstract
Logo detection is important for brand advertising and surveillance applications. The central issues of this technology are fast localization and accurate matching. Based on key traits analysis of common logos, this paper presents a two-stage detection scheme based on spatialspectral saliency (SSS) and partial spatial context (PSC). SSS speeds up logo location and avoid the impact of cluttered background. PSC filters false matching using spatial consistency of local invariant points. The integration of SSS and PSC result in faster localization and increased accuracy. Experiments on a dataset of nearly 10,000 web images containing several popular logo types are presented. The results indicate that our method is applicable and precise for different logo detection scenarios.
Ke Gao 0012, Shouxun Lin, Yongdong Zhang 0001, Sheng Tang, Dongming Zhang 0004
ICME5
2008 Object retrieval based on spatially frequent items with informative patches
abstract
Spatial relation of local image patches plays an important role in object-based image retrieval. An approach called spatial frequent items is proposed as an extension of Bag-of-Words method by introducing spatial relations between patches. Spatial frequent items are defined as frequent pairs of adjacent local image patches in polar coordinates, and exploited using data mining. Based on these frequent configurations, we develop a method to encode patches and their spatial relations for image indexing and retrieval. Besides, to avoid the interference of background patches, informative patches are filtrated based on their local entropy and self-similarity in the preprocess stage. Experimental results demonstrate that our method can be 8.6% more effective than the state-of-art object retrieval methods.
Ke Gao 0012, Shouxun Lin, Junbo Guo, Dongming Zhang 0004, Yongdong Zhang 0001
ICME4
2008 Human attention model for semantic scene analysis in movies
abstract
In this paper, we specifically propose the Weber-Fechner Law-based human attention model for semantic scene analysis in movies. Different from traditional video processing techniques, we pay more attention on bringing in the related subjects, such as psychology, physiology and cognitive informatics, for content-based video analysis. The innovation of our work has two aspects. Firstly, we originally construct the human attention model with temporal information instructed by the Weber-Fechner Law. Secondly, motivated by cognitive informatics, we formulate the computational methodology of features in visual, audio and textual modalities in the uniform metric of information quantity. With human attention analysis and semantic scene detection, we build a system for hierarchical browse and edit with semantics annotation. Large-scale experiments demonstrate the effectiveness and generality of the proposed human attention model for movie analysis.
Anan Liu, Yongdong Zhang 0001, Yan Song 0004, Dongming Zhang 0004, Jintao Li 0001, Zhaoxuan Yang
ICME4
2007 Complexity controllable DCT for real-time H.264 encoder
Dongming Zhang 0004, Shouxun Lin, Yongdong Zhang 0001, Lejun Yu
J. Vis. Commun. Image Represent.1
2005 Fast inter frame encoding based on modes pre-decision in H.264
abstract
The new video coding standard, H.264 allows motion estimation performing on tree-structured block partitioning and multiple reference frames. This feature improves the prediction accuracy significantly, but the cost of which is the complexity and computation load of video coding increase drastically. The reference software of H.264 adopts a full search scheme and its complexity increases linearly with specified the number of reference frames and the number of modes respectively. In this paper, we propose a fast mode decision algorithm in H.264 which eliminates searching unnecessary modes. The proposed algorithm takes full use of available information obtained from the previous searching process, e.g., macroblock's best mode in searching the previous reference frames. In the algorithm implementation, multiple half-stop conditions have been set in motion estimation so as to decrease encoder's complexity. Simulation results show that this proposed algorithm can effectively reduce complexity of inter-frame encoding, and the quality degradation is tiny compared with full search scheme. Furthermore, our algorithm is adaptive to variant test sequences, no need to set specific experimental threshold.
Dongming Zhang 0004, Yanfei Shen, Shouxun Lin, Yongdong Zhang 0001
ICME1
2004 Adaptive weighted prediction in video coding
abstract
An adaptive weighted prediction algorithm is proposed to improve the coding efficiency for video scenes that contain global brightness variations caused by fade in/out or local brightness variations caused by camera flashes. A two brightness-variation parameters model is used, which represent multiplier and offset components of the brightness variation. Brightness variations at frame and macroblock level, respectively, are detected; if there are brightness variations, we use weighted prediction to compensate these brightness variations; otherwise the conventional prediction method is used. Simulation results show that the proposed algorithm can improve the coding efficiency sufficiently when the coded scenes contain brightness variations. This technology is adopted by AVS.
Yanfei Shen, Dongming Zhang 0004, Jintao Li 0001
ICME2