Wenmin Wang 0001

dblp:96/7699 · DBLP profile ↗
← Back
102ranked-venue papers
0as first author
27since 2021 · last 2026
0000-0003-2664-4413ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 74 · 9 since 2021Artificial intelligence and machine learning · 28 · 19 since 2021Systems, architecture and hardware · 4Databases, data management, data science and information retrieval · 3 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MCGS: Markov Chain Gaussian Splatting for Dynamic Scenes Reconstruction
abstract
We present MCGS (Markov Chain Gaussian Splatting), a novel approach for high-fidelity dynamic scene reconstruction via combining Markov chain and 3D Gaussian splatting. Our method addresses the critical challenge of artifact-free temporal consistency in dynamic neural rendering. By integrating a Markov chain-based deformation network with multi-head temporal attention, MCGS effectively captures motion patterns and temporal dependencies, producing more accurate and stable 3D representations over time. The key innovations include: (1) a Markov Deform Network that models state transitions while preserving temporal coherence, (2) a temporal attention mechanism that adaptively weights historical states within a sliding window, and (3) strategic noise injection during training to enhance model robustness and generalization. Experiments on representative dynamic scene datasets demonstrate that MCGS outperforms previous methods in both visual quality and temporal coherence, while maintaining competitive rendering speed and efficiency. These results suggest the practical applicability of our approach to real-world dynamic scene understanding and synthesis.
Wenmin Wang 0001, Xinxing Yu, Zhongheng Chen
AAAI2
2026 DHCF-Net: Dual heterodimensional context fusion network for medical image segmentation
Yu Wang 0087, Wenzheng Ma, Wenmin Wang 0001
Expert Syst. Appl.4
2026 Insert Anyone: High-fidelity full-body photo insertion via dual-branch adapters
Yifan Zhang 0028, Zhongliang Tang, Wenmin Wang 0001
Expert Syst. Appl.4
2025 CSRP: Modeling class spatial relation with prototype network for novel class discovery
Nannan Li 0001, Jiuqing Dong, Huiwen Guo, Wenmin Wang 0001, Chuanchuan You
Appl. Intell.5
2025 Causality thinking for large-scale long-tailed video action recognition
Zhengjin Zhang, Nannan Li 0001, Wenmin Wang 0001, Huiwen Guo, Sudan Huang
Eng. Appl. Artif. Intell.3
2025 TIA2V: Video generation conditioned on triple modalities of text-image-audio
Minglu Zhao, Wenmin Wang 0001, Rui Zhang 0108, Haomei Jia
Expert Syst. Appl.2
2025 Fine-tuning feature interaction for unsupervised domain adaptive low-light object detection
Maomao Xiong, Qunshu Zhang, Dagang Li 0001, Wenmin Wang 0001, Cong Liu 0012, Da Chen 0002, Jinglin Zhang 0004
Neurocomputing4
2025 Multimodal Sensitive Adaptive Transformer for 3D medical image segmentation
abstract
Three-dimensional medical imaging segmentation presents a significant challenge within the field, with the segmentation of multiple organs and lesions in MRI images being particularly demanding. This paper introduces an innovative approach utilizing the Multimodal Sensitive Adaptive Attention (MSAA). We refer to this new structure as the Multimodal Sensitive Adaptive Transformer Network (MSAT), which incorporates downsampling and Multimodal Sensitive Adaptive Attention into the encoding phase and integrate skip connections from different layers, outputs from Multimodal Sensitive Adaptive Attention, and upsampled feature outputs into the decoding phase. The MSAT consists of two primary components. The initial component is designed to extract a richer set of high-dimensional features through an advanced network architecture. This includes integration of different layers skip connections, outputs from the MSAA, and the results of the preceding upsampling layer. The second component features a Multimodal Sensitive Adaptive Attention block, which integrates two types of attention mechanisms: Local Sensitive Adaptive Attention (LSAA) and Spatial Sensitive Adaptive Attention (SSAA). These attention mechanisms work synergistically to blend high and low-dimensional features effectively, thereby enriching the contextual information captured by the model. Our experiments, conducted across several datasets including Synapse, BTCV, ACDC, and the BraTS 2021 dataset, demonstrate that the MSAT outperforms other existing methodologies. The MSAT shows superior segmentation capabilities for 3D multi-organ, cardiac, and brain tumor segmentation tasks.
Zhibing Wang, Wenmin Wang 0001, Nannan Li 0001, Yifan Zhang 0028, Haomei Jia, Shenyong Zhang
Image Vis. Comput.2
2025 Causalseg: investigating causality modeling for semi-supervised video object segmentation
Zhengjin Zhang, Nannan Li 0001, Wenmin Wang 0001, Huiwen Guo
Multim. Syst.3
2025 SwapInpaint2: Towards high structural consistency in identity-guided inpainting via background-preserving GAN inversion
Honglei Li 0008, Yifan Zhang 0028, Wenmin Wang 0001, Shenyong Zhang
Pattern Recognit.3
2024 Span Confusion is All You Need for Chinese Spelling Correction
Dezhi Ye, Haomei Jia, Jie Liu 0075, Haijin Liang, Jin Ma 0003, Wenmin Wang 0001
CIKM7
2024 Local Information Guided Global Integration for Infrared Small Target Detection
abstract
Infrared small targets often exhibit small scale and weak semantic features, which makes it a great challenge to their detection. To address this situation, we propose a novel network for infrared small target detection that combines local details information and global contextual information. To preserve the local and high-frequency details present in infrared images, we introduce a High-frequency Aware Encoder. To extract contextual information from multi-scale feature maps, we propose a Multi-scale Context Learning Bottleneck that incorporates contextual information repeatedly and performs cross-level fusion, which enables the recognition of small targets based on their surroundings. Finally, a lightweight Transformer Decoder is employed to restore the feature map, while placing attention on the target pixels. Experimental results on the IRSTD-1k dataset demonstrate that our method outperforms other state-of-the-art approaches.
Qianchen Mao, Jinbao Wang 0001, Wenmin Wang 0001, Bingshu Wang
ICASSP5
2024 SgLFT: Semantic-guided Late Fusion Transformer for video corpus moment retrieval
Tongbao Chen, Wenmin Wang 0001, Minglu Zhao, Ruochen Li 0001
Neurocomputing2
2024 Ultrahigh-definition video quality assessment: A new dataset and benchmark
Ruochen Li 0001, Wenmin Wang 0001, Huanqiang Hu, Tongbao Chen, Minglu Zhao
Neurocomputing2
2024 Multimodal parallel attention network for medical image segmentation
Zhibing Wang, Wenmin Wang 0001, Nannan Li 0001, Shenyong Zhang
Image Vis. Comput.2
2024 Enhanced blind face inpainting via structured mask prediction
Honglei Li 0008, Yifan Zhang 0028, Wenmin Wang 0001
Pattern Recognit. Lett.3
2024 Cross-Modality Knowledge Calibration Network for Video Corpus Moment Retrieval
abstract
Video corpus moment retrieval has become a hot topic recently, which aims to localize a consequent video moments highly relevant to the given query language description from video corpus. Existing methods towards this challenging task are suffering from the cases when the visual information and textual information in the video are very different from each other or from the cases where the redundant video content is semantically irrelevant with the query language description, which make the model confused of figuring out the truly useful within- and cross-modality information. In this article, we propose a novel Cross-Modality Knowledge Calibration Network (CKCN) to solve the issue mentioned above. Specifically, a dual calibration transformer module with improved multi-head attention is proposed to simultaneously capture the within- and cross-modality features between the visual and textual modality of the video automatically compressing the redundant information, and then a query-dependent fusion module is designed to guide feature fusion of the video's multi-modal information using the prior knowledge of query which further refine more important modality features. At last, a query-guided calibration transformer module with a well-designed learnable cell is utilized to align the query and video, forming a single joint representation for moment localization. Meanwhile, we introduce transfer learning into the task of video corpus moment retrieval (VCMR) for the first time to solve the defect of insufficient labeled data. Extensive experiments have been conducted on both the widely used TVR dataset and DiDeMo dataset which have achieved new state-of-the-art, thus verifying the effectiveness of our proposed CKCN.
Tongbao Chen, Wenmin Wang 0001, Ruochen Li 0001, Bingshu Wang
IEEE Trans. Multim.2
2024 TA2V: Text-Audio Guided Video Generation
abstract
Recent conditional and unconditional video generation tasks have been accomplished mainly based on generative adversarial network (GAN), diffusion, and autoregressive models. However, in some circumstances, using only one modality cannot provide enough semantic information. Therefore, in this paper, we propose text-audio to video (TA2V) generation, a new task for generating realistic videos from two different guided modalities, text and audio, which has not been explored much thus far. Compared to image generation, video generation is a harder task because of the complexity of processing higher-dimensional data and scarcer suitable datasets, especially for multimodal video generation. To overcome these limitations, (i) we propose the Text&Audio-guided-Video-Maker (TAgVM) model, which consists of two modules: a text-guided video generator and a text&audio-guided video modifier. (ii) This model uses a 3D VQ-GAN to compress high-dimension video data to a low-dimension discrete sequence, followed by an autoregressive model to guide text-conditional generation in the latent space. Then, we apply a text&audio-guided diffusion model to the generated video scenes, providing additional semantic details corresponding to the audio and text. (iii) We introduce a newly produced music performance video dataset, the University of Rochester Multimodal Music Performance with Video-Audio-Text (URMP-VAT), and a landscape dataset, Landscape with Video-Audio-Text (Landscape-VAT), both of which include three modalities (text, audio, and video) that are aligned with each other. The results demonstrate that our model can create videos with satisfactory quality and semantic information. The source code and datasets are available athttps://github.com/Minglu58/TA2V.
Minglu Zhao, Wenmin Wang 0001, Tongbao Chen, Rui Zhang 0108, Ruochen Li 0001
IEEE Trans. Multim.2
2024 E-detector: Asynchronous Spatio-temporal for Event-based Object Detection in Intelligent Transportation System
abstract
In intelligent transportation systems, various sensors, including radar and conventional frame cameras, are used to improve system robustness in various challenging scenarios. An event camera is a novel bio-inspired sensor that has attracted the interest of several researchers. It provides a form of neuromorphic vision to capture motion information asynchronously at high speeds. Thus, it possesses advantages for intelligent transportation systems that conventional frame cameras cannot match, such as high temporal resolution, high dynamic range, as well as sparse and minimal motion blur. Therefore, this study proposes an E-detector based on event cameras that asynchronously detect moving objects. The main innovation of our framework is that the spatiotemporal domain of the event camera can be adjusted according to different velocities and scenarios. It overcomes the inherent challenges that traditional cameras face when detecting moving objects in complex environments, such as high speed, complex lighting, and motion blur. Moreover, our approach adopts filter models and transfer learning to improve the performance of event-based object detection. Experiments have shown that our method can detect high-speed moving objects better than conventional cameras using state-of-the-art detection algorithms. Thus, our proposed approach is extremely competitive and extensible, as it can be extended to other scenarios concerning high-speed moving objects. The study findings are expected to unlock the potential of event cameras in intelligent transportation system applications.
Wenmin Wang 0001, Honglei Li 0008, Shenyong Zhang
ACM Trans. Multim. Comput. Commun. Appl.2
2023 Shadow Removal of Text Document Images Using Background Estimation and Adaptive Text Enhancement
abstract
This paper proposes a simple yet effective method to re-move shadows from text document images. It mainly includes several parts. Firstly, we propose a text elimination-based background extraction strategy to estimate shadow map. It indicates the shadow regions accurately and helps to predict global background. Secondly, a binarization-based text ex-traction algorithm is designed to obtain texts from document image. By fusing texts and global background, a preparatory shadow-free image can be obtained. Thirdly, we propose an adaptive text contrast enhancement strategy to generate shadow-free results with comfortable visual perception across shadow and non-shadow regions. Quantitative and visual results performed on open datasets indicate that the proposed method can generate clear shadow-free images from text document images. Our code will be publicly available soon.
Bingshu Wang, Jiangbin Zheng 0001, Wenmin Wang 0001
ICASSP4
2023 Bounding convolutional network for refining object locations
Shenyong Zhang, Wenmin Wang 0001, Honglei Li 0008
Neural Comput. Appl.2
2023 SaGCN: Semantic-Aware Graph Calibration Network for Temporal Sentence Grounding
abstract
Temporal sentence grounding is a challenging task that aims to localize the semantic corresponding segment from the untrimmed video according to the given query language description. Existing methods either utilize a cross-modal matching architecture following a scan-and-rank pipeline or directly predict the probabilities of being the target boundary for each frame based on the entire video content. However, such methods are weak when some of the critical semantic concepts in the query are actually relevant to multiple video segments or the desired video segment contains a query-irrelevant scene due to ignoring query semantic concepts and local and global cross-modal context. In this paper, we propose a novel semantic-aware graph calibration network (SaGCN) to address the issues mentioned above. Specifically, we first introduce a semantic-aware local relational graph module to capture the inherent relationships among the specific semantic concept relevant local contextual information for fine-grained cross-modal information interactions. Then, a semantic-aware global relational graph module is derived for global contextual information integration and achieving cross-modal alignment. Finally, an attention-based calibration module is designed for eliminating the irrelevant information maintained in the visual modality under the guidance of query description. Extensive experiments verify the effectiveness of our proposed SaGCN on two widely used datasets (Charades-STA and TACoS), in which we achieve significant and consistent improvement compared to the state-of-the-art approaches.
Tongbao Chen, Wenmin Wang 0001, Kangrui Han, Huijuan Xu 0001
IEEE Trans. Circuits Syst. Video Technol.2
2022 ANGraph: attribute-interactive neighborhood-aggregative graph representation learning
Ying Shen 0001, Huizhi Li, Dagang Li 0001, Jingwei Zheng, Wenmin Wang 0001
Neural Comput. Appl.5
2022 Fast 2-step regularization on style optimization for real face morphing
Wenmin Wang 0001, Honglei Li 0008, Roberto Bugiolacchi
Neural Networks2
2022 Affective word embedding in affective explanation generation for fine art paintings
Jianhao Yan, Wenmin Wang 0001
Pattern Recognit. Lett.2
2022 Fast transformation of discriminators into encoders using pre-trained GANs
Wenmin Wang 0001
Pattern Recognit. Lett.2
2022 SwapInpaint: Identity-Specific Face Inpainting With Identity Swapping
abstract
As face editing scenarios have become popular, the face inpainting technique has become a hot topic. Although some existing methods can inpaint faces with preserved identity information, they fail to solve a more flexible inpainting problem that fills the “holes” with identity-specific content from other faces. In this work, we propose a disentangle and subject-agnostic framework that affects both full and partial-face inpainting with the guidance of a reference face image. The framework consists of an identity encoding module, a content inference module and a generative module. The identity encoding module extracts the identity embedding from the reference image, the content inference module learns to predict the content image, and the generative module integrates the content image and the reference identity embedding to generate the identity-specific inpainted result. To minimize the structure and style gap between the incomplete image and inpainted image, we use a double attribute loss to the generative module and a postprocess of blending operation to the swapped result. We compare our method with state-of-the-art works and demonstrate that our method achieves higher identity similarity and better structural correctness.
Honglei Li 0008, Wenmin Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2020 Exploring Entity-Level Spatial Relationships for Image-Text Matching
abstract
Exploring the entity-level (i.e., objects in an image, words in a text) spatial relationship contributes to understanding multimedia content precisely. The ignorance of spatial information in previous works probably leads to misunderstandings of image contents. For instance, sentences `Boats are on the water' and `Boats are under the water' describe the same objects, but correspond to different sceneries. To this end, we utilize the relative position of objects to capture entity-level spatial relationships for image-text matching. Specifically, we fuse semantic and spatial relationships of image objects in a visual intra-modal relation module. The module performs promisingly to understand image contents and improve object representation learning. It contributes to capturing entity-level latent correspondence of image-text pairs. Then the query (text) plays a role of textual context to refine the interpretable alignments of image-text pairs in the inter-modal relation module. Our proposed method achieves state-of-the-art results on MSCOCO and Flickr30K datasets.
Yaxian Xia, Lun Huang, Wenmin Wang 0001, Xiaoyong Wei, Jie Chen 0001
ICASSP3
2020 Generating Future Frames with Mask-Guided Prediction
abstract
Current approaches in video prediction tend to hallucinate the future frames directly or learn global motion transformation from the entire scene. However, it is difficult for these methods without instance-aware mechanism to learn the underlying structures, dynamics and appearances of foreground and background elements simultaneously, especially when it comes to long-term prediction. In this paper, we propose an explicit instance-level prediction approach to tackle this issue and present a novel mask-guided dual network. We utilize instance masks to extract active objects from the videos, and design two LSTM branches to predict the future dynamics and appearances for objects and backgrounds individually. Superior than most recent skeleton-aided methods that only focus on single human object with two-stage procedure, our proposed network can predict instances from other categories and be trained end-to-end with a joint loss. We evaluate our approach on KTH, Penn Action and Running Horse datasets, and achieve promising results in both quality and quantity.
Xiongtao Chen, Wenmin Wang 0001
ICME4
2020 Low Resolution Facial Manipulation Detection
abstract
Detecting manipulated images and videos is an important aspect of digital media forensics. Due to severe discriminative information loss caused by resolution degradation, the performance of most existing methods is significantly reduced on low resolution manipulated images. To address this issue, we propose an Artifacts-Focus Super-Resolution (AFSR) module and a Two-stream Feature Extractor (TFE). The AFSR recovers facial cues and manipulation artifact details using an autoencoder learned with an artifacts focus training loss. The TFE adopts a two-stream feature extractor with key points-based fusion pooling to learn discriminative facial representations. These two complementary modules are jointly trained to recover and capture distinctive manipulation artifacts in low resolution images. Extensive experiments on two benchmarks including FaceForensics++ and DeepfakeTIMIT, evidence the favorable performance of our method against other state-of-the-art methods.
Zhongyi Ji, Wenmin Wang 0001
VCIP3
2020 A Dense-Gated U-Net for Brain Lesion Segmentation
abstract
Brain lesion segmentation plays a crucial role in diagnosis and monitoring of disease progression. DenseNets have been widely used for medical image segmentation, but much redundancy arises in dense-connected feature maps and the training process becomes harder. In this paper, we address the brain lesion segmentation task by proposing a Dense-Gated U-Net (DGNet), which is a hybrid of Dense-gated blocks and U-Net. The main contribution lies in the dense-gated blocks that explicitly model dependencies among concatenated layers and alleviate redundancy. Based on dense-gated blocks, DGNet can achieve weighted concatenation and suppress useless features. Extensive experiments on MICCAI BraTS 2018 challenge and our collected intracranial hemorrhage dataset demonstrate that our approach outperforms a powerful backbone model and other state-of-the-art methods.
Zhongyi Ji, Tong Lin 0002, Wenmin Wang 0001
VCIP4
2020 Text-to-Image Generation via Semi-Supervised Training
abstract
Synthesizing images from text is an important problem and has various applications. Most of the existing studies of text-to-image generation utilize supervised methods and rely on a fully-labeled dataset, but detailed and accurate descriptions of images are onerous to obtain. In this paper, we introduce a simple but effective semi-supervised approach that considers the feature of unlabeled images as "Pseudo Text Feature". Therefore, the unlabeled data can participate in the following training process. To achieve this, we design a Modality-invariant Semantic- consistent Module which aims to make the image feature and the text feature indistinguishable and maintain their semantic information. Extensive qualitative and quantitative experiments on MNIST and Oxford-102 flower datasets demonstrate the effectiveness of our semi-supervised method in comparison to supervised ones. We also show that the proposed method can be easily plugged into other visual generation models such as image translation and performs well.
Zhongyi Ji, Wenmin Wang 0001, Baoyang Chen
VCIP2
2020 Fast and Accurate Action Detection in Videos With Motion-Centric Attention Model
abstract
A key factor that makes action detection in videos different from general video classification is human-guided clues, especially motion signals. Since not all the pixels in a video are informative for action recognition, the irrelevant and redundant parts can lead to a lot of noise and be burdensome for both feature extraction and classifier training. This encourages the researchers to seek out the design of the attentive model that can dynamically focus computations on the key spatiotemporal volumes. In this paper, we propose a motion-centric attention model for action detection in videos which imitates the human perception of saccade and fixation procedures while detecting actions in a video. Specifically, we first present a strategy to generate motion-centric locations based on the density peak of motion signals, providing reliable candidates around which actions have high possibilities to occur. Then, we introduce an attention model that conducts the saccade and fixation procedures on these candidates to observe local spatiotemporal visual information, preserve internal comprehension, and produce the action proposals on temporal bounds. Afterward, a classifier with several variants is prepared to classify the action proposals and decide which one to fixate and generate the final predictions. We show how to efficiently train our model to produce fast and accurate action detection, by scanning only a small fraction of locations in a video. The extensive experiments on three challenging datasets show promising results with both accuracy and speed.
Jinzhuo Wang, Wenmin Wang 0001, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.2
2020 Uni-and-Bi-Directional Video Prediction via Learning Object-Centric Transformation
abstract
Video prediction, including uni-directional prediction for future frames and bi-directional prediction for in-between frames, is a challenging task and a problem worth exploring in multimedia and computer vision fields. Existing practices usually make predictions by learning global motion information from the whole given image. However, humans often focus on key objects carrying vital motion information instead of the entire frame. Besides, different objects often show different movement and deformation, even in the same scene. In this connection, we build a novel model of object-centric video prediction, in which the motion signals of key objects are particularly learned. This model can predict new frames by repeatedly transforming objects into the original input images. To focus on these objects automatically, we create an attention module with substitutable strategies. Our method requires no annotated data, and we also use adversarial training to improve sharpness of the predictions. We evaluate our model through Moving MNIST, UCF101 and Penn Action datasets and achieve competitive results in both quantity and quality, compared to existing methods. The experiments demonstrate that our uni-and-bi-directional network can well predict motions for different objects and generate plausible future and in-between frames.
Xiongtao Chen, Wenmin Wang 0001
IEEE Trans. Multim.2
2019 Image Captioning with Two Cascaded Agents
abstract
Recent neural models on image captioning usually take a encoder-decoder fashion, where the decoder predicts a single word at one step recently with the encoder providing information. The encoder is a pretrained CNN model typically. Thus the decoder, the input to it, and the output from it become the most important parts of a model. We propose a pipelined image captioning framework consisting of two cascaded agents. The former is named as "semantic adaptive agent" which generates the input to the decoder by consulting the information from the current decoding process, and the latter as "caption generating agent" which select a single word of the vocabulary as the output of the decoder by taking consideration of the input and the current states of the decoder. For the framework of two cascaded agents, we design a multi-stage training procedure to train the two agents with different objectives by fully utilizing reinforcement learning. In experiments, we conduct quantitative and qualitative analysis on MS COCO dataset and our results can significantly outperform baseline methods in terms of several evaluation metrics.
Lun Huang, Wenmin Wang 0001
ICASSP2
2019 Multi-step Self-attention Network for Cross-modal Retrieval Based on a Limited Text Space
abstract
Cross-modal retrieval has been recently proposed to find an appropriate subspace where the similarity among different modalities, such as image and text, can be directly measured. In this paper, we propose Multi-step Self-Attention Network (MSAN) to perform cross-modal retrieval in a limited text space with multiple attention steps, that can selectively attend to partial shared information at each step and aggregate useful information over multiple steps to measure the final similarity. In order to achieve better retrieval results with faster training speed, we introduce global prior knowledge as the global reference information. Extensive experiments on Flickr30K and MSCOCO, show that MSAN achieves new state-of-the-art results in accuracy for cross-modal retrieval.
Wenmin Wang 0001, Ge Li 0002
ICASSP2
2019 Attention on Attention for Image Captioning
abstract
Attention mechanisms are widely used in current encoder/decoder frameworks of image captioning, where a weighted average on encoded vectors is generated at each time step to guide the caption decoding process. However, the decoder has little idea of whether or how well the attended vector and the given attention query are related, which could make the decoder give misled results. In this paper, we propose an Attention on Attention (AoA) module, which extends the conventional attention mechanisms to determine the relevance between attention results and queries. AoA first generates an information vector and an attention gate using the attention result and the current context, then adds another attention by applying element-wise multiplication to them and finally obtains the attended information, the expected useful knowledge. We apply AoA to both the encoder and the decoder of our image captioning model, which we name as AoA Network (AoANet). Experiments show that AoANet outperforms all previously published methods and achieves a new state-of-the-art performance of 129.8 CIDEr-D score on MS COCO Karpathy offline test split and 129.6 CIDEr-D (C40) score on the official online testing server. Code is available at https://github.com/husthuaan/AoANet.
Lun Huang, Wenmin Wang 0001, Jie Chen 0001, Xiaoyong Wei
ICCV2
2019 Video Prediction with Temporal-Spatial Attention Mechanism and Deep Perceptual Similarity Branch
abstract
Video prediction is a challenging but worth exploring task in computer vision. Different from image analysis, the challenge of video analysis derives from more complicated dependencies in time as well as in space. In this paper, we propose a Temporal-Spatial Attention Mechanism (TSAM) to capture not only spatial appearance dependencies but also temporal dynamic dependencies in video sequence. The TSAM is transplantable for existing networks and allows the long-range dependency modeling for various video analysis tasks (we take video prediction for specific experiments in this paper). Besides, we propose an additional Deep Perceptual Similarity Branch (DPSB) to encourage a better approximation to the ground-truth in high-level feature space, laying the foundation for frame generation. Extensive experiments on KTH, Penn Action and UCF-101 datasets demonstrate that our model performs quite competitively across diverse natural visual scenes, even in long-term video prediction.
Wenmin Wang 0001, Xiongtao Chen, Weimian Li
ICME2
2019 Adaptively Aligned Image Captioning via Adaptive Attention Time
abstract
Recent neural models for image captioning usually employ an encoder-decoder framework with an attention mechanism. However, the attention mechanism in such a framework aligns one single (attended) image feature vector to one caption word, assuming one-to-one mapping from source image regions and target caption words, which is never possible. In this paper, we propose a novel attention model, namely Adaptive Attention Time (AAT), to align the source and the target adaptively for image captioning. AAT allows the framework to learn how many attention steps to take to output a caption word at each decoding step. With AAT, an image region can be mapped to an arbitrary number of caption words while a caption word can also attend to an arbitrary number of image regions. AAT is deterministic and differentiable, and doesn't introduce any noise to the parameter gradients. In this paper, we empirically show that AAT improves over state-of-the-art methods on the task of image captioning. Code is available at https://github.com/husthuaan/AAT.
Lun Huang, Wenmin Wang 0001, Yaxian Xia, Jie Chen 0001
NeurIPS2
2019 Predicting Diverse Future Frames With Local Transformation-Guided Masking
abstract
Video prediction is the challenging task of generating the future frames of a video given a sequence of previously observed frames. This task involves the construction of an internal representation that accurately models the frame evolutions, including contents and dynamics. Video prediction is considered difficult due to the inherent compounding of errors in recursive pixel level prediction. In this paper, we present a novel video prediction system that focuses on regions of interest (ROIs) rather than on entire frames and learns frame evolutions at the transformation level rather than at the pixel level. We provide two strategies to generate high-quality ROIs that contains potential moving visual cues. The frame evolutions are modeled with a transformation generator that produces transformers and masks simultaneously, which are then combined to generate the future frame in a transformation-guided masking procedure. Compared with recent approaches, our system is able to generate more accurate predictions by modeling the visual evolutions at the transformation level rather than at the pixel level. Focusing on ROIs avoids a heavy computational burden and enables our system to generate high-quality long-term future frames without severely amplified signal loss. Moreover, our system is able to generate diverse plausible future frames, which is important in many real-world scenarios. Furthermore, we enable our system to perform video prediction conditioned on a single frame by revising the transformation generator to produce motion-centric transformers. We test our system on four datasets with different experimental settings and demonstrate its advantages over recent methods, both quantitatively and qualitatively.
Jinzhuo Wang, Wenmin Wang 0001, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.2
2018 A Motion Aided Merge Mode For Hevc
abstract
Merge prediction is a practical inter-technique in HEVC, which can significantly improve the coding efficiency, especially for homogeneous regions in video sequences. In this paper, a motion aided merge mode (MAMM) is proposed to achieve a better trade-off between the prediction accuracy and bit rate. Different from the traditional merge mode in HEVC, MAMM is accomplished by a small motion obtained by searching in a specific search region. The search range is comprised of a number of points with high occurrence possibilities. The motion vector difference (MVD) is coded by Huffman coding in MAMM and the Huffman coding table is generated according to the statistical frequency of each possible MVD value. The proposed method is implemented on top of the HEVC reference software (HM −16.15), and experimental results show that 0.6% BD-rate reduction is achieved under Random Access (RA) configuration.
Kui Fan, Ronggang Wang, Ge Li 0002, Wenmin Wang 0001
ICASSP5
2018 Local patch encoding-based method for single image super-resolution
Yang Zhao 0002, Ronggang Wang, Wei Jia 0001, Jianchao Yang, Wenmin Wang 0001, Wen Gao 0001
Inf. Sci.5
2018 MPEG Internet Video Coding Standard and Its Performance Evaluation
abstract
MPEG has produced standards that have provided the industry with the best video compression technologies. To address diverse Internet needs, MPEG issued a Call for Proposals (CfP) for Internet video coding (IVC) in July, 2011. The anticipation is that any patent declaration associated with the baseline profile of this standard will indicate that the patent owner is prepared to grant a free of charge license to an unrestricted number of applicants worldwide. Three codecs have responded to the CfP: Web video coding (WVC), video coding for browsers (VCB), and IVC. WVC is in fact the AVC baseline, and VCB uses the same coding tools as VP8. IVC has been developed in MPEG from scratch by combining well-known existing technology elements and new coding tools with royalty-free declarations. In June 2015, the IVC project was approved as ISO/IEC 14496-33 (MPEG-4 IVC). This standard can be highly beneficial for video services in the Internet domain. This paper describes the main coding tools used in IVC, and evaluates its objective and subjective performances compared with WVC, VCB, and AVC high profile (AVC HP). The experimental results show that IVC's compression performance is approximately equal to that of the AVC HP for typical operational settings, both for streaming and low-delay applications, and is superior to WVC and VCB.
Ronggang Wang, Zhenyu Wang 0002, Kui Fan, Tiejun Huang 0001, Wenmin Wang 0001, Ge Li 0002, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.5
2018 Second- and High-Order Graph Matching for Correspondence Problems
abstract
Correspondence problems are challenging due to the complexity of real-world scenes. One way to solve this problem is to improve the graph matching (GM) process, which is flexible for matching non-rigid objects. GM can be classified into three categories that correspond with the variety of object functions: first-order, second-order, and high-order matching. Graph and hypergraph matching have been proposed separately in previous works. The former is equivalent to the second-order GM, and the latter is equivalent to high-order GM, but we use the terms second- and high-order GM to unify the terminology in this paper. Second- and high-order GM fit well with different types of problems; the key goal for these processes is to find better-optimized algorithms. Because the optimal problems for second- and high-order GM are different, we propose two novel optimized algorithms for them in this paper. (1) For the second-order GM, we first introduce a$K$-nearest-neighbor-pooling matching method that integrates feature pooling into GM and reduces the complexity. Meanwhile, we evaluate each matching candidate using discriminative weights on its$k$-nearest neighbors by taking locality as well as sparsity into consideration. (2) High-order GM introduces numerous outliers, because precision is rarely considered in related methods. Therefore, we propose a sub-pattern structure to construct a robust high-order GM method that better integrates geometric information. To narrow the search space and solve the optimization problem, a new prior strategy and a cell-algorithm-based Markov Chain Monte Carlo framework are proposed. In addition, experiments demonstrate the robustness and improvements of these algorithms with respect to matching accuracy compared with other state-of-the-art algorithms.
Wenmin Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2018 Multiscale Deep Alternative Neural Network for Large-Scale Video Classification
abstract
With the rapid increase in the amount of multimedia data, video classification has become a demanding and challenging research topic. Compared with image classification, video classification requires mapping a video that contains hundreds of frames to semantic tags, which poses many challenges to the direct use of advanced models originally designed for image-oriented tasks. On the other hand, continuous frames in a video also give us more visual clues that we can leverage to achieve better classification. One of the most important clues is the context in the spatiotemporal domain. In this paper, we introduce the multiscale deep alternative neural network (DANN), a novel architecture combining the strengths of both convolutional neural network and recurrent neural networks to achieve a deep network that can collect rich context hierarchies for video classification. In particular, the DANN is stacked with alternative layers, each of which consists of a volumetric convolutional layer followed by a recurrent layer. The former acts as a local feature learner, whereas the latter is used to collect contexts. Compared with popular deep feed-forward neural networks, the DANN learns local features and their contexts from the very beginning. This setting enables preserving context evolutions, which we show to be essential for improving the accuracy of video classification. To release the full potential of the DANN, we develop a deeper version with stochastic-layer skip-connections and construct a multiscale DANN to incorporate contexts at different scales. We show how to apply the multiscale DANN for video classification with carefully designed configurations in terms of both input-output settings and training-testing methods. The DANN is shown to be robust to not only human-centric videos, but also natural videos. As there are few large-scale natural disaster video datasets, we construct a new large-scale one and make it publicly available. Experiments on four datasets show the effectiveness of our method for both human actions and natural events.
Jinzhuo Wang, Wenmin Wang 0001, Wen Gao 0001
IEEE Trans. Multim.2
2017 Beyond Monte Carlo Tree Search: Playing Go with Deep Alternative Neural Network and Long-Term Evaluation
abstract
Monte Carlo tree search (MCTS) is extremely popular in computer Go which determines each action by enormous simulations in a broad and deep search tree. However, human experts select most actions by pattern analysis and careful evaluation rather than brute search of millions of future interactions. In this paper, we propose a computer Go system that follows experts’ way of thinking and playing. Our system consists of two parts. The first part is a novel deep alternative neural network (DANN) used to generate candidates of next move. Compared with existing deep convolutional neural network (DCNN), DANN inserts recurrent layer after each convolutional layer and stacks them in an alternative manner. We show such setting can preserve more contexts of local features and its evolutions which are beneficial for move prediction. The second part is a long-term evaluation (LTE) module used to provide a reliable evaluation of candidates rather than a single probability from move predictor. This is consistent with human experts’ nature of playing since they can foresee tens of steps to give an accurate estimation of candidates. In our system, for each candidate, LTE calculates a cumulative reward after several future interactions when local variations are settled. Combining criteria from the two parts, our system determines the optimal choice of next move. For more comprehensive experiments, we introduce a new professional Go dataset (PGD), consisting of $253,233$ professional records. Experiments on GoGoD and PGD datasets show the DANN can substantially improve performance of move prediction over pure DCNN. When combining LTE, our system outperforms most relevant approaches and open engines based on MCTS.
Jinzhuo Wang, Wenmin Wang 0001, Ronggang Wang, Wen Gao 0001
AAAI2
2017 Attention-Based Two-Phase Model for Video Action Detection
Xiongtao Chen, Wenmin Wang 0001, Weimian Li, Jinzhuo Wang
CAIP (2)2
2017 A Violence Detection Approach Based on Spatio-temporal Hypergraph Transition
Jingjia Huang, Ge Li 0002, Nannan Li 0001, Ronggang Wang, Wenmin Wang 0001
CAIP (2)5
2017 Progressive Probabilistic Graph Matching with Local Consistency Regularization
Wenmin Wang 0001
CAIP (2)2
2017 A New Image Contrast Enhancement Algorithm Using Exposure Fusion Framework
Zhenqiang Ying, Ge Li 0002, Yurui Ren, Ronggang Wang, Wenmin Wang 0001
CAIP (2)5
2017 Learning a Limited Text Space for Cross-Media Retrieval
Wenmin Wang 0001, Mengdi Fan
CAIP (1)2
2017 A Multilayer Backpropagation Saliency Detection Algorithm Based on Depth Mining
Chunbiao Zhu, Ge Li 0002, Wenmin Wang 0001, Ronggang Wang
CAIP (2)4
2017 Cross-modality matching based on Fisher Vector with neural word embeddings and deep image features
abstract
Cross-modal retrieval, which aims to solve the problem that the query and the retrieved results are from different modality, becomes more and more essential with the development of the Internet. In this paper, we mainly focus on the exploration of high-level semantic representation of image and text for cross-modal matching. Deep convolutional image features and Fisher Vector with neural word embeddings are utilized as visual and textual features respectively. To further investigate the correlation among heterogeneous multimodal characteristics, we use multiclass logistic classifier for semantic matching across modalities. Experiments on Wikipedia and Pascal Sentence dataset demonstrate the robustness and effectiveness for both Img2Text and Text2Img retrieval tasks.
Wenmin Wang 0001, Mengdi Fan, Ronggang Wang
ICASSP2
2017 A joint model for action localization and classification in untrimmed video with visual attention
abstract
In this paper, we introduce a joint model that learns to directly localize the temporal bounds of actions in untrimmed videos as well as precisely classify what actions occur. Most existing approaches tend to scan the whole video to generate action instances, which are really inefficient. Instead, inspired by human perception, our model is formulated based on a recurrent neural network to observe different locations within a video over time. And, it is capable of producing temporal localizations by only observing a fixed number of fragments, and the amount of computation it performs is independent of input video size. The decision policy for determining where to look next is learned by REINFORCE which is powerful in non-differentiable settings. In addition, different from relevant ways, our model runs localization and classification serially, and possesses a strategy for extracting appropriate features to classify. We evaluate our model on ActivityNet dataset, and it greatly outperforms the baseline. Moreover, compared with a recent approach, we show that our serial design can bring about 9% increase in detection performance.
Weimian Li, Wenmin Wang 0001, Xiongtao Chen, Jinzhuo Wang, Ge Li 0002
ICME2
2017 Better deep visual attention with reinforcement learning in action recognition
abstract
Deep visual attention in computer vision has attracted much attention over the past years, which achieves great contributions especially in image classification, image caption and action recognition. However, due to taking BP training wholly or partially, they can not show the true power of attention in computational efficiency and focusing accuracy. Our intuition is that attention mechanism should be similar to the process in which human draw attention and select the next location to focus, by observing, analyzing and jumping instead of existing describing continuous features. Based on this insight, we formulate our model as a recurrent neural network-based agent that chooses attention region by reinforcement learning at each timestep. In experiments, our model explicitly outperforms baselines not only in focusing and recognizing accuracy, but also consumes much less computational resources, which can be honored as better deep visual attention.
Wenmin Wang 0001, Jingzhuo Wang, Yaohua Bu
ISCAS2
2017 Learning Object-Centric Transformation for Video Prediction
abstract
Future frame prediction for video sequences is a challenging task and worth exploring problem in computer vision. Existing methods often learn motion information for the entire image to predict next frames. However, different objects in the same scene often move and deform in different ways intuitively. Considering the human visual system, one often pays attention to the key objects that contain crucial motion signals, rather than compress an entire image into a static representation. Motivated by this property of human perception, in this work, we develop a novel object-centric video prediction model that learns local motion transformation dynamically for key object regions with visual attention. By transforming objects iteratively to the original input frames, next frame can be produced. Specifically, we design an attention module with replaceable strategies to attend to objects in video frames automatically. Our method does not require any annotated data during training procedure. To produce sharp predictions, adversarial training is adopted in our work. We evaluate our model on the Moving MNIST and UCF101 datasets and report competitive results, compared to prior methods. The generated frames demonstrate that our model can characterize motion for different objects and produce plausible future frames.
Xiongtao Chen, Wenmin Wang 0001, Jinzhuo Wang, Weimian Li
ACM Multimedia2
2017 Cross-media Retrieval by Learning Rich Semantic Embeddings of Multimedia
abstract
Cross-media retrieval aims at seeking the semantic association between different media types. Most existing methods paid much attention on learning mapping functions or finding the optimal spaces, but neglected how people accurately cognize images and texts. This paper proposes a brain inspired cross-media retrieval framework to learn rich semantic embeddings of multimedia. Different from directly using off-the-shelf image features, we combine the visual and descriptive senses for an image from the view of human perception via a joint model, called multi-sensory fusion network (MSFN). A topic model based TextNet maps texts into the same semantic space as images according to their shared ground truth labels. Moreover, in order to overcome the limitations of insufficient data for training neural networks and less complexity in text form, we introduce a large-scale image-text dataset, called Britannica dataset. Extensive experiments show the effectiveness of our framework for different lengths of texts on three benchmark datasets as well as Britannica dataset. Most of all, we report the best known average results of Img2Text and Text2Img compared with several state-of-the-art methods.
Mengdi Fan, Wenmin Wang 0001, Peilei Dong, Ronggang Wang, Ge Li 0002
ACM Multimedia2
2017 Long-term video interpolation with bidirectional predictive network
abstract
This paper considers the challenging task of long-term video interpolation. Unlike most existing methods that only generate few intermediate frames between existing adjacent ones, we attempt to speculate or imagine the procedure of an episode and further generate multiple frames between two non-consecutive frames in videos. In this paper, we present a novel deep architecture called bidirectional predictive network (BiPN) that predicts intermediate frames from two opposite directions. The bidirectional architecture allows the model to learn scene transformation with time as well as generate longer video sequences. Besides, we make attempts to extend our model to predict multiple possible procedures by sampling different noise vectors. A joint loss composed of clues in image and feature spaces and adversarial loss is designed to train our model. We demonstrate the advantages of BiPN on two benchmarks Moving 2D Shapes and UCF101 and report competitive results to recent approaches.
Xiongtao Chen, Wenmin Wang 0001, Jinzhuo Wang
VCIP2
2017 Mask-streaming CNN for pedestrian detection
abstract
Inspired by humans recognition of pedestrians, we propose a mask-streaming convolutional neural network (CNN) for pedestrian detection. The mask stream consists of one original region proposal and six masked regions. These masked regions, which aim at highlighting significantly discriminative semantic characteristic of head, body parts and contextual information, are generated through wiping off a part of the original region by `masks'. We feed the mask stream into the network to generate features of each masked region, then a concatenate layer integrates these features as the final representation of pedestrians. We evaluate the proposed model on the challenging Caltech and Eth datasets. Our approach performs better than the baseline, and achieves competitive performance compared to numerous pedestrian detection methods.
Peilei Dong, Wenmin Wang 0001, Mengdi Fan, Ronggang Wang, Ge Li 0002
VCIP2
2017 Learning multi-view embedding in joint space for bidirectional image-text retrieval
abstract
In this paper, we propose a framework for learning a joint embedding space for bidirectional image-text retrieval task, which fuses embedding spaces in multi-views. We have implemented two views currently, one is a frame-sentence view and the other is a region-phrase view. In the frame-sentence view, we project each frame of the images and each sentence of the texts into a holistic-level subspace to explore the correlation between them. In the region-phrase view, we extract each region of the frames and each phrase of the sentences and map them into a local-level subspace. We separately mine the semantic correlations in the two views, then merge them by a multi-view fusion ranking method. In each view, to embed heterogeneous data into a common space, we adopt the two-branch neural network to transform the data. Extensive experiments show that our multi-view joint space can preserve more accurate semantic correlations between images and texts in different granularities and can significantly improve the performance on the image-text retrieval task. Our method achieves better results than the state of the art on the Pascal1K and Flickr8K image-sentence datasets.
Lu Ran, Wenmin Wang 0001
VCIP2
2017 Adaptive difference modelling for background subtraction
abstract
Background subtraction plays a very important role in video analysis, especially in surveillance systems. While being straightforward, the performance based on frame differencing is unsatisfied due to its sensitiveness to issues such as camera shake and swinging objects. To address its limitations, in this paper we propose a complete adaptive difference modelling framework. First, we introduce two difference discriminators to model the evolution process of pixels. Second, we use Gaussian Mixture Models to adaptively learn the difference threshold to distinguish foreground from background. Third, three heuristics are employed to further improve the model adaptability. Experiments on real-world videos of the Background Models Challenge (BMC) demonstrate that our method performs better on global quality metric (FSD) than other state-of-the-art methods.
Xianghao Zang, Ge Li 0002, Jun Yang 0033, Wenmin Wang 0001
VCIP4
2017 Deep discriminative network with inception module for person re-identification
abstract
Convolutional neural networks have been verified to be exceptionally powerful on extracting semantic features, which contribute to a great progress in computer vision. However, focusing too much on the superiority, researchers seem to pay less attention to exploring CNNs' potential in other aspects, e.g. the ability to discriminate the difference. In this work we try to dig into the discriminative power of CNNs and introduce a deep discriminative network with inception module (DDN-IM) for person re-identification. Without individual feature extraction as prerequisite, input images from two different non-overlapping camera views are concatenated in depth at the beginning, followed by series of convolutional and nonlinear operations, etc. to predict their similarity. In addition, inception module is embedded in our network to boost the performance. We validate our proposal on several person re-identification datasets, CUHK01, QMUL GRID and PRID2011 included. We obtain competitive or superior performance compared to the state-of-the-art methods.
Wenmin Wang 0001, Jinzhuo Wang
VCIP2
2017 Iterative projection reconstruction for fast and efficient image upsampling
Yang Zhao 0002, Ronggang Wang, Wei Jia 0001, Wenmin Wang 0001, Wen Gao 0001
Neurocomputing4
2017 Color Image-Guided Boundary-Inconsistent Region Refinement for Stereo Matching
abstract
Cost computation, cost aggregation, disparity optimization, and disparity refinement are the four main steps for stereo matching. While the first three steps have been widely investigated, few efforts have been taken on disparity refinement. In this paper, we propose a color image-guided disparity refinement method to further remove the boundary-inconsistent regions on disparity map. First, the origins of boundary-inconsistent regions are analyzed. Then, these regions are detected with the proposed hybrid-superpixel-based strategy. Finally, the detected boundary-inconsistent regions are refined by a modified weighted median filtering method. Experimental results on various stereo matching conditions validate the effectiveness of the proposed method. Furthermore, depth maps obtained by active depth acquisition devices like Kinect can also be well refined with our proposed method.
Jianbo Jiao, Ronggang Wang, Wenmin Wang 0001, Dagang Li 0001, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2017 Accelerating Image-Domain-Warping Virtual View Synthesis on GPGPU
abstract
The image-domain-warping (IDW) method can effectively create high-quality virtual views. However, the IDW algorithm is very complex, and the software implementation for this method is far from real-time. In this paper, we propose an IDW-based view synthesis acceleration method on general-purpose computing on a graphics processing unit (GPGPU). Our method makes two main contributions. First, at the algorithm level, we employ FAST for sparse disparity estimation and adopt the successive over-relaxation iterative method to calculate warps. Second, at the platform level, two computation-intensive modules (data extraction and view synthesis) in IDW are offloaded to GPU using efficient data-level parallelism strategies. Experimental results demonstrate that our proposed acceleration method can speed up the original IDW algorithm by more than 110x, and HD stereo three-dimensional video can be converted to 8-view 4 K video (each view has an approximate 720P resolution) in real-time on a hybrid CPU + GPU (NVIDIA GTX980) platform.
Ronggang Wang, Jiajia Luo, Xiubao Jiang, Zhenyu Wang 0002, Wenmin Wang 0001, Ge Li 0002, Wen Gao 0001
IEEE Trans. Multim.5
2016 Tube ConvNets: Better exploiting motion for action recognition
abstract
Motion information is a key factor for action recognition and has been eagerly pursued for decades. How to effectively learn motion features in Convolutional Networks (ConvNets) remains an open issue. Prevalent ConvNets often take several full frames of video as input at a time, which can be a heavy burden for network training. In this paper, we introduce a novel framework called Tube ConvNets, by substituting action tubes for full frames to reduce this burden. Tube ConvNets focus on the regions of interest (ROI) where key motions occur, and thus eliminate the distraction of irrelevant objects. Each action tube is a fraction of spatiotemporal volumes, generated by the techniques of object detection and clustering algorithm. We demonstrate the effectiveness of Tube ConvNets for action classification on UCF-101 dataset, and illustrate its potential to support fine-grained localization on UCF-Sports dataset. Source code is available at https://github.com/wangjinzhuo/tubecnn.
Zhihao Li 0002, Wenmin Wang 0001, Nannan Li 0001, Jinzhuo Wang
ICIP2
2016 An Empirical Study of Deformable Part Model with fast feature pyramid
abstract
The performance of an object detection system relies heavily on two components: an object model to capture the compositional relationship among the object body and its parts, and a feature representation to describe object appearance. In this work, we present an empirical study of combining two state-of-the-art such components: Deformable Part Model (DPM), a proven effective and flexible part-based object model which originally adopts Histogram of Oriented Gradients (HOG) feature, and Aggregated Channel Features (ACF), a unified feature representation framework with fast pyramid calculation which is originally used in a rigid template matching scheme. DPM is known to work but slow, at the same time ACF has previously been shown to yield a massive speedup with only a minor loss in accuracy compared to competing features including HOG. By combining the two, our hope is to achieve the best of both worlds: the object structure representation power of DPM and the computational efficiency of ACF. Our experiments show that while ACF with heterogeneous feature channels could improve the accuracy of DPM, the run time benefit introduced by fast pyramid approximation is rather limited.
Jun Yang 0033, Ge Li 0002, Wenmin Wang 0001, Ronggang Wang
ICPR3
2016 An MCMC-based prior sub-hypergraph matching in presence of outliers
abstract
Correspondence problems are very challenging due to the complexity of real-world scenes. Some hypergraph matching methods have been proposed for improving the recall of the solution, but the numerous outliers are brought since the precision is rarely considered. To solve this issue, we propose a sub-hypergraph matching method, which is robust with better integration of geometric information and reduces the difficulty of NP-hard problem happened in hypergraphs. To narrow the search space and solve the optimization problem, a new prior strategy and cell-algorithm in Markov Chain Monte Carlo (MCMC) framework is proposed on sub-hypergraph matching. The experiments show that our proposed method significantly outperforms other state-of-the-art algorithms.
Wenmin Wang 0001
ICPR2
2016 Regional Subspace Projection Coding for Image Retrieval
abstract
For image retrieval task, hamming embedding, being proved to be one of the state-of-the-art methods, has been prevalently utilised. The basic idea is to project local features into orthogonal space randomly, in which the binary signature is generated based on a single partition of feature space. However, the binary signature generation process is coarse and heuristic. On the one hand, the same projection is carried out for all visual word space without consideration of difference among subspaces. On the other hand, the projection matrix is generated randomly regardless of the distribution of feature data. Therefore, the performance of hamming embedding is limited and far from the optimal. In this paper, we firstly analyse the limitation of hamming em- bedding and compare different orthogonal projection methods. Then we propose a regional subspace projection coding method that is based on the distribution of local features assigned to each visual word. Finally, our experiments on two benchmark datasets demonstrate that our proposed method outperforms current state-of-the-art methods.
Mingmin Zhen, Wenmin Wang 0001, Ronggang Wang
ICMR2
2016 A Novel Shadow-Free Feature Extractor for Real-Time Road Detection
abstract
Road detection is one of the most important research areas in driver assistance and automated driving field. However, the performance of existing methods is still unsatisfactory, especially in severe shadow conditions. To overcome those difficulties, first we propose a novel shadow-free feature extractor based on the color distribution of road surface pixels. Then we present a road detection framework based on the extractor, whose performance is more accurate and robust than that of existing extractors. Also, the proposed framework has much low-complexity, which is suitable for usage in practical systems.
Zhenqiang Ying, Ge Li 0002, Xianghao Zang, Ronggang Wang, Wenmin Wang 0001
ACM Multimedia5
2016 Deep Alternative Neural Network: Exploring Contexts as Early as Possible for Action Recognition
abstract
Contexts are crucial for action recognition in video. Current methods often mine contexts after extracting hierarchical local features and focus on their high-order encodings. This paper instead explores contexts as early as possible and leverages their evolutions for action recognition. In particular, we introduce a novel architecture called deep alternative neural network (DANN) stacking alternative layers. Each alternative layer consists of a volumetric convolutional layer followed by a recurrent layer. The former acts as local feature learner while the latter is used to collect contexts. Compared with feed-forward neural networks, DANN learns contexts of local features from the very beginning. This setting helps to preserve hierarchical context evolutions which we show are essential to recognize similar actions. Besides, we present an adaptive method to determine the temporal size for network input based on optical flow energy, and develop a volumetric pyramid pooling layer to deal with input clips of arbitrary sizes. We demonstrate the advantages of DANN on two benchmarks HMDB51 and UCF101 and report competitive or superior results to the state-of-the-art.
Jinzhuo Wang, Wenmin Wang 0001, Xiongtao Chen, Ronggang Wang, Wen Gao 0001
NIPS2
2016 An effective post quantization rate estimation for HEVC intra encoder
abstract
In high efficiency video coding (HEVC), the encoder employs a flexible quad-tree coding structure as well as a large number of prediction modes. For each size of coding unit (CU), transform unit (TU) and each prediction mode, the rate distortion optimization (RDO) is performed to select the best CU, TU and the best prediction mode. Although better coding efficiency is achieved, the computational complexity increases dramatically. In order to reduce the burden of RDO in the HEVC intra encoder, in this paper, we propose an effective approach, which is based on the generalized Gaussian distribution (GGD) model, to estimate the block level bit-rate. The weaknesses of the conventional GGD model are analyzed and relevant improvements are exploited. Our experiments show that, compared with the original RDO procedure in HM16.0, the proposed algorithm reduces RDO time by 37.7% with 0.64% BD-rate loss.
Hongbin Cao, Ronggang Wang, Zhenyu Wang 0002, Ge Li 0002, Wenmin Wang 0001
VCIP5
2016 Better region proposals for pedestrian detection with R-CNN
abstract
Region-based convolutional neural networks (R-CNNs) have achieved great success in object detection recently. These deep models depend on region proposal algorithms to hypothesize object locations. In this paper, we combine a special region proposal algorithm with R-CNN, and apply it to pedestrian detection. The special algorithm is used to generate region proposals only for pedestrian class. It is different from the popular region proposal algorithm selective search that detects generic object locations. The experimental results prove that region proposals generated by our method are more applicable than selective search for pedestrian detection. Our method performs faster training and testing than the deep model based on, and it achieves competitive performance compared to the state-of-the-arts in pedestrian detection.
Peilei Dong, Wenmin Wang 0001
VCIP2
2016 A simple but efficient way to combine VLAD with locality-constrained linear coding
abstract
The VLAD (vector of locally aggregated descriptors) representation, derived from BoF and Fisher kernel, has shown its efficiency in the field of image search. However, assigning local descriptors to a codeword is a hard voting process, which does not consider the uncertainty and the plausibility for single codeword. In this paper, we propose an approach to combine VLAD with locality-constrained linear coding, as opposed to the original one, considering several nearest neighbors when assigning local descriptors and computing weights. In order to evaluate our proposed method, experiments are conducted on several image classification benchmarks, using VLAD for comparison. The experimental results show that our method stably outperforms VLAD in terms of classification accuracy, while producing feature representation of the same dimension without much additional computational cost.
Zhenglin Tan, Wenmin Wang 0001, Yifeng Jiang 0004, Ronggang Wang
VCIP2
2016 A new video denoising method using texture metric and adaptive structure variance
abstract
In this paper, an innovated method is proposed for video denoising. The method consists of two major procedures. First, a new adaptive superpixel video texture metric is proposed. This video texture metric is calculated, which relates to different parts of a video stream. Then a new adaptive structure variance is estimated by adopting fine and coarse structures. Finally, a noise filter based on the estimated weights of different structures in a video stream is applied. By comparison, the proposed method outperforms traditional state-of-art methods, especially in block artifacts reduction.
Ge Li 0002, Wenmin Wang 0001, Ronggang Wang
VCIP4
2016 An advanced local offset matching strategy for object proposal matching
abstract
Image correspondence problem is very challenging due to the complexity of real-world scenes, especially in the presence of deformation, outliers, and other intra-class variations. Semantic flow methods are used for finding image correspondences, but the prevalent ones are mainly depended on pixel-level or regularly local-region sampled operation that are easily confused by image-specific details and element-specific scenes to individual objects, e.g. background, clutter, and they rarely care about the relationships between proposals. Therefore, we propose an advanced local offset matching strategy for object proposals, which is in term of naive Bayesian model, considering both appearance and the local spatial pattern with preprocessing. It not only builds reliable region relationships using object proposals, but also explores consistency between appearance and spatial pattern among proposals. In addition, two evaluation metrics, seven features are applied to experiments. The comparison and analysis, along with adequate experiments on standard benchmark, present that our proposed method effectively outperforms other matching strategies in diverse settings.
Wenmin Wang 0001
VCIP2
2016 Local Quantization Code histogram for texture classification
Yang Zhao 0002, Ronggang Wang, Wenmin Wang 0001, Wen Gao 0001
Neurocomputing3
2016 Spatially variant defocus blur map estimation and deblurring from a single image
Xinxin Zhang 0004, Ronggang Wang, Xiubao Jiang, Wenmin Wang 0001, Wen Gao 0001
J. Vis. Commun. Image Represent.4
2016 Multilevel Modified Finite Radon Transform Network for Image Upsampling
abstract
A local line-like feature is the most important discriminate information in the image upsampling scenario. In recent example-based upsampling methods, grayscale and gradient features are often adopted to describe the local patches, but these simple features cannot accurately characterize complex patches. In this paper, we present a feature representation of local edges by means of a multilevel filtering network, namely, multilevel modified finite Radon transform network (MMFRTN). In the proposed MMFRTN, the MFRT is utilized in the filtering layer to extract the local line-like feature; the nonlinear layer is set to be a simple local binary process; for the feature-pooling layer, we concatenate the mapped patches as the feature of local patch. Then, we propose a new example-based upsampling method by means of the MMFRTN feature. Experimental results demonstrate the effectiveness of the proposed method over some state-of-the-art methods.
Yang Zhao 0002, Ronggang Wang, Wenmin Wang 0001, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2016 CSPS: An Adaptive Pooling Method for Image Classification
abstract
This paper proposes an adaptive approach to learn class-specific pooling shapes (CSPS) for image classification. Prevalent methods for spatial pooling are often conducted on predefined grids of images, which is an ad-hoc method and, thus, lacks generalization power across different categories. In contrast, our CSPS is designed in a data-driven fashion by generating plenty of candidates and selecting the optimal subset for each class. Specifically, we establish an overcomplete spatial shape set that preserves as many geometric patterns as possible. Then, the class-specific subset is selected by training a linear classifier with structured sparsity constraints and color distribution cues. To address the high computational cost and the risk of overfitting due to the overcomplete scheme, the image representations for CSPS are first compressed according to dictionary sensitivity and shape importance. These representations are finally fed to SVMs for the classification task. We demonstrate that CSPS can learn compact yet discriminative geometric information for different classes that carries more semantic meaning than other methods. Experimental results on four datasets demonstrate the benefits of the proposed method compared with other pooling schemes and illustrate its effectiveness on both object and scene images.
Jinzhuo Wang, Wenmin Wang 0001, Ronggang Wang, Wen Gao 0001
IEEE Trans. Multim.2
2015 Learning discriminative visual dictionary for natural scene categorization
abstract
Many successful systems for scene recognition transform low-level descriptors into complex representations. This process consists of the two steps: 1) feature coding, which performs a pointwise transformation of the descriptors into a representation adapted to the task, and 2) image pooling, which summarizes the coded features. Even though these two steps have been paid so much attention, but there are still some problems in combining scene semantic with local features. The goal of this paper is threefold: to address the problem by modifying the traditional bag-of-features (BoF) framework; to show how to achieve the best performance by learning a semi-supervised discriminative dictionary; and to provide theoretical and empirical insight into the remarkable performance. By teasing apart components shared by modern scene categorization pipeline, our approach aims to facilitate the design of better scene recognition architectures.
Wenmin Wang 0001, Ronggang Wang
ICASSP2
2015 A novel integer-pixel motion estimation algorithm based on quadratic prediction
abstract
This paper presents a fast integer-pixel motion estimation (ME) algorithm for High Efficiency Video Coding (HEVC), which provides a strategy to speed up the search process significantly, while yielding the same quality performance as the Test Zone Search (TZSearch) scheme in HEVC Test Model (HM). Compared with the H.264/AVC, the ME process employs a more complex hybrid coding architecture and a larger size search window in HEVC, leading to great computational complexity. We utilize limited pixels of certain position in the current search range to build a quadratic model, and then shrink the search range repeatedly by analyzing the sum of absolute difference (SAD) distribution until the best motion vector (MV) is obtained. The proposed algorithm can be applied to various encoding conditions. Experimental results show that our method can save 56% of computations compared with the TZSearch scheme, with negligible decrease of coding quality.
Longfei Gao, Shengfu Dong, Wenmin Wang 0001, Ronggang Wang, Wen Gao 0001
ICIP3
2015 A low-light image enhancement method for both denoising and contrast enlarging
abstract
In this paper, a novel united low-light image enhancement framework for both contrast enhancement and denoising is proposed. First, the low-light image is segmented into superpixels, and the ratio between the local standard deviation and the local gradients is utilized to estimate the noise-texture level of each superpixel. Then the image is inverted to be processed in the following steps. Based on the noise-texture level, a smooth base layer is adaptively extracted by the BM3D filter, and another detail layer is extracted by the first order differential of the inverted image and smoothed with the structural filter. These two layers are adaptively combined to get a noise-free and detail-preserved image. At last, an adaptive enhancement parameter is adopt into the dark channel prior dehazing process to enlarge contrast and prevent over/under enhancement. Experimental results demonstrate that our proposed method outperforms traditional methods in both subjective and objective assessments.
Lin Li 0062, Ronggang Wang, Wenmin Wang 0001, Wen Gao 0001
ICIP3
2015 Image classification using RBM to encode local descriptors with group sparse learning
abstract
This paper proposes to employ deep learning model to encode local descriptors for image classification. Previous works using deep architectures to obtain higher representations are often operated from pixel level, which lack the power to be generalized to large-size and complex images due to computational burdens and internal essence capture. Our method slips the leash of this limitation by starting from local descriptors to leverage more semantical inputs. We investigate to use two layers of Restricted Boltzmann Machines (RBMs) to encode different local descriptors with a novel group sparse learning (GSL) inspired by the recent success of sparse coding. Besides, unlike the most existing pure unsupervised feature coding strategies, we use another RBM corresponding to semantic labels to perform supervised fine-tuning which makes our model more suitable for classification task. Experimental results on Caltech-256 and Indoor-67 datasets demonstrate the effectiveness of our method.
Jinzhuo Wang, Wenmin Wang 0001, Ronggang Wang, Wen Gao 0001
ICIP2
2015 A compact shot representation for video semantic indexing
abstract
This paper presents a compact shot representation for video semantic indexing (SIN). The proposed representation consists of visual cues from only two frames, i.e., key frame (KF) and difference frame (DF), which are both constructed with spatial pyramid. The KF describes static information while the generated DF captures non-static information. Each region of DF is derived from the same location in a selected frame, which has the most salient difference compared with the key frame in that region. We introduce a variation of DF to further enhance our model. Experimental results on TRECVID SIN demonstrate that our method obtains better accuracy than the state-of-the-art, while requiring less storage space and consuming time.
Jinzhuo Wang, Wenmin Wang 0001, Ronggang Wang, Wen Gao 0001
ICIP2
2015 Image deblurring using robust sparsity priors
abstract
In this paper, we propose a robust method to remove motion blur from a single photograph. We find that an inaccurate kernel and an unreliable final latent image reconstruction method are two main factors leading to low-quality restored images. To improve image quality, we do the following technical contributions. For robust blur kernel estimation, first, an edge mask and a smooth constraint are used to provide reliable intermediate latent images for salient structure extraction; second, we adopt an effective salient structure selection method to remove detrimental edges for kernel estimation; third, we use a gradient sparsity prior to remove kernel noise and ensure the continuity of blur kernels. For final latent image reconstruction, we combine the merits of both the TV-l2model and the hyper-Laplacian model to preserve tiny details and eliminate noise. Experimental results on synthetically blurred images and real photographs demonstrate that the proposed algorithm performs better than state-of-the-art approaches.
Xinxin Zhang 0004, Ronggang Wang, Yonghong Tian 0001, Wenmin Wang 0001, Wen Gao 0001
ICIP4
2015 Accelerating CDVS extraction on mobile platform
abstract
The extraction of MPEG-7 Compact Descriptors for Visual Search (CDVS) on most popular mobile devices is slow, which is attributed to the complexity of the algorithm and the limited computing power of the mobile platform. A feasible and straightforward way to accelerate the extraction is to excavate the potential computing power of modern mobile microprocessor. In this paper, we implement a NEON SIMD based data level parallelism and Pthread based multi-thread parallelism scheme to accelerate the CDVS extracting process on a multi-core ARM processor. Experimental results show a speed-up of 3.5x on Gaussian convolution, and 2.3x on the whole extracting process with the proposed method.
Ronggang Wang, Qiusi Wang, Wenmin Wang 0001
ICIP4
2015 Improved cluster center adaption for image classification
abstract
The feature coding algorithm, “Vector of Locally Aggregated Descriptors (VLAD)”, can be used effectively for large scale object instance retrieval. Despite its effectiveness and excellent performance, the existence of ambiguous cluster centers can reduce the performance. Though an idea to this problem has been proposed, it is not practical in fact. In this paper, we analyze possible situations that cause effect on the results and propose a novel approach to improve the VLAD method. The proposed method mainly focuses on the similarity measure between each two images. For each two images, we adapt the original cluster center to VLAD vectors. As we illustrate, our method has promising results with small vocabulary size on both datasets of 15 Scenes and VOC2007.
Mingmin Zhen, Wenmin Wang 0001, Ronggang Wang
ICIP2
2015 Learning class-specific pooling shapes for image classification
abstract
Spatial pyramid (SP) representation is an extension of bag-of-feature model which embeds spatial layout information of local features by pooling feature codes over pre-defined spatial shapes. However, the uniform style of spatial pooling shapes used in standard SP is an ad-hoc manner without theoretical motivation, thus lacking the generalization power to adapt to different distribution of geometric properties across image classes. In this paper, we propose a data-driven approach to adaptively learn class-specific pooling shapes (CSPS). Specifically, we first establish an over-complete set of spatial shapes providing candidates with more flexible geometric patterns. Then the optimal subset for each class is selected by training a linear classifier with structured sparsity constraint and color distribution cues. To further enhance the robust of our model, the representations over CSPS are compressed according to the shape importance and finally fed to SVM with a multi-shape matching kernel for classification task. Experimental results on three challenging datasets (Caltech-256, Scene-15 and Indoor-67) demonstrate the effectiveness of the proposed method on both object and scene images.
Jinzhuo Wang, Wenmin Wang 0001, Ronggang Wang, Wen Gao 0001
ICME2
2015 Fast intra mode decision algorithm based on refinement in HEVC
abstract
High Efficiency Video Coding (HEVC) is the next generation video compression standard providing significant coding performance. It adopts 35 intra prediction modes with larger CU size to improve the intra encoding efficiency, so that cause a high computational complexity. In this paper, two fast intra-prediction algorithms are proposed to reduce the number of candidate modes for rate-distortion (RD) optimization. We obtain an optimal adjacent modes (OAM) list consisting of dominant directions through the analysis of costs of several general direction modes. Furthermore, we improve the most probable mode (MPM) algorithm to make full use of the spatial correlation between neighbour prediction blocks instead of simply merging the prediction modes of neighbour prediction blocks into the candidate list. Experimental results show that the proposed algorithms can reduce about 27.3% of the encoding time compared to the HEVC test model 14.0, while the decrease of coding quality is negligible.
Longfei Gao, Shengfu Dong, Wenmin Wang 0001, Ronggang Wang, Wen Gao 0001
ISCAS3
2015 Context-adaptive fast motion estimation of HEVC
abstract
High Efficient Video Coding (HEVC) is the latest coding standard with superior compression efficiency while its encoding complexity is much higher compared with H.264/AVC. Motion estimation is one of the most time-consuming parts in video coding. In the reference software of HEVC, TZ (Test Zone) search method is adopted as the fast motion estimation method. However, its complexity is still high. There are many other fast motion estimation methods, for example, the hexagon search method, but their performance loss is larger than TZ search. In order to balance coding speed and performance, a new context-adaptive fast motion estimation algorithm is proposed in this paper. In this the proposed, motion intensity is defined in block-level, motion vectors and motion vector differences of neighbor blocks are utilized to measure the motion intensity. When motion intensity is large, TZ search method is used; otherwise, hexagon search method is used. Experimental results show that the proposed method can save 39% ~ 60% of motion estimation time with average 0.5% of BD-rate loss.
Ronggang Wang, Xiaole Cui, Wenmin Wang 0001
ISCAS4
2015 Clustering Sentences with Density Peaks for Multi-document Summarization
abstract
Multi-document Summarization (MDS) is of great value to many real world applications.Many scoring models are proposed to select appropriate sentences from documents to form the summary, in which the clustering-based methods are popular.In this work, we propose a unified sentence scoring model which measures representativeness and diversity at the same time.Experimental results on DUC04 demonstrate that our MDS method outperforms the DUC04 best method and the existing clustering-based methods, and it yields close results compared to the state-of-the-art generic MDS methods.Advantages of the proposed MDS method are two-fold: (1) The density peaks clustering algorithm is firstly adopted, which is effective and fast.(2) No external resources such as Wordnet and Wikipedia or complex language parsing algorithms is used, making reproduction and deployment very easy in real environment.
Yunqing Xia, Yi Liu 0056, Wenmin Wang 0001
HLT-NAACL4
2015 Improving VLAD with regional PCA whitening
abstract
In recent yeas, VLAD has been used to represent an image effectively and efficiently by just a few bytes in large-scale image retrieval. In spite of its remarkable performance, a series of modification methods have been presented. In addition, the redundancy between the features corresponding to the same cluster center could be improved. In this paper, a regional PCA Whitening method is proposed to decorrelate the features and reduce the dimensionality for each cluster with the consideration of mapping the descriptor into high dimensionality explicitly. Our method can also be embedded into original VLAD pipeline with global PCA very well. The experimental results on both Holidays and UKbench dataset show that our approach improves VLAD significantly.
Mingmin Zhen, Wenmin Wang 0001, Ronggang Wang
VCIP2
2015 Weighted transformable spatial pyramid and scalable query for object retrieval
abstract
Object retrieval in the large-scale image corpus is an appealing, yet challenging task. Most of existing frameworks are based on bag-of-visual-words (BoVW) model. However, BoVW has an obvious drawback, i.e. lack of spatial information. In this paper, we propose weighted transformable spatial pyramid and scalable query for object retrieval. We first break the whole image into sub-images and then make these sub-images up in a new order in the final representation. Our method has two contributions: 1) relative spatial relationships of local features instead of absolute geometric layouts of features are encoded so that translation invariance is guaranteed, 2) scaling invariance in the image representation is ensured by scalable query. The experimental results show that our approach outperforms BoVW and traditional spatial pyramid matching.
Zi'ou Zheng, Wenmin Wang 0001, Ronggang Wang
VCIP2
2015 Dynamic macroblock wavefront parallelism for parallel video coding
Zhenyu Wang 0002, Shengfu Dong, Ronggang Wang, Wenmin Wang 0001, Wen Gao 0001
J. Vis. Commun. Image Represent.4
2015 High Resolution Local Structure-Constrained Image Upsampling
abstract
With the development of ultra-high-resolution display devices, the visual perception of fine texture details is becoming more and more important. A method of high-quality image upsampling with a low cost is greatly needed. In this paper, we propose a fast and efficient image upsampling method that makes use of high-resolution local structure constraints. The average local difference is used to divide a bicubic-interpolated image into a sharp edge area and a texture area, and these two areas are reconstructed separately with specific constraints. For reconstruction of the sharp edge area, a high-resolution gradient map is estimated as an extra constraint for the recovery of sharp and natural edges; for the reconstruction of the texture area, a high-resolution local texture structure map is estimated as an extra constraint to recover fine texture details. These two reconstructed areas are then combined to obtain the final high-resolution image. The experimental results demonstrated that the proposed method recovered finer pixel-level texture details and obtained top-level objective performance with a low time cost compared with state-of-the-art methods.
Yang Zhao 0002, Ronggang Wang, Wenmin Wang 0001, Wen Gao 0001
IEEE Trans. Image Process.3
2014 HEVC decoder acceleration on multi-core X86 platform
abstract
In this paper, we propose a hybrid parallel decoding strategy for HEVC which combines task-level parallelism and datalevel parallelism based on CTUs. The data-level parallelism makes the execution time distribution of different decoding stages more balanced, and makes the task-level parallelism more efficient. Our approach imposes no constraint on bit streams that they shall be generated by optional parallel coding tools such as tiles or WPP, so it can be applied for all kinds of HEVC bit streams. Furthermore, SSE, a typical SIMD instruction set on X86 platform, is utilized to accelerate time-consuming modules, which shortens the execution time gaps between different stages and make them in favor of parallel processing. We have implemented these acceleration strategies on HM-10.0 decoder, and a great speed-up ratio is achieved.
Bingjie Han, Ronggang Wang, Zhenyu Wang 0002, Shengfu Dong, Wenmin Wang 0001, Wen Gao 0001
ICASSP5
2014 Cost-volume filtering-based stereo matching with improved matching cost and secondary refinement
abstract
Recent cost-volume filtering-based local stereo methods have achieved comparable accuracy with global methods. However, there are still some significant outliers existing in the final disparity map. In this paper, we propose a cost-volume filtering-based local stereo matching method that employs a new combined cost and a novel secondary disparity refinement mechanism. The combined cost is formulated by a modified color census transform, truncated absolute differences of color and gradients. Symmetric guided filter is used for the cost aggregation. Different from traditional stereo matching, a novel secondary disparity refinement is proposed to further remove remaining outliers. Experimental results on Mid-dlebury benchmark show that our method ranks the 5thout of the 144 submitted methods, and is the best cost-volume filtering-based local method. Furthermore, experiments on real world sequences also validate the effectiveness of our proposed method.
Jianbo Jiao, Ronggang Wang, Wenmin Wang 0001, Shengfu Dong, Zhenyu Wang 0002, Wen Gao 0001
ICME3
2014 A new frame interpolation method with pixel-level motion vector field
abstract
In this paper, a new frame interpolation method with pixel-level motion vector field (MVF) is proposed. Given that existing methods cannot handle occlusions and blocking artifacts well, there are three contributions in our method: (i) applying the pixel-level motion vectors (MVs) estimated by optical flow algorithm to eliminate blocking artifacts (ii) motion post-processing to keep spatial consistency (iii) robust warping method to address collisions and holes caused by occlusions. The method could remove blocking artifacts and alleviate the artifacts caused by occlusions. Experimental results show that the proposed method outperforms existing methods both in terms of objective and subjective performances, especially for sequences with complex motions.
Chuanxin Tang, Ronggang Wang, Wenmin Wang 0001, Wen Gao 0001
VCIP3
2013 High definition IEEE AVS decoder on ARM NEON platform
abstract
Nowadays, mobile devices are capable of displaying video up to HD resolution. In this paper, we propose two acceleration strategies for Audio Video coding Standard (AVS) software decoder on multi-core ARM NEON platform. Firstly, data level parallelism is utilized to effectively use the SIMD capability of NEON and key modules are redesigned to make them SIMD friendly. Secondly, a macroblock level wavefront parallelism is designed based on the decoding dependencies among macroblocks to utilize the processing capability of multiple cores. Experiment results show that AVS (IEEE 1857) HD video stream can be decoded in real-time by applying the proposed two acceleration strategies.
Ronggang Wang, Wenmin Wang 0001, Zhenyu Wang 0002, Shengfu Dong, Wen Gao 0001
ICIP3
2013 Adaptive motion estimation order for frame rate up-conversion
abstract
This paper proposes an adaptive motion estimation (ME) order for frame rate up-conversion (FRUC). Almost all existing FRUC methods adopt a raster scan order for ME. The ME is performed from top-left blocks to bottom-right blocks in raster scan order. Such an order can propagate some wrongly estimated motion vectors (MV) through a frame. The proposed method first detects the blocks rich in features (feature blocks) and estimates their MVs. Then ME is performed on the other blocks according to their distance to feature blocks. The closer a block to feature blocks is; the earlier the ME is performed on it. In this adaptive order, MVs of feature blocks are propagated to its neighbors. It makes the estimated motions of a frame close to the true motions. In order to demonstrate the efficiency of the proposed method, we estimate the MVs with diamond search in the proposed adaptive ME order. In the experiments, the quality of frame rate up converted videos have been significantly improved compared with the ones using traditional raster scan order. Moreover, the adaptive ME order can be easily combined with various ME methods applied in previous FRUC.
Chengzhou Tang, Ronggang Wang, Wenmin Wang 0001
ISCAS3
2013 Dynamic MB-level Scheduling for parallel video coding
abstract
MB-level parallelism is widely used in parallel video coding thanks to its merits of low latency, no performance loss and high degree of parallelism. Most of video encoders with MB-level parallelism employ MB Row Scheduling (MRS) scheme. In software video encoder, early terminate algorithms tend to cause significant difference in coding time of different MBs. Consequently, the running speeds of multiple threads are unbalanced. When the number of threads is more than that of physical cores, the running speed unbalance is further worsened by computation resources competition among multiple threads. The computation resources of multiple cores can't be fully utilized without careful handling of the above running speed unbalance. Additionally, synchronization of multiple threads can also penalize the running speed of the whole video encoder. In this paper, we analyze the running speed unbalance of multiple threads in MRS scheme, and propose a new Dynamic MB-level Scheduling (DMS) scheme for parallel video coding. DMS alleviates both the running speed unbalance and synchronization delay among multiple threads on multi-core platform. Experiment results verified that video encoder with MRS can be accelerated in average 9% by our proposed DMS, when processors are fully utilized.
Shengfu Dong, Zhenyu Wang 0002, Ronggang Wang, Wenmin Wang 0001, Wen Gao 0001
PCS4