VLDB 2026 Research / reviewers in the wild / expert
Wenmin Wang 0001
dblp:96/7699
· DBLP profile ↗
102ranked-venue papers
0as first author
27since 2021 · last 2026
0000-0003-2664-4413ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 74 · 9 since 2021Artificial intelligence and machine learning · 28 · 19 since 2021Systems, architecture and hardware · 4Databases, data management, data science and information retrieval · 3 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MCGS: Markov Chain Gaussian Splatting for Dynamic Scenes ReconstructionabstractWe present MCGS (Markov Chain Gaussian Splatting), a novel approach for high-fidelity dynamic scene reconstruction via combining Markov chain and 3D Gaussian splatting. Our method addresses the critical challenge of artifact-free temporal consistency in dynamic neural rendering. By integrating a Markov chain-based deformation network with multi-head temporal attention, MCGS effectively captures motion patterns and temporal dependencies, producing more accurate and stable 3D representations over time. The key innovations include: (1) a Markov Deform Network that models state transitions while preserving temporal coherence, (2) a temporal attention mechanism that adaptively weights historical states within a sliding window, and (3) strategic noise injection during training to enhance model robustness and generalization. Experiments on representative dynamic scene datasets demonstrate that MCGS outperforms previous methods in both visual quality and temporal coherence, while maintaining competitive rendering speed and efficiency. These results suggest the practical applicability of our approach to real-world dynamic scene understanding and synthesis. Wenmin Wang 0001, Xinxing Yu, Zhongheng Chen |
AAAI | 2 |
| 2026 | DHCF-Net: Dual heterodimensional context fusion network for medical image segmentation
Yu Wang 0087, Wenzheng Ma, Wenmin Wang 0001 |
Expert Syst. Appl. | 4 |
| 2026 | Insert Anyone: High-fidelity full-body photo insertion via dual-branch adapters
Yifan Zhang 0028, Zhongliang Tang, Wenmin Wang 0001 |
Expert Syst. Appl. | 4 |
| 2025 | CSRP: Modeling class spatial relation with prototype network for novel class discovery
Nannan Li 0001, Jiuqing Dong, Huiwen Guo, Wenmin Wang 0001, Chuanchuan You |
Appl. Intell. | 5 |
| 2025 | Causality thinking for large-scale long-tailed video action recognition
Zhengjin Zhang, Nannan Li 0001, Wenmin Wang 0001, Huiwen Guo, Sudan Huang |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | TIA2V: Video generation conditioned on triple modalities of text-image-audio
Minglu Zhao, Wenmin Wang 0001, Rui Zhang 0108, Haomei Jia |
Expert Syst. Appl. | 2 |
| 2025 | Fine-tuning feature interaction for unsupervised domain adaptive low-light object detection
Maomao Xiong, Qunshu Zhang, Dagang Li 0001, Wenmin Wang 0001, Cong Liu 0012, Da Chen 0002, Jinglin Zhang 0004 |
Neurocomputing | 4 |
| 2025 | Multimodal Sensitive Adaptive Transformer for 3D medical image segmentationabstractThree-dimensional medical imaging segmentation presents a significant challenge within the field, with the segmentation of multiple organs and lesions in MRI images being particularly demanding. This paper introduces an innovative approach utilizing the Multimodal Sensitive Adaptive Attention (MSAA). We refer to this new structure as the Multimodal Sensitive Adaptive Transformer Network (MSAT), which incorporates downsampling and Multimodal Sensitive Adaptive Attention into the encoding phase and integrate skip connections from different layers, outputs from Multimodal Sensitive Adaptive Attention, and upsampled feature outputs into the decoding phase. The MSAT consists of two primary components. The initial component is designed to extract a richer set of high-dimensional features through an advanced network architecture. This includes integration of different layers skip connections, outputs from the MSAA, and the results of the preceding upsampling layer. The second component features a Multimodal Sensitive Adaptive Attention block, which integrates two types of attention mechanisms: Local Sensitive Adaptive Attention (LSAA) and Spatial Sensitive Adaptive Attention (SSAA). These attention mechanisms work synergistically to blend high and low-dimensional features effectively, thereby enriching the contextual information captured by the model. Our experiments, conducted across several datasets including Synapse, BTCV, ACDC, and the BraTS 2021 dataset, demonstrate that the MSAT outperforms other existing methodologies. The MSAT shows superior segmentation capabilities for 3D multi-organ, cardiac, and brain tumor segmentation tasks. Zhibing Wang, Wenmin Wang 0001, Nannan Li 0001, Yifan Zhang 0028, Haomei Jia, Shenyong Zhang |
Image Vis. Comput. | 2 |
| 2025 | Causalseg: investigating causality modeling for semi-supervised video object segmentation
Zhengjin Zhang, Nannan Li 0001, Wenmin Wang 0001, Huiwen Guo |
Multim. Syst. | 3 |
| 2025 | SwapInpaint2: Towards high structural consistency in identity-guided inpainting via background-preserving GAN inversion
Honglei Li 0008, Yifan Zhang 0028, Wenmin Wang 0001, Shenyong Zhang |
Pattern Recognit. | 3 |
| 2024 | Span Confusion is All You Need for Chinese Spelling Correction
Dezhi Ye, Haomei Jia, Jie Liu 0075, Haijin Liang, Jin Ma 0003, Wenmin Wang 0001 |
CIKM | 7 |
| 2024 | Local Information Guided Global Integration for Infrared Small Target DetectionabstractInfrared small targets often exhibit small scale and weak semantic features, which makes it a great challenge to their detection. To address this situation, we propose a novel network for infrared small target detection that combines local details information and global contextual information. To preserve the local and high-frequency details present in infrared images, we introduce a High-frequency Aware Encoder. To extract contextual information from multi-scale feature maps, we propose a Multi-scale Context Learning Bottleneck that incorporates contextual information repeatedly and performs cross-level fusion, which enables the recognition of small targets based on their surroundings. Finally, a lightweight Transformer Decoder is employed to restore the feature map, while placing attention on the target pixels. Experimental results on the IRSTD-1k dataset demonstrate that our method outperforms other state-of-the-art approaches. Qianchen Mao, Jinbao Wang 0001, Wenmin Wang 0001, Bingshu Wang |
ICASSP | 5 |
| 2024 | SgLFT: Semantic-guided Late Fusion Transformer for video corpus moment retrieval
Tongbao Chen, Wenmin Wang 0001, Minglu Zhao, Ruochen Li 0001 |
Neurocomputing | 2 |
| 2024 | Ultrahigh-definition video quality assessment: A new dataset and benchmark
Ruochen Li 0001, Wenmin Wang 0001, Huanqiang Hu, Tongbao Chen, Minglu Zhao |
Neurocomputing | 2 |
| 2024 | Multimodal parallel attention network for medical image segmentation
Zhibing Wang, Wenmin Wang 0001, Nannan Li 0001, Shenyong Zhang |
Image Vis. Comput. | 2 |
| 2024 | Enhanced blind face inpainting via structured mask prediction
Honglei Li 0008, Yifan Zhang 0028, Wenmin Wang 0001 |
Pattern Recognit. Lett. | 3 |
| 2024 | Cross-Modality Knowledge Calibration Network for Video Corpus Moment RetrievalabstractVideo corpus moment retrieval has become a hot topic recently, which aims to localize a consequent video moments highly relevant to the given query language description from video corpus. Existing methods towards this challenging task are suffering from the cases when the visual information and textual information in the video are very different from each other or from the cases where the redundant video content is semantically irrelevant with the query language description, which make the model confused of figuring out the truly useful within- and cross-modality information. In this article, we propose a novel Cross-Modality Knowledge Calibration Network (CKCN) to solve the issue mentioned above. Specifically, a dual calibration transformer module with improved multi-head attention is proposed to simultaneously capture the within- and cross-modality features between the visual and textual modality of the video automatically compressing the redundant information, and then a query-dependent fusion module is designed to guide feature fusion of the video's multi-modal information using the prior knowledge of query which further refine more important modality features. At last, a query-guided calibration transformer module with a well-designed learnable cell is utilized to align the query and video, forming a single joint representation for moment localization. Meanwhile, we introduce transfer learning into the task of video corpus moment retrieval (VCMR) for the first time to solve the defect of insufficient labeled data. Extensive experiments have been conducted on both the widely used TVR dataset and DiDeMo dataset which have achieved new state-of-the-art, thus verifying the effectiveness of our proposed CKCN. Tongbao Chen, Wenmin Wang 0001, Ruochen Li 0001, Bingshu Wang |
IEEE Trans. Multim. | 2 |
| 2024 | TA2V: Text-Audio Guided Video GenerationabstractRecent conditional and unconditional video generation tasks have been accomplished mainly based on generative adversarial network (GAN), diffusion, and autoregressive models. However, in some circumstances, using only one modality cannot provide enough semantic information. Therefore, in this paper, we propose text-audio to video (TA2V) generation, a new task for generating realistic videos from two different guided modalities, text and audio, which has not been explored much thus far. Compared to image generation, video generation is a harder task because of the complexity of processing higher-dimensional data and scarcer suitable datasets, especially for multimodal video generation. To overcome these limitations, (i) we propose the Text&Audio-guided-Video-Maker (TAgVM) model, which consists of two modules: a text-guided video generator and a text&audio-guided video modifier. (ii) This model uses a 3D VQ-GAN to compress high-dimension video data to a low-dimension discrete sequence, followed by an autoregressive model to guide text-conditional generation in the latent space. Then, we apply a text&audio-guided diffusion model to the generated video scenes, providing additional semantic details corresponding to the audio and text. (iii) We introduce a newly produced music performance video dataset, the University of Rochester Multimodal Music Performance with Video-Audio-Text (URMP-VAT), and a landscape dataset, Landscape with Video-Audio-Text (Landscape-VAT), both of which include three modalities (text, audio, and video) that are aligned with each other. The results demonstrate that our model can create videos with satisfactory quality and semantic information. The source code and datasets are available athttps://github.com/Minglu58/TA2V. Minglu Zhao, Wenmin Wang 0001, Tongbao Chen, Rui Zhang 0108, Ruochen Li 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | E-detector: Asynchronous Spatio-temporal for Event-based Object Detection in Intelligent Transportation SystemabstractIn intelligent transportation systems, various sensors, including radar and conventional frame cameras, are used to improve system robustness in various challenging scenarios. An event camera is a novel bio-inspired sensor that has attracted the interest of several researchers. It provides a form of neuromorphic vision to capture motion information asynchronously at high speeds. Thus, it possesses advantages for intelligent transportation systems that conventional frame cameras cannot match, such as high temporal resolution, high dynamic range, as well as sparse and minimal motion blur. Therefore, this study proposes an E-detector based on event cameras that asynchronously detect moving objects. The main innovation of our framework is that the spatiotemporal domain of the event camera can be adjusted according to different velocities and scenarios. It overcomes the inherent challenges that traditional cameras face when detecting moving objects in complex environments, such as high speed, complex lighting, and motion blur. Moreover, our approach adopts filter models and transfer learning to improve the performance of event-based object detection. Experiments have shown that our method can detect high-speed moving objects better than conventional cameras using state-of-the-art detection algorithms. Thus, our proposed approach is extremely competitive and extensible, as it can be extended to other scenarios concerning high-speed moving objects. The study findings are expected to unlock the potential of event cameras in intelligent transportation system applications. Wenmin Wang 0001, Honglei Li 0008, Shenyong Zhang |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Shadow Removal of Text Document Images Using Background Estimation and Adaptive Text EnhancementabstractThis paper proposes a simple yet effective method to re-move shadows from text document images. It mainly includes several parts. Firstly, we propose a text elimination-based background extraction strategy to estimate shadow map. It indicates the shadow regions accurately and helps to predict global background. Secondly, a binarization-based text ex-traction algorithm is designed to obtain texts from document image. By fusing texts and global background, a preparatory shadow-free image can be obtained. Thirdly, we propose an adaptive text contrast enhancement strategy to generate shadow-free results with comfortable visual perception across shadow and non-shadow regions. Quantitative and visual results performed on open datasets indicate that the proposed method can generate clear shadow-free images from text document images. Our code will be publicly available soon. Bingshu Wang, Jiangbin Zheng 0001, Wenmin Wang 0001 |
ICASSP | 4 |
| 2023 | Bounding convolutional network for refining object locations
Shenyong Zhang, Wenmin Wang 0001, Honglei Li 0008 |
Neural Comput. Appl. | 2 |
| 2023 | SaGCN: Semantic-Aware Graph Calibration Network for Temporal Sentence GroundingabstractTemporal sentence grounding is a challenging task that aims to localize the semantic corresponding segment from the untrimmed video according to the given query language description. Existing methods either utilize a cross-modal matching architecture following a scan-and-rank pipeline or directly predict the probabilities of being the target boundary for each frame based on the entire video content. However, such methods are weak when some of the critical semantic concepts in the query are actually relevant to multiple video segments or the desired video segment contains a query-irrelevant scene due to ignoring query semantic concepts and local and global cross-modal context. In this paper, we propose a novel semantic-aware graph calibration network (SaGCN) to address the issues mentioned above. Specifically, we first introduce a semantic-aware local relational graph module to capture the inherent relationships among the specific semantic concept relevant local contextual information for fine-grained cross-modal information interactions. Then, a semantic-aware global relational graph module is derived for global contextual information integration and achieving cross-modal alignment. Finally, an attention-based calibration module is designed for eliminating the irrelevant information maintained in the visual modality under the guidance of query description. Extensive experiments verify the effectiveness of our proposed SaGCN on two widely used datasets (Charades-STA and TACoS), in which we achieve significant and consistent improvement compared to the state-of-the-art approaches. Tongbao Chen, Wenmin Wang 0001, Kangrui Han, Huijuan Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | ANGraph: attribute-interactive neighborhood-aggregative graph representation learning
Ying Shen 0001, Huizhi Li, Dagang Li 0001, Jingwei Zheng, Wenmin Wang 0001 |
Neural Comput. Appl. | 5 |
| 2022 | Fast 2-step regularization on style optimization for real face morphing
Wenmin Wang 0001, Honglei Li 0008, Roberto Bugiolacchi |
Neural Networks | 2 |
| 2022 | Affective word embedding in affective explanation generation for fine art paintings
Jianhao Yan, Wenmin Wang 0001 |
Pattern Recognit. Lett. | 2 |
| 2022 | Fast transformation of discriminators into encoders using pre-trained GANs
Wenmin Wang 0001 |
Pattern Recognit. Lett. | 2 |
| 2022 | SwapInpaint: Identity-Specific Face Inpainting With Identity SwappingabstractAs face editing scenarios have become popular, the face inpainting technique has become a hot topic. Although some existing methods can inpaint faces with preserved identity information, they fail to solve a more flexible inpainting problem that fills the “holes” with identity-specific content from other faces. In this work, we propose a disentangle and subject-agnostic framework that affects both full and partial-face inpainting with the guidance of a reference face image. The framework consists of an identity encoding module, a content inference module and a generative module. The identity encoding module extracts the identity embedding from the reference image, the content inference module learns to predict the content image, and the generative module integrates the content image and the reference identity embedding to generate the identity-specific inpainted result. To minimize the structure and style gap between the incomplete image and inpainted image, we use a double attribute loss to the generative module and a postprocess of blending operation to the swapped result. We compare our method with state-of-the-art works and demonstrate that our method achieves higher identity similarity and better structural correctness. Honglei Li 0008, Wenmin Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Exploring Entity-Level Spatial Relationships for Image-Text MatchingabstractExploring the entity-level (i.e., objects in an image, words in a text) spatial relationship contributes to understanding multimedia content precisely. The ignorance of spatial information in previous works probably leads to misunderstandings of image contents. For instance, sentences `Boats are on the water' and `Boats are under the water' describe the same objects, but correspond to different sceneries. To this end, we utilize the relative position of objects to capture entity-level spatial relationships for image-text matching. Specifically, we fuse semantic and spatial relationships of image objects in a visual intra-modal relation module. The module performs promisingly to understand image contents and improve object representation learning. It contributes to capturing entity-level latent correspondence of image-text pairs. Then the query (text) plays a role of textual context to refine the interpretable alignments of image-text pairs in the inter-modal relation module. Our proposed method achieves state-of-the-art results on MSCOCO and Flickr30K datasets. Yaxian Xia, Lun Huang, Wenmin Wang 0001, Xiaoyong Wei, Jie Chen 0001 |
ICASSP | 3 |
| 2020 | Generating Future Frames with Mask-Guided PredictionabstractCurrent approaches in video prediction tend to hallucinate the future frames directly or learn global motion transformation from the entire scene. However, it is difficult for these methods without instance-aware mechanism to learn the underlying structures, dynamics and appearances of foreground and background elements simultaneously, especially when it comes to long-term prediction. In this paper, we propose an explicit instance-level prediction approach to tackle this issue and present a novel mask-guided dual network. We utilize instance masks to extract active objects from the videos, and design two LSTM branches to predict the future dynamics and appearances for objects and backgrounds individually. Superior than most recent skeleton-aided methods that only focus on single human object with two-stage procedure, our proposed network can predict instances from other categories and be trained end-to-end with a joint loss. We evaluate our approach on KTH, Penn Action and Running Horse datasets, and achieve promising results in both quality and quantity. Xiongtao Chen, Wenmin Wang 0001 |
ICME | 4 |
| 2020 | Low Resolution Facial Manipulation DetectionabstractDetecting manipulated images and videos is an important aspect of digital media forensics. Due to severe discriminative information loss caused by resolution degradation, the performance of most existing methods is significantly reduced on low resolution manipulated images. To address this issue, we propose an Artifacts-Focus Super-Resolution (AFSR) module and a Two-stream Feature Extractor (TFE). The AFSR recovers facial cues and manipulation artifact details using an autoencoder learned with an artifacts focus training loss. The TFE adopts a two-stream feature extractor with key points-based fusion pooling to learn discriminative facial representations. These two complementary modules are jointly trained to recover and capture distinctive manipulation artifacts in low resolution images. Extensive experiments on two benchmarks including FaceForensics++ and DeepfakeTIMIT, evidence the favorable performance of our method against other state-of-the-art methods. Zhongyi Ji, Wenmin Wang 0001 |
VCIP | 3 |
| 2020 | A Dense-Gated U-Net for Brain Lesion SegmentationabstractBrain lesion segmentation plays a crucial role in diagnosis and monitoring of disease progression. DenseNets have been widely used for medical image segmentation, but much redundancy arises in dense-connected feature maps and the training process becomes harder. In this paper, we address the brain lesion segmentation task by proposing a Dense-Gated U-Net (DGNet), which is a hybrid of Dense-gated blocks and U-Net. The main contribution lies in the dense-gated blocks that explicitly model dependencies among concatenated layers and alleviate redundancy. Based on dense-gated blocks, DGNet can achieve weighted concatenation and suppress useless features. Extensive experiments on MICCAI BraTS 2018 challenge and our collected intracranial hemorrhage dataset demonstrate that our approach outperforms a powerful backbone model and other state-of-the-art methods. Zhongyi Ji, Tong Lin 0002, Wenmin Wang 0001 |
VCIP | 4 |
| 2020 | Text-to-Image Generation via Semi-Supervised TrainingabstractSynthesizing images from text is an important problem and has various applications. Most of the existing studies of text-to-image generation utilize supervised methods and rely on a fully-labeled dataset, but detailed and accurate descriptions of images are onerous to obtain. In this paper, we introduce a simple but effective semi-supervised approach that considers the feature of unlabeled images as "Pseudo Text Feature". Therefore, the unlabeled data can participate in the following training process. To achieve this, we design a Modality-invariant Semantic- consistent Module which aims to make the image feature and the text feature indistinguishable and maintain their semantic information. Extensive qualitative and quantitative experiments on MNIST and Oxford-102 flower datasets demonstrate the effectiveness of our semi-supervised method in comparison to supervised ones. We also show that the proposed method can be easily plugged into other visual generation models such as image translation and performs well. Zhongyi Ji, Wenmin Wang 0001, Baoyang Chen |
VCIP | 2 |
| 2020 | Fast and Accurate Action Detection in Videos With Motion-Centric Attention ModelabstractA key factor that makes action detection in videos different from general video classification is human-guided clues, especially motion signals. Since not all the pixels in a video are informative for action recognition, the irrelevant and redundant parts can lead to a lot of noise and be burdensome for both feature extraction and classifier training. This encourages the researchers to seek out the design of the attentive model that can dynamically focus computations on the key spatiotemporal volumes. In this paper, we propose a motion-centric attention model for action detection in videos which imitates the human perception of saccade and fixation procedures while detecting actions in a video. Specifically, we first present a strategy to generate motion-centric locations based on the density peak of motion signals, providing reliable candidates around which actions have high possibilities to occur. Then, we introduce an attention model that conducts the saccade and fixation procedures on these candidates to observe local spatiotemporal visual information, preserve internal comprehension, and produce the action proposals on temporal bounds. Afterward, a classifier with several variants is prepared to classify the action proposals and decide which one to fixate and generate the final predictions. We show how to efficiently train our model to produce fast and accurate action detection, by scanning only a small fraction of locations in a video. The extensive experiments on three challenging datasets show promising results with both accuracy and speed. Jinzhuo Wang, Wenmin Wang 0001, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Uni-and-Bi-Directional Video Prediction via Learning Object-Centric TransformationabstractVideo prediction, including uni-directional prediction for future frames and bi-directional prediction for in-between frames, is a challenging task and a problem worth exploring in multimedia and computer vision fields. Existing practices usually make predictions by learning global motion information from the whole given image. However, humans often focus on key objects carrying vital motion information instead of the entire frame. Besides, different objects often show different movement and deformation, even in the same scene. In this connection, we build a novel model of object-centric video prediction, in which the motion signals of key objects are particularly learned. This model can predict new frames by repeatedly transforming objects into the original input images. To focus on these objects automatically, we create an attention module with substitutable strategies. Our method requires no annotated data, and we also use adversarial training to improve sharpness of the predictions. We evaluate our model through Moving MNIST, UCF101 and Penn Action datasets and achieve competitive results in both quantity and quality, compared to existing methods. The experiments demonstrate that our uni-and-bi-directional network can well predict motions for different objects and generate plausible future and in-between frames. Xiongtao Chen, Wenmin Wang 0001 |
IEEE Trans. Multim. | 2 |
| 2019 | Image Captioning with Two Cascaded AgentsabstractRecent neural models on image captioning usually take a encoder-decoder fashion, where the decoder predicts a single word at one step recently with the encoder providing information. The encoder is a pretrained CNN model typically. Thus the decoder, the input to it, and the output from it become the most important parts of a model. We propose a pipelined image captioning framework consisting of two cascaded agents. The former is named as "semantic adaptive agent" which generates the input to the decoder by consulting the information from the current decoding process, and the latter as "caption generating agent" which select a single word of the vocabulary as the output of the decoder by taking consideration of the input and the current states of the decoder. For the framework of two cascaded agents, we design a multi-stage training procedure to train the two agents with different objectives by fully utilizing reinforcement learning. In experiments, we conduct quantitative and qualitative analysis on MS COCO dataset and our results can significantly outperform baseline methods in terms of several evaluation metrics. Lun Huang, Wenmin Wang 0001 |
ICASSP | 2 |
| 2019 | Multi-step Self-attention Network for Cross-modal Retrieval Based on a Limited Text SpaceabstractCross-modal retrieval has been recently proposed to find an appropriate subspace where the similarity among different modalities, such as image and text, can be directly measured. In this paper, we propose Multi-step Self-Attention Network (MSAN) to perform cross-modal retrieval in a limited text space with multiple attention steps, that can selectively attend to partial shared information at each step and aggregate useful information over multiple steps to measure the final similarity. In order to achieve better retrieval results with faster training speed, we introduce global prior knowledge as the global reference information. Extensive experiments on Flickr30K and MSCOCO, show that MSAN achieves new state-of-the-art results in accuracy for cross-modal retrieval. Wenmin Wang 0001, Ge Li 0002 |
ICASSP | 2 |
| 2019 | Attention on Attention for Image CaptioningabstractAttention mechanisms are widely used in current encoder/decoder frameworks of image captioning, where a weighted average on encoded vectors is generated at each time step to guide the caption decoding process. However, the decoder has little idea of whether or how well the attended vector and the given attention query are related, which could make the decoder give misled results. In this paper, we propose an Attention on Attention (AoA) module, which extends the conventional attention mechanisms to determine the relevance between attention results and queries. AoA first generates an information vector and an attention gate using the attention result and the current context, then adds another attention by applying element-wise multiplication to them and finally obtains the attended information, the expected useful knowledge. We apply AoA to both the encoder and the decoder of our image captioning model, which we name as AoA Network (AoANet). Experiments show that AoANet outperforms all previously published methods and achieves a new state-of-the-art performance of 129.8 CIDEr-D score on MS COCO Karpathy offline test split and 129.6 CIDEr-D (C40) score on the official online testing server. Code is available at https://github.com/husthuaan/AoANet. Lun Huang, Wenmin Wang 0001, Jie Chen 0001, Xiaoyong Wei |
ICCV | 2 |
| 2019 | Video Prediction with Temporal-Spatial Attention Mechanism and Deep Perceptual Similarity BranchabstractVideo prediction is a challenging but worth exploring task in computer vision. Different from image analysis, the challenge of video analysis derives from more complicated dependencies in time as well as in space. In this paper, we propose a Temporal-Spatial Attention Mechanism (TSAM) to capture not only spatial appearance dependencies but also temporal dynamic dependencies in video sequence. The TSAM is transplantable for existing networks and allows the long-range dependency modeling for various video analysis tasks (we take video prediction for specific experiments in this paper). Besides, we propose an additional Deep Perceptual Similarity Branch (DPSB) to encourage a better approximation to the ground-truth in high-level feature space, laying the foundation for frame generation. Extensive experiments on KTH, Penn Action and UCF-101 datasets demonstrate that our model performs quite competitively across diverse natural visual scenes, even in long-term video prediction. Wenmin Wang 0001, Xiongtao Chen, Weimian Li |
ICME | 2 |
| 2019 | Adaptively Aligned Image Captioning via Adaptive Attention TimeabstractRecent neural models for image captioning usually employ an encoder-decoder framework with an attention mechanism. However, the attention mechanism in such a framework aligns one single (attended) image feature vector to one caption word, assuming one-to-one mapping from source image regions and target caption words, which is never possible. In this paper, we propose a novel attention model, namely Adaptive Attention Time (AAT), to align the source and the target adaptively for image captioning. AAT allows the framework to learn how many attention steps to take to output a caption word at each decoding step. With AAT, an image region can be mapped to an arbitrary number of caption words while a caption word can also attend to an arbitrary number of image regions. AAT is deterministic and differentiable, and doesn't introduce any noise to the parameter gradients. In this paper, we empirically show that AAT improves over state-of-the-art methods on the task of image captioning. Code is available at https://github.com/husthuaan/AAT. Lun Huang, Wenmin Wang 0001, Yaxian Xia, Jie Chen 0001 |
NeurIPS | 2 |
| 2019 | Predicting Diverse Future Frames With Local Transformation-Guided MaskingabstractVideo prediction is the challenging task of generating the future frames of a video given a sequence of previously observed frames. This task involves the construction of an internal representation that accurately models the frame evolutions, including contents and dynamics. Video prediction is considered difficult due to the inherent compounding of errors in recursive pixel level prediction. In this paper, we present a novel video prediction system that focuses on regions of interest (ROIs) rather than on entire frames and learns frame evolutions at the transformation level rather than at the pixel level. We provide two strategies to generate high-quality ROIs that contains potential moving visual cues. The frame evolutions are modeled with a transformation generator that produces transformers and masks simultaneously, which are then combined to generate the future frame in a transformation-guided masking procedure. Compared with recent approaches, our system is able to generate more accurate predictions by modeling the visual evolutions at the transformation level rather than at the pixel level. Focusing on ROIs avoids a heavy computational burden and enables our system to generate high-quality long-term future frames without severely amplified signal loss. Moreover, our system is able to generate diverse plausible future frames, which is important in many real-world scenarios. Furthermore, we enable our system to perform video prediction conditioned on a single frame by revising the transformation generator to produce motion-centric transformers. We test our system on four datasets with different experimental settings and demonstrate its advantages over recent methods, both quantitatively and qualitatively. Jinzhuo Wang, Wenmin Wang 0001, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | A Motion Aided Merge Mode For HevcabstractMerge prediction is a practical inter-technique in HEVC, which can significantly improve the coding efficiency, especially for homogeneous regions in video sequences. In this paper, a motion aided merge mode (MAMM) is proposed to achieve a better trade-off between the prediction accuracy and bit rate. Different from the traditional merge mode in HEVC, MAMM is accomplished by a small motion obtained by searching in a specific search region. The search range is comprised of a number of points with high occurrence possibilities. The motion vector difference (MVD) is coded by Huffman coding in MAMM and the Huffman coding table is generated according to the statistical frequency of each possible MVD value. The proposed method is implemented on top of the HEVC reference software (HM −16.15), and experimental results show that 0.6% BD-rate reduction is achieved under Random Access (RA) configuration. Kui Fan, Ronggang Wang, Ge Li 0002, Wenmin Wang 0001 |
ICASSP | 5 |
| 2018 | Local patch encoding-based method for single image super-resolution
Yang Zhao 0002, Ronggang Wang, Wei Jia 0001, Jianchao Yang, Wenmin Wang 0001, Wen Gao 0001 |
Inf. Sci. | 5 |
| 2018 | MPEG Internet Video Coding Standard and Its Performance EvaluationabstractMPEG has produced standards that have provided the industry with the best video compression technologies. To address diverse Internet needs, MPEG issued a Call for Proposals (CfP) for Internet video coding (IVC) in July, 2011. The anticipation is that any patent declaration associated with the baseline profile of this standard will indicate that the patent owner is prepared to grant a free of charge license to an unrestricted number of applicants worldwide. Three codecs have responded to the CfP: Web video coding (WVC), video coding for browsers (VCB), and IVC. WVC is in fact the AVC baseline, and VCB uses the same coding tools as VP8. IVC has been developed in MPEG from scratch by combining well-known existing technology elements and new coding tools with royalty-free declarations. In June 2015, the IVC project was approved as ISO/IEC 14496-33 (MPEG-4 IVC). This standard can be highly beneficial for video services in the Internet domain. This paper describes the main coding tools used in IVC, and evaluates its objective and subjective performances compared with WVC, VCB, and AVC high profile (AVC HP). The experimental results show that IVC's compression performance is approximately equal to that of the AVC HP for typical operational settings, both for streaming and low-delay applications, and is superior to WVC and VCB. Ronggang Wang, Zhenyu Wang 0002, Kui Fan, Tiejun Huang 0001, Wenmin Wang 0001, Ge Li 0002, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2018 | Second- and High-Order Graph Matching for Correspondence ProblemsabstractCorrespondence problems are challenging due to the complexity of real-world scenes. One way to solve this problem is to improve the graph matching (GM) process, which is flexible for matching non-rigid objects. GM can be classified into three categories that correspond with the variety of object functions: first-order, second-order, and high-order matching. Graph and hypergraph matching have been proposed separately in previous works. The former is equivalent to the second-order GM, and the latter is equivalent to high-order GM, but we use the terms second- and high-order GM to unify the terminology in this paper. Second- and high-order GM fit well with different types of problems; the key goal for these processes is to find better-optimized algorithms. Because the optimal problems for second- and high-order GM are different, we propose two novel optimized algorithms for them in this paper. (1) For the second-order GM, we first introduce a$K$-nearest-neighbor-pooling matching method that integrates feature pooling into GM and reduces the complexity. Meanwhile, we evaluate each matching candidate using discriminative weights on its$k$-nearest neighbors by taking locality as well as sparsity into consideration. (2) High-order GM introduces numerous outliers, because precision is rarely considered in related methods. Therefore, we propose a sub-pattern structure to construct a robust high-order GM method that better integrates geometric information. To narrow the search space and solve the optimization problem, a new prior strategy and a cell-algorithm-based Markov Chain Monte Carlo framework are proposed. In addition, experiments demonstrate the robustness and improvements of these algorithms with respect to matching accuracy compared with other state-of-the-art algorithms. Wenmin Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Multiscale Deep Alternative Neural Network for Large-Scale Video ClassificationabstractWith the rapid increase in the amount of multimedia data, video classification has become a demanding and challenging research topic. Compared with image classification, video classification requires mapping a video that contains hundreds of frames to semantic tags, which poses many challenges to the direct use of advanced models originally designed for image-oriented tasks. On the other hand, continuous frames in a video also give us more visual clues that we can leverage to achieve better classification. One of the most important clues is the context in the spatiotemporal domain. In this paper, we introduce the multiscale deep alternative neural network (DANN), a novel architecture combining the strengths of both convolutional neural network and recurrent neural networks to achieve a deep network that can collect rich context hierarchies for video classification. In particular, the DANN is stacked with alternative layers, each of which consists of a volumetric convolutional layer followed by a recurrent layer. The former acts as a local feature learner, whereas the latter is used to collect contexts. Compared with popular deep feed-forward neural networks, the DANN learns local features and their contexts from the very beginning. This setting enables preserving context evolutions, which we show to be essential for improving the accuracy of video classification. To release the full potential of the DANN, we develop a deeper version with stochastic-layer skip-connections and construct a multiscale DANN to incorporate contexts at different scales. We show how to apply the multiscale DANN for video classification with carefully designed configurations in terms of both input-output settings and training-testing methods. The DANN is shown to be robust to not only human-centric videos, but also natural videos. As there are few large-scale natural disaster video datasets, we construct a new large-scale one and make it publicly available. Experiments on four datasets show the effectiveness of our method for both human actions and natural events. Jinzhuo Wang, Wenmin Wang 0001, Wen Gao 0001 |
IEEE Trans. Multim. | 2 |
| 2017 | Beyond Monte Carlo Tree Search: Playing Go with Deep Alternative Neural Network and Long-Term EvaluationabstractMonte Carlo tree search (MCTS) is extremely popular in computer Go which determines each action by enormous simulations in a broad and deep search tree. However, human experts select most actions by pattern analysis and careful evaluation rather than brute search of millions of future interactions. In this paper, we propose a computer Go system that follows experts’ way of thinking and playing. Our system consists of two parts. The first part is a novel deep alternative neural network (DANN) used to generate candidates of next move. Compared with existing deep convolutional neural network (DCNN), DANN inserts recurrent layer after each convolutional layer and stacks them in an alternative manner. We show such setting can preserve more contexts of local features and its evolutions which are beneficial for move prediction. The second part is a long-term evaluation (LTE) module used to provide a reliable evaluation of candidates rather than a single probability from move predictor. This is consistent with human experts’ nature of playing since they can foresee tens of steps to give an accurate estimation of candidates. In our system, for each candidate, LTE calculates a cumulative reward after several future interactions when local variations are settled. Combining criteria from the two parts, our system determines the optimal choice of next move. For more comprehensive experiments, we introduce a new professional Go dataset (PGD), consisting of $253,233$ professional records. Experiments on GoGoD and PGD datasets show the DANN can substantially improve performance of move prediction over pure DCNN. When combining LTE, our system outperforms most relevant approaches and open engines based on MCTS. Jinzhuo Wang, Wenmin Wang 0001, Ronggang Wang, Wen Gao 0001 |
AAAI | 2 |
| 2017 | Attention-Based Two-Phase Model for Video Action Detection
Xiongtao Chen, Wenmin Wang 0001, Weimian Li, Jinzhuo Wang |
CAIP (2) | 2 |
| 2017 | A Violence Detection Approach Based on Spatio-temporal Hypergraph Transition
Jingjia Huang, Ge Li 0002, Nannan Li 0001, Ronggang Wang, Wenmin Wang 0001 |
CAIP (2) | 5 |
| 2017 | Progressive Probabilistic Graph Matching with Local Consistency Regularization
Wenmin Wang 0001 |
CAIP (2) | 2 |
| 2017 | A New Image Contrast Enhancement Algorithm Using Exposure Fusion Framework
Zhenqiang Ying, Ge Li 0002, Yurui Ren, Ronggang Wang, Wenmin Wang 0001 |
CAIP (2) | 5 |
| 2017 | Learning a Limited Text Space for Cross-Media Retrieval
Wenmin Wang 0001, Mengdi Fan |
CAIP (1) | 2 |
| 2017 | A Multilayer Backpropagation Saliency Detection Algorithm Based on Depth Mining
Chunbiao Zhu, Ge Li 0002, Wenmin Wang 0001, Ronggang Wang |
CAIP (2) | 4 |
| 2017 | Cross-modality matching based on Fisher Vector with neural word embeddings and deep image featuresabstractCross-modal retrieval, which aims to solve the problem that the query and the retrieved results are from different modality, becomes more and more essential with the development of the Internet. In this paper, we mainly focus on the exploration of high-level semantic representation of image and text for cross-modal matching. Deep convolutional image features and Fisher Vector with neural word embeddings are utilized as visual and textual features respectively. To further investigate the correlation among heterogeneous multimodal characteristics, we use multiclass logistic classifier for semantic matching across modalities. Experiments on Wikipedia and Pascal Sentence dataset demonstrate the robustness and effectiveness for both Img2Text and Text2Img retrieval tasks. Wenmin Wang 0001, Mengdi Fan, Ronggang Wang |
ICASSP | 2 |
| 2017 | A joint model for action localization and classification in untrimmed video with visual attentionabstractIn this paper, we introduce a joint model that learns to directly localize the temporal bounds of actions in untrimmed videos as well as precisely classify what actions occur. Most existing approaches tend to scan the whole video to generate action instances, which are really inefficient. Instead, inspired by human perception, our model is formulated based on a recurrent neural network to observe different locations within a video over time. And, it is capable of producing temporal localizations by only observing a fixed number of fragments, and the amount of computation it performs is independent of input video size. The decision policy for determining where to look next is learned by REINFORCE which is powerful in non-differentiable settings. In addition, different from relevant ways, our model runs localization and classification serially, and possesses a strategy for extracting appropriate features to classify. We evaluate our model on ActivityNet dataset, and it greatly outperforms the baseline. Moreover, compared with a recent approach, we show that our serial design can bring about 9% increase in detection performance. Weimian Li, Wenmin Wang 0001, Xiongtao Chen, Jinzhuo Wang, Ge Li 0002 |
ICME | 2 |
| 2017 | Better deep visual attention with reinforcement learning in action recognitionabstractDeep visual attention in computer vision has attracted much attention over the past years, which achieves great contributions especially in image classification, image caption and action recognition. However, due to taking BP training wholly or partially, they can not show the true power of attention in computational efficiency and focusing accuracy. Our intuition is that attention mechanism should be similar to the process in which human draw attention and select the next location to focus, by observing, analyzing and jumping instead of existing describing continuous features. Based on this insight, we formulate our model as a recurrent neural network-based agent that chooses attention region by reinforcement learning at each timestep. In experiments, our model explicitly outperforms baselines not only in focusing and recognizing accuracy, but also consumes much less computational resources, which can be honored as better deep visual attention. Wenmin Wang 0001, Jingzhuo Wang, Yaohua Bu |
ISCAS | 2 |
| 2017 | Learning Object-Centric Transformation for Video PredictionabstractFuture frame prediction for video sequences is a challenging task and worth exploring problem in computer vision. Existing methods often learn motion information for the entire image to predict next frames. However, different objects in the same scene often move and deform in different ways intuitively. Considering the human visual system, one often pays attention to the key objects that contain crucial motion signals, rather than compress an entire image into a static representation. Motivated by this property of human perception, in this work, we develop a novel object-centric video prediction model that learns local motion transformation dynamically for key object regions with visual attention. By transforming objects iteratively to the original input frames, next frame can be produced. Specifically, we design an attention module with replaceable strategies to attend to objects in video frames automatically. Our method does not require any annotated data during training procedure. To produce sharp predictions, adversarial training is adopted in our work. We evaluate our model on the Moving MNIST and UCF101 datasets and report competitive results, compared to prior methods. The generated frames demonstrate that our model can characterize motion for different objects and produce plausible future frames. Xiongtao Chen, Wenmin Wang 0001, Jinzhuo Wang, Weimian Li |
ACM Multimedia | 2 |
| 2017 | Cross-media Retrieval by Learning Rich Semantic Embeddings of MultimediaabstractCross-media retrieval aims at seeking the semantic association between different media types. Most existing methods paid much attention on learning mapping functions or finding the optimal spaces, but neglected how people accurately cognize images and texts. This paper proposes a brain inspired cross-media retrieval framework to learn rich semantic embeddings of multimedia. Different from directly using off-the-shelf image features, we combine the visual and descriptive senses for an image from the view of human perception via a joint model, called multi-sensory fusion network (MSFN). A topic model based TextNet maps texts into the same semantic space as images according to their shared ground truth labels. Moreover, in order to overcome the limitations of insufficient data for training neural networks and less complexity in text form, we introduce a large-scale image-text dataset, called Britannica dataset. Extensive experiments show the effectiveness of our framework for different lengths of texts on three benchmark datasets as well as Britannica dataset. Most of all, we report the best known average results of Img2Text and Text2Img compared with several state-of-the-art methods. Mengdi Fan, Wenmin Wang 0001, Peilei Dong, Ronggang Wang, Ge Li 0002 |
ACM Multimedia | 2 |
| 2017 | Long-term video interpolation with bidirectional predictive networkabstractThis paper considers the challenging task of long-term video interpolation. Unlike most existing methods that only generate few intermediate frames between existing adjacent ones, we attempt to speculate or imagine the procedure of an episode and further generate multiple frames between two non-consecutive frames in videos. In this paper, we present a novel deep architecture called bidirectional predictive network (BiPN) that predicts intermediate frames from two opposite directions. The bidirectional architecture allows the model to learn scene transformation with time as well as generate longer video sequences. Besides, we make attempts to extend our model to predict multiple possible procedures by sampling different noise vectors. A joint loss composed of clues in image and feature spaces and adversarial loss is designed to train our model. We demonstrate the advantages of BiPN on two benchmarks Moving 2D Shapes and UCF101 and report competitive results to recent approaches. Xiongtao Chen, Wenmin Wang 0001, Jinzhuo Wang |
VCIP | 2 |
| 2017 | Mask-streaming CNN for pedestrian detectionabstractInspired by humans recognition of pedestrians, we propose a mask-streaming convolutional neural network (CNN) for pedestrian detection. The mask stream consists of one original region proposal and six masked regions. These masked regions, which aim at highlighting significantly discriminative semantic characteristic of head, body parts and contextual information, are generated through wiping off a part of the original region by `masks'. We feed the mask stream into the network to generate features of each masked region, then a concatenate layer integrates these features as the final representation of pedestrians. We evaluate the proposed model on the challenging Caltech and Eth datasets. Our approach performs better than the baseline, and achieves competitive performance compared to numerous pedestrian detection methods. Peilei Dong, Wenmin Wang 0001, Mengdi Fan, Ronggang Wang, Ge Li 0002 |
VCIP | 2 |
| 2017 | Learning multi-view embedding in joint space for bidirectional image-text retrievalabstractIn this paper, we propose a framework for learning a joint embedding space for bidirectional image-text retrieval task, which fuses embedding spaces in multi-views. We have implemented two views currently, one is a frame-sentence view and the other is a region-phrase view. In the frame-sentence view, we project each frame of the images and each sentence of the texts into a holistic-level subspace to explore the correlation between them. In the region-phrase view, we extract each region of the frames and each phrase of the sentences and map them into a local-level subspace. We separately mine the semantic correlations in the two views, then merge them by a multi-view fusion ranking method. In each view, to embed heterogeneous data into a common space, we adopt the two-branch neural network to transform the data. Extensive experiments show that our multi-view joint space can preserve more accurate semantic correlations between images and texts in different granularities and can significantly improve the performance on the image-text retrieval task. Our method achieves better results than the state of the art on the Pascal1K and Flickr8K image-sentence datasets. Lu Ran, Wenmin Wang 0001 |
VCIP | 2 |
| 2017 | Adaptive difference modelling for background subtractionabstractBackground subtraction plays a very important role in video analysis, especially in surveillance systems. While being straightforward, the performance based on frame differencing is unsatisfied due to its sensitiveness to issues such as camera shake and swinging objects. To address its limitations, in this paper we propose a complete adaptive difference modelling framework. First, we introduce two difference discriminators to model the evolution process of pixels. Second, we use Gaussian Mixture Models to adaptively learn the difference threshold to distinguish foreground from background. Third, three heuristics are employed to further improve the model adaptability. Experiments on real-world videos of the Background Models Challenge (BMC) demonstrate that our method performs better on global quality metric (FSD) than other state-of-the-art methods. Xianghao Zang, Ge Li 0002, Jun Yang 0033, Wenmin Wang 0001 |
VCIP | 4 |
| 2017 | Deep discriminative network with inception module for person re-identificationabstractConvolutional neural networks have been verified to be exceptionally powerful on extracting semantic features, which contribute to a great progress in computer vision. However, focusing too much on the superiority, researchers seem to pay less attention to exploring CNNs' potential in other aspects, e.g. the ability to discriminate the difference. In this work we try to dig into the discriminative power of CNNs and introduce a deep discriminative network with inception module (DDN-IM) for person re-identification. Without individual feature extraction as prerequisite, input images from two different non-overlapping camera views are concatenated in depth at the beginning, followed by series of convolutional and nonlinear operations, etc. to predict their similarity. In addition, inception module is embedded in our network to boost the performance. We validate our proposal on several person re-identification datasets, CUHK01, QMUL GRID and PRID2011 included. We obtain competitive or superior performance compared to the state-of-the-art methods. Wenmin Wang 0001, Jinzhuo Wang |
VCIP | 2 |
| 2017 | Iterative projection reconstruction for fast and efficient image upsampling
Yang Zhao 0002, Ronggang Wang, Wei Jia 0001, Wenmin Wang 0001, Wen Gao 0001 |
Neurocomputing | 4 |
| 2017 | Color Image-Guided Boundary-Inconsistent Region Refinement for Stereo MatchingabstractCost computation, cost aggregation, disparity optimization, and disparity refinement are the four main steps for stereo matching. While the first three steps have been widely investigated, few efforts have been taken on disparity refinement. In this paper, we propose a color image-guided disparity refinement method to further remove the boundary-inconsistent regions on disparity map. First, the origins of boundary-inconsistent regions are analyzed. Then, these regions are detected with the proposed hybrid-superpixel-based strategy. Finally, the detected boundary-inconsistent regions are refined by a modified weighted median filtering method. Experimental results on various stereo matching conditions validate the effectiveness of the proposed method. Furthermore, depth maps obtained by active depth acquisition devices like Kinect can also be well refined with our proposed method. Jianbo Jiao, Ronggang Wang, Wenmin Wang 0001, Dagang Li 0001, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | Accelerating Image-Domain-Warping Virtual View Synthesis on GPGPUabstractThe image-domain-warping (IDW) method can effectively create high-quality virtual views. However, the IDW algorithm is very complex, and the software implementation for this method is far from real-time. In this paper, we propose an IDW-based view synthesis acceleration method on general-purpose computing on a graphics processing unit (GPGPU). Our method makes two main contributions. First, at the algorithm level, we employ FAST for sparse disparity estimation and adopt the successive over-relaxation iterative method to calculate warps. Second, at the platform level, two computation-intensive modules (data extraction and view synthesis) in IDW are offloaded to GPU using efficient data-level parallelism strategies. Experimental results demonstrate that our proposed acceleration method can speed up the original IDW algorithm by more than 110x, and HD stereo three-dimensional video can be converted to 8-view 4 K video (each view has an approximate 720P resolution) in real-time on a hybrid CPU + GPU (NVIDIA GTX980) platform. Ronggang Wang, Jiajia Luo, Xiubao Jiang, Zhenyu Wang 0002, Wenmin Wang 0001, Ge Li 0002, Wen Gao 0001 |
IEEE Trans. Multim. | 5 |
| 2016 | Tube ConvNets: Better exploiting motion for action recognitionabstractMotion information is a key factor for action recognition and has been eagerly pursued for decades. How to effectively learn motion features in Convolutional Networks (ConvNets) remains an open issue. Prevalent ConvNets often take several full frames of video as input at a time, which can be a heavy burden for network training. In this paper, we introduce a novel framework called Tube ConvNets, by substituting action tubes for full frames to reduce this burden. Tube ConvNets focus on the regions of interest (ROI) where key motions occur, and thus eliminate the distraction of irrelevant objects. Each action tube is a fraction of spatiotemporal volumes, generated by the techniques of object detection and clustering algorithm. We demonstrate the effectiveness of Tube ConvNets for action classification on UCF-101 dataset, and illustrate its potential to support fine-grained localization on UCF-Sports dataset. Source code is available at https://github.com/wangjinzhuo/tubecnn. Zhihao Li 0002, Wenmin Wang 0001, Nannan Li 0001, Jinzhuo Wang |
ICIP | 2 |
| 2016 | An Empirical Study of Deformable Part Model with fast feature pyramidabstractThe performance of an object detection system relies heavily on two components: an object model to capture the compositional relationship among the object body and its parts, and a feature representation to describe object appearance. In this work, we present an empirical study of combining two state-of-the-art such components: Deformable Part Model (DPM), a proven effective and flexible part-based object model which originally adopts Histogram of Oriented Gradients (HOG) feature, and Aggregated Channel Features (ACF), a unified feature representation framework with fast pyramid calculation which is originally used in a rigid template matching scheme. DPM is known to work but slow, at the same time ACF has previously been shown to yield a massive speedup with only a minor loss in accuracy compared to competing features including HOG. By combining the two, our hope is to achieve the best of both worlds: the object structure representation power of DPM and the computational efficiency of ACF. Our experiments show that while ACF with heterogeneous feature channels could improve the accuracy of DPM, the run time benefit introduced by fast pyramid approximation is rather limited. Jun Yang 0033, Ge Li 0002, Wenmin Wang 0001, Ronggang Wang |
ICPR | 3 |
| 2016 | An MCMC-based prior sub-hypergraph matching in presence of outliersabstractCorrespondence problems are very challenging due to the complexity of real-world scenes. Some hypergraph matching methods have been proposed for improving the recall of the solution, but the numerous outliers are brought since the precision is rarely considered. To solve this issue, we propose a sub-hypergraph matching method, which is robust with better integration of geometric information and reduces the difficulty of NP-hard problem happened in hypergraphs. To narrow the search space and solve the optimization problem, a new prior strategy and cell-algorithm in Markov Chain Monte Carlo (MCMC) framework is proposed on sub-hypergraph matching. The experiments show that our proposed method significantly outperforms other state-of-the-art algorithms. Wenmin Wang 0001 |
ICPR | 2 |
| 2016 | Regional Subspace Projection Coding for Image RetrievalabstractFor image retrieval task, hamming embedding, being proved to be one of the state-of-the-art methods, has been prevalently utilised. The basic idea is to project local features into orthogonal space randomly, in which the binary signature is generated based on a single partition of feature space. However, the binary signature generation process is coarse and heuristic. On the one hand, the same projection is carried out for all visual word space without consideration of difference among subspaces. On the other hand, the projection matrix is generated randomly regardless of the distribution of feature data. Therefore, the performance of hamming embedding is limited and far from the optimal. In this paper, we firstly analyse the limitation of hamming em- bedding and compare different orthogonal projection methods. Then we propose a regional subspace projection coding method that is based on the distribution of local features assigned to each visual word. Finally, our experiments on two benchmark datasets demonstrate that our proposed method outperforms current state-of-the-art methods. Mingmin Zhen, Wenmin Wang 0001, Ronggang Wang |
ICMR | 2 |
| 2016 | A Novel Shadow-Free Feature Extractor for Real-Time Road DetectionabstractRoad detection is one of the most important research areas in driver assistance and automated driving field. However, the performance of existing methods is still unsatisfactory, especially in severe shadow conditions. To overcome those difficulties, first we propose a novel shadow-free feature extractor based on the color distribution of road surface pixels. Then we present a road detection framework based on the extractor, whose performance is more accurate and robust than that of existing extractors. Also, the proposed framework has much low-complexity, which is suitable for usage in practical systems. Zhenqiang Ying, Ge Li 0002, Xianghao Zang, Ronggang Wang, Wenmin Wang 0001 |
ACM Multimedia | 5 |
| 2016 | Deep Alternative Neural Network: Exploring Contexts as Early as Possible for Action RecognitionabstractContexts are crucial for action recognition in video. Current methods often mine contexts after extracting hierarchical local features and focus on their high-order encodings. This paper instead explores contexts as early as possible and leverages their evolutions for action recognition. In particular, we introduce a novel architecture called deep alternative neural network (DANN) stacking alternative layers. Each alternative layer consists of a volumetric convolutional layer followed by a recurrent layer. The former acts as local feature learner while the latter is used to collect contexts. Compared with feed-forward neural networks, DANN learns contexts of local features from the very beginning. This setting helps to preserve hierarchical context evolutions which we show are essential to recognize similar actions. Besides, we present an adaptive method to determine the temporal size for network input based on optical flow energy, and develop a volumetric pyramid pooling layer to deal with input clips of arbitrary sizes. We demonstrate the advantages of DANN on two benchmarks HMDB51 and UCF101 and report competitive or superior results to the state-of-the-art. Jinzhuo Wang, Wenmin Wang 0001, Xiongtao Chen, Ronggang Wang, Wen Gao 0001 |
NIPS | 2 |
| 2016 | An effective post quantization rate estimation for HEVC intra encoderabstractIn high efficiency video coding (HEVC), the encoder employs a flexible quad-tree coding structure as well as a large number of prediction modes. For each size of coding unit (CU), transform unit (TU) and each prediction mode, the rate distortion optimization (RDO) is performed to select the best CU, TU and the best prediction mode. Although better coding efficiency is achieved, the computational complexity increases dramatically. In order to reduce the burden of RDO in the HEVC intra encoder, in this paper, we propose an effective approach, which is based on the generalized Gaussian distribution (GGD) model, to estimate the block level bit-rate. The weaknesses of the conventional GGD model are analyzed and relevant improvements are exploited. Our experiments show that, compared with the original RDO procedure in HM16.0, the proposed algorithm reduces RDO time by 37.7% with 0.64% BD-rate loss. Hongbin Cao, Ronggang Wang, Zhenyu Wang 0002, Ge Li 0002, Wenmin Wang 0001 |
VCIP | 5 |
| 2016 | Better region proposals for pedestrian detection with R-CNNabstractRegion-based convolutional neural networks (R-CNNs) have achieved great success in object detection recently. These deep models depend on region proposal algorithms to hypothesize object locations. In this paper, we combine a special region proposal algorithm with R-CNN, and apply it to pedestrian detection. The special algorithm is used to generate region proposals only for pedestrian class. It is different from the popular region proposal algorithm selective search that detects generic object locations. The experimental results prove that region proposals generated by our method are more applicable than selective search for pedestrian detection. Our method performs faster training and testing than the deep model based on, and it achieves competitive performance compared to the state-of-the-arts in pedestrian detection. Peilei Dong, Wenmin Wang 0001 |
VCIP | 2 |
| 2016 | A simple but efficient way to combine VLAD with locality-constrained linear codingabstractThe VLAD (vector of locally aggregated descriptors) representation, derived from BoF and Fisher kernel, has shown its efficiency in the field of image search. However, assigning local descriptors to a codeword is a hard voting process, which does not consider the uncertainty and the plausibility for single codeword. In this paper, we propose an approach to combine VLAD with locality-constrained linear coding, as opposed to the original one, considering several nearest neighbors when assigning local descriptors and computing weights. In order to evaluate our proposed method, experiments are conducted on several image classification benchmarks, using VLAD for comparison. The experimental results show that our method stably outperforms VLAD in terms of classification accuracy, while producing feature representation of the same dimension without much additional computational cost. Zhenglin Tan, Wenmin Wang 0001, Yifeng Jiang 0004, Ronggang Wang |
VCIP | 2 |
| 2016 | A new video denoising method using texture metric and adaptive structure varianceabstractIn this paper, an innovated method is proposed for video denoising. The method consists of two major procedures. First, a new adaptive superpixel video texture metric is proposed. This video texture metric is calculated, which relates to different parts of a video stream. Then a new adaptive structure variance is estimated by adopting fine and coarse structures. Finally, a noise filter based on the estimated weights of different structures in a video stream is applied. By comparison, the proposed method outperforms traditional state-of-art methods, especially in block artifacts reduction. Ge Li 0002, Wenmin Wang 0001, Ronggang Wang |
VCIP | 4 |
| 2016 | An advanced local offset matching strategy for object proposal matchingabstractImage correspondence problem is very challenging due to the complexity of real-world scenes, especially in the presence of deformation, outliers, and other intra-class variations. Semantic flow methods are used for finding image correspondences, but the prevalent ones are mainly depended on pixel-level or regularly local-region sampled operation that are easily confused by image-specific details and element-specific scenes to individual objects, e.g. background, clutter, and they rarely care about the relationships between proposals. Therefore, we propose an advanced local offset matching strategy for object proposals, which is in term of naive Bayesian model, considering both appearance and the local spatial pattern with preprocessing. It not only builds reliable region relationships using object proposals, but also explores consistency between appearance and spatial pattern among proposals. In addition, two evaluation metrics, seven features are applied to experiments. The comparison and analysis, along with adequate experiments on standard benchmark, present that our proposed method effectively outperforms other matching strategies in diverse settings. Wenmin Wang 0001 |
VCIP | 2 |
| 2016 | Local Quantization Code histogram for texture classification
Yang Zhao 0002, Ronggang Wang, Wenmin Wang 0001, Wen Gao 0001 |
Neurocomputing | 3 |
| 2016 | Spatially variant defocus blur map estimation and deblurring from a single image
Xinxin Zhang 0004, Ronggang Wang, Xiubao Jiang, Wenmin Wang 0001, Wen Gao 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2016 | Multilevel Modified Finite Radon Transform Network for Image UpsamplingabstractA local line-like feature is the most important discriminate information in the image upsampling scenario. In recent example-based upsampling methods, grayscale and gradient features are often adopted to describe the local patches, but these simple features cannot accurately characterize complex patches. In this paper, we present a feature representation of local edges by means of a multilevel filtering network, namely, multilevel modified finite Radon transform network (MMFRTN). In the proposed MMFRTN, the MFRT is utilized in the filtering layer to extract the local line-like feature; the nonlinear layer is set to be a simple local binary process; for the feature-pooling layer, we concatenate the mapped patches as the feature of local patch. Then, we propose a new example-based upsampling method by means of the MMFRTN feature. Experimental results demonstrate the effectiveness of the proposed method over some state-of-the-art methods. Yang Zhao 0002, Ronggang Wang, Wenmin Wang 0001, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2016 | CSPS: An Adaptive Pooling Method for Image ClassificationabstractThis paper proposes an adaptive approach to learn class-specific pooling shapes (CSPS) for image classification. Prevalent methods for spatial pooling are often conducted on predefined grids of images, which is an ad-hoc method and, thus, lacks generalization power across different categories. In contrast, our CSPS is designed in a data-driven fashion by generating plenty of candidates and selecting the optimal subset for each class. Specifically, we establish an overcomplete spatial shape set that preserves as many geometric patterns as possible. Then, the class-specific subset is selected by training a linear classifier with structured sparsity constraints and color distribution cues. To address the high computational cost and the risk of overfitting due to the overcomplete scheme, the image representations for CSPS are first compressed according to dictionary sensitivity and shape importance. These representations are finally fed to SVMs for the classification task. We demonstrate that CSPS can learn compact yet discriminative geometric information for different classes that carries more semantic meaning than other methods. Experimental results on four datasets demonstrate the benefits of the proposed method compared with other pooling schemes and illustrate its effectiveness on both object and scene images. Jinzhuo Wang, Wenmin Wang 0001, Ronggang Wang, Wen Gao 0001 |
IEEE Trans. Multim. | 2 |
| 2015 | Learning discriminative visual dictionary for natural scene categorizationabstractMany successful systems for scene recognition transform low-level descriptors into complex representations. This process consists of the two steps: 1) feature coding, which performs a pointwise transformation of the descriptors into a representation adapted to the task, and 2) image pooling, which summarizes the coded features. Even though these two steps have been paid so much attention, but there are still some problems in combining scene semantic with local features. The goal of this paper is threefold: to address the problem by modifying the traditional bag-of-features (BoF) framework; to show how to achieve the best performance by learning a semi-supervised discriminative dictionary; and to provide theoretical and empirical insight into the remarkable performance. By teasing apart components shared by modern scene categorization pipeline, our approach aims to facilitate the design of better scene recognition architectures. Wenmin Wang 0001, Ronggang Wang |
ICASSP | 2 |
| 2015 | A novel integer-pixel motion estimation algorithm based on quadratic predictionabstractThis paper presents a fast integer-pixel motion estimation (ME) algorithm for High Efficiency Video Coding (HEVC), which provides a strategy to speed up the search process significantly, while yielding the same quality performance as the Test Zone Search (TZSearch) scheme in HEVC Test Model (HM). Compared with the H.264/AVC, the ME process employs a more complex hybrid coding architecture and a larger size search window in HEVC, leading to great computational complexity. We utilize limited pixels of certain position in the current search range to build a quadratic model, and then shrink the search range repeatedly by analyzing the sum of absolute difference (SAD) distribution until the best motion vector (MV) is obtained. The proposed algorithm can be applied to various encoding conditions. Experimental results show that our method can save 56% of computations compared with the TZSearch scheme, with negligible decrease of coding quality. Longfei Gao, Shengfu Dong, Wenmin Wang 0001, Ronggang Wang, Wen Gao 0001 |
ICIP | 3 |
| 2015 | A low-light image enhancement method for both denoising and contrast enlargingabstractIn this paper, a novel united low-light image enhancement framework for both contrast enhancement and denoising is proposed. First, the low-light image is segmented into superpixels, and the ratio between the local standard deviation and the local gradients is utilized to estimate the noise-texture level of each superpixel. Then the image is inverted to be processed in the following steps. Based on the noise-texture level, a smooth base layer is adaptively extracted by the BM3D filter, and another detail layer is extracted by the first order differential of the inverted image and smoothed with the structural filter. These two layers are adaptively combined to get a noise-free and detail-preserved image. At last, an adaptive enhancement parameter is adopt into the dark channel prior dehazing process to enlarge contrast and prevent over/under enhancement. Experimental results demonstrate that our proposed method outperforms traditional methods in both subjective and objective assessments. Lin Li 0062, Ronggang Wang, Wenmin Wang 0001, Wen Gao 0001 |
ICIP | 3 |
| 2015 | Image classification using RBM to encode local descriptors with group sparse learningabstractThis paper proposes to employ deep learning model to encode local descriptors for image classification. Previous works using deep architectures to obtain higher representations are often operated from pixel level, which lack the power to be generalized to large-size and complex images due to computational burdens and internal essence capture. Our method slips the leash of this limitation by starting from local descriptors to leverage more semantical inputs. We investigate to use two layers of Restricted Boltzmann Machines (RBMs) to encode different local descriptors with a novel group sparse learning (GSL) inspired by the recent success of sparse coding. Besides, unlike the most existing pure unsupervised feature coding strategies, we use another RBM corresponding to semantic labels to perform supervised fine-tuning which makes our model more suitable for classification task. Experimental results on Caltech-256 and Indoor-67 datasets demonstrate the effectiveness of our method. Jinzhuo Wang, Wenmin Wang 0001, Ronggang Wang, Wen Gao 0001 |
ICIP | 2 |
| 2015 | A compact shot representation for video semantic indexingabstractThis paper presents a compact shot representation for video semantic indexing (SIN). The proposed representation consists of visual cues from only two frames, i.e., key frame (KF) and difference frame (DF), which are both constructed with spatial pyramid. The KF describes static information while the generated DF captures non-static information. Each region of DF is derived from the same location in a selected frame, which has the most salient difference compared with the key frame in that region. We introduce a variation of DF to further enhance our model. Experimental results on TRECVID SIN demonstrate that our method obtains better accuracy than the state-of-the-art, while requiring less storage space and consuming time. Jinzhuo Wang, Wenmin Wang 0001, Ronggang Wang, Wen Gao 0001 |
ICIP | 2 |
| 2015 | Image deblurring using robust sparsity priorsabstractIn this paper, we propose a robust method to remove motion blur from a single photograph. We find that an inaccurate kernel and an unreliable final latent image reconstruction method are two main factors leading to low-quality restored images. To improve image quality, we do the following technical contributions. For robust blur kernel estimation, first, an edge mask and a smooth constraint are used to provide reliable intermediate latent images for salient structure extraction; second, we adopt an effective salient structure selection method to remove detrimental edges for kernel estimation; third, we use a gradient sparsity prior to remove kernel noise and ensure the continuity of blur kernels. For final latent image reconstruction, we combine the merits of both the TV-l2model and the hyper-Laplacian model to preserve tiny details and eliminate noise. Experimental results on synthetically blurred images and real photographs demonstrate that the proposed algorithm performs better than state-of-the-art approaches. Xinxin Zhang 0004, Ronggang Wang, Yonghong Tian 0001, Wenmin Wang 0001, Wen Gao 0001 |
ICIP | 4 |
| 2015 | Accelerating CDVS extraction on mobile platformabstractThe extraction of MPEG-7 Compact Descriptors for Visual Search (CDVS) on most popular mobile devices is slow, which is attributed to the complexity of the algorithm and the limited computing power of the mobile platform. A feasible and straightforward way to accelerate the extraction is to excavate the potential computing power of modern mobile microprocessor. In this paper, we implement a NEON SIMD based data level parallelism and Pthread based multi-thread parallelism scheme to accelerate the CDVS extracting process on a multi-core ARM processor. Experimental results show a speed-up of 3.5x on Gaussian convolution, and 2.3x on the whole extracting process with the proposed method. Ronggang Wang, Qiusi Wang, Wenmin Wang 0001 |
ICIP | 4 |
| 2015 | Improved cluster center adaption for image classificationabstractThe feature coding algorithm, “Vector of Locally Aggregated Descriptors (VLAD)”, can be used effectively for large scale object instance retrieval. Despite its effectiveness and excellent performance, the existence of ambiguous cluster centers can reduce the performance. Though an idea to this problem has been proposed, it is not practical in fact. In this paper, we analyze possible situations that cause effect on the results and propose a novel approach to improve the VLAD method. The proposed method mainly focuses on the similarity measure between each two images. For each two images, we adapt the original cluster center to VLAD vectors. As we illustrate, our method has promising results with small vocabulary size on both datasets of 15 Scenes and VOC2007. Mingmin Zhen, Wenmin Wang 0001, Ronggang Wang |
ICIP | 2 |
| 2015 | Learning class-specific pooling shapes for image classificationabstractSpatial pyramid (SP) representation is an extension of bag-of-feature model which embeds spatial layout information of local features by pooling feature codes over pre-defined spatial shapes. However, the uniform style of spatial pooling shapes used in standard SP is an ad-hoc manner without theoretical motivation, thus lacking the generalization power to adapt to different distribution of geometric properties across image classes. In this paper, we propose a data-driven approach to adaptively learn class-specific pooling shapes (CSPS). Specifically, we first establish an over-complete set of spatial shapes providing candidates with more flexible geometric patterns. Then the optimal subset for each class is selected by training a linear classifier with structured sparsity constraint and color distribution cues. To further enhance the robust of our model, the representations over CSPS are compressed according to the shape importance and finally fed to SVM with a multi-shape matching kernel for classification task. Experimental results on three challenging datasets (Caltech-256, Scene-15 and Indoor-67) demonstrate the effectiveness of the proposed method on both object and scene images. Jinzhuo Wang, Wenmin Wang 0001, Ronggang Wang, Wen Gao 0001 |
ICME | 2 |
| 2015 | Fast intra mode decision algorithm based on refinement in HEVCabstractHigh Efficiency Video Coding (HEVC) is the next generation video compression standard providing significant coding performance. It adopts 35 intra prediction modes with larger CU size to improve the intra encoding efficiency, so that cause a high computational complexity. In this paper, two fast intra-prediction algorithms are proposed to reduce the number of candidate modes for rate-distortion (RD) optimization. We obtain an optimal adjacent modes (OAM) list consisting of dominant directions through the analysis of costs of several general direction modes. Furthermore, we improve the most probable mode (MPM) algorithm to make full use of the spatial correlation between neighbour prediction blocks instead of simply merging the prediction modes of neighbour prediction blocks into the candidate list. Experimental results show that the proposed algorithms can reduce about 27.3% of the encoding time compared to the HEVC test model 14.0, while the decrease of coding quality is negligible. Longfei Gao, Shengfu Dong, Wenmin Wang 0001, Ronggang Wang, Wen Gao 0001 |
ISCAS | 3 |
| 2015 | Context-adaptive fast motion estimation of HEVCabstractHigh Efficient Video Coding (HEVC) is the latest coding standard with superior compression efficiency while its encoding complexity is much higher compared with H.264/AVC. Motion estimation is one of the most time-consuming parts in video coding. In the reference software of HEVC, TZ (Test Zone) search method is adopted as the fast motion estimation method. However, its complexity is still high. There are many other fast motion estimation methods, for example, the hexagon search method, but their performance loss is larger than TZ search. In order to balance coding speed and performance, a new context-adaptive fast motion estimation algorithm is proposed in this paper. In this the proposed, motion intensity is defined in block-level, motion vectors and motion vector differences of neighbor blocks are utilized to measure the motion intensity. When motion intensity is large, TZ search method is used; otherwise, hexagon search method is used. Experimental results show that the proposed method can save 39% ~ 60% of motion estimation time with average 0.5% of BD-rate loss. Ronggang Wang, Xiaole Cui, Wenmin Wang 0001 |
ISCAS | 4 |
| 2015 | Clustering Sentences with Density Peaks for Multi-document SummarizationabstractMulti-document Summarization (MDS) is of great value to many real world applications.Many scoring models are proposed to select appropriate sentences from documents to form the summary, in which the clustering-based methods are popular.In this work, we propose a unified sentence scoring model which measures representativeness and diversity at the same time.Experimental results on DUC04 demonstrate that our MDS method outperforms the DUC04 best method and the existing clustering-based methods, and it yields close results compared to the state-of-the-art generic MDS methods.Advantages of the proposed MDS method are two-fold: (1) The density peaks clustering algorithm is firstly adopted, which is effective and fast.(2) No external resources such as Wordnet and Wikipedia or complex language parsing algorithms is used, making reproduction and deployment very easy in real environment. Yunqing Xia, Yi Liu 0056, Wenmin Wang 0001 |
HLT-NAACL | 4 |
| 2015 | Improving VLAD with regional PCA whiteningabstractIn recent yeas, VLAD has been used to represent an image effectively and efficiently by just a few bytes in large-scale image retrieval. In spite of its remarkable performance, a series of modification methods have been presented. In addition, the redundancy between the features corresponding to the same cluster center could be improved. In this paper, a regional PCA Whitening method is proposed to decorrelate the features and reduce the dimensionality for each cluster with the consideration of mapping the descriptor into high dimensionality explicitly. Our method can also be embedded into original VLAD pipeline with global PCA very well. The experimental results on both Holidays and UKbench dataset show that our approach improves VLAD significantly. Mingmin Zhen, Wenmin Wang 0001, Ronggang Wang |
VCIP | 2 |
| 2015 | Weighted transformable spatial pyramid and scalable query for object retrievalabstractObject retrieval in the large-scale image corpus is an appealing, yet challenging task. Most of existing frameworks are based on bag-of-visual-words (BoVW) model. However, BoVW has an obvious drawback, i.e. lack of spatial information. In this paper, we propose weighted transformable spatial pyramid and scalable query for object retrieval. We first break the whole image into sub-images and then make these sub-images up in a new order in the final representation. Our method has two contributions: 1) relative spatial relationships of local features instead of absolute geometric layouts of features are encoded so that translation invariance is guaranteed, 2) scaling invariance in the image representation is ensured by scalable query. The experimental results show that our approach outperforms BoVW and traditional spatial pyramid matching. Zi'ou Zheng, Wenmin Wang 0001, Ronggang Wang |
VCIP | 2 |
| 2015 | Dynamic macroblock wavefront parallelism for parallel video coding
Zhenyu Wang 0002, Shengfu Dong, Ronggang Wang, Wenmin Wang 0001, Wen Gao 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2015 | High Resolution Local Structure-Constrained Image UpsamplingabstractWith the development of ultra-high-resolution display devices, the visual perception of fine texture details is becoming more and more important. A method of high-quality image upsampling with a low cost is greatly needed. In this paper, we propose a fast and efficient image upsampling method that makes use of high-resolution local structure constraints. The average local difference is used to divide a bicubic-interpolated image into a sharp edge area and a texture area, and these two areas are reconstructed separately with specific constraints. For reconstruction of the sharp edge area, a high-resolution gradient map is estimated as an extra constraint for the recovery of sharp and natural edges; for the reconstruction of the texture area, a high-resolution local texture structure map is estimated as an extra constraint to recover fine texture details. These two reconstructed areas are then combined to obtain the final high-resolution image. The experimental results demonstrated that the proposed method recovered finer pixel-level texture details and obtained top-level objective performance with a low time cost compared with state-of-the-art methods. Yang Zhao 0002, Ronggang Wang, Wenmin Wang 0001, Wen Gao 0001 |
IEEE Trans. Image Process. | 3 |
| 2014 | HEVC decoder acceleration on multi-core X86 platformabstractIn this paper, we propose a hybrid parallel decoding strategy for HEVC which combines task-level parallelism and datalevel parallelism based on CTUs. The data-level parallelism makes the execution time distribution of different decoding stages more balanced, and makes the task-level parallelism more efficient. Our approach imposes no constraint on bit streams that they shall be generated by optional parallel coding tools such as tiles or WPP, so it can be applied for all kinds of HEVC bit streams. Furthermore, SSE, a typical SIMD instruction set on X86 platform, is utilized to accelerate time-consuming modules, which shortens the execution time gaps between different stages and make them in favor of parallel processing. We have implemented these acceleration strategies on HM-10.0 decoder, and a great speed-up ratio is achieved. Bingjie Han, Ronggang Wang, Zhenyu Wang 0002, Shengfu Dong, Wenmin Wang 0001, Wen Gao 0001 |
ICASSP | 5 |
| 2014 | Cost-volume filtering-based stereo matching with improved matching cost and secondary refinementabstractRecent cost-volume filtering-based local stereo methods have achieved comparable accuracy with global methods. However, there are still some significant outliers existing in the final disparity map. In this paper, we propose a cost-volume filtering-based local stereo matching method that employs a new combined cost and a novel secondary disparity refinement mechanism. The combined cost is formulated by a modified color census transform, truncated absolute differences of color and gradients. Symmetric guided filter is used for the cost aggregation. Different from traditional stereo matching, a novel secondary disparity refinement is proposed to further remove remaining outliers. Experimental results on Mid-dlebury benchmark show that our method ranks the 5thout of the 144 submitted methods, and is the best cost-volume filtering-based local method. Furthermore, experiments on real world sequences also validate the effectiveness of our proposed method. Jianbo Jiao, Ronggang Wang, Wenmin Wang 0001, Shengfu Dong, Zhenyu Wang 0002, Wen Gao 0001 |
ICME | 3 |
| 2014 | A new frame interpolation method with pixel-level motion vector fieldabstractIn this paper, a new frame interpolation method with pixel-level motion vector field (MVF) is proposed. Given that existing methods cannot handle occlusions and blocking artifacts well, there are three contributions in our method: (i) applying the pixel-level motion vectors (MVs) estimated by optical flow algorithm to eliminate blocking artifacts (ii) motion post-processing to keep spatial consistency (iii) robust warping method to address collisions and holes caused by occlusions. The method could remove blocking artifacts and alleviate the artifacts caused by occlusions. Experimental results show that the proposed method outperforms existing methods both in terms of objective and subjective performances, especially for sequences with complex motions. Chuanxin Tang, Ronggang Wang, Wenmin Wang 0001, Wen Gao 0001 |
VCIP | 3 |
| 2013 | High definition IEEE AVS decoder on ARM NEON platformabstractNowadays, mobile devices are capable of displaying video up to HD resolution. In this paper, we propose two acceleration strategies for Audio Video coding Standard (AVS) software decoder on multi-core ARM NEON platform. Firstly, data level parallelism is utilized to effectively use the SIMD capability of NEON and key modules are redesigned to make them SIMD friendly. Secondly, a macroblock level wavefront parallelism is designed based on the decoding dependencies among macroblocks to utilize the processing capability of multiple cores. Experiment results show that AVS (IEEE 1857) HD video stream can be decoded in real-time by applying the proposed two acceleration strategies. Ronggang Wang, Wenmin Wang 0001, Zhenyu Wang 0002, Shengfu Dong, Wen Gao 0001 |
ICIP | 3 |
| 2013 | Adaptive motion estimation order for frame rate up-conversionabstractThis paper proposes an adaptive motion estimation (ME) order for frame rate up-conversion (FRUC). Almost all existing FRUC methods adopt a raster scan order for ME. The ME is performed from top-left blocks to bottom-right blocks in raster scan order. Such an order can propagate some wrongly estimated motion vectors (MV) through a frame. The proposed method first detects the blocks rich in features (feature blocks) and estimates their MVs. Then ME is performed on the other blocks according to their distance to feature blocks. The closer a block to feature blocks is; the earlier the ME is performed on it. In this adaptive order, MVs of feature blocks are propagated to its neighbors. It makes the estimated motions of a frame close to the true motions. In order to demonstrate the efficiency of the proposed method, we estimate the MVs with diamond search in the proposed adaptive ME order. In the experiments, the quality of frame rate up converted videos have been significantly improved compared with the ones using traditional raster scan order. Moreover, the adaptive ME order can be easily combined with various ME methods applied in previous FRUC. Chengzhou Tang, Ronggang Wang, Wenmin Wang 0001 |
ISCAS | 3 |
| 2013 | Dynamic MB-level Scheduling for parallel video codingabstractMB-level parallelism is widely used in parallel video coding thanks to its merits of low latency, no performance loss and high degree of parallelism. Most of video encoders with MB-level parallelism employ MB Row Scheduling (MRS) scheme. In software video encoder, early terminate algorithms tend to cause significant difference in coding time of different MBs. Consequently, the running speeds of multiple threads are unbalanced. When the number of threads is more than that of physical cores, the running speed unbalance is further worsened by computation resources competition among multiple threads. The computation resources of multiple cores can't be fully utilized without careful handling of the above running speed unbalance. Additionally, synchronization of multiple threads can also penalize the running speed of the whole video encoder. In this paper, we analyze the running speed unbalance of multiple threads in MRS scheme, and propose a new Dynamic MB-level Scheduling (DMS) scheme for parallel video coding. DMS alleviates both the running speed unbalance and synchronization delay among multiple threads on multi-core platform. Experiment results verified that video encoder with MRS can be accelerated in average 9% by our proposed DMS, when processors are fully utilized. Shengfu Dong, Zhenyu Wang 0002, Ronggang Wang, Wenmin Wang 0001, Wen Gao 0001 |
PCS | 4 |