VLDB 2026 Research / reviewers in the wild / expert
Hanli Wang
dblp:04/5757
· DBLP profile ↗
159ranked-venue papers
32as first author
58since 2021 · last 2026
0000-0002-9999-4871ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 115 · 25 first-author · 40 since 2021Artificial intelligence and machine learning · 32 · 3 first-author · 15 since 2021Databases, data management, data science and information retrieval · 8 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 2 since 2021Systems, architecture and hardware · 5 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-authorComputer networks · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CityVG: Contrastive Fine-Tuning and Reward-Based Chain-of-Thought Reasoning for Zero-Shot City-Scale 3D Visual Groundingabstract3D Visual Grounding (3DVG) locates objects in 3D scenes based on natural language descriptions. However, existing methods are primarily confined to small-scale indoor data or rely on heavy supervision, failing to generalize to the complexity of large-scale urban environments. To address this limitation, we present CityVG, the first city-scale zero-shot 3D visual grounding framework capable of localizing urban objects without manual annotations. Our approach adopts a retrieval-and-reasoning paradigm comprising two key components. Specifically, we propose a contrastive fine-tuning strategy to align textual queries with urban scene graphs. By leveraging an LLM-driven graph clustering mechanism, we automatically construct high-quality positive and negative training pairs and fine-tune the text encoder via contrastive learning, resulting in a scene-adaptive text encoder that enables efficient alignment without grounding supervision. Complementing this, we introduce a multi-trajectory reward-based Chain-of-Thought (CoT) reasoning strategy for inference. This mechanism iteratively evaluates candidate objects by aggregating reward scores across diverse reasoning trajectories, selecting the target that is most consistent with both appearance and spatial constraints. Extensive experiments on city-scale 3D grounding benchmarks demonstrate that CityVG achieves strong zero-shot localization performance and generalizes effectively to unseen urban environments. Hanli Wang |
ACL (1) | 2 |
| 2026 | Taming Image-Based Vision-Language Pre-training Model with Bootstrapped Auxiliary Tasks for Video Captioning
Hanli Wang |
MMM (1) | 2 |
| 2026 | GLM-EER: Global-Local Memory and Emotion Evaluation Refinement For Emotional Video Description
Chong Ma 0002, Shengbo Chen, Pengjie Tang, Hong Rao, Hanli Wang |
Expert Syst. Appl. | 5 |
| 2026 | Contrastive Mean Teacher for Robust Low-Light Image Enhancement
Zhangkai Ni, Menglin Han, Wenhan Yang, Hanli Wang, Lin Ma 0002, Sam Kwong |
Int. J. Comput. Vis. | 4 |
| 2026 | Visual dialog with semantic consistency: An external knowledge-driven approach
Shanshan Du, Hanli Wang |
Neural Networks | 2 |
| 2026 | Structure-Preserved Superpixel Perception for Referring Image SegmentationabstractReferring image segmentation aims to identify and segment objects in images based on linguistic expressions. Existing methods typically match individual pixels with linguistic words (pixel-word perception) and use interpolation upsampling for segmentation. However, these approaches encounter two limitations. First, as a meaningful word describes an entire entity while a pixel merely captures fragmented visual cues, the widely used pixel-word perception experiences semantic hierarchy misalignment, resulting in the misjudgment of the target entity. Second, the approximate estimation in the interpolation upsampling misleads region segmentation, struggles to preserve the target entity boundaries during feature reconstruction, and causes oversegmentation of non-target regions or undersegmentation of critical parts. To address these issues, a structure-preserved superpixel perception (SSP) framework is designed, which integrates a superpixel perception mechanism (SPM) to refine object identification and a replication upsampling strategy (RUS) to preserve appearance integrity. Specifically, SPM starts by utilizing superpixel clustering and language guidance to mine the knowledge of target object efficiently. Meanwhile, potential target regions are also extracted by SPM through traditional pixel-word associations. Then, SPM refines the recognition of target entity by establishing relationships between pixel-level representation and superpixel-level knowledge. Regarding RUS, it can directly replicate spatial and channel information for more precise feature reconstruction, serving as an alternative to interpolation upsampling. RUS consists of two complementary components: spatial replication and channel replication, which account for the assurance of target entity’s structural integrity and preservation of fine-grained details, respectively. Extensive experiments on the benchmark RefCOCO, RefCOCO+ and G-Ref datasets demonstrate that the proposed SSP framework achieves superior performance while using fewer FLOPs compared to the state-of-the-art methods. The source code of this work can be found in https://mic.tongji.edu.cn. Taiyi Su, Shanshan Du, Hanli Wang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Integrating Disparity Confidence Estimation Into Relative Depth Prior-Guided Unsupervised Stereo MatchingabstractUnsupervised stereo matching has garnered significant attention for its independence from costly disparity annotations. Typical unsupervised methods rely on the multi-view consistency assumption for training networks, which suffer considerably from stereo matching ambiguities, such as repetitive patterns and texture-less regions. A feasible solution lies in transferring 3D geometric knowledge from a relative depth map to the stereo matching networks. However, existing knowledge transfer methods learn depth ranking information from randomly built sparse correspondences, which makes inefficient utilization of 3D geometric knowledge and introduces noise from mistaken disparity estimates. This work proposes a novel unsupervised learning framework to address these challenges, which comprises a plug-and-play disparity confidence estimation algorithm and two depth prior-guided loss functions. Specifically, the local coherence consistency between neighboring disparities and their corresponding relative depths is first checked to obtain disparity confidence. Afterwards, quasi-dense correspondences are built using only confident disparity estimates to facilitate efficient depth ranking learning. Finally, a dual disparity smoothness loss is proposed to boost stereo matching performance at disparity discontinuities. Experimental results demonstrate that our method achieves state-of-the-art stereo matching accuracy on the KITTI Stereo benchmarks among all unsupervised stereo matching methods. Mingjian Sun, Cairong Zhao, Hanli Wang, Alexander V. Dvorkovich, Rui Fan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Prompted Contrastive Learning for Skeleton-Based Action RecognitionabstractLarge pre-trained vision-language models have shown great potential across various visual understanding tasks. However, in skeleton-based action recognition, it is challenging to design suitable prompts to guide the model in understanding actions depicted by skeletal structures. Moreover, the modality gap between skeleton data and text descriptions poses obstacles to effective cross-modality contrastive learning. In this work, we propose a prompted contrastive learning framework for skeleton-based action recognition. Specifically, to automatically generate natural language descriptions of actions, we introduce prompted contrast with knowledge engine (PCKE), which utilizes a pre-trained large language model as the knowledge engine to generate text prompts for the input of text encoder. Moreover, as the prompt words generated by the language model may not always be optimal, we explore the potential of the model to autonomously learn prompts. To this end, we present prompted contrast with learning engine (PCLE), a straightforward method that replaces the prompts’ context words with learnable vectors. After separately encoding skeleton data and text prompts with the skeleton and text encoder, a skeleton-language interaction process is implemented to optimize feature alignment and reduce modality gap. The interaction module contains two components: a skeleton-language bridger to connect the representation spaces of the skeleton and text encoders, and cross-modality attention to fuse information of both modalities, facilitating cross-modal knowledge transfer and refining feature alignment. Extensive experiments on four benchmark datasets of NTU RGB+D, NTU RGB+D 120, NW-UCLA, and PKU-MMD demonstrate the effectiveness of the proposed method. The source code of this work can be found in https://mic.tongji.edu.cn. Taiyi Su, Hanli Wang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Multimodal Image Representation Learning With Limited Visual-Tactile DataabstractPrevious multimodal visual-tactile image representation learning (VTL) methods have achieved significant success in object understanding through large-scale training data. However, obtaining sufficient training data is often infeasible, and the above methods struggle to effectively focus on discriminative visual and tactile features with limited data, resulting in degraded performance. To solve the above issue, we introduce a new task called visual-tactile image representation learning with limited data (VTL-L), which better facilitates real-world applications. To address the challenges of limited data and modality discrepancy in the VTL-L task, we propose a novel multi-order feature enhancement-based, alignment-free fusion network (MOA-Net). First, we introduce a multi-order feature enhancement (MFE) module to hierarchically strengthen the detailed and structural representation by aggregating the low- and high-order topological information. This approach can effectively reduce the attention noise and obtain discriminative features with limited data. Then, we propose the alignment-free visual-tactile fusion (AVTF) module to achieve representative spatial and channel features and perform the cross-modality fusion without alignment, which efficiently mitigates the modality discrepancy. Finally, we develop a dual counterfactual intervention (DCI) loss to jointly optimize fused visual-tactile feature and probability distributions, thereby improving the performance of the MOA-Net in the VTL-L task. Extensive experiments demonstrate the superiority of the proposed method across three types of tasks on four datasets under diverse limited-data settings (source code available at: https://github.com/liuxiangqiu007/MOA-Net). Liuxiang Qiu, Hui Da, Wenxi Liu, Yuzhen Niu, Hanli Wang, Tiesong Zhao |
IEEE Trans. Image Process. | 5 |
| 2025 | Generative Planning with 3D-Vision Language Pre-training for End-to-End Autonomous DrivingabstractAutonomous driving is a challenging task that requires perceiving and understanding the surrounding environment for safe trajectory planning. While existing vision-based end-to-end models have achieved promising results, these methods are still facing the challenges of vision understanding, decision reasoning and scene generalization. To solve these issues, a generative planning with 3D-vision language pre-training model named GPVL is proposed for end-to-end autonomous driving. The proposed paradigm has two significant aspects. On one hand, a 3D-vision language pre-training module is designed to bridge the gap between visual perception and linguistic understanding in the bird's eye view. On the other hand, a cross-modal language model is introduced to generate reasonable planning with perception and navigation information in an auto-regressive manner. Experiments on the challenging nuScenes dataset demonstrate that the proposed scheme achieves excellent performances compared with state-of-the-art methods. Besides, the proposed GPVL presents strong generalization ability and real-time potential when handling high-level commands in various scenarios. It is believed that the effective, robust and efficient performance of GPVL is crucial for the practical application of future autonomous driving systems. Tengpeng Li, Hanli Wang, Xianfei Li, Wenlong Liao |
AAAI | 2 |
| 2025 | MGMapNet: Multi-Granularity Representation Learning for End-to-End Vectorized HD Map ConstructionabstractThe construction of vectorized high-definition map typically requires capturing both category and geometry information of map elements. Current state-of-the-art methods often adopt solely either point-level or instance-level representation, overlooking the strong intrinsic relationship between points and instances. In this work, we propose a simple yet efficient framework named MGMapNet (multi-granularity map network) to model map elements with multi-granularity representation, integrating both coarse-grained instance-level and fine-grained point-level queries. Specifically, these two granularities of queries are generated from the multi-scale bird's eye view features using a proposed multi-granularity aggregator. In this module, instance-level query aggregates features over the entire scope covered by an instance, and the point-level query aggregates features locally. Furthermore, a point-instance interaction module is designed to encourage information exchange between instance-level and point-level queries. Experimental results demonstrate that the proposed MGMapNet achieves state-of-the-art performances, surpassing MapTRv2 by 5.3 mAP on the nuScenes dataset and 4.4 mAP on the Argoverse2 dataset, respectively. Minyue Jiang, Xiao Tan 0001, Errui Ding, Jingdong Wang 0001, Hanli Wang |
ICLR | 8 |
| 2025 | WaveCL: Wavelet Calibration Learning for Referring Video Object SegmentationabstractReferring video object segmentation (RVOS) focuses on segmenting target objects in a video based on natural language descriptions. However, existing methods typically rely on text cues that are unrelated to video content, and the target entity is only recognized in the pixel space. This often leads to ambiguous cross-modal understanding and fragmented perception across space and time, resulting in inaccurate or incomplete segmentation of the target objects. To address these challenges, a novel wavelet calibration learning (WaveCL) framework is proposed to unify cross-modal understanding and preserve spatial-temporal integrity of the target object. The WaveCL framework is built on two core components: semantic-calibrated entity perception (SEP) and wavelet-guided integrity perception (WIP). SEP aligns the textual semantics with video content, enabling more accurate and context-aware cross-modal understanding. WIP, on the other hand, leverages wavelet representations to capture fine-grained details of the target object from a global spatial-temporal perspective. By refining wavelet clues with the guidance of text queries, WIP enhances the integrity of segmentation. Through the collaboration of SEP and WIP, WaveCL enables precise, target-specific segmentation with detailed boundaries and consistent spatial-temporal perception. Extensive experiments on four benchmark datasets of Ref-YouTube-VOS, Ref-DAVIS17, A2D-Sentences, and JHMDB-Sentences show that WaveCL outperforms existing state-of-the-art methods. The source code of this work can be found in https://mic.tongji.edu.cn. Taiyi Su, Hanli Wang |
ACM Multimedia | 3 |
| 2025 | Semantic Masking with Curriculum Learning for Robust HDR Image Reconstruction
Zhangkai Ni, Kerui Ren, Wenhan Yang, Hanli Wang, Sam Kwong |
Int. J. Comput. Vis. | 5 |
| 2025 | Bootstrapping Vision-Language Models for Frequency-Centric Self-Supervised Remote Physiological Measurement
Zijie Yue, Miaojing Shi, Hanli Wang, Shuai Ding 0001, Shanlin Yang |
Int. J. Comput. Vis. | 3 |
| 2025 | SRVC-LA: Sparse regularization of visual context and latent attention based model for video description
Pengjie Tang, Jiayu Zhang 0002, Hanli Wang, Yunlan Tan, Yun Yi |
Neurocomputing | 3 |
| 2025 | MGTR-MISS: More Ground Truth Retrieving based Multimodal Interaction and Semantic Supervision for video description
Jiayu Zhang 0002, Pengjie Tang, Yunlan Tan, Hanli Wang |
Neural Networks | 4 |
| 2025 | Asynchronous Multi-Agent Collaborative Framework for Integrated Production and Procurement Optimization in RefineryabstractThe refinery industry operates as a highly complex system characterized by dynamic interactions among numerous processes, resources, and decision-making strategies. Within this context, production planning and crude oil procurement management are critical to ensuring operational efficiency and profitability. However, traditional optimization approaches often neglect the distinct decision-making time scales of these two functions, limiting their ability to address market dynamics and supply chain complexities effectively. To overcome these challenges, this study introduces an asynchronous multi-agent collaborative optimization framework that enhances the coordination between production planning and crude oil procurement in refinery operations. By enabling production and procurement agents to operate on independent time scales, the framework adapts to fluctuations in crude oil prices and variations in product demand. The study further extends the classical multi-agent reinforcement learning algorithm into an asynchronous paradigm, introducing Async-MAPPO, which allows agents to independently optimize decisions using real-time data. This approach mitigates operational delays and enhances adaptability. Experimental evaluations demonstrate that the proposed method significantly improves production efficiency, reduces operational costs, and strengthens the refinery’s resilience to market fluctuations, underscoring the critical role of asynchronous optimization in refinery operations. Kai Wang 0024, Fei Qiao, Hanli Wang |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2025 | Rethinking Artifact Mitigation in HDR Reconstruction: From Detection to OptimizationabstractArtifact remains a long-standing challenge in High Dynamic Range (HDR) reconstruction. Existing methods focus on model designs for artifact mitigation but ignore explicit detection and suppression strategies. Because artifact lacks clear boundaries, distinct shapes, and semantic consistency, and there is no existing dedicated dataset for HDR artifact, progress in direct artifact detection and recovery is impeded. To bridge the gap, we propose a unified HDR reconstruction framework that integrates artifact detection and model optimization. Firstly, we build the first HDR artifact dataset (HADataset), comprising 1,213 diverse multi-exposure Low Dynamic Range (LDR) image sets and 1,765 HDR image pairs with per-pixel artifact annotations. Secondly, we develop an effective HDR artifact detector (HADetector), a robust artifact detection model capable of accurately localizing HDR reconstruction artifact. HADetector plays two pivotal roles: (1) enhancing existing HDR reconstruction models through fine-tuning, and (2) serving as a non-reference image quality assessment (NR-IQA) metric, the Artifact Score (AS), which aligns closely with human visual perception for reliable quality evaluation. Extensive experiments validate the effectiveness and generalizability of our framework, including the HADataset, HADetector, fine-tuning paradigm, and AS metric. The code and datasets are available at: https://github.com/xinyueliii/hdr-artifact-detect-optimize. Zhangkai Ni, Wenhan Yang, Hanli Wang, Lianghua He, Sam Kwong |
IEEE Trans. Image Process. | 5 |
| 2025 | Structural Similarity-Inspired Unfolding for Lightweight Image Super-ResolutionabstractMajor efforts in data-driven image super-resolution (SR) primarily focus on expanding the receptive field of the model to better capture contextual information. However, these methods are typically implemented by stacking deeper networks or leveraging transformer-based attention mechanisms, which consequently increases model complexity. In contrast, model-driven methods based on the unfolding paradigm show promise in improving performance while effectively maintaining model compactness through sophisticated module design. Based on these insights, we propose a Structural Similarity-Inspired Unfolding (SSIU) method for efficient image SR. This method is designed through unfolding an SR optimization function constrained by structural similarity, aiming to combine the strengths of both data-driven and model-driven approaches. Our model operates progressively following the unfolding paradigm. Each iteration consists of multiple Mixed-Scale Gating Modules (MSGM) and an Efficient Sparse Attention Module (ESAM). The former implements comprehensive constraints on features, including a structural similarity constraint, while the latter aims to achieve sparse activation. In addition, we design a Mixture-of-Experts-based Feature Selector (MoE-FS) that fully utilizes multi-level feature information by combining features from different steps. Extensive experiments validate the efficacy and efficiency of our unfolding-inspired network. Our model outperforms current state-of-the-art models, boasting lower parameter counts and reduced memory consumption. Our code will be available at: https://github.com/eezkni/SSIU. Zhangkai Ni, Wenhan Yang, Hanli Wang, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Image Process. | 4 |
| 2025 | Shell-Guided Compression of Voxel Radiance FieldsabstractIn this paper, we address the challenge of significant memory consumption and redundant components in large-scale voxel-based model, which are commonly encountered in real-world 3D reconstruction scenarios. We propose a novel method called Shell-guided compression of Voxel Radiance Fields (SVRF), aimed at optimizing voxel-based model into a shell-like structure to reduce storage costs while maintaining rendering accuracy. Specifically, we first introduce a Shell-like Constraint, operating in two main aspects: 1) enhancing the influence of voxels neighboring the surface in determining the rendering outcomes, and 2) expediting the elimination of redundant voxels both inside and outside the surface. Additionally, we introduce an Adaptive Thresholds to ensure appropriate pruning criteria for different scenes. To prevent the erroneous removal of essential object parts, we further employ a Dynamic Pruning Strategy to conduct smooth and precise model pruning during training. The compression method we propose does not necessitate the use of additional labels. It merely requires the guidance of self-supervised learning based on predicted depth. Furthermore, it can be seamlessly integrated into any voxel-grid-based method. Extensive experimental results demonstrate that our method achieves comparable rendering quality while compressing the original number of voxel grids by more than 70%. Our code will be available at: https://github.com/eezkni/SVRF. Peiqi Yang, Zhangkai Ni, Hanli Wang, Wenhan Yang, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Image Process. | 3 |
| 2025 | M2Trans: Multi-Modal Regularized Coarse-to-Fine Transformer for Ultrasound Image Super-ResolutionabstractUltrasound image super-resolution (SR) aims to transform low-resolution images into high-resolution ones, thereby restoring intricate details crucial for improved diagnostic accuracy. However, prevailing methods relying solely on image modality guidance and pixel-wise loss functions struggle to capture the distinct characteristics of medical images, such as unique texture patterns and specific colors harboring critical diagnostic information. To overcome these challenges, this paper introduces the Multi-Modal Regularized Coarse-to-fine Transformer (M2Trans) for Ultrasound Image SR. By integrating the text modality, we establish joint image-text guidance during training, leveraging the medical CLIP model to incorporate richer priors from text descriptions into the SR optimization process, enhancing detail, structure, and semantic recovery. Furthermore, we propose a novel coarse-to-fine transformer comprising multiple branches infused with self-attention and frequency transforms to efficiently capture signal dependencies across different scales. Extensive experimental results demonstrate significant improvements over state-of-the-art methods on benchmark datasets, including CCA-US, US-CASE, and our newly created dataset MMUS1K, with a minimum improvement of 0.17dB, 0.30dB, and 0.28dB in terms of PSNR. Zhangkai Ni, Runyu Xiao, Wenhan Yang, Hanli Wang, Zhihua Wang 0002, Lihua Xiang |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | Similarity Shuffled Criss-Cross Transformer With Angle Loss for Image-Text MatchingabstractImage-text matching aims to retrieve images from the guidance of textual queries or retrieve text expressions with the help of images. Existing Transformer-based methods compute attention for all tokens and thus suffer from redundant information, resulting in inadequate focus on salient features. On the other hand, the widely adopted bidirectional ranking loss overlooks the importance of expanding the distance between positive and negative samples, leading to the misclassification of negative samples as positive ones. In this work, we propose similarity shuffled criss-cross Transformer (SSCT) with angle loss for image-text matching. Specifically, a grouping-shuffling operation is introduced to better distinguish salient features from redundant information, bypassing the need for fully connected mapping. The grouping-shuffling operation establishes channel dependencies across different groups of feature representations, enhancing salient features while suppressing unimportant ones. Then, a criss-cross attention mechanism that equips self-attention with a novel criss-cross convolution is designed to make isolated information cooperatively express integral semantics. Moreover, a novel angle loss is introduced to expand the distances between positive and negative samples. Extensive experiments on the benchmark datasets of MSCOCO and Flickr30K demonstrate that the proposed methods achieve superior performances compared to state-of-the-art methods. Taiyi Su, Hanli Wang, Zhangkai Ni |
IEEE Trans. Multim. | 3 |
| 2025 | Vision-Language Relational Transformer for Video-to-Text GenerationabstractVideo-to-text generation is a challenging task that involves translating video contents into accurate and expressive sentences. Existing methods often ignore the importance of establishing fine-grained semantics within visual representations and exploring textual knowledge implied by video contents, leading to difficulty in generating satisfactory sentences. To address these problems, a vision-language relational transformer model is proposed for video-to-text generation. Three key novel aspects are investigated. First, a visual relation modeling block is designed to obtain higher-order feature representations and establish semantic relationships between regional and global features. Second, a knowledge attention block is developed to explore hierarchical textual information and capture cross-modal dependencies. Third, a video-centric conversation system is constructed to complete multi-round dialogues by incorporating the proposed modules including visual relation modeling, knowledge attention and text generation. Extensive experiments on five benchmark datasets including MSVD, MSRVTT, ActivityNet, Charades and EMVPC demonstrate that the proposed scheme achieves remarkable performance compared with the state-of-the-art methods. Besides, the qualitative experiment reveals the system's favorable conversation capability and provides a valuable exemplar for future video understanding works. Tengpeng Li, Hanli Wang, Qinyu Li, Zhangkai Ni |
IEEE Trans. Multim. | 2 |
| 2025 | EIN: Exposure-Induced Network for Single-Image HDR ReconstructionabstractReconstructing high dynamic range (HDR) images from standard dynamic range (SDR) ones has received growing attention in recent years. A predominant problem of this task lies in the absence of texture and structural information in under/over-exposed regions. In this article, we propose an efficient and stable single-image HDR reconstruction method, namely exposure-induced network (EIN). More specifically, a dynamic range expansion branch (DB) is designed to expand the global dynamic range of the input SDR image. Moreover, two exposure-gated detail recovering branches for local over- (OB) and under- (UB) exposed regions are proposed to interact with the DB to progressively infer the texture and structural details with the learned confidence maps to resolve challenging ambiguities in such regions. The features from these three interactional branches are adaptively fused in the joint global–local decoder to reconstruct the final HDR image. The proposed network is trained based upon a large-scale dataset constructed with diverse content. Extensive experimental results demonstrate that the proposed model achieves consistent visual quality improvement for input SDR images with different exposures compared with state-of-the-art methods. The source code is available at: https://github.com/Yliu724/EIN . Zhangkai Ni, Peilin Chen 0001, Shiqi Wang 0001, Xinfeng Zhang 0001, Hanli Wang, Sam Kwong |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | ColNeRF: Collaboration for Generalizable Sparse Input Neural Radiance FieldabstractNeural Radiance Fields (NeRF) have demonstrated impressive potential in synthesizing novel views from dense input, however, their effectiveness is challenged when dealing with sparse input. Existing approaches that incorporate additional depth or semantic supervision can alleviate this issue to an extent. However, the process of supervision collection is not only costly but also potentially inaccurate. In our work, we introduce a novel model: the Collaborative Neural Radiance Fields (ColNeRF) designed to work with sparse input. The collaboration in ColNeRF includes the cooperation among sparse input source images and the cooperation among the output of the NeRF. Through this, we construct a novel collaborative module that aligns information from various views and meanwhile imposes self-supervised constraints to ensure multi-view consistency in both geometry and appearance. A Collaborative Cross-View Volume Integration module (CCVI) is proposed to capture complex occlusions and implicitly infer the spatial location of objects. Moreover, we introduce self-supervision of target rays projected in multiple directions to ensure geometric and color consistency in adjacent regions. Benefiting from the collaboration at the input and output ends, ColNeRF is capable of capturing richer and more generalized scene representation, thereby facilitating higher-quality results of the novel view synthesis. Our extensive experimental results demonstrate that ColNeRF outperforms state-of-the-art sparse input generalizable NeRF methods. Furthermore, our approach exhibits superiority in fine-tuning towards adapting to new scenes, achieving competitive performance compared to per-scene optimized NeRF-based methods while significantly reducing computational costs. Our code is available at: https://github.com/eezkni/ColNeRF. Zhangkai Ni, Peiqi Yang, Wenhan Yang, Hanli Wang, Lin Ma 0002, Sam Kwong |
AAAI | 4 |
| 2024 | Misalignment-Robust Frequency Distribution Loss for Image TransformationabstractThis paper aims to address a common challenge in deep learning-based image transformation methods, such as im-age enhancement and super-resolution, which heavily rely on precisely aligned paired datasets with pixel-level align-ments. However, creating precisely aligned paired images presents significant challenges and hinders the advance-ment of methods trained on such data. To overcome this challenge, this paper introduces a novel and simple frequency Distribution Loss (FDL) for computing distribution distance within the frequency domain. Specifically, we transform image features into the frequency domain using Discrete Fourier Transformation (DFT). Subsequently, frequency components (amplitude and phase) are processed separately to form the FDL loss function. Our method is empirically proven effective as a training constraint due to the thoughtful utilization of global information in the frequency domain. Extensive experimental evaluations, fo-cusing on image enhancement and super-resolution tasks, demonstrate that FDL outperforms existing misalignment-robust loss functions. Furthermore, we explore the poten-tial of our FDL for image style transfer that relies solely on completely misaligned data. Our code is available at: https://github.com/eezkni/FDL Zhangkai Ni, Juncheng Wu, Wenhan Yang, Hanli Wang, Lin Ma 0002 |
CVPR | 5 |
| 2024 | Memory-Based Contrastive Learning with Optimized Sampling for Incremental Few-Shot Semantic SegmentationabstractIncremental few-shot semantic segmentation (IFSS) aims to incrementally expand a semantic segmentation model’s ability to identify new classes based on few samples. However, it grapples with the dual challenges of catastrophic forgetting (due to feature drift in old classes) and overfitting (triggered by inadequate samples in new classes). To address these issues, a novel approach is proposed to integrate pixel-wise and region-wise contrastive learning, complemented by an optimized example and anchor sampling strategy. The proposed method incorporates a region memory and pixel memory designed to explore the high-dimensional embedding space more effectively. The memory, retaining the feature embeddings of known classes, facilitates the calibration and alignment of seen class features during the learning process of new classes. To further mitigate overfitting, the proposed approach implements an optimized example and anchor sampling strategy. Extensive experiments show the competitive performance of the proposed method. The source code of this work can be found in https://mic.tongji.edu.cn. Miaojing Shi, Taiyi Su, Hanli Wang |
ISCAS | 4 |
| 2024 | DDR: Exploiting Deep Degradation Response as Flexible Image DescriptorabstractImage deep features extracted by pre-trained networks are known to contain rich and informative representations. In this paper, we present Deep Degradation Response (DDR), a method to quantify changes in image deep features under varying degradation conditions. Specifically, our approach facilitates flexible and adaptive degradation, enabling the controlled synthesis of image degradation through text-driven prompts. Extensive evaluations demonstrate the versatility of DDR as an image descriptor, with strong correlations observed with key image attributes such as complexity, colorfulness, sharpness, and overall quality. Moreover, we demonstrate the efficacy of DDR across a spectrum of applications. It excels as a blind image quality assessment metric, outperforming existing methodologies across multiple datasets. Additionally, DDR serves as an effective unsupervised learning objective in image restoration tasks, yielding notable advancements in image deblurring and single-image super-resolution. Our code is available at: https://github.com/eezkni/DDR. Juncheng Wu, Zhangkai Ni, Hanli Wang, Wenhan Yang, Yuyin Zhou, Shiqi Wang 0001 |
NeurIPS | 3 |
| 2024 | Emotion recognition in user-generated videos with long-range correlation-aware networkabstractAbstract Emotion recognition in user‐generated videos plays an essential role in affective computing. In general, visual information directly affects human emotions, so the visual modality is significant for emotion recognition. Most classic approaches mainly focus on local temporal information of videos, which potentially restricts their capacity to encode the correlation of long‐range context. To address this issue, a novel network is proposed to recognize emotions in videos. To be specific, a spatio‐temporal correlation‐aware block is designed to depict the long‐range correlations between input tokens, where the convolutional layers are used to learn the local correlations and the inter‐image cross‐attention is designed to learn the long‐range and spatio‐temporal correlations between input tokens. To generate diverse and challenging samples, a dual‐augmentation fusion layer is devised, which fuses each frame with its corresponding frame in the temporal domain. To produce rich video clips, a long‐range sampling layer is designed, which generates clips in a wide range of spatial and temporal domains. Extensive experiments are conducted on two challenging video emotion datasets, namely VideoEmotion‐8 and Ekman‐6. The experimental results demonstrate that the proposed method obtains better performance than baseline methods. Moreover, the proposed method achieves state‐of‐the‐art results on the two datasets. The source code of the proposed network is available at: https://github.com/JinChow/LRCANet . Yun Yi, Hanli Wang, Pengjie Tang, Min Wang 0020 |
IET Image Process. | 3 |
| 2024 | Transductive Learning With Prior Knowledge for Generalized Zero-Shot Action RecognitionabstractIt is challenging to achieve generalized zero-shot action recognition. Different from the conventional zero-shot tasks which assume that the instances of the source classes are absent in the test set, the generalized zero-shot task studies the case that the test set contains both the source and the target classes. Due to the gap between visual feature and semantic embedding as well as the inherent bias of the learned classifier towards the source classes, the existing generalized zero-shot action recognition approaches are still far less effective than traditional zero-shot action recognition approaches. Facing these challenges, a novel transductive learning with prior knowledge (TLPK) model is proposed for generalized zero-shot action recognition. First, TLPK learns the prior knowledge which assists in bridging the gap between visual features and semantic embeddings, and preliminarily reduces the bias caused by the visual-semantic gap. Then, a transductive learning method that employs unlabeled target data is designed to overcome the bias problem in an effective manner. To achieve this, a target semantic-available approach and a target semantic-free approach are devised to utilize the target semantics in two different ways, where the target semantic-free approach exploits prior knowledge to produce well-performed semantic embeddings. By exploring the usage of the aforementioned prior-knowledge learning and transductive learning strategies, TLPK significantly bridges the visual-semantic gap and alleviates the bias between the source and the target classes. The experiments on the benchmark datasets of HMDB51 and UCF101 demonstrate the effectiveness of the proposed model compared to the state-of-the-art methods. The source code of this work can be found inhttps://mic.tongji.edu.cn Taiyi Su, Hanli Wang, Qiuping Qi, Bin He 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Multi-Modal Large Language Model Enhanced Pseudo 3D Perception Framework for Visual Commonsense ReasoningabstractThe visual commonsense reasoning (VCR) task is to choose an answer and provide a justifying rationale based on the given image and textural question. Representative works first recognize objects in images and then associate them with key words in texts. However, existing approaches do not consider exact positions of objects in a human-like three-dimensional (3D) manner, making them incompetent to accurately distinguish objects and understand visual relation. Recently, multi-modal large language models (MLLMs) have been used as powerful tools for several multi-modal tasks but not for VCR yet, which requires elaborate reasoning on specific visual objects referred by texts. In light of the above, an MLLM enhanced pseudo 3D perception framework is designed for VCR. Specifically, we first demonstrate that the relation between objects is relevant to object depths in images, and hence introduce object depth into VCR frameworks to infer 3D positions of objects in images. Then, a depth-aware Transformer is proposed to encode depth differences between objects into the attention mechanism of Transformer to discriminatively associate objects with visual scenes guided by depth. To further associate the answer with the depth of visual scene, each word in the answer is tagged with a pseudo depth to realize depth-aware association between answer words and objects. On the other hand, BLIP-2 as an MLLM is employed to process images and texts, and the referring expressions in texts involving specific visual objects are modified with linguistic object labels to serve as comprehensible MLLM inputs. Finally, a parameter optimization technique is devised to fully consider the quality of data batches based on multi-level reasoning confidence. Experiments on the VCR dataset demonstrate the superiority of the proposed framework over state-of-the-art approaches. The source code of this work can be found inhttps://mic.tongji.edu.cn. Jian Zhu 0006, Hanli Wang, Miaojing Shi |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Self-Supervised Video Representation Learning by Serial Restoration With Elastic ComplexityabstractSelf-supervised video representation learning leaves out heavy manual annotation by automatically excavating supervisory signals. Although contrastive learning based approaches exhibit superior performances, pretext task based approaches still deserve further study. This is because the pretext tasks exploit the nature of data and encourage feature extractors to learn spatiotemporal logic by discovering dependencies among video clips or cubes, without manual engineering on data augmentations or manual construction of contrastive pairs. To utilize chronological property more effectively and efficiently, this work proposes a novel pretext task, named serial restoration of shuffled clips (SRSC), disentangled by an elaborately designed task network composed of an order-aware encoder and a serial restoration decoder. In contrast to other order based pretext tasks that formulate clip order recognition as a one-step classification problem, the proposed SRSC task restores shuffled clips into the right order in multiple steps. Owing to the excellent elasticity of SRSC, a novel taxonomy of curriculum learning is further proposed to equip SRSC with different pre-training strategies. According to the factors that affect the complexity of solving the SRSC task, the proposed curriculum learning strategies can be categorized into task based, model based and data based. Extensive experiments are conducted on the subdivided strategies to explore their effectiveness and noteworthy laws. Compared with existing approaches, this work demonstrates that the proposed approach achieves state-of-the-art performances in pretext task based self-supervised video representation learning and a majority of the proposed strategies further boost the performance of downstream tasks. For the first time, the features pre-trained by the pretext tasks are applied to video captioning by feature-level early fusion, and enhance the input of existing approaches as a lightweight plugin. Hanli Wang, Chang Wen Chen |
IEEE Trans. Multim. | 2 |
| 2024 | Hybrid Graph Reasoning With Dynamic Interaction for Visual DialogabstractAs a pivotal branch of intelligent human-computer interaction, visual dialog is a technically challenging task that requires artificial intelligence (AI) agents to answer consecutive questions based on image content and history dialog. Despite considerable progresses, visual dialog still suffers from two major problems: (1) how to design flexible cross-modal interaction patterns instead of over-reliance on expert experience and (2) how to infer underlying semantic dependencies between dialogues effectively. To address these issues, an end-to-end framework employing dynamic interaction and hybrid graph reasoning is proposed in this work. Specifically, three major components are designed and the practical benefits are demonstrated by extensive experiments. First, a dynamic interaction module is developed to automatically determine the optimal modality interaction route for multifarious questions, which consists of three elaborate functional interaction blocks endowed with dynamic routers. Second, a hybrid graph reasoning module is designed to explore adequate semantic associations between dialogues from multiple perspectives, where the hybrid graph is constructed by aggregating a structured coreference graph and a context-aware temporal graph. Third, a unified one-stage visual dialog model with an end-to-end structure is developed to train the dynamic interaction module and the hybrid graph reasoning module in a collaborative manner. Extensive experiments on the benchmark datasets of VisDial v0.9 and VisDial v1.0 demonstrate the effectiveness of the proposed method compared to other state-of-the-art approaches. The source code of this work can be found inhttps://mic.tongji.edu.cn. Shanshan Du, Hanli Wang, Tengpeng Li, Chang Wen Chen |
IEEE Trans. Multim. | 2 |
| 2024 | Opinion-Unaware Blind Image Quality Assessment Using Multi-Scale Deep Feature StatisticsabstractDeep learning-based methods have significantly influenced the blind image quality assessment (BIQA) field, however, these methods often require training using large amounts of human rating data. In contrast, traditional knowledge-based methods are cost-effective for training but face challenges in effectively extracting features aligned with human visual perception. To bridge these gaps, we propose integrating deep features from pre-trained visual models with a statistical analysis model into a Multi-scale Deep Feature Statistics (MDFS) model for achieving opinion-unaware BIQA (OU-BIQA), thereby eliminating the reliance on human rating data and significantly improving training efficiency. Specifically, we extract patch-wise multi-scale features from pre-trained vision models, which are subsequently fitted into a multivariate Gaussian (MVG) model. The final quality score is determined by quantifying the distance between the MVG model derived from the test image and the benchmark MVG model derived from the high-quality image set. A comprehensive series of experiments conducted on various datasets show that our proposed model exhibits superior consistency with human visual perception compared to state-of-the-art BIQA models. Furthermore, it shows improved generalizability across diverse target-specific BIQA tasks. Our code is available at:https://github.com/eezkni/MDFS Zhangkai Ni, Keyan Ding, Wenhan Yang, Hanli Wang, Shiqi Wang 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | Glow in the Dark: Low-Light Image Enhancement With External MemoryabstractDeep learning-based methods have achieved remarkable success with powerful modeling capabilities. However, the weights of these models are learned over the entire training dataset, which inevitably leads to the ignorance of sample specific properties in the learned enhancement mapping. This situation causes ineffective enhancement in the testing phase for the samples that differ significantly from the training distribution. In this paper, we introduce external memory to form an external memory-augmented network (EMNet) for low-light image enhancement. The external memory aims to capture the sample specific properties of the training dataset to guide the enhancement in the testing phase. Benefiting from the learned memory, more complex distributions of reference images in the entire dataset can be “remembered” to facilitate the adjustment of the testing samples more adaptively. To further augment the capacity of the model, we take the transformer as our baseline network, which specializes in capturing long-range spatial redundancy. Experimental results demonstrate that our proposed method has a promising performance and outperforms state-of-the-art methods. It is noted that, the proposed external memory is a plug-and-play mechanism that can be integrated with any existing method to further improve the enhancement quality. More practices of integrating external memory with other image enhancement methods are qualitatively and quantitatively analyzed. The results further confirm that the effectiveness of our proposed memory mechanism when combing with existing enhancement methods. Dongjie Ye, Zhangkai Ni, Wenhan Yang, Hanli Wang, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Multim. | 4 |
| 2024 | Multi-Modal Structure-Embedding Graph Transformer for Visual Commonsense ReasoningabstractVisual commonsense reasoning (VCR) is a challenging reasoning task that aims to not only answer the question based on a given image but also provide a rationale justifying for the choice. Graph-based networks are appropriate to represent and extract the correlation between image and language for reasoning, where how to construct and learn graphs based on such multi-modal Euclidean data is a fundamental problem. Most existing graph-based methods view visual regions and linguistic words as identical graph nodes, ignoring inherent characteristics of multi-modal data. In addition, these approaches typically only have one graph-learning layer, and the performance declines as the model goes deeper. To address these issues, a novel method named Multi-modal Structure-embedding Graph Transformer (MSGT) is proposed. Specifically, an answer-vision graph and an answer-question graph are constructed to represent and model intra-modal and inter-modal correlations in VCR simultaneously, where additional multi-modal structure representations are initialized and embedded according to visual region distances and linguistic word orders for more reasonable graph representation. Then, a structure-injecting graph transformer is designed to inject embedded structure priors into the semantic correlation matrix for the evolution of node features and structure representations, which can stack more layers to make model deeper and extract more powerful features with instructive priors. To adaptively fuse graph features, a scored pooling mechanism is further developed to select valuable clues for reasoning from learnt node features. Experiments demonstrate the superiority of the proposed MSGT framework compared with state-of-the-art methods on the VCR benchmark dataset. Jian Zhu 0006, Hanli Wang, Bin He 0003 |
IEEE Trans. Multim. | 2 |
| 2023 | Adaptive Token Excitation with Negative Selection for Video-Text Retrieval
Juntao Yu, Zhangkai Ni, Taiyi Su, Hanli Wang |
ICANN (7) | 4 |
| 2023 | Structure-Aware Generative Adversarial Network for Text-to-Image GenerationabstractText-to-image generation aims at synthesizing photo-realistic images from textual descriptions. Existing methods typically align images with the corresponding texts in a joint semantic space. However, the presence of the modality gap in the joint semantic space leads to misalignment. Meanwhile, the limited receptive field of the convolutional neural network leads to structural distortions of generated images. In this work, a structure-aware generative adversarial network (SaGAN) is proposed for (1) semantically aligning multimodal features in the joint semantic space in a learnable manner; and (2) improving the structure and contour of generated images by the designed content-invariant negative samples. Experimental results show that SaGAN achieves over 30.1% and 8.2% improvements in terms of FID on the datasets of CUB and COCO when compared with the state-of-the-art approaches. Zhangkai Ni, Hanli Wang |
ICIP | 3 |
| 2023 | Knowledge-Enriched Attention Network With Group-Wise Semantic for Visual StorytellingabstractAs a technically challenging topic, visual storytelling aims at generating an imaginary and coherent story with narrative multi-sentences from a group of relevant images. Existing methods often generate direct and rigid descriptions of apparent image-based contents, because they are not capable of exploring implicit information beyond images. Hence, these schemes could not capture consistent dependencies from holistic representation, impairing the generation of reasonable and fluent stories. To address these problems, a novel knowledge-enriched attention network with group-wise semantic model is proposed. Three main novel components are designed and supported by substantial experiments to reveal practical advantages. First, a knowledge-enriched attention network is designed to extract implicit concepts from external knowledge system, and these concepts are followed by a cascade cross-modal attention mechanism to characterize imaginative and concrete representations. Second, a group-wise semantic module with second-order pooling is developed to explore the globally consistent guidance. Third, a unified one-stage story generation model with encoder-decoder structure is proposed to simultaneously train and infer the knowledge-enriched attention network, group-wise semantic module and multi-modal story generation decoder in an end-to-end fashion. Substantial experiments on the visual storytelling datasets with both objective and subjective evaluation metrics demonstrate the superior performance of the proposed scheme as compared with other state-of-the-art methods. The source code of this work can be found in https://mic.tongji.edu.cn. Tengpeng Li, Hanli Wang, Bin He 0003, Chang Wen Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Coherent Visual Storytelling via Parallel Top-Down Visual and Topic AttentionabstractVisual storytelling aims at producing a narrative paragraph for a given photo album automatically. It introduces more new challenges than individual image paragraph descriptions, mainly due to the difficulty in preserving coherent topics and in generating diverse phrases to depict the rich content of a photo album. Existing attention-based models that lack higher-level guiding information always result in a deviation between the generated sentence and the topic expressed by the image. In addition, these widely applied language generation approaches employing standard beam search tend to produce monotonous descriptions. In this work, a coherent visual storytelling (CoVS) framework is designed to address the above-mentioned problems. Specifically, in the encoding phase, an image sequence encoder is designed to efficiently extract visual features of the input photo album. Then, the novel parallel top-down visual and topic attention (PTDVTA) decoder is constructed via a topic-aware neural network, a parallel top-down attention model, and a coherent language generator. Concretely, visual attention focuses on the attributes and the relationships of the objects, while topic attention integrating a topic-aware neural network could improve the coherence of generated sentences. Eventually, a phrase beam search algorithm with$n$-gram hamming diversity is further designed to optimize the expression diversity of the generated story. To justify the proposed CoVS framework, extensive experiments are conducted on the VIST dataset, which shows that CoVS can automatically generate coherent and diverse stories in a more natural way. Moreover, CoVS obtains better performance than state-of-the-art baselines on BLEU-4 and METEOR scores, while maintaining good CIDEr and ROUGH_L scores. The source code of this work can be found inhttps://mic.tongji.edu.cn. Jinjing Gu, Hanli Wang, Ruichao Fan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | High Dynamic Range Image Quality Assessment Based on Frequency DisparityabstractIn this paper, a novel and effective image quality assessment (IQA) algorithm based on frequency disparity for high dynamic range (HDR) images is proposed, termed as local-global frequency feature-based model (LGFM). Motivated by the assumption that the human visual system (HVS) is highly adapted for extracting structural information and partial frequencies when perceiving the visual scene, the Gabor and the Butterworth filters are applied to the luminance component of the HDR image to extract the local and global frequency features, respectively. The similarity measurement and feature pooling strategy are sequentially performed on the frequency features to obtain the predicted single quality score. The experiments evaluated on four widely used benchmarks demonstrate that the proposed LGFM can provide a higher consistency with the subjective perception compared with the state-of-the-art HDR IQA methods. Our code is available at:https://github.com/eezkni/LGFM. Zhangkai Ni, Shiqi Wang 0001, Hanli Wang, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Joint Graph Attention and Asymmetric Convolutional Neural Network for Deep Image CompressionabstractRecent deep image compression methods have achieved prominent progress by using nonlinear modeling and powerful representation capabilities of neural networks. However, most existing learning-based image compression approaches employ customized convolutional neural network (CNN) to utilize visual features by treating all pixels equally, neglecting the effect of local key features. Meanwhile, the convolutional filters in CNN usually express the local spatial relationship within the receptive field and seldom consider the long-range dependencies from distant locations. This results in the long-range dependencies of latent representations not being fully compressed. To address these issues, an end-to-end image compression method is proposed by integrating graph attention and asymmetric convolutional neural network (ACNN). Specifically, ACNN is used to strengthen the effect of local key features and reduce the cost of model training. Graph attention is introduced into image compression to address the bottleneck problem of CNN in modeling long-range dependencies. Meanwhile, regarding the limitation that existing attention mechanisms for image compression hardly share information, we propose a self-attention approach which allows information flow to achieve reasonable bit allocation. The proposed self-attention approach is in compliance with the perceptual characteristics of human visual system, as information can interact with each other via attention modules. Moreover, the proposed self-attention approach takes into account channel-level relationship and positional information to promote the compression effect of rich-texture regions. Experimental results demonstrate that the proposed method achieves state-of-the-art rate-distortion performances after being optimized by MS-SSIM compared to recent deep compression models on the benchmark datasets of Kodak and Tecnick. The project page with the source code can be found inhttps://mic.tongji.edu.cn. Zhisen Tang, Hanli Wang, Xiaokai Yi, Yun Zhang 0002, Sam Kwong, C.-C. Jay Kuo |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Multi-Level Content-Aware Boundary Detection for Temporal Action Proposal GenerationabstractIt is challenging to generate temporal action proposals from untrimmed videos. In general, boundary-based temporal action proposal generators are based on detecting temporal action boundaries, where a classifier is usually applied to evaluate the probability of each temporal action location. However, most existing approaches treat boundaries and contents separately, which neglect that the context of actions and the temporal locations complement each other, resulting in incomplete modeling of boundaries and contents. In addition, temporal boundaries are often located by exploiting either local clues or global information, without mining local temporal information and temporal-to-temporal relations sufficiently at different levels. Facing these challenges, a novel approach named multi-level content-aware boundary detection (MCBD) is proposed to generate temporal action proposals from videos, which jointly models the boundaries and contents of actions and captures multi-level (i.e., frame level and proposal level) temporal and context information. Specifically, the proposed MCBD preliminarily mines rich frame-level features to generate one-dimensional probability sequences, and further exploits temporal-to-temporal proposal-level relations to produce two-dimensional probability maps. The final temporal action proposals are obtained by a fusion of the multi-level boundary and content probabilities, achieving precise boundaries and reliable confidence of proposals. The extensive experiments on the three benchmark datasets of THUMOS14, ActivityNet v1.3 and HACS demonstrate the effectiveness of the proposed MCBD compared to state-of-the-art methods. The source code of this work can be found in https://mic.tongji.edu.cn. Taiyi Su, Hanli Wang |
IEEE Trans. Image Process. | 2 |
| 2023 | CSformer: Bridging Convolution and Transformer for Compressive SensingabstractConvolutional Neural Networks (CNNs) dominate image processing but suffer from local inductive bias, which is addressed by the transformer framework with its inherent ability to capture global context through self-attention mechanisms. However, how to inherit and integrate their advantages to improve compressed sensing is still an open issue. This paper proposes CSformer, a hybrid framework to explore the representation capacity of local and global features. The proposed approach is well-designed for end-to-end compressive image sensing, composed of adaptive sampling and recovery. In the sampling module, images are measured block-by-block by the learned sampling matrix. In the reconstruction stage, the measurements are projected into an initialization stem, a CNN stem, and a transformer stem. The initialization stem mimics the traditional reconstruction of compressive sensing but generates the initial reconstruction in a learnable and efficient manner. The CNN stem and transformer stem are concurrent, simultaneously calculating fine-grained and long-range features and efficiently aggregating them. Furthermore, we explore a progressive strategy and window-based transformer block to reduce the parameters and computational complexity. The experimental results demonstrate the effectiveness of the dedicated transformer-based architecture for compressive sensing, which achieves superior performance compared to state-of-the-art methods on different datasets. Our codes is available at: https://github.com/Lineves7/CSformer. Dongjie Ye, Zhangkai Ni, Hanli Wang, Jian Zhang 0018, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Image Process. | 3 |
| 2023 | Perceptually Weighted Rate Distortion Optimization for Video-Based Point Cloud CompressionabstractDynamic point cloud is a volumetric visual data representing realistic 3D scenes for virtual reality and augmented reality applications. However, its large data volume has been the bottleneck of data processing, transmission, and storage, which requires effective compression. In this paper, we propose a Perceptually Weighted Rate-Distortion Optimization (PWRDO) scheme for Video-based Point Cloud Compression (V-PCC), which aims to minimize the perceptual distortion of reconstructed point cloud at the given bit rate. Firstly, we propose a general framework of perceptually optimized V-PCC to exploit visual redundancies in point clouds. Secondly, a multi-scale Projection based Point Cloud quality Metric (PPCM) is proposed to measure the perceptual quality of 3D point cloud. The PPCM model comprises 3D-to-2D patch projection, multi-scale structural distortion measurement, and fusion model. Approximations and simplifications of the proposed PPCM are also presented for both V-PCC integration and low complexity. Thirdly, based on the simplified PPCM model, we propose a PWRDO scheme with Lagrange multiplier adaptation, which is incorporated into the V-PCC to enhance the coding efficiency. Experimental results show that the proposed PPCM models can be used as standalone quality metrics, and they are able to achieve higher consistency with the human subjective scores than the state-of-the-art objective visual quality metrics. Also, compared with the latest V-PCC reference model, the proposed PWRDO-based V-PCC scheme achieves an average bit rate reduction of 13.52%, 8.16%, 10.56% and 9.54%, respectively, in terms of four objective visual quality metrics for point clouds. It is significantly superior to the state-of-the-art coding algorithms. The computational complexity of the proposed PWRDO increases by 1.71% and 0.05% on average to the V-PCC encoder and decoder, respectively, which is negligible. The source codes of the PPCM and PWRDO schemes are available at https://github.com/VVCodec/PPCM-PWRDO. Yun Zhang 0002, Keqin Ding, Na Li 0015, Hanli Wang, Xiaoxia Huang 0004, C.-C. Jay Kuo |
IEEE Trans. Image Process. | 4 |
| 2023 | Task-Driven Video Compression for Humans and Machines: Framework Design and OptimizationabstractLearned video compression has developed rapidly and achieved impressive progress in recent years. Despite efficient compression performance, existing signal fidelity oriented or semantic fidelity oriented video compression methods limit the capability to meet the requirements of both machine and human vision. To address this problem, a task-driven video compression framework is proposed to flexibly support vision tasks for both human vision and machine vision. Specifically, to improve the compression performance, the backbone of the video compression framework is optimized by using three novel modules, including multi-scale motion estimation, multi-frame feature fusion, and reference based in-loop filters. Then, based on the proposed efficient compression backbone, a task-driven optimization approach is designed to achieve the trade-off between signal fidelity oriented compression and semantic fidelity oriented compression. Moreover, a post-filter module is employed for the framework to further improve the performance of the human vision branch. Finally, rate-distortion performance, rate-accuracy performance, and subjective quality are employed as the evaluation metrics, and experimental results show the superiority of the proposed framework for both human vision and machine vision. The source code of this work can be found inhttps://mic.tongji.edu.cn. Xiaokai Yi, Hanli Wang, Sam Kwong, C.-C. Jay Kuo |
IEEE Trans. Multim. | 2 |
| 2022 | Learned Image Compression with Multi-Scale Spatial and Contextual Information FusionabstractAlthough learned image compression based on convolution neural network and hyperprior makes significant progress, the distinction between original and reconstructed images is still obvious. In order to reconstruct compressed image with higher quality, a novel model based on fusing multi-scale spatial and context information is proposed in this work. Since spatial information might be dropped during the forward propagation when neural networks go deeper, a multi-scale information fusion module is designed to help the encoder to retain the necessary spatial information while removing the redundancy in latent representation. Meanwhile, a multi-scale 3D context module with varying-sized masked 3D convolution kernels is devised to obtain multi-scale correlation in latent representation. The experiments demonstrate the superiority of the proposed approach over a number of state-of-the-art image compression methods and the versatile video coding. Hanli Wang, Taiyi Su |
ICIP | 2 |
| 2022 | Discretized Gaussian Mixture Hyperprior for Learned Image Compression with Mask ModuleabstractLearned image compression approaches have shown great potential with promising results. However, according to the commonly used measurement methods, there still lies a performance gap between learned compression methods and the latest compression standard versatile video coding (VVC), because of the remaining redundancy existing in contemporary algorithms. To obtain a more accurate entropy model for rate estimation, discretized Gaussian mixture hyperprior is proposed in this work to parameterize the distribution of latent codes. In addition, a proposed mask module is exploited to enhance the feature extraction ability of the encoder and adaptively allocate bit rates. In this case, the proposed learned image compression model achieves the state-of-the-art performance among the existing learned compression methods and most compression standards on both Kodak24 and CLIC Mobile Validation datasets. The proposed model also outperforms VVC under several bitrate scenarios. Shengkai Wang, Hanli Wang |
ICME | 2 |
| 2022 | Spatio-temporal Super-resolution Network: Enhance Visual Representations for Video CaptioningabstractVideo captioning is a sequence-to-sequence task of automatically generating descriptions for given videos. Due to the diversity of video scenes, learning rich representations is critical for video captioning. However, previous works mainly exploited elaborate features but neglected the loss of information caused by frame sampling and image compression. In this paper, we propose a novel spatio-temporal super-resolution (STSR) network which is jointly trained for the video captioning task and the video super-resolution task in an end-to-end fashion. Specifically, a video super-resolution task consists of two subtasks: spatial super-resolution restores high-resolution image features while temporal super-resolution reconstructs missing frame features between two adjacent sampled frames. By sharing multi-modal encoders across both of these two tasks, STSR encourages encoders to capture salient visual contents and learn context-aware representations. Experiments on two benchmark datasets demonstrate that the proposed STSR boosts video captioning performances significantly and outperforms most state-of-the-art approaches. Quanhui Cao, Pengjie Tang, Hanli Wang |
ISCAS | 3 |
| 2022 | Multi-concept Mining for Video Captioning Based on Multiple TasksabstractVideo captioning is a challenging cross-modal task that requires taking full advantage of both vision and language. To identify objects in videos, object detectors are usually employed to extract high-level object-related features, but the fine-grained knowledge from the object detectors are often neglected. Also, there is a fact that not just the task of object detection has the ability to obtain additional knowledge for video understanding. In this paper, multiple tasks are assigned to fully mine multi-concept knowledge in both vision and language, including video-to-video knowledge, video-to-text knowledge and text-to-text knowledge. Moreover, since there is a strong synergy in knowledge, both of global and local word similarities are developed based on the text-to-text knowledge to boost the robustness of the mined semantic knowledge. The mined knowledge can offer the model an extra guidance apart from linguistic prior to generate more semantically appropriate and grammatically correct sentences. The experimental results on the benchmark MSVD and MSR-VTT datasets show that the proposed method makes remarkable improvement on all metrics on MSVD and two out of four metrics on MSR-VTT. Qinyu Zhang 0005, Pengjie Tang, Hanli Wang, Jinjing Gu |
ISCAS | 3 |
| 2022 | Cycle-Interactive Generative Adversarial Network for Robust Unsupervised Low-Light EnhancementabstractGetting rid of the fundamental limitations in fitting to the paired training data, recent unsupervised low-light enhancement methods excel in adjusting illumination and contrast of images. However, for unsupervised low light enhancement, the remaining noise suppression issue due to the lacking of supervision of detailed signal largely impedes the wide deployment of these methods in real-world applications. Herein, we propose a novel Cycle-Interactive Generative Adversarial Network (CIGAN) for unsupervised low-light image enhancement, which is capable of not only better transferring illumination distributions between low/normal-light images but also manipulating detailed signals between two domains, e.g., suppressing/synthesizing realistic noise in the cyclic enhancement/degradation process. In particular, the proposed low-light guided transformation feed-forwards the features of low-light images from the generator of enhancement GAN (eGAN) into the generator of degradation GAN (dGAN). With the learned information of real low-light images, dGAN can synthesize more realistic diverse illumination and contrast in low-light images. Moreover, the feature randomized perturbation module in dGAN learns to increase the feature randomness to produce diverse feature distributions, persuading the synthesized low-light images to contain realistic noise. Extensive experiments demonstrate both the superiority of the proposed method and the effectiveness of each module in CIGAN. Zhangkai Ni, Wenhan Yang, Hanli Wang, Shiqi Wang 0001, Lin Ma 0002, Sam Kwong |
ACM Multimedia | 3 |
| 2022 | Contrastive semantic similarity learning for image captioning evaluation
Chao Zeng 0005, Sam Kwong, Tiesong Zhao, Hanli Wang |
Inf. Sci. | 4 |
| 2022 | Meta-Learning-Based Incremental Few-Shot Object DetectionabstractRecent years have witnessed meaningful progress in the task of few-shot object detection. However, most of the existing models are not capable of incremental learning with a few samples,i.e., the detector can’t detect novel-class objects by using only a few samples of novel classes (without revisiting the original training samples) while maintaining the performances on base classes. This is largely because of catastrophic forgetting, which is a general phenomenon in few-shot learning that the incorporation of the unseen information (e.g., novel-class objects) will lead to a serious loss of the knowledge learnt before (e.g., base-class objects). In this paper, a new model is proposed for incremental few-shot object detection, which takes CenterNet as the fundamental framework and redesigns it by introducing a novel meta-learning method to make the model adapted to unseen knowledge while overcoming forgetting to a great extent. Specifically, a meta-learner is trained with the base-class samples, providing the object locator of the proposed model with a good weight initialization, and thus the proposed model can be fine-tuned easily with few novel-class samples. On the other hand, the filters correlated to base classes are preserved when fine-tuning the proposed model with the few samples of novel classes, which is a simple but effective solution to mitigate the problem of forgetting. The experiments on the benchmark MS COCO and PASCAL VOC datasets demonstrate that the proposed model outperforms the state-of-the-art methods by a large margin in the detection performances on base classes and all classes while achieving best performances when detecting novel-class objects in most cases. The project page can be found inhttps://mic.tongji.edu.cn/e6/d5/c9778a190165/page.htm. Hanli Wang, Yu Long 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Emotion Expression With Fact Transfer for Video DescriptionabstractTranslating a video into natural language is a fundamental but challenging task in visual understanding, since there is a great gap between visual content and linguistic sentence. More attention has been paid to this research field and a number of state-of-the-art results are achieved in recent years. However, the emotions in videos are usually overlooked, leading to the generated description sentences being boring and colorless. In this work, we construct a new dataset for video description with emotion expression, which consists of two parts: a re-annotated subset of the MSVD dataset with emotion embedded and another subset annotated with long sentences and rich emotions based on a video emotion recognition dataset. A fact transfer based framework is designed, which incorporates a fact stream and an emotion stream to generate sentences with emotion expression for video description. In addition, we propose a novel approach for sentence evaluation by balancing facts and emotions. A group of experiments are conducted, and the experimental results demonstrate the effectiveness of the proposed methods, including the idea of dataset construction for video description with emotion expression, model training and testing, and the emotion evaluation metric. The project page (including the code and dataset) can be found inhttps://mic.tongji.edu.cn/ce/70/c9778a183920/page.htm. Hanli Wang, Pengjie Tang, Qinyu Li |
IEEE Trans. Multim. | 1 |
| 2021 | Visual Storytelling with Hierarchical BERT Semantic GuidanceabstractVisual storytelling, which aims at automatically producing a narrative paragraph for photo album, remains quite challenging due to the complexity and diversity of photo album content. In addition, open-domain photo albums cover a broad range of topics and this results in highly variable vocabularies and expression styles to describe photo albums. In this work, a novel teacher-student visual storytelling framework with hierarchical BERT semantic guidance (HBSG) is proposed to address the above-mentioned challenges. The proposed teacher module consists of two joint tasks, namely, word-level latent topic generation and semantic-guided sentence generation. The first task aims to predict the latent topic of the story. As there is no ground-truth topic information, a pre-trained BERT model based on visual contents and annotated stories is utilized to mine topics. Then the topic vector is distilled to a designed image-topic prediction model. In the semantic-guided sentence generation task, HBSG is introduced for two purposes. The first is to narrow down the language complexity across topics, where the co-attention decoder with vision and semantic is designed to leverage the latent topics to induce topic-related language models. The second is to employ sentence semantic as an online external linguistic knowledge teacher module. Finally, an auxiliary loss is devised to transform linguistic knowledge into the language generation model. Extensive experiments are performed to demonstrate the effectiveness of HBSG framework, which surpasses the state-of-the-art approaches evaluated on the VIST test set. Ruichao Fan, Hanli Wang, Jinjing Gu |
MMAsia | 2 |
| 2021 | MABAN: Multi-Agent Boundary-Aware Network for Natural Language Moment RetrievalabstractThe amount of videos over the Internet and electronic surveillant cameras is growing dramatically, meanwhile paired sentence descriptions are significant clues to select attentional contents from videos. The task of natural language moment retrieval (NLMR) has drawn great interests from both academia and industry, which aims to associate specific video moments with the text descriptions figuring complex scenarios and multiple activities. In general, NLMR requires temporal context to be properly comprehended, and the existing studies suffer from two problems: (1) limited moment selection and (2) insufficient comprehension of structural context. To address these issues, a multi-agent boundary-aware network (MABAN) is proposed in this work. To guarantee flexible and goal-oriented moment selection, MABAN utilizes multi-agent reinforcement learning to decompose NLMR into localizing the two temporal boundary points for each moment. Specially, MABAN employs a two-phase cross-modal interaction to exploit the rich contextual semantic information. Moreover, temporal distance regression is considered to deduce the temporal boundaries, with which the agents can enhance the comprehension of structural context. Extensive experiments are carried out on two challenging benchmark datasets of ActivityNet Captions and Charades-STA, which demonstrate the effectiveness of the proposed approach as compared to state-of-the-art methods. The project page can be found in https://mic.tongji.edu.cn/e5/23/c9778a189731/page.htm. Hanli Wang, Bin He 0003 |
IEEE Trans. Image Process. | 2 |
| 2021 | CaptionNet: A Tailor-made Recurrent Neural Network for Generating Image DescriptionsabstractImage captioning is a challenging task of visual understanding and has drawn more attention of researchers. In general, two inputs are required at each time step by the Long Short-Term Memory (LSTM) network used in popular attention based image captioning frameworks, including image features and previous generated words. However, error will be accumulated if the previous words are not accurate and the related semantic is not efficient enough. Facing these challenges, a novel model named CaptionNet is proposed in this work as an improved LSTM specially designed for image captioning. Concretely, only attended image features are allowed to be fed into the memory of CaptionNet through input gates. In this way, the dependency on the previous predicted words can be reduced, forcing model to focus on more visual clues of images at the current time step. Moreover, a memory initialization method called image feature encoding is designed to capture richer semantics of the target image. The evaluation on the benchmark MSCOCO and Flickr30K datasets demonstrates the effectiveness of the proposed CaptionNet model, and extensive ablation studies are performed to verify each of the proposed methods. The project page can be found in https://mic.tongji.edu.cn/3f/9c/c9778a147356/page.htm. Longyu Yang, Hanli Wang, Pengjie Tang, Qinyu Li |
IEEE Trans. Multim. | 2 |
| 2021 | Perceptual Image Compression with Block-Level Just Noticeable Difference PredictionabstractA block-level perceptual image compression framework is proposed in this work, including a block-level just noticeable difference (JND) prediction model and a preprocessing scheme. Specifically speaking, block-level JND values are first deduced by utilizing the OTSU method based on the variation of block-level structural similarity values between two adjacent picture-level JND values in the MCL-JCI dataset. After the JND value for each image block is generated, a convolutional neural network–based prediction model is designed to forecast block-level JND values for a given target image. Then, a preprocessing scheme is devised to modify the discrete cosine transform coefficients during JPEG compression on the basis of the distribution of block-level JND values of the target test image. Finally, the test image is compressed by the max JND value across all of its image blocks in the light of the initial quality factor setting. The experimental results demonstrate that the proposed block-level perceptual image compression method is able to achieve 16.75% bit saving as compared to the state-of-the-art method with similar subjective quality. The project page can be found at https://mic.tongji.edu.cn/43/3f/c9778a148287/page.htm. Tao Tian, Hanli Wang, Sam Kwong, C.-C. Jay Kuo |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2020 | Enhanced Action Tubelet Detector for Spatio-Temporal Video Action DetectionabstractCurrent spatio-temporal action detection methods usually employ a two-stream architecture, a RGB stream for raw images and an auxiliary motion stream for optical flow. Training is required individually for each stream and more efforts are necessary to improve the precision of RGB stream. To this end, a single stream network named enhanced action tubelet (EAT) detector is proposed in this work based on RGB stream. A modulation layer is designed to modulate RGB features with conditional information from the visual clues of optical flow and human pose. This network is end-to-end and the proposed layer can be easily applied into other action detectors. Experiments show that EAT detector outperforms traditional RGB stream and is competitive to existing two-stream methods while free from the trouble of training streams separately. By being embedded in a new three-stream architecture, the resulting three-stream EAT detector achieves impressive performances among the best competitors on UCF-Sports, JHMDB and UCF-101. Yutang Wu, Hanli Wang, Shuheng Wang, Qinyu Li |
ICASSP | 2 |
| 2020 | Random Occlusion Recovery with Noise Channel for Person Re-identification
Di Wu 0030, Chang-an Yuan 0001, Xiao Qin 0005, Hongjie Wu, Xingming Zhao, Yuchuan Du, Hanli Wang |
ICIC (1) | 9 |
| 2020 | Patch assembly for real-time instance segmentationabstractThe paradigm of sliding window is proven effective for the task of visual instance segmentation in many popular research works. However, it still suffers from the bottleneck of inference time. To accelerate existing instance segmentation approaches which are dense sliding window based, this work introduces a novel approach, called patch assembly, which can be integrated into bounding box detectors for segmentation without extra up-sampling computations. A well-designed detector named PAMask is proposed to verify the effectiveness of the proposed approach. Benefitting from the simple structure as well as a fusion of multiple representations, PAMask has the ability to run in real time while achieving competitive performances. Besides, another effective technique called Center-NMS is designed to reduce the number of boxes for intersection of union calculation, which can be fully parallelized on device and contributes 0.6% mAP improvement both in detection and segmentation for free. Yutao Xu, Hanli Wang |
MMAsia | 2 |
| 2020 | Wonderful Clips of Playing Basketball: A Database for Localizing Wonderful Actions
Qinyu Li, Hanli Wang |
MMM (1) | 3 |
| 2020 | Structural Pyramid Network for Cascaded Optical Flow Estimation
Zefeng Sun, Hanli Wang, Yun Yi, Qinyu Li |
MMM (1) | 2 |
| 2020 | Evolutionary recurrent neural network for image captioning
Hanli Wang, Kaisheng Xu |
Neurocomputing | 2 |
| 2020 | RepeatPadding: Balancing words and sentence length for language comprehension in visual question answering
Yu Long 0003, Pengjie Tang, Zhihua Wei 0001, Jinjing Gu, Hanli Wang |
Inf. Sci. | 5 |
| 2020 | Affective Video Content Analysis With Adaptive Fusion Recurrent NetworkabstractAffective video content analysis is an important research topic in video content analysis and has extensive applications. Intuitively, multimodal features can depict elicited emotions, and the accumulation of temporal inputs influences the viewer's emotion. Although a number of research works have been proposed for this task, the adaptive weights of modalities and the correlation of temporal inputs are still not well studied. To address these issues, a novel framework is designed to learn the weights of modalities and temporal inputs from video data. Specifically, three network layers are designed, including statistical-data layer to improve the robustness of data, temporal-adaptive-fusion layer to fuse temporal inputs, and multimodal-adaptive-fusion layer to combine multiple modalities. In particular, the feature vectors of three input modalities are respectively extracted from three pre-trained convolutional neural networks and then fed to three statistical-data layers. Then, the output vectors of these three statistical-data layers are separately connected to three recurrent layers, and the corresponding outputs are fed to a fully-connected layer which shares parameters across modalities and temporal inputs. Finally, the outputs of the fully-connected layer are fused by the temporal-adaptive-fusion layer and then combined by the multimodal-adaptive-fusion layer. To discover the correlation of both multiple modalities and temporal inputs, adaptive weights of modalities and temporal inputs are introduced into loss functions for model training, and these weights are learned by an optimization algorithm. Extensive experiments are conducted on two challenging datasets, which demonstrate that the proposed method achieves better performances than baseline and other state-of-the-art methods. Yun Yi, Hanli Wang, Qinyu Li |
IEEE Trans. Multim. | 2 |
| 2019 | P3D-CTN: Pseudo-3D Convolutional Tube Network for Spatio-Temporal Action Detection in VideosabstractThe spatial independence and temporal continuity of video data as a whole are not fully investigated for video action detection. To tackle this issue, a deep network architecture is proposed, named Pseudo-3D Convolutional Tube Network (P3D-CTN). In particular, the proposed P3D-CTN integrates the frame-based two-dimensional convolutional module with the P3D convolutional module to balance the spatial and temporal information, and generates deeper features about human actions. Evaluations on two benchmark datasets (i.e., UCF-Sports and J-HMDB) demonstrate that the proposed P3D-CTN has superior performances in the task of action label prediction and yields state-of-the-art results for spatio-temporal action detection. Jiangchuan Wei, Hanli Wang, Yun Yi, Qinyu Li, De-Shuang Huang |
ICIP | 2 |
| 2019 | Perceptual Video Coding with Block-Level Staircase Just Noticeable DistortionabstractPerceptual video coding (PVC) is able to improve video compression efficiency by employing just noticeable distortion (JND) models. However, there are limitations of conventional JND models on simulating complex human visual system. To address this issue, a novel PVC framework is proposed in this work, in which a JND model based on staircase perceptual characteristics is designed to calculate block-level JND (BLJND) levels and a convolutional neural network based predictive model is developed to predict BLJND levels for video coding. Experimental results demonstrate that the proposed PVC framework is effective and robust in terms of video compression efficiency and subjective video quality. Hanli Wang, Tao Tian |
ICIP | 2 |
| 2019 | Swell-and-Shrink: Decomposing Image Captioning by Transformation and SummarizationabstractImage captioning is currently viewed as a problem analogous to machine translation. However, it always suffers from poor interpretability, coarse or even incorrect descriptions on regional details. Moreover, information abstraction and compression, as essential characteristics of captioning, are always overlooked and seldom discussed. To overcome the shortcomings, a swell-shrink method is proposed to redefine image captioning as a compositional task which consists of two separated modules: modality transformation and text compression. The former is guaranteed to accurately transform adequate visual content into textual form while the latter consists of a hierarchical LSTM which particularly emphasizes on removing the redundancy among multiple phrases and organizing the final abstractive caption. Additionally, the order and quality of region of interest and modality processing are studied to give insights of better understanding the influence of regional visual cues on language forming. Experiments demonstrate the effectiveness of the proposed method. Hanli Wang, Kaisheng Xu |
IJCAI | 2 |
| 2019 | Data Driven Regularization for Convolutional Neural Networks on Image ClassificationabstractDeep Convolutional Neural Network has shown significant improvements in many fields of computer vision, and a series of researches are proposed to explore advanced model structures to attenuate the problem of over-fitting. In this paper, two data driven techniques are designed including SwitchNode and SwitchConnect, which employ the sparsity of deterministic data to regularize convolutional neural network models. Specifically, the proposed SwitchNode method switches from the redundant nodes which have similar activations and spatial information to new initialization nodes, while the SwitchConnect method retrains replaceable convolutional kernels. The effectiveness of the proposed data driven regularization methods has been verified by the performance gain experimented on several benchmark image classification datasets. Hanli Wang, Qinyu Li, De-Shuang Huang |
ISCAS | 2 |
| 2019 | Multi-Dilation Network for Crowd CountingabstractWith the growth of urban population, crowd analysis has become an important and necessary task in the field of computer vision. The goal of crowd counting, which is a subfield of crowd analysis, is to count the number of people in an image or a zone of a picture. Due to the problems like heavy occlusions, perspective and luminous intensity variations, it is still extremely challenging to achieve crowd counting. Recent state-of-the-art approaches are mainly designed with convolutional neural networks to generate density maps. In this work, Multi-Dilation Network (MDNet) is proposed to solve the problem of crowd counting in congested scenes. The MDNet is made up of two parts: a VGG-16 based front end for feature extraction and a back end containing multi-dilation blocks to generate density maps. Especially, a multi-dilation block has four branches which are used to collect features in different sizes. By using dilated convolutional operations, the multi-dilation block could obtain various features while the maximum kernel size is still 3 x 3. The experiments on two challenging crowd counting datasets, UCF_CC_50 and ShanghaiTech, have shown that the proposed MDNet achieves better performances than other state-of-the-art methods, with a lower mean absolute error and mean squared error. Comparing to the network with multi-scale blocks which adopt larger kernels to extract features, MDNet still gains competitive performances with fewer model parameters. Shuheng Wang, Hanli Wang, Qinyu Li |
MMAsia | 2 |
| 2019 | Multi-modal learning for affective content analysis in movies
Yun Yi, Hanli Wang |
Multim. Tools Appl. | 2 |
| 2019 | WLDISR: Weighted Local Sparse Representation-Based Depth Image Super-Resolution for 3D Video SystemabstractIn this paper, we propose a Weighted Local sparse representation based Depth Image Super-Resolution (WLDISR) schemes aiming at improving the Virtual View Image (VVI) quality of 3D video system. Different from color images, depth images are mainly used to provide geometrical information in synthesizing VVI. Due to the view synthesis characteristics difference between textural structures and smooth regions of depth images, we divide the depth images into edge and smooth patches and learn two local dictionaries, respectively. Meanwhile, the weight term is derived and incorporated explicitly in the cost function to denote different importance of edge structures and smooth regions to the VVI quality. Then, local sparse representation and weighted sparse representation are jointly used in both dictionary learning and reconstruction phases in depth image super-resolution. Based on different optimizations on learning and reconstruction modules, three WLDISR schemes, WLDISR-D, WLDISR-R, and WLDISR-ALL, are proposed. Experimental results on 3D sequences demonstrate that the proposed WLDISR-D, WLDISR-R, and WLDISR-ALL schemes can achieve more than 1.9-, 2.03-, and 2.16-dB gains on average, respectively, in terms of the VVIs' quality, as compared with the state-of-the-art schemes. In addition, the visual quality of VVIs is also improved. Huan Zhang 0008, Yun Zhang 0002, Hanli Wang, Yo-Sung Ho, Shengzhong Feng |
IEEE Trans. Image Process. | 3 |
| 2019 | A Multi-Grained Parallel Solution for HEVC Encoding on Heterogeneous PlatformsabstractTo improve the parallel processing capability of video coding, the emerging high efficiency video coding (HEVC) standard introduces two parallel techniques, i.e., Wavefront Parallel Processing (WPP) andTiles, to make it much more parallel-friendly than its predecessors. However, these two techniques are designed to explore coarse-grained parallelism in HEVC encoding on multicore Central Processing Unit (CPU) platforms. As the computing architecture undergoes a trend toward heterogeneity in the last decade, multi-grained parallel computing methods can be designed to accelerate HEVC encoding on heterogeneous systems. In this paper, a multi-grained parallel solution (MPS) is proposed to optimize HEVC encoding on a typical heterogeneous platform. A massively parallel motion estimation algorithm is employed by MPS to parallelize part of HEVC encoding on Graphic Processing Unit (GPU). Meanwhile, several other HEVC encoding modules are accelerated on CPU through the cooperation of WPP and an adaptive parallel mode decision algorithm. The parallelism between CPU and GPU is well designed and implemented to guarantee an efficient concurrent execution of HEVC encoding on multi-grained parallel levels. The effectiveness of the proposed MPS for HEVC encoding is verified on a number of experiments. Bo Xiao 0005, Hanli Wang, Jun Wu 0006, Sam Kwong, C.-C. Jay Kuo |
IEEE Trans. Multim. | 2 |
| 2019 | Rich Visual and Language Representation with Complementary Semantics for Video CaptioningabstractIt is interesting and challenging to translate a video to natural description sentences based on the video content. In this work, an advanced framework is built to generate sentences with coherence and rich semantic expressions for video captioning. A long short term memory (LSTM) network with an improved factored way is first developed, which takes the inspiration of LSTM with a conventional factored way and a common practice to feed multi-modal features into LSTM at the first time step for visual description. Then, the incorporation of the LSTM network with the proposed improved factored way and un-factored way is exploited, and a voting strategy is utilized to predict candidate words. In addition, for robust and abstract visual and language representation, residuals are employed to enhance the gradient signals that are learned from the residual network (ResNet), and a deeper LSTM network is constructed. Furthermore, three convolutional neural network based features extracted from GoogLeNet, ResNet101, and ResNet152, are fused to catch more comprehensive and complementary visual information. Experiments are conducted on two benchmark datasets, including MSVD and MSR-VTT2016, and competitive performances are obtained by the proposed techniques as compared to other state-of-the-art methods. Pengjie Tang, Hanli Wang, Qinyu Li |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2018 | Categorizing Concepts With Basic Level for Vision-to-LanguageabstractVision-to-language tasks require a unified semantic understanding of visual content. However, the information contained in image/video is essentially ambiguous on two perspectives manifested on the diverse understanding among different persons and the various understanding grains even for the same person. Inspired by the basic level in early cognition, a Basic Concept (BaC) category is proposed in this work that contains both consensus and proper level of visual content to help neural network tackle the above problems. Specifically, a salient concept category is firstly generated by intersecting the labels of ImageNet and the vocabulary of MSCOCO dataset. Then, according to the observation from human early cognition that children make fewer mistakes on the basic level, the salient category is further refined by clustering concepts with a defined confusion degree which measures the difficulty for convolutional neural network to distinguish class pairs. Finally, a pre-trained model based on GoogLeNet is produced with the proposed BaC category of 1,372 concept classes. To verify the effectiveness of the proposed categorizing method for vision-to-language tasks, two kinds of experiments are performed including image captioning and visual question answering with the benchmark datasets of MSCOCO, Flickr30k and COCO-QA. The experimental results demonstrate that the representations derived from the cognition-inspired BaC category promote representation learning of neural networks on vision-to-language tasks, and a performance improvement is gained without modifying standard models. Hanli Wang, Kaisheng Xu |
CVPR | 2 |
| 2018 | Image Captioning with Word Level AttentionabstractImage captioning is an attractive and challenging task to perform automatic image description and a number of works are designed for this task. Most of these researches are based on convolutional neural network (CNN) and recurrent neural network (RNN), where the primary input to language model for word prediction at the current time step is usually the linguistic word generated at the previous time step. In this work, a novel word level attention layer is designed to process image features with two modules for accurate word prediction. The first is a bidirectional spatial embedding module to handle feature maps, then the second module employs attention mechanism to extract word level attention which will be fed into language model. The experimental results on the benchmark MSCOCO dataset demonstrate that the proposed model achieves the state-of-the-art performances with 106.0 on CIDEr and 34.0 on B-4, respectively. Hanli Wang, Pengjie Tang |
ICIP | 2 |
| 2018 | Refining Attention: A Sequential Attention Model for Image CaptioningabstractVisual attention is widely applied to image captioning. Previous works put visual attention and linguistic word into a long short-term memory network together, but neglect the sequential relation of attention at different time steps during word prediction. Moreover, the abstraction degree of visual attention is usually different from that of linguistic word. To address these issues, a sequential attention model is proposed in this work to handle visual attention by considering the corresponding sequential relation, and hence the internal relation among attention at each word prediction step is well utilized to enhance the visual information during sentence decoding. The experimental results on the benchmark MSCOCO and Flickr30K datasets show that the proposed model achieves excellent performances with 108.1 and 34.9 respectively on the evaluation criteria of CIDEr and BLEU-4 for MSCOCO. Qinyu Li, Hanli Wang, Pengjie Tang |
ICME | 3 |
| 2018 | Large-scale video compression: recent advances and challenges
Tao Tian, Hanli Wang |
Frontiers Comput. Sci. | 2 |
| 2018 | Deep sequential fusion LSTM network for image description
Pengjie Tang, Hanli Wang, Sam Kwong |
Neurocomputing | 2 |
| 2018 | Looking deeper and transferring attention for image captioning
Hanli Wang, Pengjie Tang |
Multim. Tools Appl. | 2 |
| 2018 | Maximal granularity structure and generalized multi-view discriminant analysis for person re-identification
Cairong Zhao, Xuekuan Wang, Duoqian Miao 0001, Hanli Wang, Wei-Shi Zheng 0001, Yong Xu 0001, David Zhang 0001 |
Pattern Recognit. | 4 |
| 2018 | Real-Time Action Recognition With Deeply Transferred Motion Vector CNNsabstractThe two-stream CNNs prove very successful for video based action recognition. However the classical two-stream CNNs are time costly, mainly due to the bottleneck of calculating optical flows. In this paper, we propose a two-stream based real-time action recognition approach by using motion vector to replace optical flow. Motion vectors are encoded in video stream and can be extracted directly without extra calculation. However directly training CNN with motion vectors degrades accuracy severely due to the noise and the lack of fine details in motion vectors. In order to relieve this problem, we propose four training strategies which leverage the knowledge learned from optical flow CNN to enhance the accuracy of motion vector CNN. Our insight is that motion vector and optical flow share inherent similar structures which allows us to transfer knowledge from one domain to another. To fully utilize the knowledge learned in optical flow domain, we develop deeply transferred motion vector CNN. Experimental results on various datasets show the effectiveness of our training strategies. Our approach is significantly faster than optical flow based approaches and achieves processing speed of 390.7 frames per second, surpassing real-time requirement. We release our model and code to facilitate further research. Bowen Zhang 0002, Limin Wang 0002, Zhe Wang 0013, Yu Qiao 0001, Hanli Wang |
IEEE Trans. Image Process. | 5 |
| 2018 | A Collaborative Scheduling-Based Parallel Solution for HEVC Encoding on Multicore PlatformsabstractIn order to meet the high computational demand to achieve superior coding efficiency and to explore the parallelism of parallel processing architectures, the emerging high efficiency video coding (HEVC) standard has been designed to be more parallelizable than previous video coding standards. However, it is still desirable to design an efficient parallel HEVC encoder to fully exploit the parallelism of the increasingly powerful multicore platforms, especially when considering the amount of parallelism, the scalability of parallelization, and the coding efficiency. In this work, a performance model of HEVC encoding is first introduced to investigate the speedup and the limitations of the technique of wavefront parallel processing (WPP) under various conditions. Then, a collaborative scheduling-based parallel solution (CSPS) for HEVC encoding is proposed, which includes adaptive parallel mode decision, asynchronous frame-level pixel interpolation, and multigrained task scheduling. The goal of the proposed CSPS is to defeat the disadvantages of WPP and further improve the parallelization of HEVC encoding on multicore platforms. Extensive experimental results demonstrate the efficiency of the proposed CSPS for parallelizing HEVC encoding as the computing resources of multicore architectures can be fully utilized. Hanli Wang, Bo Xiao 0005, Jun Wu 0006, Sam Kwong, C.-C. Jay Kuo |
IEEE Trans. Multim. | 1 |
| 2018 | Motion keypoint trajectory and covariance descriptor for human action recognition
Yun Yi, Hanli Wang |
Vis. Comput. | 2 |
| 2017 | Regularization of convolutional neural networks using ShuffleNodeabstractConvolutional Neural Network (CNN) has recently achieved significant performances for visual computing, and a number of researches are made to explore advanced model structures to solve the problem of over-fitting. In this paper, a regularization technique named ShuffleNode is proposed, which shuffles feature map elements to achieve regularization functions during model training. Specifically, there are two shuffle ways including within-map shuffle and cross-map shuffle, which are suitable to be employed in convolutional layers. The method of within-map shuffle is used to provide the exchange of elements within one feature map, while the cross-map shuffle method offers the opportunity of information sharing across different feature maps. The experimental results on several benchmark image classification datasets demonstrate the efficiency of the proposed method. Hanli Wang, Yu Long 0003 |
ICME | 2 |
| 2017 | Image captioning with deep LSTM based on sequential residualabstractImage captioning is a fundamental task which requires semantic understanding of images and the ability of generating description sentences with proper and correct structure. In consideration of the problem that language models are always shallow in modern image caption frameworks, a deep residual recurrent neural network is proposed in this work with the following two contributions. First, an easy-to-train deep stacked Long Short Term Memory (LSTM) language model is designed to learn the residual function of output distributions by adding identity mappings to multi-layer LSTMs. Second, in order to overcome the over-fitting problem caused by larger-scale parameters in deeper LSTM networks, a novel temporal Dropout method is proposed into LSTM. The experimental results on the benchmark MSCOCO and Flickr30K datasets demonstrate that the proposed model achieves the state-of-the-art performances with 101.1 in CIDEr on MSCOCO and 22.9 in B-4 on Flickr30K, respectively. Kaisheng Xu, Hanli Wang, Pengjie Tang |
ICME | 2 |
| 2017 | Richer Semantic Visual and Language Representation for Video CaptioningabstractTranslating and summarizing a video into natural language is an interesting and challenging visual task. In this work, a novel framework is built to generate sentences for videos with more coherence and semantics. A long short term memory (LSTM) network with an improved factored way is first developed, which takes inspiration of the conventional factored way and a common practice of presenting multi-modal features at the first time in LSTM for video captioning. An LSTM network with the combination of improved factored and un-factored ways is exploited, and a voting strategy is employed to predict the words. Then, the residual is used to enhance the gradient signals which is learned from residual network (ResNet), and a deeper LSTM network is constructed. Furthermore, several convolutional neural network (CNN) features from deep models with different architectures are fused to catch more comprehensive and complementary visual information. Experiments are conducted on the MSR-VTT2016 and MSR-VTT2017 grand challenge datasets to demonstrate the effectiveness of each presented techniques as well as the superiority compared to other state-of-the-art methods. Pengjie Tang, Hanli Wang, Kaisheng Xu |
ACM Multimedia | 2 |
| 2017 | G-MS2F: GoogLeNet based multi-stage feature fusion of deep CNN for scene recognition
Pengjie Tang, Hanli Wang, Sam Kwong |
Neurocomputing | 2 |
| 2017 | Learning correlations for human action recognition in videos
Yun Yi, Hanli Wang, Bowen Zhang 0002 |
Multim. Tools Appl. | 2 |
| 2017 | Improving feature matching strategies for efficient image retrieval
Lei Wang 0063, Hanli Wang |
Signal Process. Image Commun. | 2 |
| 2017 | Objective Video Quality Assessment Based on Perceptually Weighted Mean Squared ErrorabstractObject quality assessment for compressed video is critical to various video compression systems that are essential in the video delivery and storage. Although mean squared error (MSE) is computationally simple, it may not be accurate to reflect the perceptual quality of compressed videos, which are also affected dramatically by the characteristics of the human visual system (HVS), such as contrast sensitivity, visual attention, and masking effect. In this paper, a video quality metric is proposed based on perceptually weighted MSE. A low-pass filter is designed to model the contrast sensitivity of the HVS with the consideration of visual attention. The imperceptible distortion is adaptively removed in the salient and nonsalient regions. To quantitatively measure the masking effect, the randomness of video content is proposed in both the spatial and temporal domains. Since the masking effect highly depends on the regularity of structure and motion in the spatial and temporal directions, the video signal is modeled as a linear dynamic system, and the prediction error of future frames from previous frames is used as randomness to measure the significance of masking. The relation is investigated between MSE and perceptual quality scores across various contents, and a masking modulation model is proposed to compensate the impact of the masking effect on the MSE. The performance of the proposed quality metric is validated on three video databases with various compression distortions. The experimental results demonstrate that the proposed algorithm outperforms other benchmark quality metrics. Sudeng Hu, Lina Jin, Hanli Wang, Yun Zhang 0002, Sam Kwong, C.-C. Jay Kuo |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | Building Correlations Between Filters in Convolutional Neural NetworksabstractIn this paper, a new optimization approach is designed for convolutional neural network (CNN) which introduces explicit logical relations between filters in the convolutional layer. In a conventional CNN, the filters' weights in convolutional layers are separately trained by their own residual errors, and the relations of these filters are not explored for learning. Different from the traditional learning mechanism, the proposed correlative filters (CFs) are initiated and trained jointly in accordance with predefined correlations, which are efficient to work cooperatively and finally make a more generalized optical system. The improvement in CNN performance with the proposed CF is verified on five benchmark image classification datasets, including CIFAR-10, CIFAR-100, MNIST, STL-10, and street view house number. The comparative experimental results demonstrate that the proposed approach outperforms a number of state-of-the-art CNN approaches. Hanli Wang, Peiqiu Chen, Sam Kwong |
IEEE Trans. Cybern. | 1 |
| 2017 | Image Quality Assessment Based on Local Linear Information and Distortion-Specific CompensationabstractImage quality assessment (IQA) is a fundamental yet constantly developing task for computer vision and image processing. Most IQA evaluation mechanisms are based on the pertinence of subjective and objective estimation. Each image distortion type has its own property correlated with human perception. However, this intrinsic property may not be fully exploited by existing IQA methods. In this paper, we make two main contributions to the IQA field. First, a novel IQA method is developed based on a local linear model that examines the distortion between the reference and the distorted images for better alignment with human visual experience. Second, a distortion-specific compensation strategy is proposed to offset the negative effect on IQA modeling caused by different image distortion types. These score offsets are learned from several known distortion types. Furthermore, for an image with an unknown distortion type, a convolutional neural network-based method is proposed to compute the score offset automatically. Finally, an integrated IQA metric is proposed by combining the aforementioned two ideas. Extensive experiments are performed to verify the proposed IQA metric, which demonstrate that the local linear model is useful in human perception modeling, especially for individual image distortion, and the overall IQA method outperforms several state-of-the-art IQA approaches. Hanli Wang, Weisi Lin, Sudeng Hu, C.-C. Jay Kuo, Lingxuan Zuo |
IEEE Trans. Image Process. | 1 |
| 2017 | Joint Compression of Near-Duplicate VideosabstractThe expanding social network and multimedia technologies encourage more and more people to store and transmit information in visual format, such as image and video. However, the cost of this convenience brings about a shock to traditional video severs and exposes them under the risk of overloading. In the huge volume of online videos, there are a large amount of near-duplicate videos (NDVs). Although quite a number of research work have been proposed to detect NDVs, little research effort is made to compress these NDVs in a more effective manner than independent video compression. In this study, we make an in-depth exploration of the data redundancy of NDVs and propose a video analysis and coding framework to jointly compress NDVs. In order to employ the proposed NDV analysis and coding framework, a graph-based similar video grouping method and a number of preprocessing functions are designed to explore the correlation of visual information among NDVs and thus suit the requirement of joint video coding. Experimental results verify that the proposed NDV analysis and coding framework is able to effectively compress NDVs and thus save video data storage. Hanli Wang, Tao Tian, Jun Wu 0006 |
IEEE Trans. Multim. | 1 |
| 2016 | Real-Time Action Recognition with Enhanced Motion Vector CNNsabstractThe deep two-stream architecture [23] exhibited excellent performance on video based action recognition. The most computationally expensive step in this approach comes from the calculation of optical flow which prevents it to be real-time. This paper accelerates this architecture by replacing optical flow with motion vector which can be obtained directly from compressed videos without extra calculation. However, motion vector lacks fine structures, and contains noisy and inaccurate motion patterns, leading to the evident degradation of recognition performance. Our key insight for relieving this problem is that optical flow and motion vector are inherent correlated. Transferring the knowledge learned with optical flow CNN to motion vector CNN can significantly boost the performance of the latter. Specifically, we introduce three strategies for this, initialization transfer, supervision transfer and their combination. Experimental results show that our method achieves comparable recognition performance to the state-of-the-art, while our method can process 390.7 frames per second, which is 27 times faster than the original two-stream method. Bowen Zhang 0002, Limin Wang 0002, Zhe Wang 0013, Yu Qiao 0001, Hanli Wang |
CVPR | 5 |
| 2016 | Blind image quality assessment for multiply distorted images via convolutional neural networksabstractThe past decade has witnessed a growing development of Image Quality Assessment (IQA) techniques. However, the researches of IQA with multiple distortion types are still limited especially on blind image quality assessment methods. In this paper, a Convolutional Neural Network (CNN) based method is proposed to predict the quality of multiply distorted images without references. Inspired by the early human visual model, the proposed CNN based method combines feature learning and regression for estimating the quality of multiply distorted images. The proposed network consists of one convolutional layer, one pooling layer with max and average pooling, two full connection layers and one softmax classification layer. With this network structure, the relationship between the accuracy of CNN and the prediction monotonicity of IQA is explored. Experimental results on the newly released LIVE multiply distorted image quality database verify the effectiveness of the proposed CNN based method. Hanli Wang, Lingxuan Zuo |
ICASSP | 2 |
| 2016 | Screen content image quality assessment via convolutional neural networkabstractA number of image quality assessment (IQA) metrics have been designed in recent years for natural images, leading to a desire to develop IQA approaches for screen content image which is composed of textual as well as pictorial regions and exhibits different visual characteristics from the natural image. In this work, a no reference IQA metric based on convolutional neural network (CNN) is proposed for screen content image, which fuses the quality scores of textual and pictorial image patches by taking the differences of their visual features into accounts. Experimental results on the benchmark screen image quality assessment database (SIQAD) verify the effectiveness of the proposed CNN based approach as compared with several state-of-the-art IQA metrics. Lingxuan Zuo, Hanli Wang |
ICIP | 2 |
| 2016 | Distortion recognition for image quality assessment with convolutional neural networkabstractThe past decades have witnessed a growing development of image quality assessment (IQA). However, there is still much room to improve the IQA performance, especially for the no reference IQA problem. In this paper, a convolutional neural network (CNN) based approach is designed to predict the distortion type of an image and assess its quality without original reference. With the proposed CNN approach, an image is divided into patches and a selective weighted average method is designed to fuse the image quality score from the image patches with the aid of recognized distortion type. Experimental results on the benchmark LIVE image quality database verify the effectiveness of the proposed CNN based approach as compared with several state-of-the-art full reference and no reference image quality metrics. Hanli Wang, Lingxuan Zuo |
ICME | 1 |
| 2016 | Optimal stopping theory based fast coding tree unit decision for high efficiency video codingabstractHigh Efficiency Video Coding (HEVC) is the most recent video coding standard aiming to further reduce the bitrate over 50% as compared to the state-of-the-art H.264/Advanced Video Coding under the same visual quality. In order to achieve this, a number of advanced coding techniques have been adopted in HEVC, including the quadtree structure of Coding Unit (CU), Prediction Unit (PU) and Transform Unit (TU), etc. However, these coding techniques lead to a tremendous increase in HEVC encoding computations. In order to reduce the HEVC encoding computational complexity, the optimal stopping theory is employed herein to design an efficient algorithm to optimize the decision making process when choosing the best coding parameters of CU, PU and TU. Extensive comparative experimental results are performed by the proposed algorithm and another two recent works, which demonstrate that the proposed algorithm is very efficient and better in reducing the HEVC encoding computations while keeping the video quality and compression efficiency almost intact. Xiuzhe Wu, Hanli Wang |
VCIP | 2 |
| 2016 | Bayesian rule based fast TU depth decision algorithm for high efficiency video codingabstractThe latest video coding standard high efficiency video coding (HEVC) has made a significant progress in compression efficiency than previous standard H.264/advanced video coding (AVC) while it has led to a tremendous increase in encoding computations. Recently, a Bayesian model based transform unit (TU) depth decision approach has been designed to accelerate TU depth decision, which requires numerous variance computations. In this work, a novel relevant feature based Bayesian model is proposed for fast TU depth decision. Experimental results demonstrate that the best performance is achieved while the depths of upper TU, left TU and co-located TU are all taken into considerations. Moreover, as compared with previous research, the proposed algorithm reduces much more encoding computations while keeping the video quality and compression efficiency more or less intact. Xiuzhe Wu, Hanli Wang |
VCIP | 2 |
| 2015 | Effectively compressing Near-Duplicate Videos in a joint wayabstractWith the increasing popularity of social network, more and more people tend to store and transmit information in visual format, such as image and video. However, the cost of this convenience brings about a shock to traditional video servers and expose them under the risk of overloading. Among the huge amount of online videos, there are quite a number of Near-Duplicate Videos (NDVs). Although many works have been proposed to detect NDVs, few researches are investigated to compress these NDVs in a more effective way than independent compression. In this work, we utilize the data redundancy of NDVs and propose a video coding method to jointly compress NDVs. In order to employ the proposed video coding method, a number of pre-processing functions are designed to explore the correlation of visual information among NDVs and to suit the video coding requirements. Experimental results verify that the proposed video coding method is able to effectively compress NDVs and thus save video data storage. Hanli Wang, Tao Tian |
ICME | 1 |
| 2015 | Twin Feature and Similarity Maximal Matching for Image RetrievalabstractIn recent years, most advanced image retrieval algorithms are built upon local features, and various up-to-date match kernels are developed to boost image retrieval performances. However, most of these image retrieval algorithms need to face up two challenging issues: (1) the locality property of local features as well as quantization noise and (2) the phenomenon of burstiness, which significantly affect image retrieval performances. In this paper, two novel techniques including Twin Feature (TF) and Similarity Maximal Matching (SMM) are proposed for image retrieval performance improvement, which can be employed with non-aggregated kernel models, for example, the Selective Match Kernel (SMK). The proposed TF employs extra information from neighboring image patches to refine visual matching. As far as SMM is concerned, it tries to control burstiness by dynamically searching the match-pair combinations to maximize the global similarity score and thus removes multiple matches. Experimental results on two benchmark image datasets including Oxford5k and Paris6k demonstrate that the new techniques SMKtf (SMK with TF) and SMKsmm (SMK with SMM) can greatly enhance image retrieval accuracy performances as compared to SMK, and their combination, i.e., SMKtf+smm, is able to achieve better image retrieval accuracies than a number of state-of-the-art approaches. Lei Wang 0063, Hanli Wang, Fengkuangtian Zhu |
ICMR | 2 |
| 2015 | Accelerating Large-scale Image Retrieval on Heterogeneous Architectures with SparkabstractApache Spark is a general-purpose cluster computing system for big data processing and has drawn much attention recently from several fields, such as pattern recognition, machine learning and so on. Unlike MapReduce, Spark is especially suitable for iterative and interactive computations. With the computing power of Spark, a utility library, referred to as IRlib, is proposed in this work to accelerate large-scale image retrieval applications by jointly harnessing the power of GPU. Similar to the built-in machine learning library of Spark, namely MLlib, IRlib fits into the Spark APIs and benefits from the powerful functionalities of Spark. The main contributions of IRlib lie in two-folds. First, IRlib provides a uniform set of APIs for the programming of image retrieval applications. Second, the computational performance of Spark equipped with multiple GPUs is dramatically boosted by developing high performance modules for common image retrieval related algorithms. Comparative experiments concerning large-scale image retrieval are carried out to demonstrate the significant performance improvement achieved by IRlib as compared with single CPU thread implementation as well as Spark without GPUs employed. Hanli Wang, Bo Xiao 0005, Lei Wang 0063, Jun Wu 0006 |
ACM Multimedia | 1 |
| 2015 | Human Action Recognition With Trajectory Based Covariance Descriptor In Unconstrained VideosabstractHuman action recognition from realistic videos plays a key role in multimedia event detection and understanding. In this paper, a novel Trajectory Based Covariance (TBC) descriptor is proposed, which is formulated along the dense trajectories. To map the descriptor matrix to vector space and trim out the redundancy of data, the TBC descriptor matrix is projected to Euclidean space by the Logarithm Principal Components Analysis (LogPCA). Our method is tested on the challenging Hollywood2 and TV Human Interaction datasets. Experimental results show that the proposed TBC descriptor outperforms three baseline descriptors (i.e., histogram of oriented gradient, histogram of optical flow and motion boundary histogram), and our method achieves better recognition performances than a number of state-of-the-art approaches. Hanli Wang, Yun Yi, Jun Wu 0006 |
ACM Multimedia | 1 |
| 2015 | Large-scale human action recognition with sparkabstractIn this paper, Apache Spark, the rising big data processing tool with in-memory computing ability, is explored to address the task of large-scale human action recognition. To achieve this, several advanced key techniques for human action recognition, such as trajectory based feature extraction, Gaussian Mixture Model, Fisher Vector, etc., are realized with parallel distributed computing power on Spark. The theory and implementation details for these distributed applications are presented in this work. The experimental results on the benchmark human action dataset Hollywood-2 show that the proposed Spark based framework which is deployed on a 9-node computer cluster can deal with large-scale video data and can dramatically accelerate the process of human action recognition. Hanli Wang, Xiaobin Zheng, Bo Xiao 0005 |
MMSP | 1 |
| 2015 | Correlative Filters for Convolutional Neural NetworksabstractThis paper introduces a regularization method called Correlative Filter (CF) for Convolutional Neural Network (CNN), which takes advantage of the relevance between the convolutional kernels belonging to the same convolutional layer. During the process of training with the proposed CF method, several pairs of filters are designed in a manner of randomness to contain opposite weights in low-level layers. Regarding higher level layers where synthetical features are processed, the relation between correlative filters is explored as translation of various directions. The proposed CF method attempts to optimize the inner structure of convolutional layers and it can work jointly with other regularization techniques, such as stochastic pooling, Dropout, etc. The experimental results on the competitive image classification benchmark dataset CIFAR-10 demonstrates the performance of the proposed CF method, additionally, it is also verified that the proposed CF method is wonderful to be employed to enhance several state-of-the-art regularization models. Peiqiu Chen, Hanli Wang, Jun Wu 0006 |
SMC | 2 |
| 2015 | Accelerating Support Vector Machine Learning with GPU-Based MapReduceabstractWith the exploding growth of data, the computational complexity required by learning Support Vector Machine (SVM) lays a heavy burden on real-world applications. To address this issue, parallel computational techniques can be employed such as the Graphics Processing Units (GPUs) and MapReduce model. As it is well known, GPUs are microprocessors on a multi-core architecture which reveal high performance in mass data parallel computing, and MapReduce allows computational tasks to be divided into a plurality of parts, distributed to various computing nodes and combined on a single node. In this paper, we propose a GPU-based MapReduce framework to accelerate SVM learning by jointly utilizing the parallel computing power of GPU and MapReduce. Extensive experimental results have verified the effectiveness and efficiency of the proposed approach. Tianyao Sun, Hanli Wang, Jun Wu 0006 |
SMC | 2 |
| 2015 | Tracking Salient Keypoints for Human Action RecognitionabstractIt plays an important role to recognize human actions from realistic videos in multimedia event detection and understanding. To this aim, a novel human tracking approach is proposed in this paper. Firstly, salient key points trajectories are generated to track human actions at multiple spatial scales. Then, camera motion elimination is utilized to further improve the robustness of motion trajectories. To depict human motions accurately and efficiently, the Histogram of Oriented Gradient (HOG), Histogram of Optical Flow (HOF) and Motion Boundary Histogram (MBH) are employed with the Fisher vector model being utilized to aggregate these three features. Extensive experimental results on four challenging human action video datasets demonstrate that the proposed approach is able to achieve better recognition performances in a more computationally efficient manner as compared with a number of state-of-the-art approaches. Hanli Wang, Yun Yi |
SMC | 1 |
| 2015 | Encoding scale into fisher vector for human action recognitionabstractIn this paper, a new kind of Fisher Vector (FV) model, named Scale FV (ScaleFV), is proposed to ameliorate visual feature encoding for human action recognition. Although several researches have been proposed for feature encoding, the temporal scale information is almost ignored. Similar to the spatial scale information which has shown to be important in extracting and encoding visual features, the temporal scale information also plays an important role in video content analysis based on our investigation. To demonstrate this, a definition of temporal scale in videos is given, and it is presented that both of the spatial and temporal scale information can be encoded into the FV model by slightly modifying the underlying Gaussian Mixture Models (GMM). Furthermore, an enhanced FV model termed as Combined FV (CombFV) is designed to capture both position and scale information for human action recognition. Comparative experiments are carried out to demonstrate the superior performance of the proposed methods. Bowen Zhang 0002, Hanli Wang |
VCIP | 2 |
| 2015 | GPU-based MapReduce for large-scale near-duplicate video retrieval
Hanli Wang, Fengkuangtian Zhu, Bo Xiao 0005, Lei Wang 0063, Yu-Gang Jiang 0001 |
Multim. Tools Appl. | 1 |
| 2015 | CHCF: A Cloud-Based Heterogeneous Computing Framework for Large-Scale Image RetrievalabstractThe last decade has witnessed a dramatic growth of multimedia content and applications, which in turn requires an increasing demand of computational resources. Meanwhile, the high-performance computing world undergoes a trend toward heterogeneity. However, it is never easy to develop domain-specific applications on heterogeneous systems while maximizing the system efficiency. In this paper, a novel framework, namely, cloud-based heterogeneous computing framework (CHCF), is proposed with a set of tools and techniques for compilation, optimization, and execution of multimedia mining applications on heterogeneous systems. With the aid of the compiler and the utility library provided by CHCF, users are able to develop multimedia mining applications rapidly and efficiently. The proposed framework employs a number of techniques, including adaptive data partitioning, knowledge-based hierarchical scheduling, and performance estimation, to achieve high computing performance. As one of the most important multimedia mining applications, large-scale image retrieval is investigated based on the proposed CHCF. The scalability, computing performance, and programmability of CHCF are studied for large-scale image retrieval by case studies and experimental evaluations. The experimental results demonstrate that CHCF can achieve good scalability and significant computing performance improvements for image retrieval. Hanli Wang, Bo Xiao 0005, Lei Wang 0063, Fengkuangtian Zhu, Yu-Gang Jiang 0001, Jun Wu 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2015 | Compressed Image Quality Metric Based on Perceptually Weighted DistortionabstractObjective quality assessment for compressed images is critical to various image compression systems that are essential in image delivery and storage. Although the mean squared error (MSE) is computationally simple, it may not be accurate to reflect the perceptual quality of compressed images, which is also affected dramatically by the characteristics of human visual system (HVS), such as masking effect. In this paper, an image quality metric (IQM) is proposed based on perceptually weighted distortion in terms of the MSE. To capture the characteristics of HVS, a randomness map is proposed to measure the masking effect and a preprocessing scheme is proposed to simulate the processing that occurs in the initial part of HVS. Since the masking effect highly depends on the structural randomness, the prediction error from neighborhood with a statistical model is used to measure the significance of masking. Meanwhile, the imperceptible signal with high frequency could be removed by preprocessing with low-pass filters. The relation is investigated between the distortions before and after masking effect, and a masking modulation model is proposed to simulate the masking effect after preprocessing. The performance of the proposed IQM is validated on six image databases with various compression distortions. The experimental results show that the proposed algorithm outperforms other benchmark IQMs. Sudeng Hu, Lina Jin, Hanli Wang, Yun Zhang 0002, Sam Kwong, C.-C. Jay Kuo |
IEEE Trans. Image Process. | 3 |
| 2014 | Predicting zero coefficients for High Efficiency Video CodingabstractSimilar to previous video coding standards, transform and quantization are used in the most recent video coding standard High Efficiency Video Coding (HEVC) and a large number of transform coefficients are quantized to zeros. In order to reduce the computations involved in transform and quantization in HEVC, a prediction approach based on Gaussian model is proposed to predict zero quantized transform coefficients within the 4 × 4, 8 × 8, 16 × 16 and 32 × 32 blocks. Extensive experiments demonstrate that the proposed algorithm is able to effectively predict zero coefficients and thus reduce redundant computations while keeping the video quality and compression efficiency almost intact. Hanli Wang, Jun Wu 0006 |
ICME | 1 |
| 2014 | A Framework of Video Coding for Compressing Near-Duplicate Videos
Hanli Wang, Yu-Gang Jiang 0001, Zhihua Wei 0001 |
MMM (1) | 1 |
| 2014 | Early detection of all-zero 4×4 blocks in High Efficiency Video Coding
Hanli Wang, Weiyao Lin, Sam Kwong, Oscar C. Au, Jun Wu 0006, Zhihua Wei 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2014 | A New Network-Based Algorithm for Human Activity Recognition in VideosabstractIn this paper, a new network-transmission-based (NTB) algorithm is proposed for human activity recognition in videos. The proposed NTB algorithm models the entire scene as an error-free network. In this network, each node corresponds to a patch of the scene and each edge represents the activity correlation between the corresponding patches. Based on this network, we further model people in the scene as packages, while human activities can be modeled as the process of package transmission in the network. By analyzing these specific package transmission processes, various activities can be effectively detected. The implementation of our NTB algorithm into abnormal activity detection and group activity recognition are described in detail in this paper. Experimental results demonstrate the effectiveness of our proposed algorithm. Weiyao Lin, Yuanzhe Chen, Jianxin Wu 0001, Hanli Wang, Bin Sheng 0001, Hongxiang Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2013 | Fast Mode Decision Based on Optimal Stopping Theory for Multiview Video Coding
Hanli Wang, Yue Heng, Tiesong Zhao, Bo Xiao 0005 |
MMM (2) | 1 |
| 2013 | Semi-supervised multi-label image classification based on nearest neighbor editing
Hanli Wang |
Neurocomputing | 2 |
| 2013 | Rate control for consistent visual quality of H.264/AVC encoding
Long Xu 0001, Sam Kwong, Hanli Wang, Debin Zhao, Wen Gao 0001 |
Signal Process. Image Commun. | 3 |
| 2013 | Multiview Coding Mode Decision With Hybrid Optimal Stopping ModelabstractIn a generic decision process, optimal stopping theory aims to achieve a good tradeoff between decision performance and time consumed, with the advantages of theoretical decision-making and predictable decision performance. In this paper, optimal stopping theory is employed to develop an effective hybrid model for the mode decision problem, which aims to theoretically achieve a good tradeoff between the two interrelated measurements in mode decision, as computational complexity reduction and rate-distortion degradation. The proposed hybrid model is implemented and examined with a multiview encoder. To support the model and further promote coding performance, the multiview coding mode characteristics, including predicted mode probability and estimated coding time, are jointly investigated with inter-view correlations. Exhaustive experimental results with a wide range of video resolutions reveal the efficiency and robustness of our method, with high decision accuracy, negligible computational overhead, and almost intact rate-distortion performance compared to the original encoder. Tiesong Zhao, Sam Kwong, Hanli Wang, Zhou Wang 0001, Zhaoqing Pan, C.-C. Jay Kuo |
IEEE Trans. Image Process. | 3 |
| 2012 | Large-scale multimedia data mining using MapReduce frameworkabstractIn this paper, the framework of MapReduce is explored for large-scale multimedia data mining. Firstly, a brief overview of MapReduce and Hadoop is presented to speed up large-scale multimedia data mining. Then, the high-level theory and low-level implementation for several key computer vision technologies involved in this work are introduced, such as 2D/3D interest point detection, clustering, bag of features, and so on. Experimental results on image classification, video event detection and near-duplicate video retrieval are carried out on a five-node Hadoop cluster to demonstrate the efficiency of the proposed MapReduce framework for large-scale multimedia data mining applications. Hanli Wang, Lei Wang 0063, Kuangtian Zhufeng, Wei Wang 0033 |
CloudCom | 1 |
| 2012 | Quality assessment for color images with tucker decompositionabstractAs an extension of the singular value decomposition based approaches, a novel metric based on Tucker decomposition for color image quality assessment is proposed in this paper. It extracts both the spacial and chromatic information of a color image with Tucker decomposition. As compared to most of the other existing quality metrics for color images, the key advantage of the proposed metric is that it treats the color image as a whole entity instead of an assembly of three independent visual channels, and thus the interrelation among the different visual channels is well involved. Experimental results on the LIVE Database Release 2 demonstrate that the proposed metric can generally achieve similar or better performance than other metrics. Hanli Wang |
ICIP | 2 |
| 2012 | Analytical study of RGB vertical stripe and RGBX square-shaped subpixel arrangementsabstractThe frequency characteristics of subpixel-based decimation with RGB vertical stripe and RGBX square-shaped subpixel arrangements are studied. To achieve higher apparent resolution than pixel-based decimation, the sampling locations are specially chosen for each of two subpixel arrangements, resulting in relatively small magnitudes of horizontal and vertical aliasing spectra in frequency domain. Thanks to 2-D RGBX square-shaped subpixel arrangement, all the horizontal, vertical, diagonal and anti-diagonal aliasing spectra merely contain low-frequency information, indicating that subpixel-based decimation with RGBX square-shaped panel is more effective in retaining original high frequency details than RGB vertical stripe subpixel arrangement. Lu Fang 0001, Oscar C. Au, Jingjing Dai, Hanli Wang, Ngai-Man Cheung |
ICIP | 4 |
| 2012 | Novel 2-D MMSE Subpixel-Based Image Down-SamplingabstractSubpixel-based down-sampling is a method that can potentially improve apparent resolution of a down-scaled image on LCD by controlling individual subpixels rather than pixels. However, the increased luminance resolution comes at price of chrominance distortion. A major challenge is to suppress color fringing artifacts while maintaining sharpness. We propose a new subpixel-based down-sampling pattern called diagonal direct subpixel-based down-sampling (DDSD) for which we design a 2-D image reconstruction model. Then, we formulate subpixel-based down-sampling as a MMSE problem and derive the optimal solution called minimum mean square error for subpixel-based down-sampling (MMSE-SD). Unfortunately, straightforward implementation of MMSE-SD is computational intensive. We thus prove that the solution is equivalent to a 2-D linear filter followed by DDSD, which is much simpler. We further reduce computational complexity using a smallk×kfilter to approximate the much larger MMSE-SD filter. To compare the performances of pixel and subpixel-based down-sampling methods, we propose two novel objective measures: normalizedl1high frequency energy for apparent luminance sharpness and PSNRU(V)for chrominance distortion. Simulation results show that both MMSE-SD and MMSE-SD(k) can give sharper images compared with conventional down-sampling methods, with little color fringing artifacts. Lu Fang 0001, Oscar C. Au, Ketan Tang, Hanli Wang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2012 | A Universal Rate Control Scheme for Video TranscodingabstractVideo transcoding is proposed for the bitrate adaption, spatial and/or temporal resolutions adaption, and video format conversion. In video streaming application, it converts videos at server to the compatible versions demanded by networks or clients' devices, so that the videos can be delivered over networks and displayed in the clients' devices successfully. This paper provides a universal rate control scheme for various video transcoding purposes. First, a new rate-distortion (R-D) model is established theoretically for better representing the real R-D feature of transcoding. Second, a window-level rate control algorithm is proposed for providing smooth visual quality with compliant buffer constraint by utilizing the two-pass R-D model and a new proposed sliding window buffer control strategy. Finally, a universal rate control scheme for transcoding is developed based on the established R-D model and the proposed window-level rate control algorithm. The extensive experimental results demonstrate that as compared to other state-of-the-art rate control algorithms for transcoding, the proposed scheme can achieve more bit control accuracy with the average mismatch below 0.2%, and much more consistent visual quality with 0.1 dB-0.3 dB peak-to-signal noise ratio improvement in average, while with low computational complexity. Long Xu 0001, Sam Kwong, Hanli Wang, Yun Zhang 0002, Debin Zhao, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2012 | Adaptive Quantization-Parameter Clip Scheme for Smooth Quality in H.264/AVCabstractIn this paper, we investigate the issues over the smooth quality and the smooth bit rate during rate control (RC) in H.264/AVC. An adaptive quantization-parameter (Q(p)) clip scheme is proposed to optimize the quality smoothness while keeping the bit-rate fluctuation at an acceptable level. First, the frame complexity variation is studied by defining a complexity ratio between two nearby frames. Second, the range of the generated bits is analyzed to prevent the encoder buffer from overflow and underflow. Third, based on the safe range of the generated bits, an optimal Q(p) clip range is developed to reduce the quality fluctuation. Experimental results demonstrate that the proposed Q(p) clip scheme can achieve excellent performance in quality smoothness and buffer regulation. Sudeng Hu, Hanli Wang, Sam Kwong |
IEEE Trans. Image Process. | 2 |
| 2012 | H.264/SVC Mode Decision Based on Optimal Stopping TheoryabstractFast mode decision algorithms have been widely used in the video encoder implementation to reduce encoding complexity yet without much sacrifice in the coding performance. Optimal stopping theory, which addresses early termination for a generic class of decision problems, is adopted in this paper to achieve fast mode decision for the H.264/Scalable Video Coding standard. A constrained model is developed with optimal stopping, and the solutions to this model are employed to initialize the candidate mode list and predict the early termination. Comprehensive simulation results are conducted to demonstrate that the proposed method strikes a good balance between low encoding complexity and high coding efficiency. Tiesong Zhao, Sam Kwong, Hanli Wang, C.-C. Jay Kuo |
IEEE Trans. Image Process. | 3 |
| 2012 | Joint Demosaicing and Subpixel-Based Down-Sampling for Bayer Images: A Fast Frequency-Domain Analysis ApproachabstractA portable device such as a digital camera with a single sensor and Bayer color filter array (CFA) requires demosaicing to reconstruct a full color image. To display a high resolution image on a low resolution LCD screen of the portable device, it must be down-sampled. The two steps, demosaicing and down-sampling, influence each other. On one hand, the color artifacts introduced in demosaicing may be magnified when followed by down-sampling; on the other hand, the detail removed in the down-sampling cannot be recovered in the demosaicing. Therefore, it is very important to consider simultaneous demosaicing and down-sampling. Lu Fang 0001, Oscar C. Au, Yan Chen 0007, Aggelos K. Katsaggelos, Hanli Wang |
IEEE Trans. Multim. | 5 |
| 2011 | Combining interpretable fuzzy rule-based classifiers via multi-objective hierarchical evolutionary algorithmabstractThe contributions of this paper are two-fold: firstly, it employs a multi-objective evolutionary hierarchical algorithm to obtain a non-dominated fuzzy rule classifier set with interpretability and diversity preservation. Secondly, a reduce-error based ensemble pruning method is utilized to decrease the size and enhance the accuracy of the combined fuzzy rule classifiers. In this algorithm, each chromosome represents a fuzzy rule classifier and compose of three different types of genes: control, parameter and rule genes. In each evolution iteration, each pair of classifiers in non-dominated solution set with the same multi-objective qualities are examined in terms of Q statistic diversity values. Then, similar classifiers are removed to preserve the diversity of the fuzzy system. Finally, experimental results on the ten UCI benchmark datasets indicate that our approach can maintain a good trade-off among accuracy, interpretability and diversity of fuzzy classifiers. Jingjing Cao, Hanli Wang, Sam Kwong, Ke Li 0001 |
SMC | 2 |
| 2011 | Hierarchical B-picture mode decision in H.264/SVC
Tiesong Zhao, Hanli Wang, Sam Kwong, C.-C. Jay Kuo, Wolfgang A. Halang |
J. Vis. Commun. Image Represent. | 2 |
| 2011 | Rate Control Optimization for Temporal-Layer Scalable Video CodingabstractA novel frame-level rate control (RC) algorithm is presented in this paper for temporal scalability of scalable video coding. First, by introducing a linear quality dependency model, the quality dependency between a coding frame and its references is investigated for the hierarchical B-picture prediction structure. Second, linear rate-quantization (R-Q) and distortion-quantization (D-Q) models are introduced based on different characteristics of temporal layers. Third, according to the proposed quality dependency model and R-Q and D-Q models for each temporal layer, adaptive weighting factors are derived to allocate bits efficiently among temporal layers. Experimental results on not only traditional quarter common intermediate format/common intermediate format but also standard definition and high definition sequences demonstrate that the proposed algorithm achieves excellent coding efficiency as compared to other benchmark RC schemes. Sudeng Hu, Hanli Wang, Sam Kwong, Tiesong Zhao, C.-C. Jay Kuo |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | Fast inter-layer mode decision in Scalable Video CodingabstractThe recently developed video encoding standard Scalable Video Coding (SVC) is designed as the extension of H.264/AVC. By employing new techniques such as hierarchal layer encoding, SVC could produce bit streams with different frame rates, bit rates and spatial resolutions; but meanwhile, the computational complexity of video encoder is also increased, which makes optimization of SVC encoding a necessity. In this paper, a novel algorithm is proposed, which utilizes the encoding mode of Base Layer (BL) macroblocks (MBs) to initialize the potential optimal candidate encoding mode list for Enhancement Layer (EL) MBs, by mapping encoding modes onto a Two-Dimensional (2D) map. Experimental results demonstrate that, the proposed algorithm could achieve significantly better and more robust time saving than two recent works while maintaining the video performances. Tiesong Zhao, Hanli Wang, Sam Kwong |
ICIP | 2 |
| 2010 | Probability-based coding mode prediction for H.264/AVCabstractFast mode decision (FMD) algorithm aims to reduce the computational complexity of video encoders, especially for H.264/AVC, and meanwhile keep the coding efficiency. It is necessary for coding high resolution sequences, such as standard definition (SD) or high definition (HD) sequences. In this paper, an efficient algorithm is proposed to speed up inter-mode decision for H.264/AVC. Firstly, for a macroblock (MB) being coded, with modes of neighboring blocks, a probability model is utilized to build up an optimal encoding mode list; after that, rate-distortion (RD) cost-based early termination is employed to skip unnecessary inter-modes. Simulation results demonstrate that, the proposed algorithm could achieve significant encoding time saving for QCIF/CIF and SD/HD sequences, meanwhile with almost the same RD performance as the original encoder. Tiesong Zhao, Hanli Wang, Sam Kwong, Sudeng Hu |
ICIP | 2 |
| 2010 | Frame level rate control for H.264/AVC with novel Rate-Quantization modelabstractIn this paper, a frame level rate control algorithm is proposed with a novel Rate-Quantization (R-Q) model for H.264/AVC. Firstly, a two-stage rate control scheme is adopted to decouple the inter-dependency between Rate Distortion Optimization (RDO) and rate control. Secondly, in order to predict the frame complexity accurately, instead of the Mean Absolute Difference (MAD) of the residual signal, bits information in the RDO-based mode decision process is employed to predict the frame complexity. Thirdly, a self-adaptive exponential R-Q model is proposed for rate control. Experimental results reveal that the proposed R-Q model can estimate the actual output bits very well, and the novel rate control scheme has excellent performance both in bit rate accuracy and coding efficiency as compared to JVT-W043 and the FixedQp tool in the Joint Scalable Video Model reference software. Sudeng Hu, Hanli Wang, Sam Kwong, Tiesong Zhao |
ICME | 2 |
| 2010 | Fast Mode Decision Based on Mode AdaptationabstractThis paper proposes an efficient algorithm for fast mode decision in H.264/advanced video coding by adaptively predicting the optimal mode for each macroblock (MB) to be coded. Firstly, encoding modes are projected as points onto a 2-D map, and an optimal 2-D point of the MB to be coded is predicted based on the encoding information of spatial-temporal neighboring blocks. Then, a priority-based mode candidate list with a descending order to be the best mode is constructed based on the optimal 2-D point. Finally, mode decision is performed according to the priority-based mode candidate list in the checking order, from the most important mode to the least one, with early termination conditions. Extensive experimental results demonstrate that the proposed algorithm is superior to three recent fast mode decision algorithms, with the entire encoding time being reduced by about 60% for quarter common intermediate format/common intermediate format/standard-definition sequences on average and the rate distortion performance being kept almost intact. Tiesong Zhao, Hanli Wang, Sam Kwong, C.-C. Jay Kuo |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2009 | Decision-based median filter using k-nearest noise-free pixelsabstractTraditional median filter replaces each pixel in an image with the median value of their k-nearest pixels (commonly known as pixels in 2-D window). The problem associated with this approach is that the restored pixel is noise if median value of their k-nearest pixels is a corrupted pixel. To mitigate the above problem, this paper proposes a novel decision-based median filter that replaces each corrupted pixel with the me-dian value of their k-nearest noise-free pixels. Advantages of the median filter using k-nearest noise-free pixels instead of k-nearest pixels are two facets: first, it guarantees that pix-els after being restored must be noise-free, because the me-dian filter operator is executed on noise-free pixels; second, the median filter using k-nearest noise-free pixels adaptively adjusts its window size for each pixel such that the number of noise-free pixels locating in the window increases up to k. To realize it, the median filter using k-nearest noise-free pixels firstly detects noise-free pixels in an image, then re-places each corrupted pixel with the median value of their k-nearest noise-free pixels. The proposed median filter is tested on four real images corrupted by different levels of salt-and-pepper noise. Experimental results confirm the effectiveness of decision-based median filter using k-nearest noise-free pix-els. Index Terms — Decision-based median filter, image restoration, impulse noise, median filter, salt-and-pepper noise. 1. Yi Hong 0002, Sam Kwong, Hanli Wang |
ICASSP | 3 |
| 2009 | Resampling-based selective clustering ensembles
Yi Hong 0002, Sam Kwong, Hanli Wang, Qingsheng Ren |
Pattern Recognit. Lett. | 3 |
| 2009 | Early Determination of Zero-Quantized 8 , ×, 8 DCT CoefficientsabstractThis paper proposes a novel approach to early determination of zero-quantized 8 × 8 discrete cosine transform (DCT) coefficients for fast video encoding. First, with the dynamic range analysis of DCT coefficients at different frequency positions, several sufficient conditions are derived to early determine whether a prediction error block (8 × 8) is an all-zero or a partial-zero block, i.e., the DCT coefficients within the block are all or partially zero-quantized. Being different from traditional methods that utilize the sum of absolute difference (SAD) of the entire prediction error block, the sufficient conditions are derived based on the SAD of each row of the prediction error block. For partial-zero blocks, fast DCT/IDCT algorithms are further developed by pruning conventional 8-point butterfly-based DCT/IDCT algorithms. Experimental results exhibit that the proposed early determination algorithm greatly reduces computational complexity in terms of DCT/IDCT, quantization, and inverse quantization, as compared with existing algorithms. Xiangyang Ji, Sam Kwong, Debin Zhao, Hanli Wang, C.-C. Jay Kuo, Qionghai Dai |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2008 | Probabilistic and Graphical Model based Genetic Algorithm Driven Clustering with Instance-level ConstraintsabstractClustering is traditionally viewed as an unsupervised method for data analysis. However, several recent studies have shown that some limited prior instance-level knowledge can significantly improve the performance of clustering algorithm. This paper proposes a semi-supervised clustering algorithm termed as the Probabilistic and Graphical Model based Genetic Algorithm Driven Clustering with Instance-level Constraints (Cop-CGA). In Cop-CGA, all prior knowledge about pairs of instances that should or should not be classified into the same groups is denoted as a graph and all candidate clustering solutions are sampled from this graph with different orders to assign instances into a certain number of groups. We illustrate how to design the Cop-CGA to guarantee that all candidate solutions satisfy the given constraints and demonstrate the usefulness of background knowledge for genetic algorithm driven clustering algorithm through experiments on several real data sets with artificial hard constraints. One advantage of Cop-CGA is both positive and negative instance-level constraints can be easily incorporated. Moreover, the performance of Cop-CGA is not sensitive to the order of assignment of instances to groups. Yi Hong 0002, Sam Kwong, Hanli Wang, Qingsheng Ren, Yuchou Chang |
IEEE Congress on Evolutionary Computation | 3 |
| 2008 | SVPCGA: Selection on virtual population based compact genetic algorithmabstractThis paper describes a novel virtual population based truncation selection operator that extends our previously proposed virtual population based tournament selection operator. Moreover, two extensions of compact genetic algorithm (CGA) that make use of virtual population based selection operators are presented in this paper: one is the tournament selection on virtual population based compact genetic algorithm (SVPCGA-TO); the other is the truncation selection on virtual population based compact genetic algorithm (SVPCGA-TR). Both SVPCGA-TO and SVPCGA-TR are tested on several benchmark problems and their results are compared with those obtained by CGA and ne-CGA. Some superiorities of SVPCGA in search reliability can be achieved. Yi Hong 0002, Sam Kwong, Hanli Wang, Qingsheng Ren |
IEEE Congress on Evolutionary Computation | 3 |
| 2008 | A Novel Approach for Service Capabilities Representation Based on Statistical Study on WSDLabstractWeb services (WS) are becoming more and more popular nowadays and service-oriented architecture (SOA) have been widely used in the construction of information systems. But most of the SOA applications are not brand new and usually evolved from legacy systems. Our research group is building a service identification framework used for the SOA reengineering of existing large-scale information applications. Service capabilities have become an important issue we care about which are the actions performed or the information delivered by a service. But the semantic description for service capabilities has not been fully addressed in the current approach. In this paper, we analyze the contributions and limitations of current approach. In order to find the practical features of existing web services, we conducted a statistical study on more than four hundred WSDL documents collected from XMethods.net, Amazon and Google. Then we propose a new approach for Semantic Web Services description. Service capabilities are represented by informative entities and standard actions are specified for each informative entity. Yan Liu 0011, Mingguang Zhuang, Qingling Wang, Hanli Wang |
ICIW | 4 |
| 2008 | Rate-Distortion Optimization of Rate Control for H.264 With Adaptive Initial Quantization Parameter DeterminationabstractA rate-distortion (R-D) optimization rate control (RC) algorithm with adaptive initialization is presented for H.264. First, a linear distortion-quantization (D-Q) model is introduced and thus a close-form solution is developed to derive optimal quantization parameters (Qp)for encoding each macroblock. Then we exploit to determine the initial Qpefficiently and adaptively according to the content of video sequences. The experimental results demonstrate that the proposed algorithm can achieve better R-D performance than that of other two RC algorithms including the algorithm JVT-G012 which is the current recommended RC scheme implemented in the H.264 reference software JM9.5. Hanli Wang, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2008 | Prediction of Zero Quantized DCT Coefficients in H.264/AVC Using Hadamard Transformed InformationabstractThis paper presents an efficient approach for detecting zero quantized discrete cosine transform (ZQDCT) coefficients using the sum of absolute transformed difference (SATD). Previously, all the ZQDCT prediction approaches employ the sum of absolute difference (SAD) available ahead of DCT and quantization (Q) for early detection. However, when the Hadamard transform is enabled for H.264/AVC encoding, only the SATD instead of SAD is available before DCT and Q, and all the prediction approaches can not be directly applied. To solve this problem, the Gaussian distribution is applied to study the integer 4x4 DCT coefficients in H.264/AVC and hence an adaptive scheme with multiple thresholds against SATD is derived to realize different types of DCT and Q implementations. In addition, another two SATD based sufficient conditions are proposed for early detecting zero quantized DC coefficients for the luma components encoded with the intra 16 times 16 mode and the chroma components. The experimental results demonstrate that the proposed approach can greatly reduce the DCT and Q computations and obtain almost the same rate-distortion performance as the original encoder. Hanli Wang, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2007 | Gaussian Model Based Approach to Reducing DCT Computations in H.264abstractThis paper presents an efficient method to predict zero quantized DCT coefficients in order to reduce redundant computations in H.264 encoding. The Gaussian distribution is firstly applied to study the quantized integer DCT coefficients in H.264 and then an adaptive scheme with multiple thresholds is proposed to realize different types of DCT and quantization implementations. Compared with other approaches in the literature, the experimental results demonstrate that the proposed method can achieve the best performance in reducing computations and obtain almost the same rate-distortion performance as the original encoder. Hanli Wang, Sam Kwong |
ICASSP (1) | 1 |
| 2007 | A Rate-Distortion Optimization Algorithm for Rate Control in H.264abstractThis paper presents a novel rate-distortion (R-D) joint optimization rate control (RC) algorithm for H.264 encoding. For RC in H.264, one of the most important topics is to model the R-D characteristics accurately. To achieve this, an efficient linear model is proposed to model the distortion-quantization (D-Q) relation. With the proposed linear D-Q model, a closed-form solution is derived to calculate the optimal quantization parameter for encoding each macroblock (MB). The proposed RC algorithm can be applied to both P-frames and B-frames. It is shown by experimental results that the proposed algorithm can control the bit rates accurately with the R-D performance better than that of the RC algorithm JVT-G012 implemented in the H.264 reference software JM9.5. Hanli Wang, Sam Kwong |
ICASSP (1) | 1 |
| 2007 | Efficient predictive model of zero quantized DCT coefficients for fast video encoding
Hanli Wang, Sam Kwong, Chi-Wah Kok |
Image Vis. Comput. | 1 |
| 2007 | Genetic-fuzzy rule mining approach and evaluation of feature selection techniques for anomaly intrusion detection
Chi-Ho Tsang, Sam Kwong, Hanli Wang |
Pattern Recognit. | 3 |
| 2007 | Hybrid Model to Detect Zero Quantized DCT Coefficients in H.264abstractIn H.264 coding, there are a large number of discrete cosine transform (DCT) coefficients of the prediction residue which are quantized to zeros. Therefore, it is desired to design a method which can early detect zero quantized DCT coefficients (ZQDCT) before implementing DCT and quantization (Q) and thus reduce redundant computations for H.264 coding. To achieve this, a hybrid model is proposed in this paper in order to predict ZQDCT coefficients. First, the Gaussian distribution is applied to study the integer DCT coefficients in H.264 and hence an adaptive scheme with multiple thresholds is derived to realize different types of DCT and Q implementations. Then the adaptive scheme is further optimized by considering a more efficient condition to sufficiently detect all-zero DCT blocks. As a result, a hybrid model is developed. Compared with other methods in the literature, the proposed hybrid model is able to detect more ZQDCT coefficients and hence reduce more computations for H.264 encoding. It is shown by experimental results that the proposed hybrid model can achieve the best performance in reducing computations and obtain almost the same rate-distortion (R-D) performance as the original encoder in the H.264 reference software JM9.5 Hanli Wang, Sam Kwong |
IEEE Trans. Multim. | 1 |
| 2007 | An Efficient Mode Decision Algorithm for H.264/AVC Encoding OptimizationabstractThe H.264 video coding standard significantly outperforms previous standards in terms of coding efficiency. However, this comes as a cost of extremely high computational complexity due to mode decision where variable block size motion estimation (ME) is employed. In this paper, we propose an efficient algorithm to jointly optimize mode decision and ME. A theoretical analysis is performed to study the sufficient condition to detect all-zero blocks in H.264, and thus adaptive thresholds are derived to early terminate mode decision and ME. Besides the aforementioned early termination technique, the proposed algorithm also introduces temporal-spatial checking, thresholds based prediction and monotonic error surface based prediction methods to skip checking unnecessary modes. Experimental results demonstrate that the proposed algorithm can significantly reduce the computational complexity of H.264 encoding while maintaining almost the same rate distortion (RD) performance as the original encoder Hanli Wang, Sam Kwong, Chi-Wah Kok |
IEEE Trans. Multim. | 1 |
| 2006 | Effectively Detecting All-Zero DCT Blocks for H.264 OptimizationabstractThis paper presents a novel efficient algorithm to reduce redundant DCT and quantization computations for H.264 encoding. A theoretical analysis is performed to study the sufficient condition for DCT coefficients to be quantized to zeros in H.264. As a result, a more tight sufficient condition is derived to early detect all-zero 4x4 DCT blocks. Compared with other algorithms, the proposed algorithm provides a more precise and efficient condition to predict all-zero DCT blocks. The experimental results demonstrate that the proposed algorithm outperforms other algorithms in terms of all the evaluated performances. Hanli Wang, Sam Kwong, Chi-Wah Kok |
ICIP | 1 |
| 2006 | Analytical Model of Zero Quantized DCT Coefficients for Video Encoder OptimizationabstractThis paper proposes a novel analytical model to predict zero quantized DCT coefficients for fast video encoding. The dynamic range of quantized DCT coefficients are analyzed and a threshold scheme is derived in order to determine DCT and quantization computations to be skipped without video quality degradation. The proposed model is compared with other models in the literature. Experimental results demonstrate that the proposed analytical model can greatly reduce the computational complexity of video encoding without any performance degradation, and outperforms other models Hanli Wang, Sam Kwong, Chi-Wah Kok |
ICME | 1 |
| 2006 | Fast video coding based on Gaussian model of DCT coefficientsabstractDiscrete cosine transform (DCT), quantization (Q), inverse quantization (IQ) and inverse DCT (IDCT) are the building blocks in video coding standards. A lot of computations are required to perform the DCT, Q, IQ, and IDCT operations. With this concern, a novel statistical model based on Gaussian distribution is proposed to predict zero quantized DCT (ZQDCT) coefficients in this paper to reduce the computational complexity of video encoding. Compared with other widely used models in the literature, the proposed model can achieve the best real time performance. Experimental results demonstrate that the proposed statistical model is superior to others in terms of processing speed at the expense of negligible degradation of video quality. Hanli Wang, Sam Kwong, Chi-Wah Kok |
ISCAS | 1 |
| 2006 | Novel quantized DCT for video encoder optimizationabstractA novel technique is proposed to reduce the computational complexity of discrete cosine transform (DCT)-based video encoders. The proposed method merges the DCT and quantization into a single procedure, which is referred to as the novel quantized DCT (NQDCT), such that the DCT output does not need to be explicitly quantized. Thus, a lot of computations related to quantization can be saved. The video encoder's performance using NQDCT is evaluated by comparing it with those using the traditional separate DCT and quantization method and the quantized DCT (QDCT) method. Simulation results demonstrate that the proposed NQDCT outperforms the other two methods in improving the real-time performance for video encoding. Hanli Wang, Ming-Yan Chan, Sam Kwong, Chi-Wah Kok |
IEEE Signal Process. Lett. | 1 |
| 2006 | Efficient prediction algorithm of integer DCT coefficients for H.264/AVC optimizationabstractThis paper presents a novel efficient prediction algorithm to reduce redundant discrete cosine transform (DCT) and quantization computations for H.264 encoding optimization. A theoretical analysis is performed to study the sufficient condition for DCT coefficients to be quantized to zeros. As a result, three sufficient conditions corresponding to three types of transform and quantization methods in H.264 are proposed. Compared with other algorithms in the literature, the proposed algorithm derives more precise and efficient conditions to predict zero quantized DCT coefficients. Both the theoretical analysis and experimental results demonstrate that the proposed algorithm is superior to other algorithms in terms of the computational complexity reduction, encoded video quality, false acceptance rate, and false rejection rate. Hanli Wang, Sam Kwong, Chi-Wah Kok |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2005 | Anomaly Intrusion Detection Using Multi-Objective Genetic Fuzzy System and Agent-Based Evolutionary Computation FrameworkabstractIn this paper, we present a multi-objective genetic fuzzy system for anomaly intrusion detection. The proposed system extracts accurate and interpret able fuzzy rule-based knowledge from network data using an agent-based evolutionary computation framework. The experimental results on KDD-Cup99 intrusion detection benchmark data demonstrate that our system can achieve high detection rate for intrusion attacks and low false positive rate for normal network traffic. Chi-Ho Tsang, Sam Kwong, Hanli Wang |
ICDM | 3 |
| 2005 | Multi-objective hierarchical genetic algorithm for interpretable fuzzy rule-based knowledge extraction
Hanli Wang, Sam Kwong, Yaochu Jin, Kim-Fung Man |
Fuzzy Sets Syst. | 1 |
| 2005 | Agent-based evolutionary approach for interpretable rule-based knowledge extractionabstractAn agent-based evolutionary approach is proposed to extract interpretable rule-based knowledge. In the multiagent system, each fuzzy set agent autonomously determines its own fuzzy sets information, such as the number and distribution of the fuzzy sets. It can further consider the interpretability of fuzzy systems with the aid of hierarchical chromosome formulation and interpretability-based regulation method. Based on the obtained fuzzy sets, the Pittsburgh-style approach is applied to extract fuzzy rules that take both the accuracy and interpretability of fuzzy systems into consideration. In addition, the fuzzy set agents can cooperate with each other to exchange their fuzzy sets information and generate offspring agents. The parent agents and their offspring compete with each other through the arbitrator agent based on the criteria associated with the accuracy and interpretability to allow them to remain competitive enough to move into the next population. The performance with emphasis upon both the accuracy and interpretability based on the agent-based evolutionary approach is studied through some benchmark problems reported in the literature. Simulation results show that the proposed approach can achieve a good tradeoff between the accuracy and interpretability of fuzzy systems. Hanli Wang, Sam Kwong, Yaochu Jin, Kim-Fung Man |
IEEE Trans. Syst. Man Cybern. Part C | 1 |
| 2004 | Optimization of Gaussian Mixture Model Parameters for Speaker Identification
Q. Y. Hong, Sam Kwong, Hanli Wang |
GECCO (2) | 3 |