Wenjun Zeng 0001

dblp:57/145 · also Wenjun Kevin Zeng · DBLP profile ↗
← Back
251ranked-venue papers
24as first author
80since 2021 · last 2026
0000-0003-2531-3137ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 187 · 22 first-author · 47 since 2021Artificial intelligence and machine learning · 99 · 59 since 2021Computer networks · 11 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 6 since 2021Systems, architecture and hardware · 4 · 1 since 2021Security and privacy · 3Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Restoration Adaptation for Semantic Segmentation on Low Quality Images
abstract
Abstract In real-world scenarios, the performance of semantic segmentation often deteriorates when processing low-quality (LQ) images, which may lack clear semantic structures and high-frequency details. Although image restoration techniques offer a promising direction for enhancing degraded visual content, conventional real-world image restoration (Real-IR) models primarily focus on pixel-level fidelity and often fail to recover task-relevant semantic cues, limiting their effectiveness when directly applied to downstream vision tasks. Conversely, existing segmentation models trained on high-quality data lack robustness under real-world degradations. In this paper, we propose Restoration Adaptation for Semantic Segmentation (RASS), which effectively integrates semantic image restoration into the segmentation process, enabling high-quality semantic segmentation on the LQ images directly. Specifically, we first propose a Semantic-Constrained Restoration (SCR) model, which injects segmentation priors into the restoration model by aligning its cross-attention maps with segmentation masks, encouraging semantically faithful image reconstruction. Then, RASS transfers semantic restoration knowledge into segmentation through LoRA-based module merging and task-specific fine-tuning, thereby enhancing the model’s robustness to LQ images. To validate the effectiveness of our framework, we construct a real-world LQ image segmentation dataset with high-quality annotations, and conduct extensive experiments on both synthetic and real-world LQ benchmarks. The results show that SCR and RASS significantly outperform state-of-the-art methods in segmentation and restoration tasks. Code, models, and datasets will be available at https://github.com/Ka1Guan/RASS.git .
Rongyuan Wu, Shuai Li 0014, Wentao Zhu 0001, Wenjun Zeng 0001, Lei Zhang 0006
Int. J. Comput. Vis.5
2026 Hierarchical Context Alignment With Disentangled Geometric and Temporal Modeling for Semantic Occupancy Prediction
abstract
Camera-based 3D Semantic Occupancy Prediction (SOP) is crucial for understanding complex 3D scenes from limited 2D image observations. Existing SOP methods typically aggregate contextual features to assist the occupancy representation learning, alleviating issues like occlusion or ambiguity. However, these solutions often face misalignment issues wherein the corresponding features at the same position across different frames may have different semantic meanings during the aggregation process, which leads to unreliable contextual fusion results and an unstable representation learning process. To address this problem, we introduce a new Hierarchical context alignment paradigm for a more accurate SOP (Hi-SOP). Hi-SOP first disentangles the geometric and temporal context for separate alignment, which two branches are then composed to enhance the reliability of SOP. This parsing of the visual input into a local-global alignment hierarchy includes: (I) disentangled geometric and temporal separate alignment, within each leverages depth confidence and camera pose as prior for relevant feature matching respectively; (II) global alignment and composition of the transformed geometric and temporal volumes based on semantics consistency. Our method outperforms SOTAs for semantic scene completion on the SemanticKITTI & NuScenes-Occupancy datasets and LiDAR semantic segmentation on the NuScenes dataset.
Bohan Li 0015, Xin Jin 0014, Jiajun Deng, Yasheng Sun, Wenjun Zeng 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2026 A Survey of Behavior Foundation Model: Next-Generation Whole-Body Control System of Humanoid Robots
abstract
Humanoid robots are drawing significant attention as versatile platforms for complex motor control, human-robot interaction, and general-purpose physical intelligence. However, achieving efficient whole-body control (WBC) in humanoids remains a fundamental challenge due to sophisticated dynamics, underactuation, and diverse task requirements. While learning-based controllers have shown promise for complex tasks, their reliance on labor-intensive and costly retraining for new scenarios limits real-world applicability. To address these limitations, behavior(al) foundation models (BFMs) have emerged as a new paradigm that leverages large-scale pre-training to learn reusable primitive skills and broad behavioral priors, enabling zero-shot or rapid adaptation to a wide range of downstream tasks. In this paper, we present a comprehensive overview of BFMs for humanoid WBC, tracing their development across diverse pre-training pipelines. Furthermore, we discuss real-world applications, current limitations, urgent challenges, and future opportunities, positioning BFMs as a key approach toward scalable and general-purpose humanoid intelligence. Finally, we provide a curated and regularly updated collection of BFM papers and projects to facilitate further research, which is available at https://github.com/yuanmingqi/awesome-bfm-papers.
Mingqi Yuan, Tao Yu 0012, Wenqi Ge, Xiuyong Yao, Huijiang Wang, Jiayu Chen 0006, Bo Li 0037, Wei Zhang 0262, Wenjun Zeng 0001, Hua Chen 0007, Xin Jin 0014
IEEE Trans. Pattern Anal. Mach. Intell.10
2025 RLLTE: Long-Term Evolution Project of Reinforcement Learning
abstract
We present RLLTE: a long-term evolution, extremely modular, and open-source framework for reinforcement learning (RL) research and application. Beyond delivering top-notch algorithm implementations, RLLTE also serves as a toolkit for developing algorithms. More specifically, RLLTE decouples the RL algorithms completely from the exploitation-exploration perspective, providing a large number of components to accelerate algorithm development and evolution. In particular, RLLTE is the first RL framework to build a comprehensive ecosystem, which includes model training, evaluation, deployment, benchmark hub, and large language model (LLM)-empowered copilot. RLLTE is expected to set standards for RL engineering practice and be highly stimulative for industry and academia. Our documentation, examples, and source code are available at https://github.com/RLE-Foundation/rllte.
Mingqi Yuan, Zequn Zhang, Shihao Luo, Bo Li 0037, Xin Jin 0014, Wenjun Zeng 0001
AAAI7
2025 UniScene: Unified Occupancy-centric Driving Scene Generation
abstract
Generating high-fidelity, controllable, and annotated training data is critical for autonomous driving. Existing methods typically generate a single data form directly from a coarse scene layout, which not only fails to output rich data forms required for diverse downstream tasks but also struggles to model the direct layout-to-data distribution. In this paper, we introduce UniScene, the first unified framework for generating three key data forms — semantic occupancy, video, and LiDAR — in driving scenes. UniScene employs a progressive generation process that decomposes the complex task of scene generation into two hierarchical steps: (a) first generating semantic occupancy from a customized scene layout as a meta scene representation rich in both semantic and geometric information, and then (b) conditioned on occupancy, generating video and LiDAR data, respectively, with two novel transfer strategies of Gaussian-based Joint Rendering and Prior-guided Sparse Modeling. This occupancy-centric approach reduces the generation burden, especially for intricate scenes, while providing detailed intermediate representations for the subsequent generation stages. Extensive experiments demonstrate that UniScene outperforms previous SOTAs in the occupancy, video, and LiDAR generation, which also indeed benefits downstream driving tasks. The Project is available at https://arlo0o.github.io/uniscene/.
Bohan Li 0015, Jiazhe Guo, Hongsi Liu, Yingshuang Zou, Yikang Ding, Xiwu Chen, Hu Zhu, Feiyang Tan, Tiancai Wang, Shuchang Zhou 0001, Li Zhang 0040, Xiaojuan Qi 0001, Hao Zhao 0002, Mu Yang, Wenjun Zeng 0001, Xin Jin 0014
CVPR16
2025 Disentangled World Models: Learning to Transfer Semantic Knowledge from Distracting Videos for Reinforcement Learning
abstract
Training visual reinforcement learning (RL) in practical scenarios presents a significant challenge, $\textit{i.e.,}$ RL agents suffer from low sample efficiency in environments with variations. While various approaches have attempted to alleviate this issue by disentangled representation learning, these methods usually start learning from scratch without prior knowledge of the world. This paper, in contrast, tries to learn and understand underlying semantic variations from distracting videos via offline-to-online latent distillation and flexible disentanglement constraints. To enable effective cross-domain semantic knowledge transfer, we introduce an interpretable model-based RL framework, dubbed Disentangled World Models (DisWM). Specifically, we pretrain the action-free video prediction model offline with disentanglement regularization to extract semantic knowledge from distracting videos. The disentanglement capability of the pretrained model is then transferred to the world model through latent distillation. For finetuning in the online environment, we exploit the knowledge from the pretrained model and introduce a disentanglement constraint to the world model. During the adaptation phase, the incorporation of actions and rewards from online environment interactions enriches the diversity of the data, which in turn strengthens the disentangled representation learning. Experimental results validate the superiority of our approach on various benchmarks.
Qi Wang 0080, Baao Xie, Xin Jin 0014, Yunbo Wang, Liaomo Zheng, Xiaokang Yang 0001, Wenjun Zeng 0001
ICCV9
2025 Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions
Liang Xu 0012, Chengqun Yang, Zili Lin, Fei Xu 0008, Congsheng Xu, Yiyi Zhang 0002, Jie Qin 0004, Xingdong Sheng, Yunhui Liu 0006, Xin Jin 0014, Yichao Yan, Wenjun Zeng 0001, Xiaokang Yang 0001
ICCV13
2025 ULTHO: Ultra-Lightweight Yet Efficient Hyperparameter Optimization in Deep Reinforcement Learning
abstract
Hyperparameter optimization (HPO) is a billion-dollar problem in machine learning, which significantly impacts the training efficiency and model performance. However, achieving efficient and robust HPO in deep reinforcement learning (RL) is consistently challenging due to its high non-stationarity and computational cost. To tackle this problem, existing approaches attempt to adapt common HPO techniques (e.g., population-based training or Bayesian optimization) to the RL scenario. However, they remain sample-inefficient and computationally expensive, which cannot facilitate a wide range of applications. In this paper, we propose ULTHO, an ultra-lightweight yet powerful framework for fast HPO in deep RL within single runs. Specifically, we formulate the HPO process as a multi-armed bandit with clustered arms (MABC) and link it directly to long-term return optimization. ULTHO also provides a quantified and statistical perspective to filter the HPs efficiently. We test ULTHO on benchmarks including ALE, Procgen, MiniGrid, and PyBullet. Extensive experiments demonstrate that the ULTHO can achieve superior performance with a simple architecture, contributing to the development of advanced and automated RL systems.
Mingqi Yuan, Bo Li 0037, Xin Jin 0014, Wenjun Zeng 0001
ICCV4
2025 Hybrid-Grained Feature Aggregation with Coarse-to-Fine Language Guidance for Self-Supervised Monocular Depth Estimation
abstract
Current self-supervised monocular depth estimation (MDE) approaches encounter performance limitations due to insufficient semantic-spatial knowledge extraction. To address this challenge, we propose Hybrid-depth, a novel framework that systematically integrates foundation models (e.g., CLIP and DINO) to extract visual priors and acquire sufficient contextual information for MDE. Our approach introduces a coarse-to-fine progressive learning framework: 1) Firstly, we aggregate multi-grained features from CLIP (global semantics) and DINO (local spatial details) under contrastive language guidance. A proxy task comparing close-distant image patches is designed to enforce depth-aware feature alignment using text prompts; 2) Next, building on the coarse features, we integrate camera pose information and pixel-wise language alignment to refine depth predictions. This module seamlessly integrates with existing self-supervised MDE pipelines (e.g., Monodepth2, ManyDepth) as a plug-and-play depth encoder, enhancing continuous depth estimation. By aggregating CLIP's semantic context and DINO's spatial details through language guidance, our method effectively addresses feature granularity mismatches. Extensive experiments on the KITTI benchmark demonstrate that our method significantly outperforms SOTA methods across all metrics, which also indeed benefits downstream tasks like BEV perception. Code is available at https://github.com/Zhangwenyao1/Hybrid-depth.
Hongsi Liu, Bohan Li 0015, Jiawei He 0002, Zekun Qi, Yunnan Wang, Shengyang Zhao, Xinqiang Yu, Wenjun Zeng 0001, Xin Jin 0014
ICCV9
2025 Open-World Reinforcement Learning over Long Short-Term Imagination
abstract
Training visual reinforcement learning agents in a high-dimensional open world presents significant challenges. While various model-based methods have improved sample efficiency by learning interactive world models, these agents tend to be “short-sighted”, as they are typically trained on short snippets of imagined experiences. We argue that the primary challenge in open-world decision-making is improving the exploration efficiency across a vast state space, especially for tasks that demand consideration of long-horizon payoffs. In this paper, we present LS-Imagine, which extends the imagination horizon within a limited number of state transition steps, enabling the agent to explore behaviors that potentially lead to promising long-term feedback. The foundation of our approach is to build a $\textit{long short-term world model}$. To achieve this, we simulate goal-conditioned jumpy state transitions and compute corresponding affordance maps by zooming in on specific areas within single images. This facilitates the integration of direct long-term values into behavior learning. Our method demonstrates significant improvements over state-of-the-art techniques in MineDojo.
Jiajian Li, Qi Wang 0080, Yunbo Wang, Xin Jin 0014, Yang Li 0041, Wenjun Zeng 0001, Xiaokang Yang 0001
ICLR6
2025 Representation Disentanglement for Semantic Coding
abstract
The learned image compression methods have achieved advances in both human perception and machine vision. However, previous methods focus on transmitting visual symbols losslessly instead of precisely conveying semantic meaning, resulting in bandwidth waste, especially for AI applications. In this work, we study "Semantic Coding" and propose a novel compression method based on representation disentanglement, which understands images at the attribute level and separates semantic factors into different parts, achieving a semantically structured bitstream for transmission. Specifically, we first leverage a conditional generative diffusion procedure for a disentangled representation learning, which learns meaningful semantic attribute factors in the latent space of the image assisted by the extra language inductive bias. Furthermore, we employ a learned codec to compress the inner disentangled representations as a bitstream, where each part represents a specific semantic and can be used for purposely decoding. Experiments show that our semantic coding method could reconstruct high-quality images and enable encryption by shifting the inner semantics.
Jinming Liu 0001, Junhao Geng, Lexiang Lv, Wenjun Zeng 0001, Xin Jin 0014
ICME4
2025 Knowledge Regularized Negative Feature Tuning of Vision-Language Models for Out-of-Distribution Detection
abstract
Out-of-distribution (OOD) detection is crucial for building reliable machine learning models. Although negative prompt tuning has enhanced the OOD detection capabilities of vision-language models, these tuned models often suffer from reduced generalization performance on unseen classes and styles. To address this challenge, we propose a novel method called Knowledge Regularized Negative Feature Tuning (KR-NFT), which integrates an innovative adaptation architecture termed Negative Feature Tuning (NFT) and a corresponding knowledge-regularization (KR) optimization strategy. Specifically, NFT applies distribution-aware transformations to pre-trained text features, effectively separating positive and negative features into distinct spaces. This separation maximizes the distinction between in-distribution (ID) and OOD images. Additionally, we introduce image-conditional learnable factors through a lightweight meta-network, enabling dynamic adaptation to individual images and mitigating sensitivity to class and style shifts. Compared to traditional negative prompt tuning, NFT demonstrates superior efficiency and scalability. To optimize this adaptation architecture, the KR optimization strategy is designed to enhance the discrimination between ID and OOD sets while mitigating pre-trained knowledge forgetting. This enhances OOD detection performance on trained ID classes while simultaneously improving OOD detection on unseen ID datasets. Notably, when trained with few-shot samples from ImageNet dataset, KR-NFT not only improves ID classification accuracy and OOD detection but also significantly reduces the FPR95 by 5.44% under an unexplored generalization setting with unseen ID categories. Codes can be found at https://github.com/ZhuWenjie98/KRNFT.
Wenjie Zhu 0003, Yabin Zhang 0001, Xin Jin 0014, Wenjun Zeng 0001, Lei Zhang 0006
ACM Multimedia4
2025 DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
abstract
Recent advances in vision-language-action (VLA) models have shown promise in integrating image generation with action prediction to improve generalization and reasoning in robot manipulation. However, existing methods are limited to challenging image-based forecasting, which suffers from redundant information and lacks comprehensive and critical world knowledge, including dynamic, spatial and semantic information. To address these limitations, we propose DreamVLA, a novel VLA framework that integrates comprehensive world knowledge forecasting to enable inverse dynamics modeling, thereby establishing a perception-prediction-action loop for manipulation tasks. Specifically, DreamVLA introduces a dynamic-region-guided world knowledge prediction, integrated with the spatial and semantic cues, which provide compact yet comprehensive representations for action planning. This design aligns with how humans interact with the world by first forming abstract multimodal reasoning chains before acting. To mitigate interference among the dynamic, spatial and semantic information during training, we adopt a block-wise structured attention mechanism that masks their mutual attention, preventing information leakage and keeping each representation clean and disentangled. Moreover, to model the conditional distribution over future actions, we employ a diffusion-based transformer that disentangles action representations from shared latent features. Extensive experiments on both real-world and simulation environments demonstrate that DreamVLA achieves 76.7 success rate on real robot tasks and 4.44 average length on the CALVIN ABC-D benchmarks.
Hongsi Liu, Zekun Qi, Yunnan Wang, Xinqiang Yu, Jiazhao Zhang, Runpei Dong, Jiawei He 0002, He Wang 0010, Zhizheng Zhang 0011, Li Yi 0001, Wenjun Zeng 0001, Xin Jin 0014
NeurIPS12
2025 Cross-Modal World Models for Offline Visual Reinforcement Learning
abstract
Offline reinforcement learning (RL) with visual pixels encounters two primary challenges: overfitting in representation learning induced by limited data, and value overestimation of out-of-distribution states. Recent work has adopted accessible simulators to mitigate these issues, but rendering images introduces additional computational costs and training challenges. In this paper, we study the problem of anti-exploration with accessible state-based simulators for effective learning of offline visual control. To address these challenges, we introduce a model-based RL framework, dubbed Cross-Modal World Models (X-MWM). Concretely, we build two independent agents: a source model trained on low-dimensional states and a target model that learns from high-dimensional images. Initially, in light of the reward function discrepancy between the domains, we pretrain the source agent with latent disagreement-based intrinsic rewards. Subsequently, to prevent overfitting in offline representation learning, cross-modal latent alignment is employed to close the distance of latent state distributions. In this way, during the target agent training phase, the value of the source critic serves as an anti-exploration constraint to adjust the learning target of the offline RL agent, which encourages more conservative behavior, effectively alleviating value overestimation induced by out-of-distribution states. Experimental results of various robotic manipulation tasks on MetaWorld validate the superiority of our approach.
Qi Wang 0080, Xin Jin 0014, Baao Xie, Xiaokang Yang 0001, Wenjun Zeng 0001
SMC5
2025 Hierarchical Procedural Framework for Low-latency Robot-Assisted Hand-Object Interaction
abstract
Advances in robotics have been driving the development of human-robot interaction (HRI) technologies. However, accurately perceiving human actions and achieving adaptive control remains a challenge in facilitating seamless coordination between human and robotic movements. In this paper, we propose a hierarchical procedural framework to enable dynamic robot-assisted hand-object interaction (HOI). An open-loop hierarchy leverages the RGB-based 3D reconstruction of the human hand, based on which motion primitives have been designed to translate hand motions into robotic actions. The low-level coordination hierarchy fine-tunes the robot’s action by using the continuously updated 3D hand models. Experimental validation demonstrates the effectiveness of the hierarchical control architecture. The adaptive coordination between human and robot behavior has achieved a delay of ≤ 0.3 seconds in the tele-interaction scenario. A case study of ring-wearing tasks indicates the potential application of this work in assistive technologies such as healthcare and manufacturing.
Mingqi Yuan, Huijiang Wang, Kai-Fung Chu, Fumiya Iida, Bo Li 0037, Wenjun Zeng 0001
SMC6
2025 Standard Codec is Enough: A Training-Free 4D Gaussian Compression with Dynamic UV Mapping
abstract
4D Gaussian Splatting (4DGS) has demonstrated advances in the dynamic scene representation. However, the time-varying attributes across frames introduce considerable storage and transmission costs, making 4DGS challenging to widely deploy. Existing compression methods struggle to obtain inter-frame residuals due to the unstructured nature of Gaussian representations, making explicit motion estimation and residual modeling inherently challenging. To address these, we propose a Training-Free 4D Gaussian Compression framework, TF4DGC, which transforms 4D Gaussian into a well-structured 2D representation, easy to estimate motion for coding, via a UV mapping. Specifically, we project 3D Gaussians onto a canonical sphere to obtain temporally consistent UV coordinates, and organize per-frame Gaussian attributes into multi-channel video sequences. This design enables the direct use of standard video codecs (e.g., AVC, HEVC) for compression, which is compatible with widespread hardware decoder support on laptops and mobile devices. Experimental results show that our method efficiently compresses both reconstructed and generated Gaussian scenarios, highlighting its general applicability. Our method offers a scalable and practical solution for 4DGS compression and facilitates real-time deployment in bandwidth constrained environments.
Jinming Liu 0001, Shengyang Zhao, Qiang Hu 0003, Zhibo Chen 0001, Wenjun Zeng 0001, Xin Jin 0014
VCIP6
2025 Quadtree Partitioning-based Visual Token Pruning for MLLMs Considering Information Density
abstract
Multimodal Large Language Models (MLLMs) excel at comprehensive understanding by integrating visual and textual information. However, their inference speed is often bottlenecked by redundant visual token inputs. Existing methods tend to alleviate this issue with a heuristic pruning strategy based on token importance, tailored to certain commonly adopted vision encoders like CLIP. In this paper, we propose a novel training-free token pruning method based on a well-designed metric of information density, where we decide which tokens are retained according to their entropy, following the classic information theory. Based on that, we further propose a quadtree partitioning strategy, in which we retain these tokens with higher entropy so as to preserve the visual spatial structure while allocating more tokens to more informative regions. Experiments on LLaVA-v1.5-7B and 13B across six benchmarks show our method achieves state-of-the-art performance—retaining over 90% of full-token accuracy even at a 6.25% token budget—while cutting TFLOPs by up to 20% compared to FastV and by 81% compared to the original LLaVA-v1.5.
Yuntao Wei, Jinming Liu 0001, Shengyang Zhao, Zhibo Chen 0001, Wenjun Zeng 0001, Xin Jin 0014
VCIP5
2025 Diffusion Models for Image Restoration and Enhancement: A Comprehensive Survey
Xin Li 0082, Yulin Ren, Xin Jin 0014, Cuiling Lan, Xingrui Wang, Wenjun Zeng 0001, Xinchao Wang, Zhibo Chen 0001
Int. J. Comput. Vis.6
2025 Canvas: Compositional Generation for Art Painting With Seamless Subject-Driven Infusion
Yunnan Wang, Lexiang Lv, Zequn Zhang, Xiaoyu Shen 0001, Xin Jin 0014, Wenjun Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.8
2025 Exploring Contrastive Pre-Training for Domain Connections in Medical Image Segmentation
abstract
Unsupervised domain adaptation (UDA) in medical image segmentation aims to improve the generalization of deep models by alleviating domain gaps caused by inconsistency across equipment, imaging protocols, and patient conditions. However, existing UDA works remain insufficiently explored and present great limitations: 1) Exhibit cumbersome designs that prioritize aligning statistical metrics and distributions, which limits the model's flexibility and generalization while also overlooking the potential knowledge embedded in unlabeled data; 2) More applicable in a certain domain, lack the generalization capability to handle diverse shifts encountered in clinical scenarios. To overcome these limitations, we introduce MedCon, a unified framework that leverages general unsupervised contrastive pre-training to establish domain connections, effectively handling diverse domain shifts without tailored adjustments. Specifically, it initially explores a general contrastive pre-training to establish domain connections by leveraging the rich prior knowledge from unlabeled images. Thereafter, the pre-trained backbone is fine-tuned using source-based images to ultimately identify per-pixel semantic categories. To capture both intra- and inter-domain connections of anatomical structures, we construct positive-negative pairs from a hybrid aspect of both local and global scales. In this regard, a shared-weight encoder-decoder is employed to generate pixel-level representations, which are then mapped into hyper-spherical space using a non-learnable projection head to facilitate positive pair matching. Comprehensive experiments on diverse medical image datasets confirm that MedCon outperforms previous methods by effectively managing a wide range of domain shifts and showcasing superior generalization capabilities.
Zequn Zhang, Yunnan Wang, Baao Xie, Yuhang Li 0005, Zhen Chen 0013, Xin Jin 0014, Wenjun Zeng 0001
IEEE Trans. Medical Imaging9
2025 Unleash the Power of Vision-Language Models by Visual Attention Prompt and Multimodal Interaction
abstract
Pre-trained vision-language models (VLMs), equipped with parameter-efficient tuning (PET) methods like prompting, have shown impressive knowledge transferability on new downstream tasks, but they are still prone to be limited by catastrophic forgetting and overfitting dilemma due to large gaps among tasks. Furthermore, the underlying physical mechanisms of prompt-based tuning methods (especially for visual prompting) remain largely unexplored. It is unclear why these methods work solely based on learnable parameters as prompts for adaptation. To address the above challenges, we present a new prompt-based framework for vision-language models, termed Uni-prompt. Our framework transfers VLMs to downstream tasks by designing visual prompts from an attention perspective that reduces the transfer/solution space, which enables the vision model to focus on task-relevant regions of the input image while also learning task-specific knowledge. Additionally, Uni-prompt further aligns visual-text prompts learning through a pretext task with masked representation modeling interactions, which implicitly learns a global cross-modal matching between visual and language concepts for consistency. We conduct extensive experiments on the few-shot classification task and achieve significant improvement using our Uni-prompt method while requiring minimal extra parameters cost.
Letian Wu, Zequn Zhang, Tao Yu 0012, Chao Ma 0004, Xin Jin 0014, Xiaokang Yang 0001, Wenjun Zeng 0001
IEEE Trans. Multim.8
2024 One at a Time: Progressive Multi-Step Volumetric Probability Learning for Reliable 3D Scene Perception
abstract
Numerous studies have investigated the pivotal role of reliable 3D volume representation in scene perception tasks, such as multi-view stereo (MVS) and semantic scene completion (SSC). They typically construct 3D probability volumes directly with geometric correspondence, attempting to fully address the scene perception tasks in a single forward pass. However, such a single-step solution makes it hard to learn accurate and convincing volumetric probability, especially in challenging regions like unexpected occlusions and complicated light reflections. Therefore, this paper proposes to decompose the complicated 3D volume representation learning into a sequence of generative steps to facilitate fine and reliable scene perception. Considering the recent advances achieved by strong generative diffusion models, we introduce a multi-step learning framework, dubbed as VPD, dedicated to progressively refining the Volumetric Probability in a Diffusion process. Specifically, we first build a coarse probability volume from input images with the off-the-shelf scene perception baselines, which is then conditioned as the basic geometry prior before being fed into a 3D diffusion UNet, to progressively achieve accurate probability distribution modeling. To handle the corner cases in challenging areas, a Confidence-Aware Contextual Collaboration (CACC) module is developed to correct the uncertain regions for reliable volumetric learning based on multi-scale contextual contents. Moreover, an Online Filtering (OF) strategy is designed to maintain representation consistency for stable diffusion sampling. Extensive experiments are conducted on scene perception tasks including multi-view stereo (MVS) and semantic scene completion (SSC), to validate the efficacy of our method in learning reliable volumetric representations. Notably, for the SSC task, our work stands out as the first to surpass LiDAR-based methods on the SemanticKITTI dataset.
Bohan Li 0015, Yasheng Sun, Jingxin Dong 0002, Jinming Liu 0001, Xin Jin 0014, Wenjun Zeng 0001
AAAI7
2024 Consistency Prior Matters: Biomedical-Prompting Dual Augmentation for Domain Adaptive Medical Image Segmentation
abstract
Existing domain adaptive medical image segmentation works typically rely on style transfer techniques to mitigate the unexpected domain gap, which inevitably suffers from synthesized artifacts or unreasonable stylization. In this paper, we propose to inject biomedical-related prior knowledge (i.e., intensity and anatomical consistency) as regularization in a prompting manner, bridging the domain gap across modalities. Technically, we develop an efficient scheme called Biomedical-Prompting Dual Augmentation (BPDA) to learn domain-invariant representations by enforcing consistent model predictions across different augmented views. BPDA augments unpaired source and target images from intensity and anatomical aspects in a dual manner, while prompting the framework to fully understand the anatomical structure-invariant features. In this way, our method captures discriminative inherent representations on cross-modality scenarios. Furthermore, we also introduce a Cross-Domain Prototype Denoising (CDPD) in BPDA to refine pseudo-labeling results with the class centroids for a reliable augmentation. Extensive experiments on the cross-modality abdominal and cardiac segmentation benchmarks demonstrate the superiority of our method over state-of-the-art alternatives.
Yunnan Wang, Zequn Zhang, Xin Jin 0014, Wenjun Zeng 0001
BIBM4
2024 Inter-X: Towards Versatile Human-Human Interaction Analysis
abstract
The analysis of the ubiquitous human-human interactions is pivotal for understanding humans as social beings. Existing human-human interaction datasets typically suffer from inaccurate body motions, lack of hand gestures and fine- grained textual descriptions. To better perceive and generate human-human interactions, we propose Inter-X, a currently largest human-human interaction dataset with accurate body movements and diverse interaction patterns, together with detailed hand gestures. The dataset includes
Liang Xu 0012, Xintao Lv, Yichao Yan, Xin Jin 0014, Shuwen Wu, Congsheng Xu, Yizhou Zhou, Fengyun Rao, Xingdong Sheng, Yunhui Liu 0006, Wenjun Zeng 0001, Xiaokang Yang 0001
CVPR12
2024 ReGenNet: Towards Human Action-Reaction Synthesis
abstract
Humans constantly interact with their surrounding environments. Current human-centric generative models mainly focus on synthesizing humans plausibly interacting with static scenes and objects, while the dynamic human action-reaction synthesis for ubiquitous causal human-human interactions is less explored. Human-human interactions can be regarded as asymmetric with actors and reactors in atomic interaction periods. In this paper, we compre-hensively analyze the asymmetric, dynamic, synchronous, and detailed nature of human-human interactions and propose the first multi-setting human action-reaction synthe-sis benchmark to generate human reactions conditioned on given human actions. To begin with, we propose to an-notate the actor-reactor order of the interaction sequences for the NTU120, InterHuman, and Chi3D datasets. Based on them, a diffusion-based generative model with a Trans-former decoder architecture called ReGenNet together with an explicit distance-based interaction loss is proposed to predict human reactions in an online manner, where the future states of actors are unavailable to reactors. Quantitative and qualitative results show that our method can gener-ate instant and plausible human reactions compared to the baselines, and can generalize to unseen actor motions and viewpoint changes.
Liang Xu 0012, Yizhou Zhou, Yichao Yan, Xin Jin 0014, Wenhan Zhu, Fengyun Rao, Xiaokang Yang 0001, Wenjun Zeng 0001
CVPR8
2024 Closed-Loop Unsupervised Representation Disentanglement with β-VAE Distillation and Diffusion Probabilistic Feedback
Xin Jin 0014, Bohan Li 0015, Baao Xie, Jinming Liu 0001, Tao Yang 0032, Wenjun Zeng 0001
ECCV (45)8
2024 Hierarchical Temporal Context Learning for Camera-Based Semantic Scene Completion
Bohan Li 0015, Jiajun Deng, Zhujin Liang, Dalong Du, Xin Jin 0014, Wenjun Zeng 0001
ECCV (4)7
2024 Rate-Distortion-Cognition Controllable Versatile Neural Image Compression
Jinming Liu 0001, Ruoyu Feng 0001, Yunpeng Qi, Qiuyu Chen, Zhibo Chen 0001, Wenjun Zeng 0001, Xin Jin 0014
ECCV (56)6
2024 HIMO: A New Benchmark for Full-Body Human Interacting with Multiple Objects
Xintao Lv, Liang Xu 0012, Yichao Yan, Xin Jin 0014, Congsheng Xu, Shuwen Wu, Lincheng Li, Mengxiao Bi, Wenjun Zeng 0001, Xiaokang Yang 0001
ECCV (4)10
2024 Bridging Stereo Geometry and BEV Representation with Reliable Mutual Interaction for Semantic Scene Completion
Bohan Li 0015, Yasheng Sun, Zhujin Liang, Dalong Du, Zhuanghui Zhang, Yunnan Wang, Xin Jin 0014, Wenjun Zeng 0001
IJCAI9
2024 Scene Graph Disentanglement and Composition for Generalizable Complex Image Generation
abstract
There has been exciting progress in generating images from natural language or layout conditions. However, these methods struggle to faithfully reproduce complex scenes due to the insufficient modeling of multiple objects and their relationships. To address this issue, we leverage the scene graph, a powerful structured representation, for complex image generation. Different from the previous works that directly use scene graphs for generation, we employ the generative capabilities of variational autoencoders and diffusion models in a generalizable manner, compositing diverse disentangled visual clues from scene graphs. Specifically, we first propose a Semantics-Layout Variational AutoEncoder (SL-VAE) to jointly derive (layouts, semantics) from the input scene graph, which allows a more diverse and reasonable generation in a one-to-many mapping. We then develop a Compositional Masked Attention (CMA) integrated with a diffusion model, incorporating (layouts, semantics) with fine-grained attributes as generation guidance. To further achieve graph manipulation while keeping the visual content consistent, we introduce a Multi-Layered Sampler (MLS) for an "isolated" image editing effect. Extensive experiments demonstrate that our method outperforms recent competitors based on text, layout, or scene graph, in terms of generation rationality and controllability.
Yunnan Wang, Zequn Zhang, Baao Xie, Xihui Liu, Wenjun Zeng 0001, Xin Jin 0014
NeurIPS7
2024 Making Offline RL Online: Collaborative World Models for Offline Visual Reinforcement Learning
abstract
Training offline RL models using visual inputs poses two significant challenges, *i.e.*, the overfitting problem in representation learning and the overestimation bias for expected future rewards. Recent work has attempted to alleviate the overestimation bias by encouraging conservative behaviors. This paper, in contrast, tries to build more flexible constraints for value estimation without impeding the exploration of potential advantages. The key idea is to leverage off-the-shelf RL simulators, which can be easily interacted with in an online manner, as the “*test bed*” for offline policies. To enable effective online-to-offline knowledge transfer, we introduce CoWorld, a model-based RL approach that mitigates cross-domain discrepancies in state and reward spaces. Experimental results demonstrate the effectiveness of CoWorld, outperforming existing RL approaches by large margins.
Qi Wang 0080, Yunbo Wang, Xin Jin 0014, Wenjun Zeng 0001, Xiaokang Yang 0001
NeurIPS5
2024 Graph-based Unsupervised Disentangled Representation Learning via Multimodal Large Language Models
abstract
Disentangled representation learning (DRL) aims to identify and decompose underlying factors behind observations, thus facilitating data perception and generation. However, current DRL approaches often rely on the unrealistic assumption that semantic factors are statistically independent. In reality, these factors may exhibit correlations, which off-the-shelf solutions have yet to properly address. To tackle this challenge, we introduce a bidirectional weighted graph-based framework, to learn factorized attributes and their interrelations within complex data. Specifically, we propose a $\beta$-VAE based module to extract factors as the initial nodes of the graph, and leverage the multimodal large language model (MLLM) to discover and rank latent correlations, thereby updating the weighted edges. By integrating these complementary modules, our model successfully achieves fine-grained, practical and unsupervised disentanglement. Experiments demonstrate our method's superior performance in disentanglement and reconstruction. Furthermore, the model inherits enhanced interpretability and generalizability from MLLMs.
Baao Xie, Qiuyu Chen, Yunnan Wang, Zequn Zhang, Xin Jin 0014, Wenjun Zeng 0001
NeurIPS6
2024 Tell Codec What Worth Compressing: Semantically Disentangled Image Coding for Machine with LMMs
abstract
We present a new image compression paradigm to achieve "intelligently coding for machine" by cleverly leveraging the common sense of Large Multimodal Models (LMMs). We are motivated by the evidence that large language/multimodal models are powerful general-purpose semantics predictors for understanding the real world. Different from traditional image compression typically optimized for human eyes, the image coding for machines (ICM) framework we focus on requires the compressed bitstream to more comply with different downstream intelligent analysis tasks. To this end, we employ LMM to${\text{tell codec what to compress}}$: 1) first utilize the powerful semantic understanding capability of LMMs w.r.t object grounding, identification, and importance ranking via prompts, to disentangle image content before compression, 2) and then based on these semantic priors we accordingly encode and transmit objects of the image in order with a structured bitstream. In this way, diverse vision benchmarks including image classification, object detection, instance segmentation, etc., can be well supported with such a semantically structured bitstream. We dub our method "SDComp" for "Semantically Disentangled Compression", and compare it with state-of-the-art codecs on a wide variety of different vision tasks. SDComp codec leads to more flexible reconstruction results, promised decoded visual quality, and a more generic/satisfactory intelligent task-supporting ability.
Jinming Liu 0001, Yuntao Wei, Junyan Lin, Shengyang Zhao, Heming Sun, Zhibo Chen 0001, Wenjun Zeng 0001, Xin Jin 0014
VCIP7
2024 Understanding mobile GUI: From pixel-words to screen-sentences
Jingwen Fu, Yuwang Wang, Wenjun Zeng 0001, Nanning Zheng 0001
Neurocomputing4
2024 Correlation-Embedded Transformer Tracking: A Single-Branch Framework
abstract
Developing robust and discriminative appearance models has been a long-standing research challenge in visual object tracking. In the prevalent Siamese-based paradigm, the features extracted by the Siamese-like networks are often insufficient to model the tracked targets and distractor objects, thereby hindering them from being robust and discriminative simultaneously. While most Siamese trackers focus on designing robust correlation operations, we propose a novel single-branch tracking framework inspired by the transformer. Unlike the Siamese-like feature extraction, our tracker deeply embeds cross-image feature correlation in multiple layers of the feature network. By extensively matching the features of the two images through multiple layers, it can suppress non-target features, resulting in target-aware feature extraction. The output features can be directly used to predict target locations without additional correlation steps. Thus, we reformulate the two-branch Siamese tracking as a conceptually simple, fully transformer-based Single-Branch Tracking pipeline, dubbed SBT. After conducting an in-depth analysis of the SBT baseline, we summarize many effective design principles and propose an improved tracker dubbed SuperSBT. SuperSBT adopts a hierarchical architecture with a local modeling layer to enhance shallow-level features. A unified relation modeling is proposed to remove complex handcrafted layer pattern designs. SuperSBT is further improved by masked image modeling pre-training, integrating temporal modeling, and equipping with dedicated prediction heads. Thus, SuperSBT outperforms the SBT baseline by 4.7%,3.0%, and 4.5% AUC scores in LaSOT, TrackingNet, and GOT-10K. Notably, SuperSBT greatly raises the speed of SBT from 37 FPS to 81 FPS. Extensive experiments show that our method achieves superior results on eight VOT benchmarks.
Wankou Yang, Chunyu Wang 0001, Yue Cao 0001, Chao Ma 0004, Wenjun Zeng 0001
IEEE Trans. Pattern Anal. Mach. Intell.7
2024 Domain Prompt Tuning via Meta Relabeling for Unsupervised Adversarial Adaptation
abstract
Unsupervised adversarial domain adaptation (ADA) aims to learn domain-invariant features by confusing a domain discriminator. As training goes on, the feature distributions of source and target samples are increasingly aligned/indistinguishable. The discrimination capability of the domain discriminator w.r.t. those aligned samples deteriorates due to the domain label of each sample is still fixed all through the learning process, which thus cannot effectively further drive the feature learning. A recently proposed method named Re-enforceable Adversarial Domain Adaptation (RADA) [1] tend to re-energize the domain discriminator during the training by using dynamic domain labels. Specifically, RADA sets up a heuristic criterion and uses it to relabel the well aligned target domain samples as source domain samples on the fly. In our study, we identify a critical problem of RADA: it is a kind of heuristic domain data re-partition solution without explicitly serving the adaptation task itself, suggesting that the criteria of RADA on which sample should be relabeled is hard to decide. To address the problem, we revisit domain relabeling process from a perspective of prompt tuning, and introduce a meta-optimized learnable prompts into RADA to replace some hand-craft designs in dynamic relabeling process, which scheme is named as RADA-prompt. Particularly, we employ a module of meta-prompter, which learns to adaptively relabel the samples based on the objective of serving UDA task. To train the meta-prompter, we leverage a domain alignment measurement and a classification measurement as the meta optimization objective. Extensive experiments on multiple unsupervised domain adaptation benchmarks demonstrate the effectiveness and superiority of RADA-prompt, this scheme also achieves state-of-the-art performance.
Xin Jin 0014, Cuiling Lan, Wenjun Zeng 0001, Zhibo Chen 0001
IEEE Trans. Multim.3
2023 NaviNeRF: NeRF-based 3D Representation Disentanglement by Latent Semantic Navigation
abstract
3D representation disentanglement aims to identify, decompose, and manipulate the underlying explanatory factors of 3D data, which helps AI fundamentally understand our 3D world. This task is currently under-explored and poses great challenges: (i) the 3D representations are complex and in general contains much more information than 2D image; (ii) many 3D representations are not well suited for gradient-based optimization, let alone disentanglement. To address these challenges, we use NeRF as a differentiable 3D representation, and introduce a self-supervised Navigation to identify interpretable semantic directions in the latent space. To our best knowledge, this novel method, dubbed NaviNeRF, is the first work to achieve fine-grained 3D disentanglement without any priors or supervisions. Specifically, NaviNeRF is built upon the generative NeRF pipeline, and equipped with an Outer Navigation Branch and an Inner Refinement Branch. They are complementary —— the outer navigation is to identify global-view semantic directions, and the inner refinement dedicates to fine-grained attributes. A synergistic loss is further devised to coordinate two branches. Extensive experiments demonstrate that NaviNeRF has a superior fine-grained 3D disentanglement ability than the previous 3D-aware models. Its performance is also comparable to editing-oriented models relying on semantic or geometry priors.*
Baao Xie, Bohan Li 0015, Zequn Zhang, Junting Dong, Xin Jin 0014, Jing-Yu Yang 0002, Wenjun Zeng 0001
ICCV7
2023 ActFormer: A GAN-based Transformer towards General Action-Conditioned 3D Human Motion Generation
abstract
We present a GAN-based Transformer for general action-conditioned 3D human motion generation, including not only single-person actions but also multi-person interactive actions. Our approach consists of a powerful Action-conditioned motion TransFormer (ActFormer) under a GAN training scheme, equipped with a Gaussian Process latent prior. Such a design combines the strong spatio-temporal representation capacity of Transformer, superiority in generative modeling of GAN, and inherent temporal correlations from the latent prior. Furthermore, ActFormer can be naturally extended to multi-person motions by alternately modeling temporal correlations and human interactions with Transformer encoders. To further facilitate research on multi-person motion generation, we introduce a new synthetic dataset of complex multi-person combat behaviors. Extensive experiments on NTU-13, NTU RGB+D 120, BABEL and the proposed combat dataset show that our method can adapt to various human motion representations and achieve superior performance over the state-of-the-art methods on both single-person and multi-person motion generation tasks, demonstrating a promising step towards a general human motion generator. The project website can be found at https://liangxuy.github.io/actformer/.
Liang Xu 0012, Jing Su 0005, Zhicheng Fang, Chenjing Ding, Weihao Gan, Yichao Yan, Xin Jin 0014, Xiaokang Yang 0001, Wenjun Zeng 0001, Wei Wu 0021
ICCV11
2023 Automatic Intrinsic Reward Shaping for Exploration in Deep Reinforcement Learning
abstract
We present AIRS: **A**utomatic **I**ntrinsic **R**eward **S**haping that intelligently and adaptively provides high-quality intrinsic rewards to enhance exploration in reinforcement learning (RL). More specifically, AIRS selects shaping function from a predefined set based on the estimated task return in real-time, providing reliable exploration incentives and alleviating the biased objective problem. Moreover, we develop an intrinsic reward toolkit to provide efficient and reliable implementations of diverse intrinsic reward approaches. We test AIRS on various tasks of MiniGrid, Procgen, and DeepMind Control Suite. Extensive simulation demonstrates that AIRS can outperform the benchmarking schemes and achieve superior performance with simple architecture.
Mingqi Yuan, Bo Li 0037, Xin Jin 0014, Wenjun Zeng 0001
ICML4
2023 WEDGE: Web-Image Assisted Domain Generalization for Semantic Segmentation
abstract
Domain generalization for semantic segmentation is highly demanded in real applications, where a trained model is expected to work well in previously unseen domains. One challenge lies in the lack of data which could cover the diverse distributions of the possible unseen domains for training. In this paper, we propose a WEb-image assisted Domain GEneralization (WEDGE) scheme, which is the first to exploit the diversity of web-crawled images for generalizable semantic segmentation. To explore and exploit the real-world data distributions, we collect web-crawled images which present large diversity in terms of weather conditions, sites, lighting, camera styles, etc. We also present a method which injects styles of the web-crawled images into training images on-the-fly during training, which enables the network to experience images of diverse styles with reliable labels for effective training. Moreover, we use the web-crawled images with their predicted pseudo labels for training to further enhance the capability of the network. Extensive experiments demonstrate that our method clearly outperforms existing domain generalization techniques.
Namyup Kim, Taeyoung Son, Jaehyun Pahk, Cuiling Lan, Wenjun Zeng 0001, Suha Kwak
ICRA5
2023 Composable Image Coding for Machine via Task-oriented Internal Adaptor and External Prior
abstract
Traditional image coding standards are typically optimized with a focus on human perception, which conflicts with the fact that most of the images are now analyzed by machines. To enable a variety of downstream intelligent tasks, contemporary approaches either utilize traditional codecs for image compression which are then used for task analysis, or develop a unified feature compression paradigm with deep learning techniques. However, they might suffer from accumulative errors and poor compatibility/generalization due to the conflict between standardized codecs and diverse machine tasks. We argue that a favorable image coding for machine (ICM) framework should have highly efficient adaptation capability, and take the ultimate task goals into account. Oriented at this, we propose a composable ICM solution dubbed Com-ICM, which develops plug-and-play lightweight internal adaptors injected into the codec architecture for efficient task transfer, and leverages off-the-shelf (large) models to provide external prior information for further task-oriented semantics learning. The internal adaptors (from the architectural aspect) and external priors (from the precondition aspect) complement each other, resulting in a mutually beneficial effect. We evaluate Com-ICM on diverse vision benchmarks, including image classification, object detection, and semantic segmentation, demonstrating its effectiveness and superiority. We are also actively submitting Com-ICM as a technical proposal to the international organization for standardization.
Jinming Liu 0001, Xin Jin 0014, Ruoyu Feng 0001, Zhibo Chen 0001, Wenjun Zeng 0001
VCIP5
2023 VoxelTrack: Multi-Person 3D Human Pose Estimation and Tracking in the Wild
abstract
We present VoxelTrack for multi-person 3D pose estimation and tracking from a few cameras which are separated by wide baselines. It employs a multi-branch network to jointly estimate 3D poses and re-identification (Re-ID) features for all people in the environment. In contrast to previous efforts which require to establish cross-view correspondence based on noisy 2D pose estimates, it directly estimates and tracks 3D poses from a 3D voxel-based representation constructed from multi-view images. We first discretize the 3D space by regular voxels and compute a feature vector for each voxel by averaging the body joint heatmaps that are inversely projected from all views. We estimate 3D poses from the voxel representation by predicting whether each voxel contains a particular body joint. Similarly, a Re-ID feature is computed for each voxel which is used to track the estimated 3D poses over time. The main advantage of the approach is that it avoids making any hard decisions based on individual images. The approach can robustly estimate and track 3D poses even when people are severely occluded in some cameras. It outperforms the state-of-the-art methods by a large margin on four public datasets including Shelf, Campus, Human3.6 M and CMU Panoptic.
Chunyu Wang 0001, Xinggang Wang, Wenyu Liu 0001, Wenjun Zeng 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Extracting 3-D Structural Lines of Building From ALS Point Clouds Using Graph Neural Network Embedded With Corner Information
abstract
The representation quantifies the geometric shape and topology of a building is a necessary procedure for many urban planning applications. A sharp line framework is a high-level structural cue providing a compact building representation. However, accurate and efficient structural line extraction remains a challenging task given the variety and complexity of buildings. This study proposes a general 3-D structural line extraction method from point clouds. The building points are extracted and further divided into various single-building units. In the proposed 3-D structural line extraction method, individual building point cloud is the input. First, the corners are detected by an associative learning module. Next, the curve connection is implemented by a link prediction block based on the graph neural network (GNN) embedded with corner information. After that, the obtained curves are subsequently converted into a topological graph. Finally, the corner points are optimized to achieve precise fitting of the structural lines. The experiments and comparisons on two airborne laser scanning (ALS) point cloud datasets demonstrate the effectiveness of the proposed method and the ability to retrieve ideal structural line results for building point clouds. Furthermore, without reprocessing, the proposed method yielded better results for various dataset types (outdoor building, indoor scene, and furniture point clouds) than the prevalent published methods (i.e., EC-Net, PIE-Net, and PC2WF), verifying its strength and efficacy. To further verify the accuracy of the obtained structural lines, we also introduce a line-based model reconstruction method that employ these lines for building reconstruction.
Tengping Jiang, Zequn Zhang, Yongchao Yang, Xin Jin 0014, Wenjun Zeng 0001
IEEE Trans. Geosci. Remote. Sens.8
2023 RailSeg: Learning Local-Global Feature Aggregation With Contextual Information for Railway Point Cloud Semantic Segmentation
abstract
Incomplete or outdated inventories of railway infrastructures may disrupt the railway sector’s administration and maintenance of transportation infrastructure, thus posing potential threats to the safety of traffic networks. Previous studies have adopted point clouds to accelerate inventory and inspection automation procedures. However, owing to the complexity of the railway scenes, previous studies reveal an imbalance between semantic richness, segmentation accuracy, and processing efficiency. This study aims to advance our understanding by providing a deep-learning framework for railway point cloud semantic segmentation. The proposed framework, named RailSeg, encompasses point cloud downsampling, integrated local-global feature extraction, spatial context aggregation, and semantic regularization. The proposed method, validated using point clouds collected in suburban and rural scenes, generates a point-level railway furniture inventory of 11 categories and achieves competitive performance in overall accuracy and mean intersection over union. In addition, RailSeg achieves better results than the baseline for additional types of point clouds (i.e., plateau railway mobile laser scanning (MLS) point clouds, street MLS point clouds, and urban-scale photogrammetric point clouds), demonstrating the superior generalization capabilities of RailSeg. This study may contribute to the development of 3D semantic segmentation, digital railway, and intelligent transportation.
Tengping Jiang, Bisheng Yang, Qinyu Zhang 0008, Xin Jin 0014, Wenjun Zeng 0001
IEEE Trans. Geosci. Remote. Sens.10
2023 Generalizing to Unseen Domains: A Survey on Domain Generalization
abstract
Machine learning systems generally assume that the training and testing distributions are the same. To this end, a key requirement is to develop models that can generalize to unseen distributions. Domain generalization (DG), i.e., out-of-distribution generalization, has attracted increasing interests in recent years. Domain generalization deals with a challenging setting where one or several different but related domain(s) are given, and the goal is to learn a model that can generalize to an unseen test domain. Great progress has been made in the area of domain generalization for years. This paper presents the first review of recent advances in this area. First, we provide a formal definition of domain generalization and discuss several related fields. We then thoroughly review the theories related to domain generalization and carefully analyze the theory behind generalization. We categorize recent algorithms into three classes: data manipulation, representation learning, and learning strategy, and present several popular algorithms in detail for each category. Third, we introduce the commonly used datasets, applications, and our open-sourced codebase for fair evaluation. Finally, we summarize existing literature and present some potential research topics for the future.
Jindong Wang 0001, Cuiling Lan, Chang Liu 0030, Yidong Ouyang, Tao Qin 0001, Wang Lu 0003, Yiqiang Chen 0001, Wenjun Zeng 0001, Philip S. Yu
IEEE Trans. Knowl. Data Eng.8
2023 Skeleton-Based Mutually Assisted Interacted Object Localization and Human Action Recognition
abstract
Skeleton data carries valuable motion information and is widely explored in human action recognition. However, not only the motion information but also the interaction with the environment provides discriminative cues to recognize the action of persons. In this paper, we propose a joint learning framework for mutually assisted “interacted object localization” and “human action recognition” based on skeleton data. The two tasks are serialized together and collaborate to promote each other, where preliminary action type derived from skeleton alone helps improve interacted object localization, which in turn provides valuable cues for the final human action recognition. Besides, we explore the temporal consistency of interacted object as constraint to better localize the interacted object with the absence of ground-truth labels. Extensive experiments on the datasets of SYSU-3D, NTU60 RGB+D, Northwestern-UCLA and UAV-Human show that our method achieves the best or competitive performance with the state-of-the-art methods for human action recognition. Visualization results show that our method can also provide reasonable interacted object localization results.
Liang Xu 0012, Cuiling Lan, Wenjun Zeng 0001, Cewu Lu
IEEE Trans. Multim.3
2022 Sparse MLP for Image Recognition: Is Self-Attention Really Necessary?
abstract
Transformers have sprung up in the field of computer vision. In this work, we explore whether the core self-attention module in Transformer is the key to achieving excellent performance in image recognition. To this end, we build an attention-free network called sMLPNet based on the existing MLP-based vision models. Specifically, we replace the MLP module in the token-mixing step with a novel sparse MLP (sMLP) module. For 2D image tokens, sMLP applies 1D MLP along the axial directions and the parameters are shared among rows or columns. By sparse connection and weight sharing, sMLP module significantly reduces the number of model parameters and computational complexity, avoiding the common over-fitting problem that plagues the performance of MLP-like models. When only trained on the ImageNet-1K dataset, the proposed sMLPNet achieves 81.9% top-1 accuracy with only 24M parameters, which is much better than most CNNs and vision Transformers under the same model size constraint. When scaling up to 66M parameters, sMLPNet achieves 83.4% top-1 accuracy, which is on par with the state-of-the-art Swin Transformer. The success of sMLPNet suggests that the self-attention mechanism is not necessarily a silver bullet in computer vision. The code and models are publicly available at https://github.com/microsoft/SPACH.
Chuanxin Tang, Guangting Wang, Chong Luo 0001, Wenxuan Xie, Wenjun Zeng 0001
AAAI6
2022 When Shift Operation Meets Vision Transformer: An Extremely Simple Alternative to Attention Mechanism
abstract
Attention mechanism has been widely believed as the key to success of vision transformers (ViTs), since it provides a flexible and powerful way to model spatial relationships. However, is the attention mechanism truly an indispensable part of ViT? Can it be replaced by some other alternatives? To demystify the role of attention mechanism, we simplify it into an extremely simple case: ZERO FLOP and ZERO parameter. Concretely, we revisit the shift operation. It does not contain any parameter or arithmetic calculation. The only operation is to exchange a small portion of the channels between neighboring features. Based on this simple operation, we construct a new backbone network, namely ShiftViT, where the attention layers in ViT are substituted by shift operations. Surprisingly, ShiftViT works quite well in several mainstream tasks, e.g., classification, detection, and segmentation. The performance is on par with or even better than the strong baseline Swin Transformer. These results suggest that the attention mechanism might not be the vital factor that makes ViT successful. It can be even replaced by a zero-parameter operation. We should pay more attentions to the remaining parts of ViT in the future work. Code is available at github.com/microsoft/SPACH.
Guangting Wang, Chuanxin Tang, Chong Luo 0001, Wenjun Zeng 0001
AAAI5
2022 Lifelong Unsupervised Domain Adaptive Person Re-identification with Coordinated Anti-forgetting and Adaptation
abstract
Unsupervised domain adaptive person re-identification (ReID) has been extensively investigated to mitigate the adverse effects of domain gaps. Those works assume the target domain data can be accessible all at once. However, for the real-world streaming data, this hinders the timely adaptation to changing data statistics and sufficient exploitation of increasing samples. In this paper, to address more practical scenarios, we propose a new task, Lifelong Un-supervised Domain Adaptive (LUDA) person ReID. This is challenging because it requires the model to continuously adapt to unlabeled data in the target environments while alleviating catastrophic forgetting for such a fine-grained person retrieval task. We design an effective scheme for this task, dubbed CLUDA-ReID, where the anti-forgetting is harmoniously coordinated with the adaptation. Specifically, a meta-based Coordinated Data Replay strategy is proposed to replay old data and update the network with a coordinated optimization direction for both adaptation and memorization. Moreover, we propose Relational Consistency Learning for old knowledge distillation/inheritance in line with the objective of retrieval-based tasks. We set up two evaluation settings to simulate the practical application scenarios. Extensive experiments demonstrate the effectiveness of our CLUDA-ReID for both scenarios with stationary target streams and scenarios with dynamic target streams.
Zhipeng Huang 0014, Zhizheng Zhang 0004, Cuiling Lan, Wenjun Zeng 0001, Peng Chu, Quanzeng You, Jiang Wang 0012, Zicheng Liu 0001, Zhengjun Zha
CVPR4
2022 ReSTR: Convolution-free Referring Image Segmentation Using Transformers
abstract
Referring image segmentation is an advanced semantic segmentation task where target is not a predefined class but is described in natural language. Most of existing methods for this task rely heavily on convolutional neural networks, which however have trouble capturing long-range dependencies between entities in the language expression and are not flexible enough for modeling interactions between the two different modalities. To address these issues, we present the first convolution-free model for referring image segmentation using transformers, dubbed ReSTR. Since it extracts features of both modalities through transformer encoders, it can capture long-range dependencies between entities within each modality. Also, ReSTR fuses features of the two modalities by a self-attention encoder, which enables flexible and adaptive interactions between the two modalities in the fusion process. The fused features are fed to a segmentation module, which works adaptively according to the image and language expression in hand. ReSTR is evaluated and compared with previous work on all public benchmarks, where it outperforms all existing models.
Namyup Kim, Suha Kwak, Cuiling Lan, Wenjun Zeng 0001
CVPR5
2022 Correlation-Aware Deep Tracking
abstract
Robustness and discrimination power are two fundamental requirements in visual object tracking. In most tracking paradigms, we find that the features extracted by the popular Siamese-like networks cannot fully discriminatively model the tracked targets and distractor objects, hindering them from simultaneously meeting these two requirements. While most methods focus on designing robust correlation operations, we propose a novel target-dependent feature network inspired by the self-/cross-attention scheme. In contrast to the Siamese-like feature extraction, our network deeply embeds cross-image feature correlation in multiple layers of the feature network. By extensively matching the features of the two images through multiple layers, it is able to suppress non-target features, resulting in instance-varying feature extraction. The output features of the search image can be directly used for predicting target locations without extra correlation step. Moreover, our model can be flexibly pre-trained on abundant unpaired images, leading to notably faster convergence than the existing methods. Extensive experiments show our method achieves the state-of-the-art results while running at real-time. Our feature networks also can be applied to existing tracking pipelines seamlessly to raise the tracking performance.
Chunyu Wang 0001, Guangting Wang, Yue Cao 0001, Wankou Yang, Wenjun Zeng 0001
CVPR6
2022 VirtualPose: Learning Generalizable 3D Human Pose Models from Virtual Data
Jiajun Su, Chunyu Wang 0001, Xiaoxuan Ma 0001, Wenjun Zeng 0001, Yizhou Wang 0001
ECCV (6)4
2022 Robust Multi-object Tracking by Marginal Inference
Chunyu Wang 0001, Xinggang Wang, Wenjun Zeng 0001, Wenyu Liu 0001
ECCV (22)4
2022 Learning Disentangled Representation by Exploiting Pretrained Generative Models: A Contrastive Learning View
Xuanchi Ren, Tao Yang 0032, Yuwang Wang, Wenjun Zeng 0001
ICLR4
2022 Towards Building A Group-based Unsupervised Representation Disentanglement Framework
Tao Yang 0032, Xuanchi Ren, Yuwang Wang, Wenjun Zeng 0001, Nanning Zheng 0001
ICLR4
2022 Retriever: Learning Content-Style Representation as a Token-Level Bipartite Graph
Dacheng Yin, Xuanchi Ren, Chong Luo 0001, Yuwang Wang, Zhiwei Xiong, Wenjun Zeng 0001
ICLR6
2022 Gaze- and Spacing-flow Unveil Intentions: Hidden Follower Discovery
abstract
We raise a new and challenging multimedia application in video surveillance system, i.e., Hidden Follower Discovery (HFD). In contrast to the common abnormal behaviors that are occurring, hidden following is not an ongoing activity, but a preparatory action. Hidden following behavior does not have salient features, making it hard to be discovered. Fortunately, from a socio-cognitive perspective, we found and verified the phenomena that the gaze-flow pattern and the spacing-flow pattern between hidden and normal followers are different. To promote HFD research, we construct two pioneering datasets and devise an HFD baseline network based on the recognition of both gaze-flow and spacing-flow patterns from surveillance videos. Extensive experiments demonstrate their effectiveness.
Danni Xu, Ruimin Hu, Zheng Wang 0007, Linbo Luo 0001, Dengshi Li, Wenjun Zeng 0001
ACM Multimedia6
2022 FPCR-Net: Feature pyramidal correlation and residual reconstruction for optical flow estimation
Jing-Yu Yang 0002, Cuiling Lan, Wenjun Zeng 0001
Neurocomputing5
2022 APANet: Auto-Path Aggregation for Future Instance Segmentation Prediction
abstract
Despite the remarkable progress achieved in conventional instance segmentation, the problem of predicting instance segmentation results for unobserved future frames remains challenging due to the unobservability of future data. Existing methods mainly address this challenge by forecasting features of future frames. However, these methods always treat features of multiple levels (e.g., coarse-to-fine pyramid features) independently and do not exploit them collaboratively, which results in inaccurate prediction for future frames; and moreover, such a weakness can partially hinder self-adaption of a future segmentation prediction model for different input samples. To solve this problem, we propose an adaptive aggregation approach called Auto-Path Aggregation Network (APANet), where the spatio-temporal contextual information obtained in the features of each individual level is selectively aggregated using the developed "auto-path". The "auto-path" connects each pair of features extracted at different pyramid levels for task-specific hierarchical contextual information aggregation, which enables selective and adaptive aggregation of pyramid features in accordance with different videos/frames. Our APANet can be further optimized jointly with the Mask R-CNN head as a feature decoder and a Feature Pyramid Network (FPN) feature encoder, forming a joint learning system for future instance segmentation prediction. We experimentally show that the proposed method can achieve state-of-the-art performance on three video-based instance segmentation benchmarks for future instance segmentation prediction.
Jianfang Hu, Jiangxin Sun, Zihang Lin, Jian-Huang Lai, Wenjun Zeng 0001, Wei-Shi Zheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Style Normalization and Restitution for Domain Generalization and Adaptation
abstract
For many computer vision applications, the learned models usually have high performance on the training datasets but suffer from significant performance degradation when deployed in new environments, where there are usually style differences between the training images and the testing images. For high-level vision tasks, an effective domain generalizable model is expected to be able to learn feature representations that are both generalizable and discriminative. In this paper, we design a novel Style Normalization and Restitution module (SNR) to simultaneously ensure high generalization and discrimination capability of the networks. In SNR, particularly, we filter out the style variations (e.g., illumination, color contrast) by performing Instance Normalization (IN) to obtain style normalized features, where the discrepancy among different samples/domains is reduced. However, such a process is task-ignorant and inevitably removes some task-relevant discriminative information, which may hurt the performance. To remedy this, we propose to distill task-relevant discriminative features from the residual (i.e., the difference between the original feature and the style normalized feature) and add them back to the network to ensure high discrimination. Moreover, for better disentanglement, we enforce a dual restitution loss constraint to encourage the better separation of task-relevant and task-irrelevant features. We validate the effectiveness of our SNR on different vision tasks, including classification, semantic segmentation, and object detection. Experiments demonstrate that our SNR is capable of improving the performance of networks for domain generalization (DG) and unsupervised domain adaptation (UDA).
Xin Jin 0014, Cuiling Lan, Wenjun Zeng 0001, Zhibo Chen 0001
IEEE Trans. Multim.3
2022 Beyond Triplet Loss: Meta Prototypical N-Tuple Loss for Person Re-identification
abstract
Person Re-identification (ReID) aims at matching a person of interest across images. In convolutional neural network (CNN) based approaches, loss design plays a vital role in pulling closer features of the same identity and pushing far apart features of different identities. In recent years, triplet loss achieves superior performance and is predominant in ReID. However, triplet loss considers only three instances of two classes in per-query optimization (with an anchor sample as query) and it is actually equivalent to a two-class classification. There is a lack of loss design which enables the joint optimization of multiple instances (of multiple classes) within per-query optimization for person ReID. In this paper, we introduce a multi-class classification loss,i.e., N-tuple loss, to jointly consider multiple ($N$) instances for per-query optimization. This in fact aligns better with the ReID test/inference process, which conducts the ranking/comparisons among multiple instances. Furthermore, for more efficient multi-class classification, we propose a new meta prototypical N-tuple loss. With the multi-class classification incorporated, our model achieves the state-of-the-art performance on the benchmark person ReID datasets
Zhizheng Zhang 0004, Cuiling Lan, Wenjun Zeng 0001, Zhibo Chen 0001, Shih-Fu Chang
IEEE Trans. Multim.3
2022 Neighborhood Geometric Structure-Preserving Variational Autoencoder for Smooth and Bounded Data Sources
abstract
Many data sources, such as human poses, lie on low-dimensional manifolds that are smooth and bounded. Learning low-dimensional representations for such data is an important problem. One typical solution is to utilize encoder-decoder networks. However, due to the lack of effective regularization in latent space, the learned representations usually do not preserve the essential data relations. For example, adjacent video frames in a sequence may be encoded into very different zones across the latent space with holes in between. This is problematic for many tasks such as denoising because slightly perturbed data have the risk of being encoded into very different latent variables, leaving output unpredictable. To resolve this problem, we first propose a neighborhood geometric structure-preserving variational autoencoder (SP-VAE), which not only maximizes the evidence lower bound but also encourages latent variables to preserve their structures as in ambient space. Then, we learn a set of small surfaces to approximately bound the learned manifold to deal with holes in latent space. We extensively validate the properties of our approach by reconstruction, denoising, and random image generation experiments on a number of data sources, including synthetic Swiss roll, human pose sequences, and facial expression images. The experimental results show that our approach learns more smooth manifolds than the baselines. We also apply our approach to the tasks of human pose refinement and facial expression image interpolation where it gets better results than the baselines.
Xingyu Chen 0001, Chunyu Wang 0001, Xuguang Lan, Nanning Zheng 0001, Wenjun Zeng 0001
IEEE Trans. Neural Networks Learn. Syst.5
2021 Very Important Person Localization in Unconstrained Conditions: A New Benchmark
abstract
This paper presents a new high-quality dataset for Very Important Person Localization (VIPLoc), named Unconstrained-7k. Generally, current datasets: 1) are limited in scale; 2) built under simple and constrained conditions, where the number of disturbing non-VIPs is not large, the scene is relatively simple, and the face of VIP is always in frontal view and salient. To tackle these problems, the proposed Unconstrained-7k dataset is featured in two aspects. First, it contains over 7,000 annotated images, making it the largest VIPLoc dataset under unconstrained conditions to date. Second, our dataset is collected freely on the Internet, including multiple scenes, where images are in unconstrained conditions. VIPs in the new dataset are in different settings, e.g., large view variation, varying sizes, occluded, and complex scenes. Meanwhile, each image has more persons (> 20), making the dataset more challenging. As a minor contribution, motivated by the observation that VIPs are highly related to not only neighbors but also iconic objects, this paper proposes a Joint Social Relation and Individual Interaction Graph Neural Networks (JSRII-GNN) for VIPLoc. Experiments show that the JSRII-GNN yields competitive accuracy on NCAA (National Collegiate Athletic Association), MS (Multi-scene), and Unconstrained-7k datasets. https://github.com/xiaowang1516/VIPLoc.
Xiao Wang 0029, Zheng Wang 0007, Toshihiko Yamasaki, Wenjun Zeng 0001
AAAI4
2021 Exploiting Sample Uncertainty for Domain Adaptive Person Re-Identification
abstract
Many unsupervised domain adaptive (UDA) person ReID approaches combine clustering-based pseudo-label prediction with feature fine-tuning. However, because of domain gap, the pseudo-labels are not always reliable and there are noisy/incorrect labels. This would mislead the feature representation learning and deteriorate the performance. In this paper, we propose to estimate and exploit the credibility of the assigned pseudo-label of each sample to alleviate the influence of noisy labels, by suppressing the contribution of noisy samples. We build our baseline framework using the mean teacher method together with an additional contrastive loss. We have observed that a sample with a wrong pseudo-label through clustering in general has a weaker consistency between the output of the mean teacher model and the student model. Based on this finding, we propose to exploit the uncertainty (measured by consistency levels) to evaluate the reliability of the pseudo-label of a sample and incorporate the uncertainty to re-weight its contribution within various ReID losses, including the ID classification loss per sample, the triplet loss, and the contrastive loss. Our uncertainty-guided optimization brings significant improvement and achieves the state-of-the-art performance on benchmark datasets.
Kecheng Zheng, Cuiling Lan, Wenjun Zeng 0001, Zhizheng Zhang 0004, Zhengjun Zha
AAAI3
2021 S2R-DepthNet: Learning a Generalizable Depth-Specific Structural Representation
abstract
Human can infer the 3D geometry of a scene from a sketch instead of a realistic image, which indicates that the spatial structure plays a fundamental role in understanding the depth of scenes. We are the first to explore the learning of a depth-specific structural representation, which captures the essential feature for depth estimation and ignores irrelevant style information. Our S2R-DepthNet (Synthetic to Real DepthNet) can be well generalized to un-seen real-world data directly even though it is only trained on synthetic data. S2R-DepthNet consists of: a) a Structure Extraction (STE) module which extracts a domain-invariant structural representation from an image by dis-entangling the image into domain-invariant structure and domain-specific style components, b) a Depth-specific Attention (DSA) module, which learns task-specific knowledge to suppress depth-irrelevant structures for better depth estimation and generalization, and c) a depth prediction module (DP) to predict depth from the depth-specific representation. Without access of any real-world images, our method even outperforms the state-of-the-art unsupervised domain adaptation methods which use real-world images of the tar-get domain for training. In addition, when using a small amount of labeled real-world data, we achieve the state-of-the-art performance under the semi-supervised setting.
Xiaotian Chen, Yuwang Wang, Xuejin Chen, Wenjun Zeng 0001
CVPR4
2021 Unsupervised Visual Representation Learning by Tracking Patches in Video
abstract
Inspired by the fact that human eyes continue to develop tracking ability in early and middle childhood, we propose to use tracking as a proxy task for a computer vision system to learn the visual representations. Modelled on the Catch game played by the children, we design a Catch-the-Patch (CtP) game for a 3D-CNN model to learn visual representations that would help with video-related tasks. In the proposed pretraining framework, we cut an image patch from a given video and let it scale and move according to a pre-set trajectory. The proxy task is to estimate the position and size of the image patch in a sequence of video frames, given only the target bounding box in the first frame. We discover that using multiple image patches simultaneously brings clear benefits. We further increase the difficulty of the game by randomly making patches invisible. Extensive experiments on mainstream benchmarks demonstrate the superior performance of CtP against other video pretraining methods. In addition, CtP-pretrained features are less sensitive to domain gaps than those trained by a supervised action recognition task. When both trained on Kinetics-400, we are pleasantly surprised to find that CtP-pretrained representation achieves much higher action classification accuracy than its fully supervised counterpart on Something-Something dataset.
Guangting Wang, Yizhou Zhou, Chong Luo 0001, Wenxuan Xie, Wenjun Zeng 0001, Zhiwei Xiong
CVPR5
2021 MetaAlign: Coordinating Domain Alignment and Classification for Unsupervised Domain Adaptation
abstract
For unsupervised domain adaptation (UDA), to alleviate the effect of domain shift, many approaches align the source and target domains in the feature space by adversarial learning or by explicitly aligning their statistics. However, the optimization objective of such domain alignment is generally not coordinated with that of the object classification task itself such that their descent directions for optimization may be inconsistent. This will reduce the effectiveness of domain alignment in improving the performance of UDA. In this paper, we aim to study and alleviate the optimization inconsistency problem between the domain alignment and classification tasks. We address this by proposing an effective meta-optimization based strategy dubbed MetaAlign, where we treat the domain alignment objective and the classification objective as the meta-train and meta-test tasks in a meta-learning scheme. MetaAlign encourages both tasks to be optimized in a coordinated way, which maximizes the inner product of the gradients of the two tasks during training. Experimental results demonstrate the effectiveness of our proposed method on top of various alignment-based baseline approaches, for tasks of object classification and object detection. MetaAlign helps achieve the state-of-the-art performance.
Guoqiang Wei, Cuiling Lan, Wenjun Zeng 0001, Zhibo Chen 0001
CVPR3
2021 Re-energizing Domain Discriminator with Sample Relabeling for Adversarial Domain Adaptation
abstract
Many unsupervised domain adaptation (UDA) methods exploit domain adversarial training to align the features to reduce domain gap, where a feature extractor is trained to fool a domain discriminator in order to have aligned feature distributions. The discrimination capability of the domain classifier w.r.t. the increasingly aligned feature distributions deteriorates as training goes on, thus cannot effectively further drive the training of feature extractor. In this work, we propose an efficient optimization strategy named Re-enforceable Adversarial Domain Adaptation (RADA) which aims to re-energize the domain discriminator during the training by using dynamic domain labels. Particularly, we relabel the well aligned target domain samples as source domain samples on the fly. Such relabeling makes the less separable distributions more separable, and thus leads to a more powerful domain classifier w.r.t. the new data distributions, which in turn further drives feature alignment. Extensive experiments on multiple UDA benchmarks demonstrate the effectiveness and superiority of our RADA.
Xin Jin 0014, Cuiling Lan, Wenjun Zeng 0001, Zhibo Chen 0001
ICCV3
2021 An Empirical Study of the Collapsing Problem in Semi-Supervised 2D Human Pose Estimation
abstract
Most semi-supervised learning models are consistency-based, which leverage unlabeled images by maximizing the similarity between different augmentations of an image. But when we apply them to human pose estimation that has extremely imbalanced class distribution, they often collapse and predict every pixel in unlabeled images as background. We find this is because the decision boundary passes the high-density areas of the minor class so more and more pixels are gradually misclassified as background. In this work, we present a surprisingly simple approach to drive the model to learn in the correct direction. For each image, it composes a pair of easy-hard augmentations and uses the more accurate predictions on the easy image to teach the network to learn pose information of the hard one. The accuracy superiority of teaching signals allows the network to be "monotonically" improved which effectively avoids collapsing. We apply our method to the state-of-the-art pose estimators and it further improves their performance on three public datasets.
Rongchang Xie, Chunyu Wang 0001, Wenjun Zeng 0001, Yizhou Wang 0001
ICCV3
2021 Self-Supervised Visual Representations Learning by Contrastive Mask Prediction
abstract
Advanced self-supervised visual representation learning methods rely on the instance discrimination (ID) pretext task. We point out that the ID task has an implicit semantic consistency (SC) assumption, which may not hold in unconstrained datasets. In this paper, we propose a novel contrastive mask prediction (CMP) task for visual representation learning and design a mask contrast (MaskCo) framework to implement the idea. MaskCo contrasts region-level features instead of view-level features, which makes it possible to identify the positive sample without any assumptions. To solve the domain gap between masked and unmasked features, we design a dedicated mask prediction head in MaskCo. This module is shown to be the key to the success of the CMP. We evaluated MaskCo on training datasets beyond ImageNet and compare its performance with MoCo V2 [4]. Results show that MaskCo achieves comparable performance with MoCo V2 using ImageNet training dataset, but demonstrates a stronger performance across a range of downstream tasks when COCO or Conceptual Captions are used for training. MaskCo provides a promising alternative to the ID-based methods for self-supervised learning in the wild.
Guangting Wang, Chong Luo 0001, Wenjun Zeng 0001, Zhengjun Zha
ICCV4
2021 Uncertainty-Aware Few-Shot Image Classification
abstract
Few-shot image classification learns to recognize new categories from limited labelled data. Metric learning based approaches have been widely investigated, where a query sample is classified by finding the nearest prototype from the support set based on their feature similarities. A neural network has different uncertainties on its calculated similarities of different pairs. Understanding and modeling the uncertainty on the similarity could promote the exploitation of limited samples in few-shot optimization. In this work, we propose Uncertainty-Aware Few-Shot framework for image classification by modeling uncertainty of the similarities of query-support pairs and performing uncertainty-aware optimization. Particularly, we exploit such uncertainty by converting observed similarities to probabilistic representations and incorporate them to the loss for more effective optimization. In order to jointly consider the similarities between a query and the prototypes in a support set, a graph-based model is utilized to estimate the uncertainty of the pairs. Extensive experiments show our proposed method brings significant improvements on top of a strong baseline and achieves the state-of-the-art performance.
Zhizheng Zhang 0004, Cuiling Lan, Wenjun Zeng 0001, Zhibo Chen 0001, Shih-Fu Chang
IJCAI3
2021 Zero-Shot Text-to-Speech for Text-Based Insertion in Audio Narration
abstract
Given a piece of speech and its transcript text, text-based speech editing aims to generate speech that can be seamlessly inserted into the given speech by editing the transcript.Existing methods adopt a two-stage approach: synthesize the input text using a generic text-to-speech (TTS) engine and then transform the voice to the desired voice using voice conversion (VC).A major problem of this framework is that VC is a challenging problem which usually needs a moderate amount of parallel training data to work satisfactorily.In this paper, we propose a one-stage context-aware framework to generate natural and coherent target speech without any training data of the target speaker.In particular, we manage to perform accurate zero-shot duration prediction for the inserted text.The predicted duration is used to regulate both text embedding and speech embedding.Then, based on the aligned cross-modality input, we directly generate the mel-spectrogram of the edited speech with a transformer-based decoder.Subjective listening tests show that despite the lack of training data for the speaker, our method has achieved satisfactory results.It outperforms a recent zero-shot TTS engine by a large margin.
Chuanxin Tang, Chong Luo 0001, Zhiyuan Zhao 0001, Dacheng Yin, Wenjun Zeng 0001
Interspeech6
2021 Pose-Guided Feature Learning with Knowledge Distillation for Occluded Person Re-Identification
abstract
Occluded person re-identification (ReID) aims to match person images with occlusion. It is fundamentally challenging because of the serious occlusion which aggravates the misalignment problem between images. At the cost of incorporating a pose estimator, many works introduce pose information to alleviate the misalignment in both training and testing. To achieve high accuracy while preserving low inference complexity, we propose a network named Pose-Guided Feature Learning with Knowledge Distillation (PGFL-KD), where the pose information is exploited to regularize the learning of semantics aligned features but is discarded in testing. PGFL-KD consists of a main branch (MB), and two pose-guided branches, e.g., a foreground-enhanced branch (FEB), and a body part semantics aligned branch (SAB). The FEB intends to emphasise the features of visible body parts while excluding the interference of obstructions and background (e.g., foreground feature alignment). The SAB encourages different channel groups to focus on different body parts to have body part semantics aligned representation. To get rid of the dependency on pose information when testing, we regularize the MB to learn the merits of the FEB and SAB through knowledge distillation and interaction-based training. Extensive experiments on occluded, partial, and holistic ReID tasks show the effectiveness of our proposed network.
Kecheng Zheng, Cuiling Lan, Wenjun Zeng 0001, Jiawei Liu 0001, Zhizheng Zhang 0004, Zhengjun Zha
ACM Multimedia3
2021 ToAlign: Task-Oriented Alignment for Unsupervised Domain Adaptation
abstract
Unsupervised domain adaptive classifcation intends to improve the classifcation performance on unlabeled target domain. To alleviate the adverse effect of domain shift, many approaches align the source and target domains in the feature space. However, a feature is usually taken as a whole for alignment without explicitly making domain alignment proactively serve the classifcation task, leading to sub-optimal solution. In this paper, we propose an effective Task-oriented Alignment (ToAlign) for unsupervised domain adaptation (UDA). We study what features should be aligned across domains and propose to make the domain alignment proactively serve classifcation by performing feature decomposition and alignment under the guidance of the prior knowledge induced from the classifcation task itself. Particularly, we explicitly decompose a feature in the source domain into a task-related/discriminative feature that should be aligned, and a task-irrelevant feature that should be avoided/ignored, based on the classifcation meta-knowledge. Extensive experimental results on various benchmarks (e.g., Offce-Home, Visda-2017, and DomainNet) under different domain adaptation settings demonstrate the effectiveness of ToAlign which helps achieve the state-of-the-art performance. The code is publicly available at https://github.com/microsoft/UDA.
Guoqiang Wei, Cuiling Lan, Wenjun Zeng 0001, Zhizheng Zhang 0004, Zhibo Chen 0001
NeurIPS3
2021 PlayVirtual: Augmenting Cycle-Consistent Virtual Trajectories for Reinforcement Learning
abstract
Learning good feature representations is important for deep reinforcement learning (RL). However, with limited experience, RL often suffers from data inefficiency for training. For un-experienced or less-experienced trajectories (i.e., state-action sequences), the lack of data limits the use of them for better feature learning. In this work, we propose a novel method, dubbed PlayVirtual, which augments cycle-consistent virtual trajectories to enhance the data efficiency for RL feature representation learning. Specifically, PlayVirtual predicts future states in a latent space based on the current state and action by a dynamics model and then predicts the previous states by a backward dynamics model, which forms a trajectory cycle. Based on this, we augment the actions to generate a large amount of virtual state-action trajectories. Being free of groudtruth state supervision, we enforce a trajectory to meet the cycle consistency constraint, which can significantly enhance the data efficiency. We validate the effectiveness of our designs on the Atari and DeepMind Control Suite benchmarks. Our method achieves the state-of-the-art performance on both benchmarks. Our code is available at https://github.com/microsoft/Playvirtual.
Tao Yu 0012, Cuiling Lan, Wenjun Zeng 0001, Mingxiao Feng, Zhizheng Zhang 0004, Zhibo Chen 0001
NeurIPS3
2021 AdaFuse: Adaptive Multiview Fusion for Accurate Human Pose Estimation in the Wild
Zhe Zhang 0045, Chunyu Wang 0001, Weichao Qiu, Wenhu Qin, Wenjun Zeng 0001
Int. J. Comput. Vis.5
2021 FairMOT: On the Fairness of Detection and Re-identification in Multiple Object Tracking
Chunyu Wang 0001, Xinggang Wang, Wenjun Zeng 0001, Wenyu Liu 0001
Int. J. Comput. Vis.4
2021 CASINet: Content-Adaptive Scale Interaction Networks for scene parsing
Xin Jin 0014, Cuiling Lan, Wenjun Zeng 0001, Zhizheng Zhang 0004, Zhibo Chen 0001
Neurocomputing3
2021 AttributeNet: Attribute enhanced vehicle re-identification
Rodolfo Quispe, Cuiling Lan, Wenjun Zeng 0001, Hélio Pedrini
Neurocomputing3
2020 Uncertainty-Aware Multi-Shot Knowledge Distillation for Image-Based Object Re-Identification
abstract
Object re-identification (re-id) aims to identify a specific object across times or camera views, with the person re-id and vehicle re-id as the most widely studied applications. Re-id is challenging because of the variations in viewpoints, (human) poses, and occlusions. Multi-shots of the same object can cover diverse viewpoints/poses and thus provide more comprehensive information. In this paper, we propose exploiting the multi-shots of the same identity to guide the feature learning of each individual image. Specifically, we design an Uncertainty-aware Multi-shot Teacher-Student (UMTS) Network. It consists of a teacher network (T-net) that learns the comprehensive features from multiple images of the same object, and a student network (S-net) that takes a single image as input. In particular, we take into account the data dependent heteroscedastic uncertainty for effectively transferring the knowledge from the T-net to S-net. To the best of our knowledge, we are the first to make use of multi-shots of an object in a teacher-student learning manner for effectively boosting the single image based re-id. We validate the effectiveness of our approach on the popular vehicle re-id and person re-id datasets. In inference, the S-net alone significantly outperforms the baselines and achieves the state-of-the-art performance.
Xin Jin 0014, Cuiling Lan, Wenjun Zeng 0001, Zhibo Chen 0001
AAAI3
2020 Semantics-Aligned Representation Learning for Person Re-Identification
abstract
Person re-identification (reID) aims to match person images to retrieve the ones with the same identity. This is a challenging task, as the images to be matched are generally semantically misaligned due to the diversity of human poses and capture viewpoints, incompleteness of the visible bodies (due to occlusion), etc. In this paper, we propose a framework that drives the reID network to learn semantics-aligned feature representation through delicate supervision designs. Specifically, we build a Semantics Aligning Network (SAN) which consists of a base network as encoder (SA-Enc) for re-ID, and a decoder (SA-Dec) for reconstructing/regressing the densely semantics aligned full texture image. We jointly train the SAN under the supervisions of person re-identification and aligned texture generation. Moreover, at the decoder, besides the reconstruction loss, we add Triplet ReID constraints over the feature maps as the perceptual losses. The decoder is discarded in the inference and thus our scheme is computationally efficient. Ablation studies demonstrate the effectiveness of our design. We achieve the state-of-the-art performances on the benchmark datasets CUHK03, Market1501, MSMT17, and the partial person reID dataset Partial REID.
Xin Jin 0014, Cuiling Lan, Wenjun Zeng 0001, Guoqiang Wei, Zhibo Chen 0001
AAAI3
2020 PHASEN: A Phase-and-Harmonics-Aware Speech Enhancement Network
abstract
Time-frequency (T-F) domain masking is a mainstream approach for single-channel speech enhancement. Recently, focuses have been put to phase prediction in addition to amplitude prediction. In this paper, we propose a phase-and-harmonics-aware deep neural network (DNN), named PHASEN, for this task. Unlike previous methods which directly use a complex ideal ratio mask to supervise the DNN learning, we design a two-stream network, where amplitude stream and phase stream are dedicated to amplitude and phase prediction. We discover that the two streams should communicate with each other, and this is crucial to phase prediction. In addition, we propose frequency transformation blocks to catch long-range correlations along the frequency axis. Visualization shows that the learned transformation matrix implicitly captures the harmonic correlation, which has been proven to be helpful for T-F spectrogram reconstruction. With these two innovations, PHASEN acquires the ability to handle detailed phase patterns and to utilize harmonic patterns, getting 1.76dB SDR improvement on AVSpeech + AudioSet dataset. It also achieves significant gains over Google's network on this dataset. On Voice Bank + DEMAND dataset, PHASEN outperforms previous methods by a large margin on four metrics.
Dacheng Yin, Chong Luo 0001, Zhiwei Xiong, Wenjun Zeng 0001
AAAI4
2020 Posterior-Guided Neural Architecture Search
abstract
The emergence of neural architecture search (NAS) has greatly advanced the research on network design. Recent proposals such as gradient-based methods or one-shot approaches significantly boost the efficiency of NAS. In this paper, we formulate the NAS problem from a Bayesian perspective. We propose explicitly estimating the joint posterior distribution over pairs of network architecture and weights. Accordingly, a hybrid network representation is presented which enables us to leverage the Variational Dropout so that the approximation of the posterior distribution becomes fully gradient-based and highly efficient. A posterior-guided sampling method is then presented to sample architecture candidates and directly make evaluations. As a Bayesian approach, our posterior-guided NAS (PGNAS) avoids tuning a number of hyper-parameters and enables a very effective architecture sampling in posterior probability space. Interestingly, it also leads to a deeper insight into the weight sharing used in the one-shot NAS and naturally alleviates the mismatch between the sampled architecture and weights caused by the weight sharing. We validate our PGNAS method on the fundamental image classification task. Results on Cifar-10, Cifar-100 and ImageNet show that PGNAS achieves a good trade-off between precision and speed of search among NAS methods. For example, it takes 11 GPU days to search a very competitive architecture with 1.98% and 14.28% test errors on Cifar10 and Cifar100, respectively.
Yizhou Zhou, Xiaoyan Sun 0001, Chong Luo 0001, Zhengjun Zha, Wenjun Zeng 0001
AAAI5
2020 Semantics-Guided Neural Networks for Efficient Skeleton-Based Human Action Recognition
abstract
Skeleton-based human action recognition has attracted great interest thanks to the easy accessibility of the human skeleton data. Recently, there is a trend of using very deep feedforward neural networks to model the 3D coordinates of joints without considering the computational efficiency. In this paper, we propose a simple yet effective semantics-guided neural network (SGN) for skeleton-based action recognition. We explicitly introduce the high level semantics of joints (joint type and frame index) into the network to enhance the feature representation capability. In addition, we exploit the relationship of joints hierarchically through two modules, i.e., a joint-level module for modeling the correlations of joints in the same frame and a framelevel module for modeling the dependencies of frames by taking the joints in the same frame as a whole. A strong baseline is proposed to facilitate the study of this field. With an order of magnitude smaller model size than most previous works, SGN achieves the state-of-the-art performance on the NTU60, NTU120, and SYSU datasets.
Pengfei Zhang 0005, Cuiling Lan, Wenjun Zeng 0001, Junliang Xing, Jianru Xue, Nanning Zheng 0001
CVPR3
2020 Style Normalization and Restitution for Generalizable Person Re-Identification
abstract
Existing fully-supervised person re-identification (ReID) methods usually suffer from poor generalization capability caused by domain gaps. The key to solving this problem lies in filtering out identity-irrelevant interference and learning domain-invariant person representations. In this paper, we aim to design a generalizable person ReID framework which trains a model on source domains yet is able to generalize/perform well on target domains. To achieve this goal, we propose a simple yet effective Style Normalization and Restitution (SNR) module. Specifically, we filter out style variations (e.g., illumination, color contrast) by Instance Normalization (IN). However, such a process inevitably removes discriminative information. We propose to distill identity-relevant feature from the removed information and restitute it to the network to ensure high discrimination. For better disentanglement, we enforce a dual causal loss constraint in SNR to encourage the separation of identity-relevant features and identity-irrelevant features. Extensive experiments demonstrate the strong generalization capability of our framework. Our models empowered by the SNR modules significantly outperform the state-of-the-art domain generalization approaches on multiple widely-used person ReID benchmarks, and also show superiority on unsupervised domain adaptation.
Xin Jin 0014, Cuiling Lan, Wenjun Zeng 0001, Zhibo Chen 0001, Li Zhang 0040
CVPR3
2020 Tracking by Instance Detection: A Meta-Learning Approach
abstract
We consider the tracking problem as a special type of object detection problem, which we call instance detection. With proper initialization, a detector can be quickly converted into a tracker by learning the new instance from a single image. We find that model-agnostic meta-learning (MAML) offers a strategy to initialize the detector that satisfies our needs. We propose a principled three-step approach to build a high-performance tracker. First, pick any modern object detector trained with gradient descent. Second, conduct offline training (or initialization) with MAML. Third, perform domain adaptation using the initial frame. We follow this procedure to build two trackers, named Retina-MAML and FCOS-MAML, based on two modern detectors RetinaNet and FCOS. Evaluations on four benchmarks show that both trackers are competitive against state-of-the-art trackers. On OTB-100, Retina-MAML achieves the highest ever AUC of 0.712. On TrackingNet, FCOS-MAML ranks the first on the leader board with an AUC of 0.757 and the normalized precision of 0.822. Both trackers run in real-time at 40 FPS.
Guangting Wang, Chong Luo 0001, Xiaoyan Sun 0001, Zhiwei Xiong, Wenjun Zeng 0001
CVPR5
2020 Multi-Granularity Reference-Aided Attentive Feature Aggregation for Video-Based Person Re-Identification
abstract
Video-based person re-identification (reID) aims at matching the same person across video clips. It is a challenging task due to the existence of redundancy among frames, newly revealed appearance, occlusion, and motion blurs. In this paper, we propose an attentive feature aggregation module, namely Multi-Granularity Reference-aided Attentive Feature Aggregation (MG-RAFA), to delicately aggregate spatio-temporal features into a discriminative video-level feature representation. In order to determine the contribution/importance of a spatial-temporal feature node, we propose to learn the attention from a global view with convolutional operations. Specifically, we stack its relations, \ieno, pairwise correlations with respect to a representative set of reference feature nodes (S-RFNs) that represents global video information, together with the feature itself to infer the attention. Moreover, to exploit the semantics of different levels, we propose to learn multi-granularity attentions based on the relations captured at different granularities. Extensive ablation studies demonstrate the effectiveness of our attentive feature aggregation module MG-RAFA. Our framework achieves the state-of-the-art performance on three benchmark datasets.
Zhizheng Zhang 0004, Cuiling Lan, Wenjun Zeng 0001, Zhibo Chen 0001
CVPR3
2020 Relation-Aware Global Attention for Person Re-Identification
abstract
For person re-identification (re-id), attention mechanisms have become attractive as they aim at strengthening discriminative features and suppressing irrelevant ones, which matches well the key of re-id, i.e., discriminative feature learning. Previous approaches typically learn attention using local convolutions, ignoring the mining of knowledge from global structure patterns. Intuitively, the affinities among spatial positions/nodes in the feature map provide clustering-like information and are helpful for inferring semantics and thus attention, especially for person images where the feasible human poses are constrained. In this work, we propose an effective Relation-Aware Global Attention (RGA) module which captures the global structural information for better attention learning. Specifically, for each feature position, in order to compactly grasp the structural information of global scope and local appearance information, we propose to stack the relations, i.e., its pairwise correlations/affinities with all the feature positions (e.g., in raster scan order), and the feature itself together to learn the attention with a shallow convolutional model. Extensive ablation studies demonstrate that our RGA can significantly enhance the feature representation power and help achieve the state-of-the-art performance on several popular benchmarks. The source code is available at https://github.com/microsoft/Relation-Aware-Global-Attention-Networks.
Zhizheng Zhang 0004, Cuiling Lan, Wenjun Zeng 0001, Xin Jin 0014, Zhibo Chen 0001
CVPR3
2020 Fusing Wearable IMUs With Multi-View Images for Human Pose Estimation: A Geometric Approach
abstract
We propose to estimate 3D human pose from multi-view images and a few IMUs attached at person's limbs. It operates by firstly detecting 2D poses from the two signals, and then lifting them to the 3D space. We present a geometric approach to reinforce the visual features of each pair of joints based on the IMUs. This notably improves 2D pose estimation accuracy especially when one joint is occluded. We call this approach Orientation Regularized Network (ORN). Then we lift the multi-view 2D poses to the 3D space by an Orientation Regularized Pictorial Structure Model (ORPSM) which jointly minimizes the projection error between the 3D and 2D poses, along with the discrepancy between the 3D pose and IMU orientations. The simple two-step approach reduces the error of the state-of-the-art by a large margin on a public dataset. Our code will be released at https://github.com/microsoft/imu-human-pose-estimation-pytorch.
Zhe Zhang 0045, Chunyu Wang 0001, Wenhu Qin, Wenjun Zeng 0001
CVPR4
2020 Spatiotemporal Fusion in 3D CNNs: A Probabilistic View
abstract
Despite the success in still image recognition, deep neural networks for spatiotemporal signal tasks (such as human action recognition in videos) still suffers from low efficacy and inefficiency over the past years. Recently, human experts have put more efforts into analyzing the importance of different components in 3D convolutional neural networks (3D CNNs) to design more powerful spatiotemporal learning backbones. Among many others, spatiotemporal fusion is one of the essentials. It controls how spatial and temporal signals are extracted at each layer during inference. Previous attempts usually start by ad-hoc designs that empirically combine certain convolutions and then draw conclusions based on the performance obtained by training the corresponding networks. These methods only support network-level analysis on limited number of fusion strategies. In this paper, we propose to convert the spatiotemporal fusion strategies into a probability space, which allows us to perform network-level evaluations of various fusion strategies without having to train them separately. Besides, we can also obtain fine-grained numerical information such as layer-level preference on spatiotemporal fusion within the probability space. Our approach greatly boosts the efficiency of analyzing spatiotemporal fusion. Based on the probability space, we further generate new fusion strategies which achieve the state-of-the-art performance on four well-known action recognition datasets.
Yizhou Zhou, Xiaoyan Sun 0001, Chong Luo 0001, Zhengjun Zha, Wenjun Zeng 0001
CVPR5
2020 Global Distance-Distributions Separation for Unsupervised Person Re-identification
Xin Jin 0014, Cuiling Lan, Wenjun Zeng 0001, Zhibo Chen 0001
ECCV (7)3
2020 VoxelPose: Towards Multi-camera 3D Human Pose Estimation in Wild Environment
Hanyue Tu, Chunyu Wang 0001, Wenjun Zeng 0001
ECCV (1)3
2020 Joint Time-Frequency and Time Domain Learning for Speech Enhancement
abstract
For single-channel speech enhancement, both time-domain and time-frequency-domain methods have their respective pros and cons. In this paper, we present a cross-domain framework named TFT-Net, which takes time-frequency spectrogram as input and produces time-domain waveform as output. Such a framework takes advantage of the knowledge we have about spectrogram and avoids some of the drawbacks that T-F-domain methods have been suffering from. In TFT-Net, we design an innovative dual-path attention block (DAB) to fully exploit correlations along the time and frequency axes. We further discover that a sample-independent DAB (SDAB) achieves a good tradeoff between enhanced speech quality and complexity. Ablation studies show that both the cross-domain design and the SDAB block bring large performance gain. When logarithmic MSE is used as the training criteria, TFT-Net achieves the highest SDR and SSNR among state-of-the-art methods on two major speech enhancement benchmarks.
Chuanxin Tang, Chong Luo 0001, Zhiyuan Zhao 0001, Wenxuan Xie, Wenjun Zeng 0001
IJCAI5
2020 Beyond Intra-modality: A Survey of Heterogeneous Person Re-identification
abstract
An efficient and effective person re-identification (ReID) system relieves the users from painful and boring video watching and accelerates the process of video analysis. Recently, with the explosive demands of practical applications, a lot of research efforts have been dedicated to heterogeneous person re-identification (Hetero-ReID). In this paper, we provide a comprehensive review of state-of-the-art Hetero-ReID methods that address the challenge of inter-modality discrepancies. According to the application scenario, we classify the methods into four categories --- low-resolution, infrared, sketch, and text. We begin with an introduction of ReID, and make a comparison between Homogeneous ReID (Homo-ReID) and Hetero-ReID tasks. Then, we describe and compare existing datasets for performing evaluations, and survey the models that have been widely employed in Hetero-ReID. We also summarize and compare the representative approaches from two perspectives, i.e., the application scenario and the learning pipeline. We conclude by a discussion of some future research directions. Follow-up updates are available at https://github.com/lightChaserX/Awesome-Hetero-reID
Zheng Wang 0007, Zhixiang Wang 0001, Yinqiang Zheng, Yang Wu 0001, Wenjun Zeng 0001, Shin'ichi Satoh 0001
IJCAI5
2020 Multi-Scale Group Transformer for Long Sequence Modeling in Speech Separation
abstract
In this paper, we introduce Transformer to the time-domain methods for single-channel speech separation. Transformer has the potential to boost speech separation performance because of its strong sequence modeling capability. However, its computational complexity, which grows quadratically with the sequence length, has made it largely inapplicable to speech applications. To tackle this issue, we propose a novel variation of Transformer, named multi-scale group Transformer (MSGT). The key ideas are group self-attention, which significantly reduces the complexity, and multi-scale fusion, which retains Transform's ability to capture long-term dependency. We implement two versions of MSGT with different complexities, and apply them to a well-known time-domain speech separation method called Conv-TasNet. By simply replacing the original temporal convolutional network (TCN) with MSGT, our approach called MSGT-TasNet achieves a large gain over Conv-TasNet on both WSJ0-2mix and WHAM! benchmarks. Without bells and whistles, the performance of MSGT-TasNet is already on par with the SOTA methods.
Chong Luo 0001, Zhengjun Zha, Wenjun Zeng 0001
IJCAI4
2020 Object Detection in Videos by High Quality Object Linking
abstract
Compared with object detection in static images, object detection in videos is more challenging due to degraded image qualities. An effective way to address this problem is to exploit temporal contexts by linking the same object across video to form tubelets and aggregating classification scores in the tubelets. In this paper, we focus on obtaining high quality object linking results for better classification. Unlike previous methods that link objects by checking boxes between neighboring frames, we propose to link in the same frame. To achieve this goal, we extend prior methods in following aspects: (1) a cuboid proposal network that extracts spatio-temporal candidate cuboids which bound the movement of objects; (2) a short tubelet detection network that detects short tubelets in short video segments; (3) a short tubelet linking algorithm that links temporally-overlapping short tubelets to form long tubelets. Experiments on the ImageNet VID dataset show that our method outperforms both the static image detector and the previous state of the art. In particular, our method improves results by 8.8 percent over the static image detector for fast moving objects.
Peng Tang 0005, Chunyu Wang 0001, Xinggang Wang, Wenyu Liu 0001, Wenjun Zeng 0001, Jingdong Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2020 Temporal-Spatial Mapping for Action Recognition
abstract
Deep learning models have enjoyed great success for image related computer vision tasks such as image classification and object detection. For video related tasks such as human action recognition, however, the advancements are not as significant yet. The main challenge is the lack of effective and efficient models in modeling the rich temporal-spatial information in a video. We introduce a simple yet effective operation, termed temporal-spatial mapping, for capturing the temporal evolution of the frames by jointly analyzing all the frames of a video. We propose a video level 2D feature representation by transforming the convolutional features of all frames to a 2D feature map, referred to as VideoMap. With each row being the vectorized feature representation of a frame, the temporal-spatial features are compactly represented, while the temporal dynamic evolution is also well embedded. Based on the VideoMap representation, we further propose a temporal attention model within a shallow convolutional neural network to efficiently exploit the temporal-spatial dynamics. The experiment results show that the proposed scheme achieves state-of-the-art performance, with 4.2% accuracy gain over the temporal segment network, a competing baseline method, on the challenging human action benchmark dataset HMDB51.
Cuiling Lan, Wenjun Zeng 0001, Junliang Xing, Xiaoyan Sun 0001, Jing-Yu Yang 0002
IEEE Trans. Circuits Syst. Video Technol.3
2020 View Invariant 3D Human Pose Estimation
abstract
The recent success of neural networks has significantly advanced the performance of 3D human pose estimation from 2D input images. However, the diversity of capturing viewpoints and the flexibility of the human poses remain some significant challenges. In this paper, we propose a view-invariant 3D human pose estimation module to alleviate the effects of viewpoint diversity. The proposed framework consists of a base network, which provides an initial estimation of a 3D pose, a view-invariant hierarchical correction network (VI-HC) on top of that to learn the 3D pose refinement under consistent views, and a view-invariant discriminative network (VID) to enforce high-level constraints over body configurations. In VI-HC, the initial 3D pose inputs are automatically transformed to consistent views for further refinements at the global body and local body parts level, respectively. For the VID, under consistent viewpoints, we use adversarial learning to differentiate between estimated 3D poses and real 3D poses to avoid implausible results. The experimental results demonstrate that the constraint on viewpoint consistency can dramatically enhance the performance of 3D human pose estimation. Our module shows robustness for different 3D pose base networks and achieves a significant improvement (about 9%) over a powerful baseline on the public 3D pose estimation benchmark Human3.6M.
Guoqiang Wei, Cuiling Lan, Wenjun Zeng 0001, Zhibo Chen 0001
IEEE Trans. Circuits Syst. Video Technol.3
2020 EleAtt-RNN: Adding Attentiveness to Neurons in Recurrent Neural Networks
abstract
Recurrent neural networks (RNNs) are capable of modeling temporal dependencies of complex sequential data. In general, current available structures of RNNs tend to concentrate on controlling the contributions of current and previous information. However, the exploration of different importance levels of different elements within an input vector is always ignored. We propose a simple yet effective Element-wise-Attention Gate (EleAttG), which can be easily added to an RNN block (e.g. all RNN neurons in an RNN layer), to empower the RNN neurons to have attentiveness capability. For an RNN block, an EleAttG is used for adaptively modulating the input by assigning different levels of importance, i.e., attention, to each element/dimension of the input. We refer to an RNN block equipped with an EleAttG as an EleAtt-RNN block. Instead of modulating the input as a whole, the EleAttG modulates the input at fine granularity, i.e., element-wise, and the modulation is content adaptive. The proposed EleAttG, as an additional fundamental unit, is general and can be applied to any RNN structures, e.g., standard RNN, Long Short-Term Memory (LSTM), or Gated Recurrent Unit (GRU). We demonstrate the effectiveness of the proposed EleAtt-RNN by applying it to different tasks including the action recognition, from both skeleton-based data and RGB videos, gesture recognition, and sequential MNIST classification. Experiments show that adding attentiveness through EleAttGs to RNN blocks significantly improves the power of RNNs.
Pengfei Zhang 0005, Jianru Xue, Cuiling Lan, Wenjun Zeng 0001, Zhanning Gao, Nanning Zheng 0001
IEEE Trans. Image Process.4
2019 Learning to Refine 3D Human Pose Sequences
abstract
We present a basis approach to refine noisy 3D human pose sequences by jointly projecting them onto a non-linear pose manifold, which is represented by a number of basis dictionaries with each covering a small manifold region. We learn the dictionaries by jointly minimizing the distance between the original poses and their projections on the dictionaries, along with the temporal jittering of the projected poses. During testing, given a sequence of noisy poses which are probably off the manifold, we project them to the manifold using the same strategy as in training for refinement. We apply our approach to the monocular 3D pose estimation and the long term motion prediction tasks. The experimental results on the benchmark dataset shows the estimated 3D poses are notably improved in both tasks. In particular, the smoothness constraint helps generate more robust refinement results even when some poses in the original sequence have large errors.
Jieru Mei, Xingyu Chen 0001, Chunyu Wang 0001, Alan L. Yuille, Xuguang Lan, Wenjun Zeng 0001
3DV6
2019 Detect or Track: Towards Cost-Effective Video Object Detection/Tracking
abstract
State-of-the-art object detectors and trackers are developing fast. Trackers are in general more efficient than detectors but bear the risk of drifting. A question is hence raised – how to improve the accuracy of video object detection/tracking by utilizing the existing detectors and trackers within a given time budget? A baseline is frame skipping – detecting every N-th frames and tracking for the frames in between. This baseline, however, is suboptimal since the detection frequency should depend on the tracking quality. To this end, we propose a scheduler network, which determines to detect or track at a certain frame, as a generalization of Siamese trackers. Although being light-weight and simple in structure, the scheduler network is more effective than the frame skipping baselines and flow-based approaches, as validated on ImageNet VID dataset in video object detection/tracking.
Wenxuan Xie, Xinggang Wang, Wenjun Zeng 0001
AAAI4
2019 Learning Basis Representation to Refine 3D Human Pose Estimations
abstract
Estimating 3D human poses from 2D joint positions is an illposed problem, and is further complicated by the fact that the estimated 2D joints usually have errors to which most of the 3D pose estimators are sensitive. In this work, we present an approach to refine inaccurate 3D pose estimations. The core idea of the approach is to learn a number of bases to obtain tight approximations of the low-dimensional pose manifold where a 3D pose is represented by a convex combination of the bases. The representation requires that globally the refined poses are close to the pose manifold thus avoiding generating illegitimate poses. Second, the designed bases also have the property to guarantee that the distances among the body joints of a pose are within reasonable ranges. Experiments on benchmark datasets show that our approach obtains more legitimate poses over the baselines. In particular, the limb lengths are closer to the ground truth.
Chunyu Wang 0001, Haibo Qiu, Alan L. Yuille, Wenjun Zeng 0001
AAAI4
2019 SPM-Tracker: Series-Parallel Matching for Real-Time Visual Object Tracking
abstract
The greatest challenge facing visual object tracking is the simultaneous requirements on robustness and discrimination power. In this paper, we propose a SiamFC-based tracker, named SPM-Tracker, to tackle this challenge. The basic idea is to address the two requirements in two separate matching stages. Robustness is strengthened in the coarse matching (CM) stage through generalized training while discrimination power is enhanced in the fine matching (FM) stage through a distance learning network. The two stages are connected in series as the input proposals of the FM stage are generated by the CM stage. They are also connected in parallel as the matching scores and box location refinements are fused to generate the final results. This innovative series-parallel structure takes advantage of both stages and results in superior performance. The proposed SPM-Tracker, running at 120 fps on GPU, achieves an AUC of 0.687 on OTB-100 and an EAO of 0.434 on VOT-16, exceeding other real-time trackers by a notable margin.
Guangting Wang, Chong Luo 0001, Zhiwei Xiong, Wenjun Zeng 0001
CVPR4
2019 Densely Semantically Aligned Person Re-Identification
abstract
We propose a densely semantically aligned person re-identification (re-ID) framework. It fundamentally addresses the body misalignment problem caused by pose/viewpoint variations, imperfect person detection, occlusion, etc.. By leveraging the estimation of the dense semantics of a person image, we construct a set of densely semantically aligned part images (DSAP-images), where the same spatial positions have the same semantics across different person images. We design a two-stream network that consists of a main full image stream (MF-Stream) and a densely semantically-aligned guiding stream (DSAG-Stream). The DSAG-Stream, with the DSAP-images as input, acts as a regulator to guide the MF-Stream to learn densely semantically aligned features from the original image. In the inference, the DSAG-Stream is discarded and only the MF-Stream is needed, which makes the inference system computationally efficient and robust. To our best knowledge, we are the first to make use of fine grained semantics for addressing misalignment problems for re-ID. Our method achieves rank-1 accuracy of 78.9% (new protocol) on the CUHK03 dataset, 90.4% on the CUHK01 dataset, and 95.7% on the Market1501 dataset, outperforming state-of-the-art methods.
Zhizheng Zhang 0004, Cuiling Lan, Wenjun Zeng 0001, Zhibo Chen 0001
CVPR3
2019 Context-Reinforced Semantic Segmentation
abstract
Recent efforts have shown the importance of context on deep convolutional neural network based semantic segmentation. Among others, the predicted segmentation map (p-map) itself which encodes rich high-level semantic cues (e.g. objects and layout) can be regarded as a promising source of context. In this paper, we propose a dedicated module, Context Net, to better explore the context information in p-maps. Without introducing any new supervisions, we formulate the context learning problem as a Markov Decision Process and optimize it using reinforcement learning during which the p-map and Context Net are treated as environment and agent, respectively. Through adequate explorations, the Context Net selects the information which has long-term benefit for segmentation inference. By incorporating the Context Net with a baseline segmentation scheme, we then propose a Context-reinforced Semantic Segmentation network (CiSS-Net), which is fully end-to-end trainable. Experimental results show that the learned context brings 3.9% absolute improvement on mIoU over the baseline segmentation method, and the CiSS-Net achieves the state-of-the-art segmentation performance on ADE20K, PASCAL-Context and Cityscapes.
Yizhou Zhou, Xiaoyan Sun 0001, Zhengjun Zha, Wenjun Zeng 0001
CVPR4
2019 Content-Aware Personalised Rate Adaptation for Adaptive Streaming via Deep Video Analysis
abstract
Adaptive bitrate (ABR) streaming is the de facto solution for achieving smooth viewing experiences under unstable network conditions. However, most of the existing rate adaptation approaches for ABR are content-agnostic, without considering the semantic information of the video content. Nevertheless, semantic information largely determines the informativeness and interestingness of the video content, and consequently affects the QoE for video streaming. One common case is that the user may expect higher quality for the parts of video content that are more interesting or informative so as to reduce overall subjective quality loss. This creates two main challenges for such a problem: First, how to determine which parts of the video content are more interesting? Second, how to allocate bitrate budgets for different parts of the video content with different significances? To address these challenges, we propose a Content-of-Interest (CoI) based rate adaptation scheme for ABR. We first design a deep learning approach for recognizing the interestingness of the video content, and then design a Deep Q-Network (DQN) approach for rate adaptation by incorporating video interestingness information. The experimental results show that our method can recognize video interestingness precisely, and the bitrate allocation for ABR can be aligned with the interestingness of video content while not compromising the performances on objective QoE metrics.
Guanyu Gao, Linsen Dong, Huaizheng Zhang, Yonggang Wen 0001, Wenjun Zeng 0001
ICC5
2019 Cross View Fusion for 3D Human Pose Estimation
abstract
We present an approach to recover absolute 3D human poses from multi-view images by incorporating multi-view geometric priors in our model. It consists of two separate steps: (1) estimating the 2D poses in multi-view images and (2) recovering the 3D poses from the multi-view 2D poses. First, we introduce a cross-view fusion scheme into CNN to jointly estimate 2D poses for multiple views. Consequently, the 2D pose estimation for each view already benefits from other views. Second, we present a recursive Pictorial Structure Model to recover the 3D pose from the multi-view 2D poses. It gradually improves the accuracy of 3D pose with affordable computational cost. We test our method on two public datasets H36M and Total Capture. The Mean Per Joint Position Errors on the two datasets are 26mm and 29mm, which outperforms the state-of-the-arts remarkably (26mm vs 52mm, 29mm vs 35mm).
Haibo Qiu, Chunyu Wang 0001, Jingdong Wang 0001, Naiyan Wang, Wenjun Zeng 0001
ICCV5
2019 Unsupervised High-Resolution Depth Learning From Videos With Dual Networks
abstract
Unsupervised depth learning takes the appearance difference between a target view and a view synthesized from its adjacent frame as supervisory signal. Since the supervisory signal only comes from images themselves, the resolution of training data significantly impacts the performance. High-resolution images contain more fine-grained details and provide more accurate supervisory signal. However, due to the limitation of memory and computation power, the original images are typically down-sampled during training, which suffers heavy loss of details and disparity accuracy. In order to fully explore the information contained in high-resolution data, we propose a simple yet effective dual networks architecture, which can directly take high-resolution images as input and generate high-resolution and high-accuracy depth map efficiently. We also propose a Self-assembled Attention (SA-Attention) module to handle low-texture region. The evaluation on the benchmark KITTI and Make3D datasets demonstrates that our method achieves state-of-the-art results in the monocular depth estimation task.
Junsheng Zhou, Yuwang Wang, Kaihuai Qin, Wenjun Zeng 0001
ICCV4
2019 Moving Indoor: Unsupervised Video Depth Learning in Challenging Environments
abstract
Recently unsupervised learning of depth from videos has made remarkable progress and the results are comparable to fully supervised methods in outdoor scenes like KITTI. However, there still exist great challenges when directly applying this technology in indoor environments, e.g., large areas of non-texture regions like white wall, more complex ego-motion of handheld camera, transparent glasses and shiny objects. To overcome these problems, we propose a new optical-flow based training paradigm which reduces the difficulty of unsupervised learning by providing a clearer training target and handles the non-texture regions. Our experimental evaluation demonstrates that the result of our method is comparable to fully supervised methods on the NYU Depth V2 benchmark. To the best of our knowledge, this is the first quantitative result of purely unsupervised learning method reported on indoor datasets.
Junsheng Zhou, Yuwang Wang, Kaihuai Qin, Wenjun Zeng 0001
ICCV4
2019 Quality-Gated Convolutional Lstm for Enhancing Compressed Video
abstract
The past decade has witnessed great success in applying deep learning to enhance the quality of compressed video. However, the existing approaches aim at quality enhancement on a single frame, or only using fixed neighboring frames. Thus they fail to take full advantage of the inter-frame correlation in the video. This paper proposes the Quality-Gated Convolutional Long Short-Term Memory (QG-ConvLSTM) network with bi-directional recurrent structure to fully exploit the advantageous information in a large range of frames. More importantly, due to the obvious quality fluctuation among compressed frames, higher quality frames can provide more useful information for other frames to enhance quality. Therefore, we propose learning the "forget" and "'input" gates in the ConvLSTM cell from quality-related features. As such, the frames with various quality contribute to the memory in ConvLSTM with different importance, making the information of each frame reasonably and adequately used. Finally, the experiments validate the effectiveness of our QG-ConvLSTM approach in advancing the state-of-the-art quality enhancement of compressed video, and the ablation study shows that our QG-ConvLSTM approach is learnt to make a trade-off between quality and correlation when leveraging multi-frame information. The project page: https://github.com/ryangchn/QG-ConvLSTM.git.
Xiaoyan Sun 0001, Mai Xu, Wenjun Zeng 0001
ICME4
2019 Predicting Future Instance Segmentation with Contextual Pyramid ConvLSTMs
abstract
Despite the remarkable progress in instance segmentation, the problem of predicting future instance segmentation remains challenging due to the unobservability of future data. Existing methods mainly address this challenge by forecasting pyramid features to represent unobserved future frames. However, they mainly predict features for each pyramid level independently, and ignore the underlying structural relationship between features of different levels.
Jiangxin Sun, Jiafeng Xie, Jianfang Hu, Zihang Lin, Jian-Huang Lai, Wenjun Zeng 0001, Wei-Shi Zheng 0001
ACM Multimedia6
2019 High-Speed Hyperspectral Video Acquisition By Combining Nyquist and Compressive Sampling
abstract
We propose a novel hybrid imaging system to acquire 4D high-speed hyperspectral (HSHS) videos with high spatial and spectral resolution. The proposed system consists of two branches: one branch performs Nyquist sampling in the temporal dimension while integrating the whole spectrum, resulting in a high-frame-rate panchromatic video; the other branch performs compressive sampling in the spectral dimension with longer exposures, resulting in a low-frame-rate hyperspectral video. Owing to the high light throughput and complementary sampling, these two branches jointly provide reliable measurements for recovering the underlying HSHS video. Moreover, the panchromatic video can be used to learn an over-complete 3D dictionary to represent each band-wise video sparsely, thanks to the inherent structural similarity in the spectral dimension. Based on the joint measurements and the self-adaptive dictionary, we further propose a simultaneous spectral sparse (3S) model to reinforce the structural similarity across different bands and develop an efficient computational reconstruction algorithm to recover the HSHS video. Both simulation and hardware experiments validate the effectiveness of the proposed approach. To the best of our knowledge, this is the first time that hyperspectral videos can be acquired at a frame rate up to 100fps with commodity optical elements and under ordinary indoor illumination.
Lizhi Wang 0001, Zhiwei Xiong, Hua Huang 0001, Guangming Shi, Feng Wu 0001, Wenjun Zeng 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2019 View Adaptive Neural Networks for High Performance Skeleton-Based Human Action Recognition
abstract
Skeleton-based human action recognition has recently attracted increasing attention thanks to the accessibility and the popularity of 3D skeleton data. One of the key challenges in action recognition lies in the large variations of action representations when they are captured from different viewpoints. In order to alleviate the effects of view variations, this paper introduces a novel view adaptation scheme, which automatically determines the virtual observation viewpoints over the course of an action in a learning based data driven manner. Instead of re-positioning the skeletons using a fixed human-defined prior criterion, we design two view adaptive neural networks, i.e., VA-RNN and VA-CNN, which are respectively built based on the recurrent neural network (RNN) with the Long Short-term Memory (LSTM) and the convolutional neural network (CNN). For each network, a novel view adaptation module learns and determines the most suitable observation viewpoints, and transforms the skeletons to those viewpoints for the end-to-end recognition with a main classification network. Ablation studies find that the proposed view adaptive models are capable of transforming the skeletons of various views to much more consistent virtual viewpoints. Therefore, the models largely eliminate the influence of the viewpoints, enabling the networks to focus on the learning of action-specific features and thus resulting in superior performance. In addition, we design a two-stream scheme (referred to as VA-fusion) that fuses the scores of the two networks to provide the final prediction, obtaining enhanced performance. Moreover, random rotation of skeleton sequences is employed to improve the robustness of view adaptation models and alleviate overfitting during training. Extensive experimental evaluations on five challenging benchmarks demonstrate the effectiveness of the proposed view-adaptive networks and superior performance over state-of-the-art approaches.
Pengfei Zhang 0005, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Jianru Xue, Nanning Zheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2019 Skeleton-Based Action Recognition With Gated Convolutional Neural Networks
abstract
For skeleton-based action recognition, most of the existing works used recurrent neural networks. Using convolutional neural networks (CNNs) is another attractive solution considering their advantages in parallelization, effectiveness in feature learning, and model base sufficiency. Besides these, skeleton data are low-dimensional features. It is natural to arrange a sequence of skeleton features chronologically into an image, which retains the original information. Therefore, we solve the sequence learning problem as an image classification task using CNNs. For better learning ability, we build a classification network with stacked residual blocks and having a special design called linear skip gated connection which can benefit information propagation across multiple residual blocks. When arranging the coordinates of body joints in one frame into a skeleton feature, we systematically investigate the performance of part-based, chain-based, and traversal-based orders. Furthermore, a fully convolutional permutation network is designed to learn an optimized order for data rearrangement. Without any bells and whistles, our proposed model achieves state-of-the-art performance on two challenging benchmark datasets, outperforming existing methods significantly.
Congqi Cao, Cuiling Lan, Yifan Zhang 0001, Wenjun Zeng 0001, Hanqing Lu, Yanning Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2019 Multi-Modality Multi-Task Recurrent Neural Network for Online Action Detection
abstract
Online action detection is a brand new challenge and plays a critical role in visual surveillance analytics. It goes one step further than a conventional action recognition task, which recognizes human actions from well-segmented clips. Online action detection is desired to identify the action type and localize action positions on the fly from the untrimmed stream data. In this paper, we propose a multi-modality multi-task recurrent neural network, which incorporates both RGB and Skeleton networks. We design different temporal modeling networks to capture specific characteristics from various modalities. Then, a deep long short-term memory subnetwork is utilized effectively to capture the complex long-range temporal dynamics, naturally avoiding the conventional sliding window design and thus ensuring high computational efficiency. Constrained by a multi-task objective function in the training phase, this network achieves superior detection performance and is capable of automatically localizing the start and end points of actions more accurately. Furthermore, embedding subtask of regression provides the ability to forecast the action prior to its occurrence. We evaluate the proposed method and several other methods in action detection and forecasting on the online action detection data set and gaming action data set datasets. Experimental results demonstrate that our model achieves the state-of-the-art performance on both tasks.
Jiaying Liu 0001, Yanghao Li, Sijie Song, Junliang Xing, Cuiling Lan, Wenjun Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.6
2019 High-Order Statistical Modeling Based on a Decision Tree for Distributed Video Coding
abstract
Aiming at low-complexity encoding, distributed video coding (DVC) based on the Wyner-Ziv theorem has attracted significant attention. However, there is still a compression performance gap between the state-of-the-art DVC and the conventional video coding. One of the most important factors is the efficient estimation of the source correlation statistics. The first-order Laplacian distribution has been widely used for source correlation modeling, but is not effective enough; high-order statistical modeling is necessary, but needs more context features for the estimation of a source symbol's conditional probability. How to analyze the strength of the correlation between the source symbol and the context features in order to utilize the features effectively is crucial for such modeling. In this paper, the estimation of the source statistical distribution is first treated as a classification problem. The symbols of the source can be classified into different classes when the relevant context features are given. Then decision tree learning is introduced to analyze the strength of the correlation between the source symbol and the context features. Specifically, by constructing the decision trees composed of the selected context features, the selected context features can be organized effectively to derive the rules to estimate the current symbol's conditional probability, upon which the high-order statistical modeling is designed. Experimental results show that the proposed model can achieve significant coding gain over existing DVC systems, especially for natural videos with high motion intensity.
Linbo Qing, Wenjun Zeng 0001, Xiaohai He
IEEE Trans. Circuits Syst. Video Technol.3
2019 Benchmarking Single-Image Dehazing and Beyond
abstract
In this paper, we present a comprehensive study and evaluation of existing single image dehazing algorithms, using a new large-scale benchmark consisting of both synthetic and real-world hazy images, called REalistic Single Image DEhazing (RESIDE). RESIDE highlights diverse data sources and image contents, and is divided into five subsets, each serving different training or evaluation purposes. We further provide a rich variety of criteria for dehazing algorithm evaluation, ranging from full-reference metrics, to no-reference metrics, to subjective evaluation and the novel task-driven evaluation. Experiments on RESIDE shed light on the comparisons and limitations of stateof- the-art dehazing algorithms, and suggest promising future directions.
Boyi Li 0001, Wenqi Ren, Dengpan Fu, Dacheng Tao, Dan Feng 0001, Wenjun Zeng 0001, Zhangyang Wang
IEEE Trans. Image Process.6
2019 Learning to Update for Object Tracking With Recurrent Meta-Learner
abstract
Model update lies at the heart of object tracking. Generally, model update is formulated as an online learning problem where a target model is learned over the online training set. Our key innovation is to formulate the model update problem in the meta-learning framework and learn the online learning algorithm itself using large numbers of offline videos, i.e., learning to update. The learned updater takes as input the online training set and outputs an updated target model. As a first attempt, we design the learned updater based on recurrent neural networks (RNNs) and demonstrate its application in a template-based tracker and a correlation filter-based tracker. Our learned updater consistently improves the base trackers and runs faster than realtime on GPU while requiring small memory footprint during testing. Experiments on standard benchmarks demonstrate that our learned updater outperforms commonly used update baselines including the efficient exponential moving average (EMA)-based update and the well-designed stochastic gradient descent (SGD)-based update. Equipped with our learned updater, the template-based tracker achieves state-of-the-art performance among realtime trackers on GPU.
Bi Li 0005, Wenxuan Xie, Wenjun Zeng 0001, Wenyu Liu 0001
IEEE Trans. Image Process.3
2019 Learning Attentional Recurrent Neural Network for Visual Tracking
abstract
Existing visual tracking methods face many challenges: 1) the changed size and number of targets over time, occlusion in discrete frames, and mis-identification for crossing targets. Long short-term memory (LSTM) has the advantage of modeling long-term tasks and is suitable for tracking. We propose a novel online attentional recurrent neural network (ARNN) model for visual tracking, whose core component is a two-layer bidirectional LSTM along the x-and y-axes. Several bidirectional LSTMs can be cascaded or parallelly connected together to exploit multiscale target features and can give more precise tracked object locations. Each bidirectional LSTM utilizes the convolutional features of a convolutional neural network inside two bounding boxes from two frames to check whether the target in the current frame is the one in previous frames. An attention mechanism is also adopted to enhance the proposed model to better express the patch-level features of the tracking targets. Interattention and intra-attention models are proposed to imitate the temporal and spatial tracking mechanism of primate visual cortex. Interattention learns to overcome the occlusion problem, and intra-attention is able to mark important regions to better trace the target. The bidirectional LSTM and the attention mechanism are jointly trained. The combination of them further improves the accuracy of target tracking in videos. The outstanding performances in the experiments demonstrate the effectiveness of our proposed online method ARNN and yield competitive results compared with the state-of-the-art tracking methods.
Qiurui Wang, Chun Yuan 0003, Jingdong Wang 0001, Wenjun Zeng 0001
IEEE Trans. Multim.4
2018 A Twofold Siamese Network for Real-Time Object Tracking
abstract
Observing that Semantic features learned in an image classification task and Appearance features learned in a similarity matching task complement each other, we build a twofold Siamese network, named SA-Siam, for real-time object tracking. SA-Siam is composed of a semantic branch and an appearance branch. Each branch is a similaritylearning Siamese network. An important design choice in SA-Siam is to separately train the two branches to keep the heterogeneity of the two types of features. In addition, we propose a channel attention mechanism for the semantic branch. Channel-wise weights are computed according to the channel activations around the target position. While the inherited architecture from SiamFC [3] allows our tracker to operate beyond real-time, the twofold design and the attention mechanism significantly improve the tracking performance. The proposed SA-Siam outperforms all other real-time trackers by a large margin on OTB-2013/50/100 benchmarks.
Anfeng He, Chong Luo 0001, Xinmei Tian 0001, Wenjun Zeng 0001
CVPR4
2018 MiCT: Mixed 3D/2D Convolutional Tube for Human Action Recognition
abstract
Human actions in videos are three-dimensional (3D) signals. Recent attempts use 3D convolutional neural networks (CNNs) to explore spatio-temporal information for human action recognition. Though promising, 3D CNNs have not achieved high performance on this task with respect to their well-established two-dimensional (2D) counterparts for visual recognition in still images. We argue that the high training complexity of spatio-temporal fusion and the huge memory cost of 3D convolution hinder current 3D CNNs, which stack 3D convolutions layer by layer, by outputting deeper feature maps that are crucial for high-level tasks. We thus propose a Mixed Convolutional Tube (MiCT) that integrates 2D CNNs with the 3D convolution module to generate deeper and more informative feature maps, while reducing training complexity in each round of spatio-temporal fusion. A new end-to-end trainable deep 3D network, MiCT-Net, is also proposed based on the MiCT to better explore spatio-temporal information in human actions. Evaluations on three well-known benchmark datasets (UCF101, Sport-1M and HMDB-51) show that the proposed MiCT-Net significantly outperforms the original 3D CNNs. Compared with state-of-the-art approaches for action recognition on UCF101 and HMDB51, our MiCT-Net yields the best performance.
Yizhou Zhou, Xiaoyan Sun 0001, Zhengjun Zha, Wenjun Zeng 0001
CVPR4
2018 Online Dictionary Learning for Approximate Archetypal Analysis
Jieru Mei, Chunyu Wang 0001, Wenjun Zeng 0001
ECCV (3)3
2018 Adding Attentiveness to the Neurons in Recurrent Neural Networks
Pengfei Zhang 0005, Jianru Xue, Cuiling Lan, Wenjun Zeng 0001, Zhanning Gao, Nanning Zheng 0001
ECCV (9)4
2018 Cooperative Hybrid Digital-Analog Video Transmission in D2D Networks
abstract
In this paper, we propose a cooperative video transmission scheme in D2D networks. This research is motivated by the growing interests in hybrid digital-analog video transmissions and device-to-device (D2D) communications. The framework of D2D communications can be generally modeled as a three-node network. In this network, coset coding is used to allow the destination to exploit the correlations between the video signals received in two phases. We have done some work of further optimization to improve the video quality at destination in this network. First, we derive a closed form of the reconstruction error at the destination. This provides a theoretical foundation for finding the optimal quantization step size in coset coding. Then, based on the accurate analysis on the coset coding we design a new power allocation algorithm. Experimental results verify that our scheme outperforms the recently proposed WCVC and DCVC.
Jian Shen 0002, Chong Luo 0001, Houqiang Li, Wenjun Zeng 0001
ICIP5
2018 Skeleton-Indexed Deep Multi-Modal Feature Learning for High Performance Human Action Recognition
abstract
This paper presents a new framework for action recognition with multi-modal data. A skeleton-indexed feature learning procedure is developed to further exploit the detailed local features from RGB and optical flow videos. In particular, the proposed framework is built based on a deep Convolutional Network (ConvNet) and a Recurrent Neural Network (RNN) with Long Short Term Memory (LSTM). A skeleton-indexed transform layer is designed to automatically extract visual features around key joints, and a part-aggregated pooling is developed to uniformly regulate the visual features from different body parts and actors. Besides, several fusion schemes are explored to take advantage of multi-modal data. The proposed deep architecture is end-to-end trainable and can better incorporate different modalities to learn effective feature representations. Quantitative experiment results on two datasets, the NTU RGB+D dataset and the MSR dataset, demonstrate the excellent performance of our scheme over other state-of-the-arts. To our knowledge, the performance obtained by the proposed framework is currently the best on the challenging NTU RGB+D dataset.
Sijie Song, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Jiaying Liu 0001
ICME4
2018 Fast Discrete Cross-modal Hashing With Regressing From Semantic Labels
abstract
Hashing has recently received great attention in cross-modal retrieval. Cross-modal retrieval aims at retrieving information across heterogeneous modalities (e.g., texts vs. images). Cross-modal hashing compresses heterogeneous high-dimensional data into compact binary codes with similarity preserving, which provides efficiency and facility in both retrieval and storage. In this study, we propose a novel fast discrete cross-modal hashing (FDCH) method with regressing from semantic labels to take advantage of supervised labels to improve retrieval performance. In contrast to existing methods that learn the projection from hash codes to semantic labels, the proposed FDCH regresses the semantic labels of training examples to the corresponding hash codes with a drift. It not only accelerates the hash learning process, but also helps generate stable hash codes. Furthermore, the drift can adjust the regression and enhance the discriminative capability of hash codes. Especially in the case of training efficiency, FDCH is much faster than existing methods. Comparisons with several state-of-the-art techniques on three benchmark datasets have demonstrated the superiority of FDCH under various cross-modal retrieval scenarios.
Xingbo Liu, Xiushan Nie, Wenjun Zeng 0001, Chaoran Cui, Lei Zhu 0002, Yilong Yin
ACM Multimedia3
2018 Real-Time Object Tracking with Motion Information
abstract
Motion is a vital information for object tracking. However, most existing methods, including the classic Siamese FC network [1], only consider the object appearance, and ignore the vital motion feature. In this paper, we design a dual-network object tracker, which is called DOT for short, to effectively combine the appearance and motion information. Our method employs two branches, S-net and M-net, to exploit the appearance and motion information respectively. Moreover, an attention fusion module is also introduced to effectively integrate these two aspects. The experiments carried out on OTB-2013 demonstrate the improvement on object tracking by the integration of motion information with our dual-network and attention fusion.
Chaoqun Wang 0011, Xiaoyan Sun 0001, Xuejin Chen, Wenjun Zeng 0001
VCIP4
2018 A Practical Hybrid Digital-Analog Scheme for Wireless Video Transmission
abstract
We propose a hybrid digital-analog framework for wireless video transmission, which benefits from both the high distortion-power performance of digital systems and the graceful performance degradation of analog systems. The proposed framework models video frames as a parallel Gaussian source, which is separated into digital and analog parts through scalar quantization. It features entropy coding and channel coding in digital transmission and power scaling in analog transmission. The key challenge in this framework is how to allocate the constrained power and bandwidth resources between and among digital and analog components to achieve minimal distortion at the receiver. Given the worst-case channel signal-to-noise ratio, we are able to derive a closed-form expression of the overall distortion. However, minimizing it is a mixed-integer non-linear programming problem, which is generally non-deterministic polynomial-time hard. By making reasonable and justified simplifications, we approach the optimal solution through a practical scheme. Evaluations show that the proposed scheme outperforms the state-of-the-art analog scheme SoftCast by a large margin. The gain in received video peak signal-to-noise ratio is up to 5.0 dB for various types of videos.
Cuiling Lan, Chong Luo 0001, Wenjun Zeng 0001, Feng Wu 0001
IEEE Trans. Circuits Syst. Video Technol.3
2018 Variable Block-Sized Signal-Dependent Transform for Video Coding
abstract
Transform, as one of the most important modules of mainstream video coding systems, seems very stable over the past several decades. However, recent developments indicate that bringing more options for transform can lead to coding efficiency benefits. In this paper, we go further to investigate how the coding efficiency can be improved over the state-of-the-art method by adapting a transform for each block. We present a variable block-sized signal-dependent transforms (SDTs) design based on the High Efficiency Video Coding (HEVC) framework. For a coding block ranged from $4\times4$ to $32\times32$ , we collect a quantity of similar blocks from the reconstructed area and use them to derive the Karhunen-Loève transform. We avoid sending overhead bits to denote the transform by performing the same procedure at the decoder. In this way, the transform for every block is tailored according to its statistics, to be signal-dependent. To make the large block-sized SDTs feasible, we present a fast algorithm for transform derivation. Experimental results show the effectiveness of the SDTs for different block sizes, which leads to up to 23.3% bit-saving. On average, we achieve BD-rate saving of 2.2%, 2.4%, 3.3%, and 7.1% under AI-Main10, RA-Main10, RA-Main10, and LP-Main10 configurations, respectively, compared with the test model HM-12 of HEVC. The proposed scheme has also been adopted into the joint exploration test model for the exploration of potential future video coding standard.
Cuiling Lan, Jizheng Xu, Wenjun Zeng 0001, Guangming Shi, Feng Wu 0001
IEEE Trans. Circuits Syst. Video Technol.3
2018 Superimposed Modulation for Soft Video Delivery With Hidden Resources
abstract
Analog-transmission-based soft video delivery suffers from the leveling-off effect when the allocated channel bandwidth is severely insufficient. Fortunately, with superimposed modulation, it is possible for analog traffic to share bandwidth with digital traffic. In this paper, we design and analyze such a hybrid digital-analog superimposed modulation (HDA-SIM) scheme for soft video delivery. Unlike previous work, we treat the bandwidth of competing digital traffic as hidden resources for the video delivery system. The key problem in this scheme is how to allocate the bandwidth and power resources among various modulation symbols so that we can improve the performance of video delivery without sacrificing the throughput of existing digital traffic. The resource allocation problem is formulated and the optimal solution under any given channel signal-to-noise ratio is derived. Based on the results, the sufficient and necessary condition for the video delivery system to achieve performance gain is given. In addition, we implement the proposed scheme for two state-of-the-art soft video delivery systems known as SoftCast and SharpCast. Both simulations and testbed evaluations show that the HDA-SIM version can achieve significant gains in the received video quality over their original designs.
Chong Luo 0001, Ruiqin Xiong, Wenjun Zeng 0001, Feng Wu 0001
IEEE Trans. Circuits Syst. Video Technol.4
2018 Simultaneous Depth and Spectral Imaging With a Cross-Modal Stereo System
abstract
This letter presents a novel approach for simultaneous depth and spectral imaging with a cross-modal stereo system. Two images of the target scene are captured at the same time: one compressively sampled hyperspectral measurement and one panchromatic measurement. The underlying hyperspectral cube is first reconstructed by leveraging the compressive sensing theory, during which a self-adaptive dictionary is learned from the panchromatic measurement to facilitate the reconstruction. The depth information of the scene is then recovered by estimating a disparity map between the hyperspectral cube and the panchromatic measurement through stereo matching. This disparity map, once obtained, is used to align the hyperspectral and panchromatic measurements to boost the hyperspectral reconstruction in an iterative manner. Through hardware experiments, for the first time to our knowledge, we demonstrate a snapshot system that allows for simultaneous depth and spectral imaging. The proposed system is capable of recording depth and spectral videos of dynamic scenes.
Lizhi Wang 0001, Zhiwei Xiong, Guangming Shi, Wenjun Zeng 0001, Feng Wu 0001
IEEE Trans. Circuits Syst. Video Technol.4
2018 Spatio-Temporal Attention-Based LSTM Networks for 3D Action Recognition and Detection
abstract
Human action analytics has attracted a lot of attention for decades in computer vision. It is important to extract discriminative spatio-temporal features to model the spatial and temporal evolutions of different actions. In this paper, we propose a spatial and temporal attention model to explore the spatial and temporal discriminative features for human action recognition and detection from skeleton data. We build our networks based on the recurrent neural networks with long short-term memory units. The learned model is capable of selectively focusing on discriminative joints of skeletons within each input frame and paying different levels of attention to the outputs of different frames. To ensure effective training of the network for action recognition, we propose a regularized cross-entropy loss to drive the learning process and develop a joint training strategy accordingly. Moreover, based on temporal attention, we develop a method to generate the action temporal proposals for action detection. We evaluate the proposed method on the SBU Kinect Interaction data set, the NTU RGB + D data set, and the PKU-MMD data set, respectively. Experiment results demonstrate the effectiveness of our proposed model on both action recognition and action detection.
Sijie Song, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Jiaying Liu 0001
IEEE Trans. Image Process.4
2018 Optimizing Quality of Experience for Adaptive Bitrate Streaming via Viewer Interest Inference
abstract
Rate adaptation is widely adopted in video streaming to improve the quality of experience (QoE). However, most of the existing rate adaptation approaches neglect the underlying video semantic information. In fact, influenced by video semantics and viewer preferences, the viewer may have different degrees of interest on different parts of a video. The interesting parts of a video can draw more visual attention from the viewer and have higher visual importance. As such, delivering the parts of a video that are interesting to the viewer in a higher quality can improve the perceptual video quality, compared with the semantics-agnostic approaches that treat each part of a video equally. Thus, it is natural to wonder: how to allocate bitrate budgets temporally over a video session under time-varying bandwidth while considering viewer interest? As an exploratory study, we propose an interest-aware rate adaptation approach for improving QoE by inferring viewer interest based on video semantics. We adopt the deep learning method to recognize the scenes of video frames and leverage the term frequency-inverse document frequency method to analyze the degrees of an individual viewer's interest on different types of scenes. The bandwidth, buffer occupancy, and viewer interest are jointly considered under the model predictive control framework for selecting appropriate bitrates for maximizing QoE. The objective and subjective evaluations measured in a real environment show that our method can achieve a higher QoE compared with the semantics-agnostic approaches.
Guanyu Gao, Huaizheng Zhang, Han Hu 0003, Yonggang Wen 0001, Jianfei Cai 0001, Chong Luo 0001, Wenjun Zeng 0001
IEEE Trans. Multim.7
2018 Hybrid Digital-Analog Video Delivery With Shannon-Kotel'nikov Mapping
abstract
Hybrid digital-analog (HDA) transmission is becoming an attractive solution for mobile video delivery because it not only has graceful degradation with channel variations but also yields high power efficiency. However, the heavy bandwidth demand of analog transmission is still an unsolved problem, limiting the received video quality when the bandwidth is not sufficient. To address this problem, we propose adopting Shannon-Kotel'nikov (SK) mapping for HDA video transmission and design an HDA scheme called SK-Cast. SK-Cast consists of a digital and an analog branch. In the digital branch, SK-Cast compresses the video sequence using an high efficiency video coding digital encoder to produce a base layer. The base layer is transmitted through digital methods with strong protection. The residual signals are then decorrelated using three-dimensional discrete cosine transform transform. The SK mapping is exploited to transmit these coefficients, as they can achieve efficient bandwidth compression. We address the resource allocation problems in SK-Cast, including the allocation between digital and analog branches and the allocation among analog symbols. The simulation results show that the SK-Cast outperforms the state-of-the-art HDA systems, including WSVC and SharpCast, and a digital scalable video coding system.
Chong Luo 0001, Ruiqin Xiong, Wenjun Zeng 0001, Feng Wu 0001
IEEE Trans. Multim.4
2018 Photo Stylistic Brush: Robust Style Transfer via Superpixel-Based Bipartite Graph
abstract
With the rapid development of social network and multimedia technology, customized image and video stylization have been widely used for various social-media applications. In this paper, we explore the problem of exemplar-based photo style transfer, which provides a flexible and convenient way to invoke fantastic visual impression. Rather than investigating some fixed artistic patterns to represent certain styles as was done in some previous works, our work emphasizes styles related to a series of visual effects in the photograph (e.g., color, tone, and contrast). We propose a photo stylistic brush, an automatic robust style transfer approach based on Super pixel-based BIpartite Graph (SuperBIG). A two-step bipartite graph algorithm with different granularity levels is employed to aggregate pixels into superpixels and find their correspondences. In the first step, with the extracted hierarchical features, a bipartite graph is constructed to describe the content similarity for pixel partition to produce superpixels. In the second step, superpixels in the input/reference image are rematched to form a new superpixel-based bipartite graph, and superpixel-level correspondences are generated by bipartite matching. Finally, the refined correspondence guides SuperBIG to perform the transformation in a decorrelated color space. Extensive experimental results demonstrate the effectiveness and robustness of the proposed method for transferring various styles of exemplar images, even for some challenging cases, such as night images.
Jiaying Liu 0001, Wenhan Yang, Xiaoyan Sun 0001, Wenjun Zeng 0001
IEEE Trans. Multim.4
2017 An End-to-End Spatio-Temporal Attention Model for Human Action Recognition from Skeleton Data
abstract
Human action recognition is an important task in computer vision. Extracting discriminative spatial and temporal features to model the spatial and temporal evolutions of different actions plays a key role in accomplishing this task. In this work, we propose an end-to-end spatial and temporal attention model for human action recognition from skeleton data. We build our model on top of the Recurrent Neural Networks (RNNs) with Long Short-Term Memory (LSTM), which learns to selectively focus on discriminative joints of skeleton within each frame of the inputs and pays different levels of attention to the outputs of different frames. Furthermore, to ensure effective training of the network, we propose a regularized cross-entropy loss to drive the model learning process and develop a joint training strategy accordingly. Experimental results demonstrate the effectiveness of the proposed model, both on the small human action recognition dataset of SBU and the currently largest NTU dataset.
Sijie Song, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Jiaying Liu 0001
AAAI4
2017 A CNN-Based Approach for Automatic License Plate Recognition in the Wild
Meng Dong, Dongliang He, Chong Luo 0001, Dong Liu 0002, Wenjun Zeng 0001
BMVC5
2017 Dropout Prediction in Home Care Training
Wenjun Zeng 0001, Si-Chi Chin, Brenda Zeimet, Rui Kuang, Chih-Lin Chi
EDM1
2017 Human Pose Estimation Using Global and Local Normalization
abstract
In this paper, we address the problem of estimating the positions of human joints, i.e., articulated pose estimation. Recent state-of-the-art solutions model two key issues, joint detection and spatial configuration refinement, together using convolutional neural networks. Our work mainly focuses on spatial configuration refinement by reducing variations of human poses statistically, which is motivated by the observation that the scattered distribution of the relative locations of joints (e.g., the left wrist is distributed nearly uniformly in a circular area around the left shoulder) makes the learning of convolutional spatial models hard. We present a two-stage normalization scheme, human body normalization and limb normalization, to make the distribution of the relative joint locations compact, resulting in easier learning of convolutional spatial models and more accurate pose estimation. In addition, our empirical results show that incorporating multi-scale supervision and multi-scale fusion into the joint detection network is beneficial. Experiment results demonstrate that our method consistently outperforms state-of-the-art methods on the benchmarks.
Ke Sun 0009, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Dong Liu 0002, Jingdong Wang 0001
ICCV4
2017 View Adaptive Recurrent Neural Networks for High Performance Human Action Recognition from Skeleton Data
abstract
Skeleton-based human action recognition has recently attracted increasing attention due to the popularity of 3D skeleton data. One main challenge lies in the large view variations in captured human actions. We propose a novel view adaptation scheme to automatically regulate observation viewpoints during the occurrence of an action. Rather than re-positioning the skeletons based on a human defined prior criterion, we design a view adaptive recurrent neural network (RNN) with LSTM architecture, which enables the network itself to adapt to the most suitable observation viewpoints from end to end. Extensive experiment analyses show that the proposed view adaptive RNN model strives to (1) transform the skeletons of various views to much more consistent viewpoints and (2) maintain the continuity of the action rather than transforming every frame to the same position with the same body orientation. Our model achieves significant improvement over the state-of-the-art approaches on three benchmark datasets.
Pengfei Zhang 0005, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Jianru Xue, Nanning Zheng 0001
ICCV4
2017 Personalized long-term prediction of cognitive function: Using sequential assessments to improve model performance
Chih-Lin Chi, Wenjun Zeng 0001, Wonsuk Oh, Soo Borson, Tatiana Lenskaia, Xinpeng Shen, Peter J. Tonellato
J. Biomed. Informatics2
2017 Adaptive Nonlocal Sparse Representation for Dual-Camera Compressive Hyperspectral Imaging
abstract
Leveraging the compressive sensing (CS) theory, coded aperture snapshot spectral imaging (CASSI) provides an efficient solution to recover 3D hyperspectral data from a 2D measurement. The dual-camera design of CASSI, by adding an uncoded panchromatic measurement, enhances the reconstruction fidelity while maintaining the snapshot advantage. In this paper, we propose an adaptive nonlocal sparse representation (ANSR) model to boost the performance of dual-camera compressive hyperspectral imaging (DCCHI). Specifically, the CS reconstruction problem is formulated as a 3D cube based sparse representation to make full use of the nonlocal similarity in both the spatial and spectral domains. Our key observation is that, the panchromatic image, besides playing the role of direct measurement, can be further exploited to help the nonlocal similarity estimation. Therefore, we design a joint similarity metric by adaptively combining the internal similarity within the reconstructed hyperspectral image and the external similarity within the panchromatic image. In this way, the fidelity of CS reconstruction is greatly enhanced. Both simulation and hardware experimental results show significant improvement of the proposed method over the state-of-the-art.
Lizhi Wang 0001, Zhiwei Xiong, Guangming Shi, Feng Wu 0001, Wenjun Zeng 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2017 Fully Reversible Privacy Region Protection for Cloud Video Surveillance
abstract
Privacy becomes one of the major concerns of cloud-based multimedia applications such as cloud video surveillance. Privacy protection of surveillance videos aims to protect privacy information without hampering normal processing tasks of the cloud. Privacy Region Protection only protects the privacy region while keeping the non-privacy region visually intact to facilitate processing in the cloud. However, full reversibility, i.e. the complete recovery of the original video which is critical to digital investigation and law enforcement has not been properly addressed in privacy region protection. In this paper, we introduce fully reversible privacy region protection into cloud video surveillance and propose a novel fully reversible privacy protection method for H.264/AVC compressed video. All the operations are performed in the compressed domain and avoid lossy re-encoding, so the original H.264/AVC compressed video can be fully recovered. To our best knowledge, the proposed scheme is the first fully reversible one for privacy region protection. Experimental results and performance comparison demonstrate the effectiveness and efficiency of the proposed approach.
Xiaojing Ma 0002, Laurence T. Yang, Yang Xiang 0001, Wenjun Zeng 0001, Deqing Zou, Hai Jin 0001
IEEE Trans. Cloud Comput.4
2017 Guest Editorial Special Issue on Visual Computing in the Cloud: Mobile Computing
abstract
Recent advances in mobile devices (e.g., smartphones and wearables) and wireless technologies are fueling a new wave of user demands for an improved user experience. Indeed, users are not only expecting ubiquitous network connections for traditional services (e.g., messaging and calling), but also demanding extensive access to a wealth of video contents and services. However, this growing demand is seriously hindered by the fact that the onboard resources with mobile devices are inherently limited and their growth rate falls behind that of their desktop counterparts. It follows that new solutions should be in order to resolve this fundamental tussle. Fortunately, the emerging cloud computing offers a natural solution to extend the desktop visual experience to mobile devices. It actually provides both computational and storage support for media-rich applications with both front-end and back-end functionalities.
Yonggang Wen 0001, Jacob Chakareski, Pascal Frossard, Di Wu 0001, Wenjun Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.5
2017 Progressive Pseudo-analog Transmission for Mobile Video Streaming
abstract
We propose a progressive pseudo-analog video transmission scheme that simultaneously handles SNR and bandwidth variations with graceful quality degradation for mobile video streaming. With the inherited SNR-adaptability from pseudo-analog transmission, the proposed progressive solution acquires bandwidth adaptability through an innovative scheduling algorithm with optimal power allocation. The basic idea is to aggressively transmit or retransmit important coefficients so that distortion is minimized at the receiver after each received packet. We derive the closed-form expression of reduced distortion for each packet under given transmission power and known channel conditions, and show that the optimal solution can be obtained with a water-filling algorithm. We also illustrate through analyses and simulations that a near-optimal solution can be found through approximation when only statistical channel information is available. Simulations show that our solution approaches the performance upper bound of pseudo-analog transmission in an additive white Gaussian noise channel and significantly outperforms existing pseudo-analog solutions in a fast Rayleigh fading channel. Trace-driven emulations are also carried out to demonstrate the advantage of the proposed solution over the state-of-the-art digital and pseudo-analog solutions under a real dramatically varying wireless environment.
Dongliang He, Cuiling Lan, Chong Luo 0001, Enhong Chen, Feng Wu 0001, Wenjun Zeng 0001
IEEE Trans. Multim.6
2016 Co-Occurrence Feature Learning for Skeleton Based Action Recognition Using Regularized Deep LSTM Networks
abstract
Skeleton based action recognition distinguishes human actions using the trajectories of skeleton joints, which provide a very good representation for describing actions. Considering that recurrent neural networks (RNNs) with Long Short-Term Memory (LSTM) can learn feature representations and model long-term temporal dependencies automatically, we propose an end-to-end fully connected deep LSTM network for skeleton based action recognition. Inspired by the observation that the co-occurrences of the joints intrinsically characterize human actions, we take the skeleton as the input at each time slot and introduce a novel regularization scheme to learn the co-occurrence features of skeleton joints. To train the deep LSTM network effectively, we propose a new dropout algorithm which simultaneously operates on the gates, cells, and output responses of the LSTM neurons. Experimental results on three human action recognition datasets consistently demonstrate the effectiveness of the proposed model.
Wentao Zhu 0001, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Yanghao Li, Li Shen 0005, Xiaohui Xie
AAAI4
2016 Online Human Action Detection Using Joint Classification-Regression Recurrent Neural Networks
Yanghao Li, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Chunfeng Yuan, Jiaying Liu 0001
ECCV (7)4
2016 A super-fast online face tracking system for video surveillance
abstract
In this paper, we propose a novel and practical system for robust online face tracking in surveillance videos. The proposed system has two contributions: 1) sustained high performance for long-term tracking even when faces come in and out of the view frequently, and 2) extremely low complexity which allows for real-time deployment on various platforms. These advantages are achieved by designing a regular update framework based on a state-of-the-art face detector and a new histogram-assisted KLT (HAKLT) tracker. Experimental results demonstrate a superior and super-fast (>100fps) practical face tracking system.
Xiaosong Lan, Zhiwei Xiong, Wei Zhang 0262, Shuxiao Li, Hongxing Chang, Wenjun Zeng 0001
ISCAS6
2016 Compressive hyperspectral imaging with complementary RGB measurements
abstract
Coded aperture snapshot spectral imaging (CASSI) has been demonstrated as a feasible solution to recover a 3D hyperspectral image by using a single 2D measurement. In this paper, we propose a new hybrid camera design for CASSI to capture high quality hyperspectral images while maintaining the snapshot advantage. Specifically, we employ a complementary RGB camera in conjunction with the CASSI system. The recorded RGB image can provide reliable spectral clue of the scene. By combining the coded hyperspectral information from the CASSI branch and the uncoded color information from the RGB branch, hyperspectral images can be reconstructed with high fidelity. Furthermore, by conducting demosaicing on the raw RGB image as a preprocessing procedure, even better performance can be achieved. Both theoretical analysis and simulation results show improved accuracy of the proposed method compared to the state-of-the-arts.
Lizhi Wang 0001, Zhiwei Xiong, Guangming Shi, Wenjun Zeng 0001, Feng Wu 0001
VCIP4
2016 Lossless Compression of JPEG Coded Photo Collections
abstract
The explosion of digital photos has posed a significant challenge to photo storage and transmission for both personal devices and cloud platforms. In this paper, we propose a novel lossless compression method to further reduce the size of a set of JPEG coded correlated images without any loss of information. The proposed method jointly removes inter/intra image redundancy in the feature, spatial, and frequency domains. For each collection, we first organize the images into a pseudo video by minimizing the global prediction cost in the feature domain. We then present a hybrid disparity compensation method to better exploit both the global and local correlations among the images in the spatial domain. Furthermore, the redundancy between each compensated signal and the corresponding target image is adaptively reduced in the frequency domain. Experimental results demonstrate the effectiveness of the proposed lossless compression method. Compared with the JPEG coded image collections, our method achieves average bit savings of more than 31%.
Hao Wu 0042, Xiaoyan Sun 0001, Jing-Yu Yang 0002, Wenjun Zeng 0001, Feng Wu 0001
IEEE Trans. Image Process.4
2015 High-speed hyperspectral video acquisition with a dual-camera architecture
abstract
We propose a novel dual-camera design to acquire 4D high-speed hyperspectral (HSHS) videos with high spatial and spectral resolution. Our work has two key technical contributions. First, we build a dual-camera system that simultaneously captures a panchromatic video at a high frame rate and a hyperspectral video at a low frame rate, which jointly provide reliable projections for the underlying HSHS video. Second, we exploit the panchromatic video to learn an over-complete 3D dictionary to represent each band-wise video sparsely, and a robust computational reconstruction is then employed to recover the HSHS video based on the joint videos and the self-learned dictionary. Experimental results demonstrate that, for the first time to our knowledge, the hyperspectral video frame rate reaches up to 100fps with decent quality, even when the incident light is not strong.
Lizhi Wang 0001, Zhiwei Xiong, Dahua Gao, Guangming Shi, Wenjun Zeng 0001, Feng Wu 0001
CVPR5
2015 Removing camera fingerprint to disguise photograph source
abstract
Sensor-based camera source identification (CSI) is believed to be an effective tool for linking a photograph to its source camera. In this paper, we propose a photo response non-uniformity (PRNU) removing attack on CSI. The PRNU fingerprint is a kind of multiplicative noise, and is believed to be difficult to be removed from an image. In this work, both the pattern and magnitude of the PRNU fingerprint of an unaltered image are estimated first, then it is subtracted from the target image. Theoretical analysis and experimental results show that the PRNU fingerprint can be removed using only common processing techniques without introducing visual artefact, therefore CSI can be easily defeated.
Hui Zeng 0002, Xiangui Kang, Wenjun Zeng 0001
ICIP4
2015 Compound image compression using lossless and lossy LZMA in HEVC
abstract
We present a compound image compression scheme based on the dictionary-based Lempel-Ziv-Markov chain algorithm (LZMA), under the framework of High Efficiency Video Coding (HEVC). Through matching strings from the sliding window dictionary, LZMA exploits the characteristics of the repeated patterns over the text and graphics regions of compound images, and represents them compactly. To obtain high compression efficiency even for noisy text and graphics contents, we have modified LZMA to support both lossless and lossy compression. We develop and treat it as a new intramode of HEVC. Experimental results show that the proposed scheme achieves significant coding gains for compound image compression. Thanks to the introduction of the lossy LZMA, the compression performance for noisy compound images is improved for more than 5dB in terms of PSNR in comparison with the lossless LZMA scheme.
Cuiling Lan, Jizheng Xu, Wenjun Zeng 0001, Feng Wu 0001
ICME3
2015 Swift: A Hybrid Digital-Analog Scheme for Low-Delay Transmission of Mobile Stereo Video
abstract
Efficient and robust wireless stereo video delivery is an enabling technology for various mobile 3D applications. Existing digital solutions have high source coding efficiency but are not robust to channel variations, while analog solutions have the opposite characteristics. In this paper, we design a novel hybrid digital-analog (HDA) solution to embrace the advantages of both solutions and avoid their drawbacks. Basically, in each pair of stereo frames, one frame is digitally encoded to ensure basic quality and the other is analogly processed to opportunistically utilize good channels for better quality. To improve the system efficiency, we design a zigzag coding structure such that both intra-view and inter-view correlations can be explored through prediction in the frames to be analogly coded. A reference selection mechanism is proposed to further improve the coding efficiency. In addition, we address the problem of optimal power and bandwidth allocation between digital and analog streams. We implement a system, named Swift, and perform extensive trace-driven evaluations based on a software-defined radio platform. We show that Swift outperforms an omniscient digital scheme under the same bandwidth and power constraints, or can have around 2x power saving in order to achieve comparable performance. Subjective quality assessment evidences that Swift provides significantly better visual quality than a straightforward HDA extension of SoftCast.
Dongliang He, Chong Luo 0001, Feng Wu 0001, Wenjun Zeng 0001
MSWiM4
2015 Progressive pseudo-analog transmission for mobile video live streaming
abstract
Mobile video live streaming is facing great challenges in offering high quality of experience (QoE) under varying channel conditions. In this paper, we propose a progressive pseudo-analog transmission scheme in which the received video quality gracefully adapts to both SNR and bandwidth variations. Building upon the emerging pseudoanalog video transmission, the proposed scheme further adopts a greedy approach to improve the received video quality with each allocated bandwidth share. The optimal scheduling and power allocation are derived under the mean squared error (MSE) criterion. Testbed evaluations show that the proposed scheme outperforms the state-of-the-art digital and analog transmission schemes by a notable margin.
Cuiling Lan, Dongliang He, Chong Luo 0001, Feng Wu 0001, Wenjun Zeng 0001
VCIP5
2015 Graph-based video fingerprinting using double optimal projection
Xiushan Nie, Wenjun Zeng 0001
J. Vis. Commun. Image Represent.4
2015 Structure-Preserving Hybrid Digital-Analog Video Delivery in Wireless Networks
abstract
Hybrid digital-analog (HDA) transmission has gained increasing attention recently in the context of wireless video delivery , for its ability to simultaneously achieve high transmission efficiency and smooth quality adaptation. However, previous systems are optimized solely based on the mean squared error criterion without taking the perceptual video quality into consideration. In this work, we propose a structure-preserving HDA video delivery system, named SharpCast, to improve both the objective and subjective visual quality. SharpCast decomposes a video into a content part and structure part. The latter is important to the human perception and therefore is protected with a robust digital transmission scheme. Then, the energy-intensive part in the content information is extracted and transmitted in digital for energy efficiency while the residual is transmitted in analog to achieve the desired smooth adaptation. We formulate the resource (power and bandwidth) allocation problem in SharpCast and solve the problem with a greedy strategy. Evaluations over nine standard 720p video sequences show that the proposed SharpCast system outperforms the state-of-the-art digital, analog, and HDA schemes by a notable margin in both peak signal-to-noise ratio (PSNR) and structural similarity (SSIM).
Dongliang He, Chong Luo 0001, Cuiling Lan, Feng Wu 0001, Wenjun Zeng 0001
IEEE Trans. Multim.5
2014 Improving distributed video coding by exploiting context-adaptive modeling
abstract
The statistical model of the bits to be encoded is crucial for the coding performance of distributed video coding (DVC). In this paper, a bit-level context-adaptive correlation model is proposed to exploit high-order statistical correlation for better channel coding performance, which consequently improves the video coding efficiency. In the proposed scheme, the wavelet domain DVC is considered and the coefficients are coded in a bit-plane fashion. The context for each bit to be coded is first formed. Then the probability distribution of each bit is estimated by using previously available data with the same context. For magnitude coding, the significant state of the following elements are included in the context, (1) the side information, (2) the local neighborhood, (3) the parent coefficients (if applicable). The condition of side information is considered as well. For sign coding, the context consists of the sign and the quality of the side information. The proposed model is implemented within a recently proposed DVC framework with decoderside multi-resolution motion refinement (MRMR). Experimental results show the effectiveness of the proposed scheme with significant coding gain over the original MRMR based DVC system, especially for videos with high motion intensity and for lower bit rates.
Linbo Qing, Wenjun Zeng 0001
ICME2
2014 Delta interpolation for upsampling imaging solutions
abstract
The advances in imaging optics, sensors, and camera processing capabilities in today's digital cameras have resulted in photos with over 10 megapixels. The high resolution may pose a challenge to the computational complexity of many imaging solutions, such as color enhancement. To reduce the complexity of these imaging solutions and enable real time operations, it has been proposed to employ an upsampling imaging solution approach, such as the celebrated joint bilateral upsampling algorithm. In this paper we present a delta interpolation scheme, which reconstructs a high resolution solution by explicitly employing knowledge from both the high and low resolution images and operating on a delta image. This method achieves substantially enhanced visual quality on color enhancement, and its performance is validated by experimental results.
Wenjun Zeng 0001
ICME2
2014 Compressive sensing based secure multiparty privacy preserving framework for collaborative data-mining and signal processing
abstract
In many real-world applications, multiple parties who provide data need to collaboratively perform certain data-mining and signal processing tasks. Security and privacy protection is a critical issue in such application scenarios. In this paper, we propose a compressive sensing (CS) based privacy preserving framework for collaborative data-mining and signal processing using secure multiparty computation (MPC) in which the data-mining and the signal processing are performed in the compressive sensing domain. In our framework, the MPC protocols are used only for compressive sensing transformation and reconstruction while the data-mining/signal processing tasks are de-coupled from MPC operations. So our framework enjoys a great deal of flexibility and scalability when compared to the prior works because the decoupling allows CS transformed data to be reused and many data processing algorithms can be applied in such CS domain. Our framework also enables privacy preserving data storage in the cloud at the same time. Additionally, we develop a MPC based orthogonal matching pursuit algorithm and its corresponding MPC protocol for the CS reconstruction. Our analysis and experimental results demonstrate that the proposed framework is effective in enabling efficient privacy preserving data-mining/signal processing and storage.
Qia Wang 0002, Wenjun Zeng 0001
ICME2
2014 Structural similarity-based video fingerprinting for video copy detection
abstract
The authors propose a video fingerprinting method based on structural similarity and a graph model. Structural similarity‐based fingerprint generation and double‐layer matching are the two main contributions. The video is mapped to a graph with frames as its vertices, and structural similarity is proposed to compute the weights of the edges. Then, the fingerprint consisting of a match tag (a coarse fingerprint) and a fine fingerprint is generated by this graph. The match tag is generated by an independent set of this graph, and the fine fingerprint is generated by the weight matrix of the graph based on the two‐block‐dimensional discrete cosine transform. During the matching, the video can be matched at the first‐layer using the match tag to obtain a candidate set, whereas the second‐layer matching is performed in this candidate set using the fine fingerprint to find a final match. The proposed video fingerprinting method is shown to be resistant to geometric attacks on frames and impairment of transmission channels.
Xiushan Nie, Wenjun Zeng 0001, Jiande Sun 0001
IET Image Process.2
2014 Secure and robust image hashing via compressive sensing
Wenjun Zeng 0001
Multim. Tools Appl.2
2014 A Compressive Sensing Based Secure Watermark Detection and Privacy Preserving Storage Framework
abstract
Privacy is a critical issue when the data owners outsource data storage or processing to a third party computing service, such as the cloud. In this paper, we identify a cloud computing application scenario that requires simultaneously performing secure watermark detection and privacy preserving multimedia data storage. We then propose a compressive sensing (CS)-based framework using secure multiparty computation (MPC) protocols to address such a requirement. In our framework, the multimedia data and secret watermark pattern are presented to the cloud for secure watermark detection in a CS domain to protect the privacy. During CS transformation, the privacy of the CS matrix and the watermark pattern is protected by the MPC protocols under the semi-honest security model. We derive the expected watermark detection performance in the CS domain, given the target image, watermark pattern, and the size of the CS matrix (but without the CS matrix itself). The correctness of the derived performance has been validated by our experiments. Our theoretical analysis and experimental results show that secure watermark detection in the CS domain is feasible. Our framework can also be extended to other collaborative secure signal processing and data-mining applications in the cloud.
Qia Wang 0002, Wenjun Zeng 0001
IEEE Trans. Image Process.2
2013 SDNAN: Software-defined networking in ad hoc networks of smartphones
abstract
In this paper, SDNAN, a first attempt to implement software-defined networking (SDN) over a wireless ad hoc network of smartphones, is presented. Its modular ad hoc network management structure can be easily modified and extended. Its abstractions and interfaces allow components to communicate without knowing how other components work. Third-party applications can use the interfaces to access the ad hoc network, significantly reducing development time and program complexity. A prototype system has been implemented on Android smartphones over Wi-Fi and achieved good preliminary results.
Paul Baskett, Yi Shang, Wenjun Zeng 0001, Brandon Guttersohn
CCNC3
2013 Mainstream media vs. social media for trending topic prediction - an experimental study
abstract
In the recent years, we have witnessed social networks blossom. Social networking reshaped worldwide communication significantly increased the speed of news spread, and connected the world stronger than ever. Although social networking has been such a revolutionary invention for the society, and many researchers have turned towards social media to explore trending topics, mainstream media still remains as the origin of the majority of the news discussed in social networking sites. Social stream mining to make video recommendations based on the trending topics has been an active direction in the research community. Understanding the trending topics and its impact on video sharing sites is very interesting for network traffic engineers. Quality of service can be significantly improved if we can predict what kind of video content will generate large traffic. The focus of this paper is to study which type of media, mainstream or social, can contribute better towards identifying trending topics. We present the experimental study of the story development process in mainstream and social media based on the real-world data. The study helps us properly identify which media source is more appropriate for the video recommendation and network traffic prediction systems. Through our findings, we discovered mainstream media could significantly improve the trend detection.
Alex Lobzhanidze, Wenjun Zeng 0001, Paige Gentry, Angelique Taylor
CCNC2
2013 A hybrid approach for tree classification in airborne LIDAR data
abstract
In this paper we propose a hybrid approach for tree classification in airborne LIDAR (Light Detection and Ranging) data by integrating the point based supervised classification with region-based unsupervised clustering method. Furthermore we propose a novel 3D robust statistics-based shape feature that can overcome the limitations of existing methods in separating building boundary points from tree points. Experimental results show the new algorithm is very effective and can achieve very high accuracy.
Wenjun Zeng 0001, Ye Duan
ICASSP2
2013 Geometry based airborne LIDAR data compression
abstract
Airborne LIDAR data often consumes hundreds of gigabytes. Existing LIDAR data compression schemes can compress the file to 5%-23% of the original size. Even after compression, the compressed data size is still in the order of gigabyte, which makes it impractical for many applications. This paper proposes a novel geometry based compression scheme. It first introduces a LIDAR classification method that accurately classifies airborne LIDAR data into tree and non-tree points; different geometry based compression schemes are then applied for different types of data. The proposed method can not only compress LIDAR data significantly, but also extract useful semantic information from the data. Experimental results show that the new approach achieves very high compression ratio, making applications that were not practical before feasible.
Wenjun Zeng 0001, Ye Duan
ICME2
2013 Proactive caching of online video by mining mainstream media
abstract
Online video sharing is becoming more and more popular and demands a large amount of network resources. Caching content has been an effective method for providing better quality of service for video applications. In the age of social network, information sharing and consumption has significantly changed. More and more viewers of online videos are influenced by the trends in social and mainstream media. Given the rich information about the trends in the media, we propose a cross-platform, proactive video caching scheme in this paper. We employ a combination of the topic modeling tool Latent Dirichlet Allocation (LDA) and frequent pattern mining algorithm Apriori to effectively detect and connect trending topics from mainstream media to online videos. Furthermore, we design a reputation-based video-ranking algorithm to select candidates for caching at geographically relevant proxy nodes to reduce the delay and improve the overall traffic in the network. Simulation results show that our cross-domain proactive caching method performs significantly better than classical caching methods that are based on the historical popularity.
Alex Lobzhanidze, Wenjun Zeng 0001
ICME2
2013 Towards Cross-Domain Learning for Social Video Popularity Prediction
abstract
Previous research on online media popularity prediction concluded that the rise in popularity of online videos maintains a conventional logarithmic distribution. However, recent studies have shown that a significant portion of online videos exhibit bursty/sudden rise in popularity, which cannot be accounted for by video domain features alone. In this paper, we propose a novel transfer learning framework that utilizes knowledge from social streams (e.g., Twitter) to grasp sudden popularity bursts in online content. We develop a transfer learning algorithm that can learn topics from social streams allowing us to model the social prominence of video content and improve popularity predictions in the video domain. Our transfer learning framework has the ability to scale with incoming stream of tweets, harnessing physical world event information in real-time. Using data comprising of 10.2 million tweets and 3.5 million YouTube videos, we show that social prominence of the video topic (context) is responsible for the sudden rise in its popularity where social trends have a ripple effect as they spread from the Twitter domain to the video domain. We envision that our cross-domain popularity prediction model will be substantially useful for various media applications that could not be previously solved by traditional multimedia techniques alone.
Suman Deb Roy, Tao Mei 0001, Wenjun Zeng 0001, Shipeng Li 0001
IEEE Trans. Multim.3
2013 Introduction to the special section of best papers of ACM multimedia 2012
abstract
No abstract available.
Ioannis Kompatsiaris, Wenjun Zeng 0001, Gang Hua 0001, Liangliang Cao
ACM Trans. Multim. Comput. Commun. Appl.2
2012 Nest: Networked smartphones for target localization
abstract
This paper presents Nest, a novel system using wirelessly connected smartphones to localize remote targets based on sound and image inputs. The system has four major components: image-based localization, acoustics-based localization, wireless ad-hoc networking, and middle-ware services for time synchronization and secure communication. Single-image, two-image, and TDOA-acoustics based methods have been developed and a prototype system has been implemented on Google Nexus One smartphones running Android. Experimental results show that the localization accuracies of the single-image-based and two-image-based method are around 94.6% and up to 92.1%, respectively. The localization errors of the acoustics-based method are within 80 centimeters.
Yi Shang, Wenjun Zeng 0001, K. C. Ho 0001, Qia Wang 0002, Yue Wang 0022, Tiancheng Zhuang, Alex Lobzhanidze, Liyang Rui
CCNC2
2012 A Computational Cognitive Model for Semantic Sub-Network Extraction from Natural Language Queries
Suman Deb Roy, Wenjun Zeng 0001
COLING2
2012 Scalable Lossy Compression for Pixel-Value Encrypted Images
abstract
Compression of encrypted data draws much attention in recent years due to the security concerns in a service oriented environment such as cloud computing. We propose a scalable lossy compression scheme for images having their pixel value encrypted with a standard stream cipher. The encrypted data are simply compressed by transmitting a uniformly sub sampled portion of the encrypted data and some bit-planes of another uniformly sub sampled portion of the encrypted data. With a proposed content adaptive interpolation prediction method with side information, at the receiver side, a decoder performs content adaptive interpolation based on the decrypted partial information, where the received bit-plane information serves as the side information that reflects the image edge information, making the image reconstruction more precise. When more bit-planes are transmitted, higher quality of the decompressed image can be achieved. The experimental results show that our proposed scheme achieves much better performance than the existing lossy compression scheme for pixel value encrypted images, and also similar performance as the state-of-the-art lossy compression for pixel permutation based encrypted images. In addition, our proposed scheme has the following advantages: at the decoder side, no computationally intensive iteration and no additional public orthogonal matrix is needed. It works well for both smooth and texture-rich images.
Xiangui Kang, Xianyu Xu, Anjie Peng, Wenjun Zeng 0001
DCC4
2012 Empowering Cross-Domain Internet Media with Real-Time Topic Learning from Social Streams
abstract
This paper aims to connect social media from disparate sources on the Internet by building a common topic space in-between, using which cross domain media recommendations can be realized on the web. The topic space is built and updated in real time by extending the Latent Dirichlet Allocation (LDA) model to cater to streaming online data. Our topical model, named Online Streaming LDA (OSLDA), is able to extract, learn, populate, and update the topic space in real time, scaling with streaming tweets. Based on the proposed topic space learned in real time, we present media recommendation applications that cannot be achieved by conventional media analysis techniques: (1) tweet enrichment by recommending related videos, and (2) popular video recommendation for featuring socially trending topical videos. We conduct experiments over a collection of 3.6 million tweets and 1.2 million click-through data from a video search engine. Our results show that the learned topic model plays a natural role connecting cross-domain social media, leading to a better user experience consuming social media.
Suman Deb Roy, Tao Mei 0001, Wenjun Zeng 0001, Shipeng Li 0001
ICME3
2012 Video Based Real-World Remote Target Tracking on Smartphones
abstract
Smartphone's increasing computing power and built-in sensors such as digital camera and GPS have created a new and open platform for developing compelling mobile multimedia tools and systems. In this paper, a novel, accurate and computationally efficient video-based remote target positioning and tracking system on Android smart phones is proposed and presented, with the assumption that the physical size of the remote target is known. Smartphone's user-friendly interface is exploited in our system to facilitate accurate and efficient video object tracking. With the help of some initial user input, our algorithm generates tight bounding boxes around the moving targets across video frames with high accuracy and low complexity. Unlike most existing video tracking systems that are limited to tracking on the image plane, we re-project the accurately detected object locations on the image plane to real world positions to estimate the remote target's moving trajectory. In addition, the Extended Kalman Filtering model is applied to provide a smooth remote target trajectory and velocity estimation. Our experimental results show that the proposed mobile multimedia system works accurately and efficiently for tracking remote moving targets.
Qia Wang 0002, Alex Lobzhanidze, Hyun Ik Jang, Wenjun Zeng 0001, Yi Shang
ICME4
2012 Expert Talk for Time Machine Session: High Order Entropy Coding - From Conventional Video Coding to Distributed Video Coding
abstract
High order entropy coding has proved to be critical for improving the coding efficiency of conventional image/video coding. What role will it play in distributed video coding? This talk intends to shed some light on that.
Wenjun Zeng 0001
ICME1
2012 Kinect-like depth denoising
abstract
Accuracy and stability of Kinect-like depth data is limited by its generating principle. In order to serve further applications with high quality depth, the preprocessing on depth data is essential. In this paper, we analyze the characteristics of the Kinect-like depth data by examing its generation principle and propose a spatial-temporal denoising algorithm taking into account its special properties. Both the intra-frame spatial correlation and the inter-frame temporal correlation are exploited to fill the depth hole and suppress the depth noise. Moreover, a divisive normalization approach is proposed to assist the noise filtering process. The 3D rendering results of the processed depth demonstrates that the lost depth is recovered in some hole regions and the noise is suppressed with depth features preserved.
Jingjing Fu, Shiqi Wang 0001, Yan Lu 0001, Shipeng Li 0001, Wenjun Zeng 0001
ISCAS5
2012 SocialTransfer: cross-domain transfer learning from social streams for media applications
abstract
The usage and applications of social media have become pervasive. This has enabled an innovative paradigm to solve multimedia problems (e.g., recommendation and popularity prediction), which are otherwise hard to address purely by traditional approaches. In this paper, we investigate how to build a mutual connection among the disparate social media on the Internet, using which cross-domain media recommendation can be realized. We accomplish this goal through SocialTransfer---a novel cross-domain real-time transfer learning framework. While existing transfer learning methods do not address how to utilize the real time social streams, our proposed SocialTransfer is able to effectively learn from social streams to help multimedia applications, assuming an intermediate topic space can be built across domains. It is characterized by two key components: 1) a topic space learned in real time from social streams via Online Streaming Latent Dirichlet Allocation (OSLDA), and 2) a real-time cross-domain graph spectra analysis based transfer learning method that seamlessly incorporates learned topic models from social streams into the transfer learning framework. We present as use cases of \emph{SocialTransfer} two video recommendation applications that otherwise can hardly be achieved by conventional media analysis techniques: 1) socialized query suggestion for video search, and 2) socialized video recommendation that features socially trending topical videos. We conduct experiments on a real-world large-scale dataset, including 10.2 million tweets and 5.7 million YouTube videos and show that \emph{SocialTransfer} outperforms traditional learners significantly, and plays a natural and interoperable connection across video and social domains, leading to a wide variety of cross-domain applications.
Suman Deb Roy, Tao Mei 0001, Wenjun Zeng 0001, Shipeng Li 0001
ACM Multimedia3
2012 Multiple description coded video streaming in peer-to-peer networks
Yuanyuan Xu 0001, Ce Zhu, Wenjun Zeng 0001, Xue Jun Li
Signal Process. Image Commun.3
2011 Network Assisted Media Streaming in Multi-Hop Wireless Networks
abstract
Delivery of high-quality streaming services over multi-hop wireless mesh networks (WMNs) is a challenging research problem because of quality fluctuation and interference of wireless links in WMNs, as well as strict throughput, delay, and reliability requirements of streaming applications. In this paper, we propose a Network Assisted Peer-to-Peer (NAP2P) system for file-based media streaming services such as video-on-demand in WMNs. In NAP2P, mesh routers dynamically cache content and form a P2P network with end user devices. This architecture enables several efficient and scalable communication mechanisms, for example, automatic caching and multi-source multi-path streaming for optimizing system performance. We address the design issues in NAP2P. Especially we design a multi-source multi-path routing mechanism to meet the QoS requirements of streaming sessions. A mathematical formulation for such source selection and routing optimization problem is presented, which determines the optimal content sources and streaming paths for multiple concurrent flows in the presence of wireless interference and subject to flow QoS constraints. By leveraging the unique architecture of NAP2P, we investigate the performance gain by applying mesh cache routers as network peers in P2P networks, as well as the benefits of allowing end user devices to provide content to their peers.
Yingnan Zhu, Hang Liu 0003, Yang Guo 0001, Wenjun Zeng 0001
ICCCN4
2011 A multi-layer key stream based approach for joint encryption and compression of H.264 video
abstract
Encryption of multimedia video data has been extensively studied. However, a pipelined method of compression followed by encryption is not always feasible and springs new challenges, namely the effect of visual degradation and the computational cost of the encryption process. Recently, joint encryption and compression techniques, which aim to encrypt the video as it is being compressed by the encoder, have gained popularity. This work aims to develop a lightweight yet effective multi-layer key stream based technique to encrypt video data by performing joint encryption and compression. Our joint encryptor-encoder performs almost as fast as a normal H.264 compression-only encoder. The multi-layer key stream we use for encryption is extracted from the underlying video structure itself, thus thwarting (piracy) key generators. It also facilitates efficient handing of packet loss in real-time transmission. Results show our method produces almost complete obscuration of the video, a consistent degradation for diverse scenes, and format compliant encrypted H.264 video.
Suman Deb Roy, Hong Heather Yu, Wenjun Zeng 0001
ICME4
2011 Positionit: an image-based remote target localization system on smartphones
abstract
We present PositionIt, a novel, low-cost, image-based remote target localization system using commodity smartphones. By leveraging smartphones' built-in sensors such as camera, digital compass, GPS, etc., the system can estimate the distance/position of remote targets. The system also takes advantage of the user-friendly interface of the smartphones to facilitate low complexity implementation (e.g., to locate target of interest and restrict the search area for matching).The feasibility of both single-image and two-image based systems are demonstrated.
Qia Wang 0002, Alex Lobzhanidze, Suman Deb Roy, Wenjun Zeng 0001, Yi Shang
ACM Multimedia4
2010 Cross-Site Request Forgery: Attack and Defense
abstract
Cross Site Request Forgery (CSRF) has emerged as a potent threat to Web 2.0 applications. Because of the stateless nature of the HTTP protocol, a malicious Website can force the user's browser to send unauthorized requests to a trusted site. This demo provides hands on exposure of the various ways in which some popular Web applications are exploited using CSRF, in addition to demonstrating techniques by which CSRF signatures can be detected and attacks effectively resisted even before initiation. The user needs only to install a simple extension to get notified about potential CSRF vulnerabilities. Because validating the Referer Header is a common CSRF prevention method, a novel solution to the Referer Privacy issue will also be demonstrated.
Tatiana Alexenko, Mark Jenne, Suman Deb Roy, Wenjun Zeng 0001
CCNC4
2010 iImage: An Image Based Information Retrieval Application for the iPhone
abstract
Image processing is a powerful technology that can be used to analyze an image for many useful purposes. However, software like this is often out of the typical user's reach, ilmage directly confronts this problem. The ilmage application takes the sophisticated technologies of image analysis and identification based on MPEG-7 image feature tools and makes them readily available on the go. With the use of the iPhone Source Development Kit, ilmage has become an application that is available on the world's most popular mobile device. The user simply takes a photo using the iPhone's built in camera, and in seconds appear similar images with relevant information from within our database.
Stephen Bock, S. Newsome, Wenjun Zeng 0001
CCNC4
2010 Design Issues of the prTorrent File Sharing Protocol
abstract
The absence of piece rarity in the BitTorrent Unchoking Algorithm exposes it to various exploitations and hinders optimized performances. prTorrent is a file sharing protocol based on BitTorrent which considers rarity of pieces swapped during unchoking for optimized performance within the swarm. This paper discusses the philosophies, issues and structure behind development of the prTorrent p2p protocol from scratch. The protocol was developed in Java to promote portability in simulation as well as emulation. We hope this work is a guide to students and developers planning to build a p2p simulator from the very base and alleviate design differences between simulation and emulation.
Suman Deb Roy, Tyler Knierim, Shanieka Churchman, Wenjun Zeng 0001
CCNC4
2010 Overview of the OMA Secure Content IDentification Mechanism
abstract
With the wide availability of digital content and proliferation of Web2.0 applications, supporting secure and efficient identification of digital content becomes a more and more important issue. In early 2010, the Open Mobile Alliance (OMA) ratified a new standard known as Secure Content IDentification Mechanism (SCIDM) to facilitate the management of digital content. As the main contributors to the SCIDM standard, the authors provide an overview of this standard as well as a use case in this paper.
Wenjun Zeng 0001, Qia Wang 0002, Stephen Bock
ICME1
2010 Motion Refinement Based Progressive Side-Information Estimation for Wyner-Ziv Video Coding
abstract
During the past ten years, Wyner-Ziv video coding (WZVC) has gained a lot of research interests because of its unique characteristics of "simple encoding, complex de coding." However, the performance gap between WZVC and conventional video coding has never been closed to the point promised by the information theory. In this paper, we illustrate the chicken-and-egg dilemma encountered in WZVC: high efficiency WZVC requires good estimation of side information (SI); however, good SI estimation is not possible for the decoder without access to the decoded current frame. To resolve such a dilemma, we present and advocate a framework that explores an important concept of decoder-side progressive-learning. More specifically, a decoder-side multi-resolution motion refinement (MRMR) scheme is proposed, where the decoder is able to learn from the already-decoded lower-resolution data to refine the motion estimation (ME), which in turn greatly improves the SI quality as well as the coding efficiency for the higher resolution data. Theoretical analysis shows that at high rates, decoder-side MRMR outperforms motion extrapolation by as much as 5 dB, while falling behind conventional encoder-side inter-frame ME by only about 1.5 dB. In addition, since decoder-side ME does not suffer from the bit-rate overhead in transmitting the motion information, further performance gain can be achieved for decoder-side MRMR by incorporating fractional-pel motion search, block matching with smaller block sizes, and multiple hypothesis prediction. We also present a practical WZVC implementation with MRMR, which shows comparable coding performance as H.264 at very high bit rates.
Wei Liu 0025, Lina Dong, Wenjun Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.3
2010 Efficient general print-scanning resilient data hiding based on uniform log-polar mapping
abstract
This paper proposes an efficient, blind, and robust data hiding scheme which is resilient to both geometric distortion and the general print-scan process, based on a near uniform log-polar mapping (ULPM). In contrast to performing inverse log-polar mapping (a mapping from the log-polar system to the Cartesian system) to the watermark signal or its index as done in the prior works, we apply ULPM to the frequency index (u,v) in the Cartesian system to obtain the discrete log-polar coordinate (l1,l2), then embed one watermark bitw(l1,l2) in the corresponding discrete Fourier transform coefficientc(u,v). This mapping of index from the Cartesian system to the log-polar system but embedding the corresponding watermark directly in the Cartesian domain not only completely removes the interpolation distortion and the interference distortion introduced to the watermark signal as observed in some prior works, but also largely expands the cardinality of watermark in the log-polar mapping domain. Both theoretical analysis and experimental results show that the proposed watermarking scheme achieves excellent robustness to geometric distortion, normal signal processing, and the general print-scan process. Compared to existing watermarking schemes, our algorithm offers significant improvement in terms of robustness against general print-scan, receiver operating characteristic (ROC) performance, and efficiency of blind resynchronization.
Xiangui Kang, Jiwu Huang, Wenjun Zeng 0001
IEEE Trans. Inf. Forensics Secur.3
2010 Efficient Compression of Encrypted Grayscale Images
abstract
Lossless compression of encrypted sources can be achieved through Slepian-Wolf coding. For encrypted real-world sources, such as images, the key to improve the compression efficiency is how the source dependency is exploited. Approaches in the literature that make use of Markov properties in the Slepian-Wolf decoder do not work well for grayscale images. In this correspondence, we propose a resolution progressive compression scheme which compresses an encrypted image progressively in resolution, such that the decoder can observe a low-resolution version of the image, study local statistics based on it, and use the statistics to decode the next resolution level. Good performance is observed both theoretically and experimentally.
Wei Liu 0025, Wenjun Zeng 0001, Lina Dong, Qiuming Yao
IEEE Trans. Image Process.2
2009 High Performance Adaptive Video Services Based on Bitstream Switching for IPTV Systems
abstract
As the next generation network gains increasing attention, IPTV has been recognized as an efficient way to provide key IP multimedia services. Compared to conventional TV services, IPTV is sensitive to the IP network characteristics such as packet loss, delay and jittering. In addition, the potential long delay in channel switching is another significant challenge for IPTV systems. In this paper, to improve the QoS (quality of service) and QoE (quality of experience), we propose a framework which includes (1) introducing server-side bit-rate adaptation to deal with bandwidth variation; (2) multicasting low bit-rate videos for preview of multi-channel programs and Picture-in-Picture (PiP); and (3) using a two-step P-frame switching mechanism to reduce the delay in channel switching. Our results from a real-world test-bed show that the proposed framework greatly improves the perceived video quality and user experience at the client side.
Yingnan Zhu, Wei Liu 0025, Lina Dong, Wenjun Zeng 0001, Hong Heather Yu
CCNC4
2009 Estimating side-information for Wyner-Ziv video coding using resolution-progressive decoding and extensive motion exploration
abstract
In Wyner-Ziv video coding (WZVC), the quality of side information (SI) has a critical impact on the coding efficiency. Most existing WZVC schemes generate SI using decoder-side motion estimation (ME), where the current frame is unavailable and the ME accuracy is greatly impaired. In this paper, without introducing additional bitrate overhead, we incorporate fractional-pel motion search, reduced block sizes and multiple hypothesis prediction into our previously proposed decoder-side multi-resolution motion refinement framework, where the current frame is progressively decoded, and based on which the decoder iteratively refines the motion to improve its accuracy. Theoretical analysis shows significant gain of the combination of these advanced techniques. A practical SI estimator is implemented and provides prediction performance comparable to H.264/AVC.
Wei Liu 0025, Lina Dong, Wenjun Zeng 0001
ICASSP3
2009 Multi-resolution based hybrid spatiotemporal compression of encrypted videos
abstract
Compression of encrypted data can be viewed as a special case of distributed source coding and can be achieved by applying Slepian-Wolf coding. However, how to compress the encrypted video efficiently remains a challenging problem especially for those videos with irregular high motion. This paper proposes a novel multi-resolution based approach which makes it possible not only to effectively derive the temporal side information from previous frames, but also to generate the spatial side information by having partial access to the current frame. The spatial and temporal side information can then be integrated adaptively to facilitate the compression. Simulation results show that the proposed scheme significantly outperforms other schemes, especially for those video clips with irregular high motion.
Qiuming Yao, Wenjun Zeng 0001, Wei Liu 0025
ICASSP2
2009 Challenges and opportunities in supporting video streaming over infrastructure wireless mesh networks
abstract
Multi-hop wireless mesh networks (WMNs) are emerging as a promising technology with applications in last mile Internet access, public safety, disaster response, battlefield, etc. The unique characteristics of WMNs such as limited and dynamic bandwidth, error prone nature, interferences, and user mobility, however, pose a number of significant challenges for high quality video streaming. In this paper, we discuss the major challenges, review some general approaches, and present a Unified Cache and Peer-to-Peer (UNICAP) framework we developed recently to support high quality streaming over WMNs. Issues addressed in UNICAP include the architecture, cache server/peer/route selection with cross-layer optimization, and admission control. Some future research directions are also discussed.
Yingnan Zhu, Hang Liu 0003, Yang Guo 0001, Wenjun Zeng 0001
ICME4
2009 Throughput and Delay Analysis of the IEEE 802.15.3 CSMA/CA Mechanism
abstract
Unlike in IEEE 802.11, the CSMA/CA traffic conditions in IEEE 802.15.3 are typically unsaturated. This paper presents an extended analytical model based on Bianchi's model in IEEE 802.11, by taking into account the device suspending events, unsaturated traffic conditions, as well as the effects of error-prone channels. Based on this model we re-derive a closed form expression of the average service time. The accuracy of the model is validated through extensive simulations. The analysis is also instructional for IEEE 802.11 networks under limited load.
Wenjun Zeng 0001
MASS2
2009 prTorrent: On Establishment of Piece Rarity in the BitTorrent Unchoking Algorithm
abstract
BitTorrent is an extensively adopted p2p Content Distribution System on the Internet. In spite of its proincentive approach and ease of implementation, recent research has empirically shown BitTorrent to be vulnerable to strategic manipulation by its constituent peers in a swarm. Moreover, Honest Piece Revelation and Free-Riding is becoming an increasing concern. Our findings indicate that till date, it is the orthogonal treatment of piece rarity and unchoking, that has encouraged strategic manipulation enabling unfair maximization of incentives in p2p systems. In this paper, we propose that solution to such concerns lies in unifying a Piece Rarity factor with the BitTorrent Unchoking Algorithm. We also discuss a new Discount Parameter attack that compromises most Tit-for-Tat mechanisms. Our analysis demonstrates how underreporting, as a Piece Revelation Strategy in auction based choking algorithms, could result in Starvation. prTorrent, a novel approach based on BitTorrent shows how strategic formulation of the Piece Rarity parameter can optimize incentives in a swarm, and help its constituent peers in achieving the equilibrium facilitating truly co-operative behavior.
Suman Deb Roy, Wenjun Zeng 0001
Peer-to-Peer Computing2
2009 Non-ambiguity of blind watermarking: a revisit with analytical resolution
Xiangui Kang, Jiwu Huang, Wenjun Zeng 0001, Yun Q. Shi 0001
Sci. China Ser. F Inf. Sci.3
2009 Path-Diversity P2P Overlay Retransmission for Reliable IP-Multicast
abstract
IP-multicast is a bandwidth efficient transmission mechanism for group communications. Reliability in IP-multicast, however, poses a set of significant challenges. To address the reliability and scalability issues in IP-multicast, this paper proposes a novel, highly distributed, and lightweight overlay peer-to-peer retransmission architecture that exploits path-diversity by taking advantages of both IP-multicast and an overlay network. An approach that leverages both disjoint-path-finding and periodic selective probing to take into account peer's recent packet loss probability, retransmission delay and recent retransmission success rate is proposed to effectively construct an efficient and dynamic overlay retransmission network. We show that the proposed path diversity overlay retransmission architecture has the potential to significantly reduce the retransmission delay, improve the reliability, playback quality, and scalability of IP-multicast based multimedia applications. Given a deployed IP-multicast network, the proposed overlay retransmission architecture is practical, scalable, and easy to deploy, requiring no change to the existing network infrastructure.
Wenjun Zeng 0001, Yingnan Zhu, Haibin Lu, Xinhua Zhuang
IEEE Trans. Multim.1
2008 Power-efficient rate allocation for Slepian-Wolf coding over wireless sensor networks
abstract
Power consumption is a critical concern in communications over wireless sensor networks (WSN). In this paper, we address the rate-allocation problem for Slepian-Wolf coding of multiple correlated sources. The goal is to find the optimal rate-point that allows lossless reconstruction of the sources, while minimizing the overall transmission power consumption of the WSN under an exponential cost model. A novel water-filling algorithm to be performed by the receiver is proposed to solve the problem in a recursive manner. The feasibility and optimality of the proposed solution are analyzed mathematically and verified experimentally. Compared to the conventional Lagrangian-multiplier approach, our algorithm achieves dramatic reduction in computational complexity.
Wei Liu 0025, Lina Dong, Wenjun Zeng 0001
ICASSP3
2008 Supporting Video Streaming Services in Infrastructure Wireless Mesh Networks: Architecture and Protocols
abstract
In this paper, we present UPAC, a unified peer-to-peer (P2P) and cache framework for high quality video-on-demand services over infrastructure multi-hop wireless mesh networks. Streaming video in multi-hop wireless networks faces many challenges, e.g., the varying available path bandwidth, interference due to shared medium, the impact of multiple relay nodes, etc. To increase the capacity of streaming services and ensure high video quality, in UPAC, the video content is cached at selected wireless mesh access points (MAPs) inside the mesh network. Furthermore peers help each other on video downloading in a best effort manner to reduce the workload imposed on the servers and networks. This unified framework has the advantages of both the content distribution network approach and peer-to-peer network approach. To obtain optimal video quality, a user device can form the P2P relationship with MAP content cache servers and other user devices. Meanwhile, the client-server relationship is established between the device and the MAP content cache servers. In addition, we propose and compare several methods to select the content cache servers for serving a device and to establish the end-to-end routes between the device and the selected servers. Our preliminary simulation results show that the proposed framework increases the number of users that can be concurrently served in the infrastructure wireless mesh network and improves the received video quality.
Yingnan Zhu, Wenjun Zeng 0001, Hang Liu 0003, Yang Guo 0001, Saurabh Mathur 0001
ICC2
2008 An efficient print-scanning resilient data hiding scheme based on a novel LPM
abstract
Print-scan resilient data hiding has not been extensively researched. This paper presents an efficient multi-bit blind watermarking scheme based on a novel Fourier log-polar mapping (LPM). The watermark resynchronization after print-scanning is efficiently solved by an embedded tracking pattern which cannot be removed by template removing attacks and is not detectable for a malicious part. Experimental results show that the proposed watermarking scheme has excellent robustness to print-scanning, cropping, geometric distortion and JPEG compression etc. The obtained success ratios of extraction 60 bits message without error from the combination attack of JPEG compressed with quality factor of 50–100 and then print-scanning were at least 95%.
Xiangui Kang, Xiong Zhong, Jiwu Huang, Wenjun Zeng 0001
ICIP4
2008 Resolution-progressive compression of encrypted grayscale images
abstract
Compression of encrypted data can be achieved by employing Slepian-Wolf coding (SWC). However, how to efficiently exploit the source dependency in an encrypted colored signal such as an image remains a challenging issue. Previous works incorporate 2-D Markov models in the SWC, which is not accurate enough for natural grayscale images; as a result, the compression performance is usually poor. In this paper, we propose to compress the image progressively, such that the decoder can observe a low-resolution version of the image, from which local statistics is learned and used for the decoding of the next resolution level. Good performance is observed both theoretically and experimentally.
Wei Liu 0025, Wenjun Zeng 0001, Lina Dong, Qiuming Yao
ICIP2
2008 Improving Robustness of Quantization-Based Image Watermarking via Adaptive Receiver
abstract
In this paper, the watermarking channel is modeled as a generalized channel with fading andnonzeromeanadditive noise. In order to improve the watermark robustness against the generalized channel, we present an optimized watermark extraction scheme by using an adaptive receiver for quantization-based watermarking. In the proposed extraction scheme, we adaptively estimate the decision zone of the binary data bits and the quantization step size. A training sequence is embedded into the original image together with the informative watermark. The estimation of the decision zone takes advantage of the response function of the training sequence. Compared to those watermarking schemes without receiver adaptation, the main improvement is the enhanced robustness against median filtering, image intensity Direct Current (DC) change, histogram equalization, color reduction, image intensity linear scaling, image intensity nonlinear scaling such as Gamma correction etc.
Xiangui Kang, Jiwu Huang, Wenjun Zeng 0001
IEEE Trans. Multim.3
2007 Optimum Detection for Spread-Spectrum Watermarking that Employs Self-masking
abstract
Digital watermarking is an effective and promising approach to protect intellectual property rights of digital media. Spread spectrum (SS) is one of the most widely used image watermarking schemes. In SS watermarking, the watermark signal is usually modulated by the just-noticeable difference (JND) of the host image. The JND is measured by advanced perceptual models as a non-linear function of local image features. In this paper, the optimum detection scheme for such non-linearly embedded watermarks is addressed. Closed-form detectors are found for arbitrary JND models that exploit the self-masking property of the human visual system.
Wei Liu 0025, Lina Dong, Wenjun Zeng 0001
ICIP (5)3
2007 Exploiting Overlay Path-Diversity for Scalable Reliable Multicast
abstract
IP-multicast is a bandwidth efficient transmission mechanism for multimedia communication. Reliability in IP-multicast, however, remains a significant challenge. This paper addresses the reliability and scalability issues in IP-multicast by exploring a novel, highly distributed overlay peer-to-peer retransmission architecture that exploits path-diversity. A simple unicast-based "tracert" tool is proposed to help to identify peers with disjoint path to the sender as potential retransmission nodes. Probing can help to adapt to the dynamics of the network. We show that a hybrid system with both "tracert" and probing can perform better than probing only or "tracert" only approach. In addition, we identify the proper probing interval that does not introduce significant probing overhead yet can collect sufficiently updated information to help recover most of the lost packets. The proposed hybrid multicast system is practical, scalable, and easy to deploy, requiring no change to the existing network infrastructure.
Yingnan Zhu, Wenjun Zeng 0001, Haibin Lu
ICME2
2007 Optimum Detection for Spread-Spectrum Watermarking That Employs Self-Masking
abstract
Digital watermarking is an efficient and promising approach to protect intellectual property rights of digital media. Spread spectrum (SS) is one of the most widely used image watermarking schemes because of its robustness against attacks and its support for the exploitation of the properties of the human visual system (HVS). To maximize the watermark strength without introducing visual artifacts, in SS watermarking, the watermark signal is usually modulated by the just-noticeable difference (JND) of the host image. In advanced perceptual models, the JND is characterized as a nonlinear function of local image features. The optimum detection scheme for such nonlinearly embedded watermarks, however, has rarely been studied. In this paper, we address this problem and propose a novel approach that transforms the test signal to a perceptually uniform domain and then performs Bayesian hypothesis testing in that domain. Locally optimum detectors for arbitrary host signal distributions and arbitrary JND models that exploit the self-masking property of the HVS are derived in closed forms, in which the test signal is first nonlinearly preprocessed before a linear correlator is applied. The optimality of the proposed detector is justified mathematically according to the Neyman–Pearson criterion. Simulation results demonstrate the superior performances of the proposed detector over the conventional linear correlation detector.
Wei Liu 0025, Lina Dong, Wenjun Zeng 0001
IEEE Trans. Inf. Forensics Secur.3
2007 Fast Bitstream Switching Algorithms for Real-Time Adaptive Video Multicasting
abstract
Bitstream switching among multiple bitstreams encoded at different bit rates is an effective way to address the bandwidth variation issue in transmitting multimedia over the Internet or wireless networks. This paper proposes two new fast real-time bitstream switching algorithms that aim to minimize the drifting error, while avoiding the problems of long delay, high complexity and bit-rate overhead for storage and transmission that often occur in prior solutions. The basic idea is to choose a switching point in a neighborhood with the highest encoding quality, within a switching window determined by the switching delay constraint. We show that they can significantly outperform a simple switching algorithm, and achieve performance that is closer to an offline mean-square-error-optimized bitstream switching solution, when compared to our previous work based on the similarity of the reference frames. The proposed schemes are especially useful in the scenario of real-time multicasting over dynamic heterogeneous networks, where multiple bitstreams with different bit rates are generated on the fly and dynamic bitstream switching is required for individual clients
Bo Xie 0003, Wenjun Zeng 0001
IEEE Trans. Multim.2
2007 Joint Design of Source Rate Control and QoS-Aware Congestion Control for Video Streaming Over the Internet
abstract
Multimedia streaming over the Internet has been a very challenging issue due to the dynamic uncertain nature of the channels. This paper proposes an algorithm for the joint design of source rate control and congestion control for video streaming over the Internet. With the incorporation of a virtual network buffer management mechanism (VB), the quality of service (QoS) requirements of the application can be translated into the constraints of the source rate and the sending rate. Then at the application layer, the source rate control is implemented based on the derived constraints, and at the transport layer, aQoS-awarecongestion control mechanism is proposed that strives to meet the send rate constraint derived from VB, by allowing temporary violation of transport control protocol (TCP)-friendliness when necessary. Long-term TCP-friendliness, nevertheless, is preserved by introducing a rate-compensation algorithm. Simulation results show that compared with traditional source rate/congestion control algorithms, this cross-layer design approach can better support the QoS requirements of the application, and significantly improve the playback quality by reducing the overflow and underflow of the decoder buffer, and improving quality smoothness, while maintaining good long-term TCP-friendliness.
Wenjun Zeng 0001, Chunwen Li
IEEE Trans. Multim.2
2006 Rate-Distortion Optimized Transmission Power Adaptation for Video Streaming over Wireless Channels
abstract
An important characteristic of video transmission over wireless channels is that the channel is error-prone and the compressed video data is highly sensitive to errors. The transmission errors will cause decoding failure which will distort the reconstructed pictures. One of the challenging issues in wireless video communication system design is to minimize the overall energy consumption and maximize the operational lifetime of the devices. In this work, we focus on the energy consumption in wireless video transmission. Based on a transmission distortion model, we develop a rate-distortion optimized transmission power adaptation scheme for video streaming over wireless channels. Our experimental results demonstrate that this scheme is able to significantly improve the video quality under the transmission power constraints.
Zhiquan He, Wenjun Zeng 0001
ICIP2
2006 Optimum Detection of Image-Adaptive Watermarking in the DCT Domain
abstract
Digital watermarking is an efficient and promising means to protect intellectual properties. Spread spectrum (SS) is one of the most widely used watermarking schemes, and can be further facilitated by adapting the watermark strength to the host signal. In this paper we address the problem of optimum blind detection of SS watermarks in the DCT domain, where Watson's non-linear perceptual model is used for the modulation of the watermarks. The basic idea is to transform the DCT coefficients to a perceptually uniform domain and perform Bayesian hypothesis testing in that domain. Simulation results show the superior performance of the proposed scheme over the conventional linear correlation detector (LCD).
Wei Liu 0025, Lina Dong, Wenjun Zeng 0001
ICIP3
2006 Path-Diversity Overlay Retransmission Architecture for Reliable Multicast
abstract
IP-multicast is a bandwidth efficient transmission mechanism for group communications. Reliability in IP-multicast, however, poses a set of significant challenges. To address the reliability and scalability issues in IP-multicast, this paper proposes a novel overlay retransmission architecture that exploits path-diversity by taking advantages of both IP multicast and an overlay network. We show that the proposed path diversity overlay retransmission architecture has the potential to significantly improve the reliability, delay, playback quality, and scalability of IP-multicast based multimedia applications. The general concept of using P2P overlay networks to help improve the QoS performance of multimedia applications as illustrated in this paper is expected to have significant impact on the deployment of next generation multimedia services
Wenjun Zeng 0001, Yingnan Zhu, Haibin Lu, Hongbing Jiang
ICME1
2006 Cross-Layer Design of Source Rate Control and Qos-Aware Congestion Control for Wireless Video Streaming
abstract
Cross-layer design has been used in streaming video over the wireless channels to optimize the overall system performance. In this paper we extend our previous work (i.e. joint design of source rate control and congestion control for video streaming over the Internet) in [1] and propose a cross-layer design approach for wireless video streaming. By jointly designing the source rate control at the application layer and congestion control at the transport layer, and taking advantage of MAC layer information, our approach can avoid the throughput degradation caused by transmission error of the wireless channel, and better support the QoS requirements of the application. Simulation results show that the proposed mechanism can significantly improve the playback quality of the application, while maintaining good performance of the transport protocol.
Wenjun Zeng 0001, Chunwen Li
ICME2
2006 A sequence-based rate control framework for consistent quality real-time video
abstract
Most model-based rate control solutions have the generally questionable assumption that video sequence is stationary. In addition, they often suffer from the fundamental problem of model parameter misestimation. In this paper, we propose a sequence-based frame-level bit allocation framework employing a rate-complexity model that has the capability of tracking the nonstationary characteristics in the video source without look-ahead encoding. In addition, a new nonlinear model parameter estimation approach is proposed to overcome the existing problems in previous model parameter estimation schemes where quantization parameter (QP) is determined to achieve the allocated bits for a frame. Furthermore, a general concept of bit allocation guarantee is discussed and its importance is highlighted. The proposed rate control solution can achieve smoother video quality with less quality flicker and motion jerkiness. Both a complete solution where requantization is employed to guarantee the achievement of the allocated bits, and a simplified solution without requantization are studied. Experimental results show that they both provide significantly better performance, in terms of average peak-signal-to-noise ratio and quality smoothness, than the MPEG-4 Annex L frame-level rate control solution.
Bo Xie 0003, Wenjun Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.2
2005 A protocol for simultaneous real time playback and full quality storage of streaming media
abstract
In this paper, we introduce the new problem of simultaneous streaming of a single media bitstream to multiple devices with different quality of service (QoS) requirements. In particular, we address simultaneous streaming of a single video stream for both real time playback and full quality storage, where the QoS requirements of the two targets are different. We design a joint streaming protocol to fully exploit the available bandwidth to deliver both real-time and retransmitted packets simultaneously, as bandwidth allows. Preliminary results show that the proposed joint streaming protocol can simultaneously address the requirements of both real-time playback and less time-critical, higher quality storage of streaming media.
Mike Sullivan, Wenjun Zeng 0001
ICC2
2005 On the rate-distortion performance of dynamic bitstream switching mechanisms
abstract
Bitstream switching is an effective way to deal with bandwidth variation in transmitting multimedia over time-varying channels. A number of bitstream switching mechanisms have been proposed to reduce the potential error drift introduced by bitstream switching, often without explicitly taking into account the impact of the switching on the bandwidth and delay requirements in their performance analysis. This paper studies the relative rate-distortion (R-D) performance of some of these bitstream streaming mechanisms by extending a delay-aware R-D optimized dynamic bitstream switching framework we previously proposed. By considering different switching mechanisms, we show that the proposed extended R-D based framework can adaptively choose the right switching mechanism and effectively switch at the right time to the right bitstream to achieve optimized rate-distortion performance.
Bo Xie 0003, Wenjun Zeng 0001
ICME2
2005 Multi-band Wavelet Based Digital Watermarking Using Principal Component Analysis
Xiangui Kang, Yun Q. Shi 0001, Jiwu Huang, Wenjun Zeng 0001
IWDW4
2005 Scalable Non-Binary Distributed Source Coding Using Gray Codes
abstract
In this paper we propose a bit-plane based, scalable distributed source coding (DSC) scheme for non-binary correlated sources. In the source coding part of DSC, the symbols are quantized and the quantization indices are converted into binary forms using Gray codes. The channel coding part encodes each bit-plane independently using turbo codes. The turbo decoder that works in binary domain is facilitated by an outer loop that captures the symbol-domain correlation. Simulation results show that, in asymmetric DSC of joint Gaussian sources, the overall system performance is greatly improved by using Gray codes. It is also demonstrated that progressive transmission is well supported
Wei Liu 0025, Wenjun Zeng 0001
MMSP2
2005 Joint Design of Source Rate Control and QoS-Aware Congestion Control for Video Streaming over the Internet
abstract
This paper proposes an algorithm for joint design of source rate control and congestion control for video streaming over the Internet. At the application layer, with the incorporation of a virtual network buffer management mechanism(VB), the QoS requirements of the application can be translated into the constraints of the source rate and the send rate. At the transport layer, a QoS-aware congestion control mechanism is proposed that strives to meet the send rate constraint derived from VB, by allowing temporary violation of TCP-friendliness when necessary. Long-term TCP-friendliness, nevertheless, is preserved by introducing a rate compensation algorithm. Simulation results show that compared with traditional source rate/congestion control algorithms, this cross-layer design approach can significantly improve the playback quality by reducing the overflow and underflow of the decoder buffer, and improving quality smoothness, while maintaining long-term TCP-friendliness.
Wenjun Zeng 0001, Chunwen Li
MMSP2
2005 Operational distortion-quantization curve-based bit allocation for smooth video quality
Junqiang Lan, Wenjun Zeng 0001, Xinhua Zhuang
J. Vis. Commun. Image Represent.2
2005 Adaptive spatial-temporal error concealment with embedded side information
Wenjun Zeng 0001
J. Vis. Commun. Image Represent.1
2005 Low-pass filtering of rate-distortion functions for quality smoothing in real-time video communication
abstract
In variable-bit-rate video coding, the video is preprocessed to collect sequence-level statistics, which are used for global bit allocation in the actual encoding stage to obtain a smoothed video presentation quality. However, in real-time video recording and network streaming, this type of two-pass encoding scheme is not allowed because the access to future frames and global statistics is not available. To address this issue, we introduce the concept of low-pass filtering of rate-distortion functions and develop a smoothed rate control (SRC) framework for real-time video recording and streaming. Theoretically, we prove that, using a geometric averaging filter, the SRC algorithm is able to maintain a smoothed video presentation quality while achieving the target bit rate automatically. We also analyze the buffer requirement of the SRC algorithm in real-time video streaming, and propose a scheme to seamlessly integrate robust buffer control into the SRC framework. The proposed SRC algorithm has very low computational complexity and implementation cost. Our extensive experimental results demonstrate that the SRC algorithm significantly reduces the picture quality variation in the encoded video clips.
Zhihai He, Wenjun Zeng 0001, Chang Wen Chen
IEEE Trans. Circuits Syst. Video Technol.2
2005 Tile-boundary artifact reduction using odd tile size and the low-pass first convention
abstract
It is well known that tile-boundary artifacts occur in wavelet-based lossy image coding. However, until now, their cause has not been understood well. In this paper, we show that boundary artifacts are an inescapable consequence of the usual methods used to choose tile size and the type of symmetric extension employed in a wavelet-based image decomposition system. This paper presents a novel method for reducing these tile-boundary artifacts. The method employs odd tile sizes (2N + 1 samples) rather than the conventional even tile sizes (2N samples). It is shown that, for the same bit rate, an image compressed using an odd tile length low-pass first (OTLPF) convention has significantly less boundary artifacts than an image compressed using even tile sizes. The OTLPF convention can also be incorporated into the JPEG 2000 image compression algorithm using extensions defined in Part 2 of this standard.
Jianxin Wei 0002, Mark R. Pickering, Michael R. Frater, John F. Arnold, John A. Boman, Wenjun Zeng 0001
IEEE Trans. Image Process.6
2004 Two fast bitstream switching algorithms for real-time adaptive multicasting of video
abstract
Bitstream switching is an effective way to deal with the bandwidth variation issue in transmitting multimedia over the Internet or wireless networks. This paper proposes two new fast intelligent bitstream switching algorithms that aim to minimize the drifting error while avoiding the problem of long delay and high complexity that often occurred in prior solutions. The basic ideas are to choose a switching point in the most stationary neighborhood or in a neighborhood with the highest encoding quality, within a switching window determined by the switching delay constraint. We show that they can achieve performance that is closer or even identical to an off-line MSE-optimized bitstream switching solution, when compared to the algorithm we previously presented in B. Xie et al. (2003). They are especially useful in the scenario of real-time multicasting over heterogeneous networks, where multiple bitstreams with different bitrates are generated on the fly and dynamic bitstream switching is required for individual clients.
Bo Xie 0003, Wenjun Zeng 0001
ICC2
2004 An improved rate-quantization model for rate control in real-time video encoding
Bo Xie 0003, Wenjun Zeng 0001
ICIP2
2004 Network friendly media security: rationales, solutions, and open issues
abstract
Network friendly media security refers to the security technologies that are specifically designed to cope with existing and future multimedia networking infrastructures and technologies so as to ease the deployment and maintain or improve the quality of service performance of multimedia applications. It is especially useful for streaming and mobile multimedia applications where content adaptation is a necessity. In this paper, we analyze the various motivations behind network-friendly security solutions, review some of the most recent approaches, discuss some open issues, and suggest some potential solutions.
Wenjun Zeng 0001, Xinhua Zhuang, Junqiang Lan
ICIP1
2004 Single-pass frame-level constant distortion bit allocation for smooth video quality
abstract
Quality fluctuation has a major negative effect on subjective video quality. A desirable single-pass frame-level constant-distortion bit allocation scheme is proposed in the paper for smooth video quality throughout the video sequence. The average distortion of previous coded frames is taken as the target distortion for the current frame. According to the linear rate control algorithm, the distortion and rate are exponential and linear function of the number of zero coefficients, respectively. Based on the target distortion and the distribution of the absolute DCT coefficients, we then derive the closed-form formulas for estimating the number of zero coefficients and the slope /spl theta/ in the linear rate control function as well as the bit budget for the current frame. Experimental results show that the proposed constant-distortion (but variable rate) bit allocation scheme provides much smoother video quality on all testing video sequences than constant bit allocation scheme.
Junqiang Lan, Xinhua Zhuang, Wenjun Zeng 0001
ICME3
2004 Rate-distortion optimized dynamic bitstream switching for scalable video streaming
abstract
Bitstream switching is an effective way to deal with bandwidth variation in transmitting multimedia over time-varying channels such as the Internet or wireless networks. Most previous work has focused on how to switch to reduce the potential error drift, without necessarily taking into account the impact of the switching on the bandwidth and delay requirements. This paper proposes a rate-distortion optimized dynamic bitstream switching algorithm that explicitly takes into account the impact of the switching on both drifting effect and bandwidth requirement. The rate control (or switching decision) of the proposed algorithm explicitly takes into account the delay constraint of multimedia applications, using an integrated end-to-end virtual network buffer management approach. Initial results show that it can achieve good performance with well-controlled quality drift and delay, and provides a promising research direction.
Bo Xie 0003, Wenjun Zeng 0001
ICME2
2003 Source characteristics based fast bitstream switching
abstract
Bitstream switching is an effective way to deal with the bandwidth variation issue in transmitting multimedia over the Internet or wireless networks. This paper proposes a fast intelligent bitstream switching algorithm that avoids the problem of long delay and high complexity that often occurred in prior solutions. We show that it can achieve very close performance as an off-line MSE-optimized bitstream switching solution, with well-controlled quality drift. It is especially useful in the scenario of real-time multicasting over heterogeneous networks, where multiple bitstreams with different bitrates are generated on the fly and dynamic bitstream switching is required for individual clients.
Bo Xie 0003, Wenjun Zeng 0001
ICME2
2003 Spatial-temporal error concealment with side information for standard video codecs
abstract
Error concealment (EC) is important for transmitting video over error prone networks such as the Internet or wireless networks. Different EC strategies have their own advantages for different scenarios. However, it is generally difficult for the decoder to figure out which strategy works the best for a specific case. This paper proposes to use data embedding to convey the necessary high level side information in a standard compliant way to help improve the decoder's EC performance. We show that the EC "mode" information (i.e.., whether spatial EC or temporal EC should be used) is critical for a spatial-temporal EC approach.
Wenjun Zeng 0001
ICME1
2003 Efficient frequency domain selective scrambling of digital video
abstract
Multimedia data security is very important for multimedia commerce on the Internet such as video-on-demand and real-time video multicast. Traditional cryptographic algorithms/systems for data security are often not fast enough to process the vast amount of data generated by multimedia applications to meet real-time constraints. This paper presents a joint encryption and compression framework in which video data are scrambled efficiently in the frequency domain by employing selective bit scrambling, block shuffling and block rotation of the transform coefficients and motion vectors. The new approach is very simple to implement, yet provides considerable levels of security and different levels of transparency, and has a very limited adverse impact on compression efficiency and no adverse impact on error resiliency. Furthermore, it allows transcodability/scalability, and other content processing functionalities without having to access the cryptographic key and perform decryption and re-encryption.
Wenjun Zeng 0001, Shawmin Lei
IEEE Trans. Multim.1
2002 Sequence-based rate control for constant quality video
abstract
Most model based rate control solutions have the generally questionable assumption that video sequence is stationary, and also suffer from the fundamental problems of model parameter mis-estimation. In this paper, we propose a sequence based bit allocation solution with the capability of tracking the nonstationary characteristics in the video source without look-ahead encoding and thus the stationarity assumption is no longer needed. In addition, a new model parameter estimation approach is provided to solve the problems in the existing model parameter estimation schemes. Moreover, a general concept of bit allocation guarantee is presented to achieve the allocated bits in a deterministic way. The proposed rate control solution can achieve constant quality video with less quality flicker and motion jerkiness.
Bo Xie 0003, Wenjun Zeng 0001
ICIP (1)2
2002 3G wireless multimedia: technologies and practical issues
abstract
This paper provides an overview of the emerging wireless communication standards, end-to-end wireless streaming systems, and relevant wireless multimedia technologies. It highlights some of the challenges in the deployment of 3G wireless multimedia services, using PacketVideo's solutions as an example.
Wenjun Zeng 0001, Jiangtao Wen
ICIP (1)1
2002 Fast self-synchronous content scrambling by spatially shuffling codewords of compressed bitstreams
abstract
This paper presents a content access control method by spatially shuffling codewords of the compressed bitstream. This approach is lightweight and incurs no bit overhead. One important advantage of the approach is that the resulting scrambled bitstream can be made compliant to the compression format, thus providing some level of scalability, error resiliency, network friendliness and capability of performing signal processing directly on the encrypted bitstream. In addition, we propose a method for generating or updating the shuffling tables on the fly, based on encrypting some local-content-specific bits using a standard cipher. This local-content-specific bits based table generation process is self-synchronous, which is critical in the presence of packet loss. It also enhances the resistance of this spatial shuffling approach to plain-text attack.
Wenjun Zeng 0001, Jiangtao Wen, Mike Severa
ICIP (3)1
2002 An overview of the visual optimization tools in JPEG 2000
Wenjun Zeng 0001, Scott Daly, Shawmin Lei
Signal Process. Image Commun.1
2002 A format-compliant configurable encryption framework for access control of video
abstract
We introduce new methods of performing selective encryption and spatial/frequency shuffling of compressed digital content that maintain syntax compliance after content has been secured. The tools described have been proposed to the MPEG-4 Intellectual Property Management and Protection (IPMP) standardization group and have been adopted into the MPEG-4 IPMP Final Proposed Draft Amendment (FPDAM). We describe the application of the new methods to the protection of MPEG-4 video content in the wireless environment, and illustrate how they are used to leverage established encryption algorithms for the protection of only the information fields in the bitstream that are critical to the reconstructed video quality, while maintaining compliance to the syntax of MPEG-4 video, and thereby reduces the amount of data to be encrypted and guarantees the inheritance of many of the good properties of the unprotected bitstreams that have been carefully studied and built, such as error resiliency and network friendliness. The encrypted content bitstream works with many existing random access, network bandwidth adaptation, and error control techniques that have been developed for standard-compliant compressed video, thus making it especially suitable for wireless multimedia applications. Standard compliance also allows subsequent signal processing techniques to be applied to the encrypted bitstream.
Jiangtao Wen, Mike Severa, Wenjun Zeng 0001, Max H. Luttrell, Weiyin Jin
IEEE Trans. Circuits Syst. Video Technol.3
2002 3G wireless multimedia: technologies and practical issues
abstract
Abstract This paper provides overviews of the emerging wireless communication standards, end‐to‐end wireless streaming systems, and relevant wireless multimedia technologies. It highlights some of the challenges in the deployment of 3G wireless multimedia services, using PacketVideo's solutions as an example. Copyright © 2002 John Wiley & Sons, Ltd.
Wenjun Zeng 0001, Jiangtao Wen
Wirel. Commun. Mob. Comput.1
2001 Scalable streaming of JPEG2000 images using hypertext transfer protocol
abstract
This paper describes a scalable architecture for streaming of JPEG2000 images, using Hypertext Transfer Protocol (HTTP). JPEG2000 is a new image compression standard. One of the goals of JPEG2000 is to support large images. For a large image, even the compressed image file size can be very big. Thus downloading the entire image at its full resolution can take a long time depending upon the user's connection speed. Thus we propose to use streaming of JPEG2000 images. We use Hypertext transfer protocol (HTTP) for streaming. The use of HTTP to stream JPEG2000 images will ease its deployment, because of the widespread availability and accessibility of web servers. Thus JPEG2000 images can be hosted by web server and can be streamed to the client helper application. Our solution is scalable. The clients can view image at a variety of resolutions and quality levels. It is also possible to stream only selected region of the image at a particular resolution and decode this partial stream and display it at the client. Thus client devices with different capabilities, variety of screen resolutions, heterogeneous bandwidths can all achieve a scalable viewing of the same content stored in a single file.
Sachin Deshpande, Wenjun Zeng 0001
ACM Multimedia2
2001 A format-compliant configurable encryption framework for access control of multimedia
abstract
We present a framework for access control of standard-compliant video bitstreams for entertainment purposes. The approach leverages well-known encryption algorithms and maintains standard-compliance of the encrypted bitstream. The standard compliance feature guarantees the inheritance of the error resiliency properties of the video compression standards, and works with many existing network bandwidth adaptation and error control techniques that have been developed for standards-compliant compressed video, thus making it especially suitable for wireless multimedia applications. Standards compliance also allows subsequent signal processing techniques to be applied to the encrypted bitstream. The approach can be applied to any of the common video coding standards. It is also capable of providing layered compromises between security, complexity, delay, bit overhead, etc.
Jiangtao Wen, Mike Severa, Wenjun Zeng 0001, Max H. Luttrell, Weiyin Jin
MMSP3
2000 Point-Wise Extended Visual Masking for JPEG-2000 Image Compression
abstract
One common visual optimization strategy for image compression is to exploit the visual masking effect where artifacts are locally masked by the image acting as a background signal. In this paper, we present a point-wise extended visual masking approach that nonlinearly maps the wavelet coefficients to a perceptually uniform domain prior to quantization by taking advantages of both self-contrast masking and neighborhood masking effects, thus achieving very good visual quality. It is essentially a coefficient-wise adaptive quantization without any overhead. It allows bitstream scalability, as opposed to many previous works. The proposed scheme has been adopted into the working draft of JPEG-2000 Part II.
Wenjun Zeng 0001, Scott Daly, Shawmin Lei
ICIP1
2000 Visual Optimization Tools in JPEG 2000
abstract
We review the various tools in JPEG 2000 that allow the users to take advantages of the various properties of the human visual system such as spatial frequency sensitivity and the visual masking effect. We show that the visual tool sets in JPEG 2000 are much richer than what was available in JPEG, where only locally invariant frequency weighting can be exploited.
Wenjun Zeng 0001, Scott Daly, Shawmin Lei
ICIP1
2000 An Efficient Color Re-Indexing Scheme for Palette-Based Compression
abstract
This paper presents a fast and efficient way for color re-indexing that tends to maximize the compression performance of a palette-based compression system. The proposed scheme relates the index difference of neighboring pixels to the potential cost of bits. It optimizes the assignment of index values to colors in a one-step look-ahead greedy fashion. Experimental results suggest that the proposed re-indexing scheme can reduce the bit rate by up to 43%, when compared to a previously proposed intensity-based color-indexing scheme. Furthermore, we show that with the proposed color re-indexing scheme, the palette-based JPEG-LS and palette-based JPEG-2000 can often outperform the graphical interchange format (GIF) significantly.
Wenjun Zeng 0001, Shawmin Lei
ICIP1
1999 Efficient frequency domain video scrambling for content access control
abstract
Multimedia data security is very important for multimedia commerce on the Internet such as video-on-demand and real-time video multicast. Traditional cryptographic algorithms for data security are often not fast enough to process the vast amount of data generated by the multimedia applications to meet the real-time constraints. This paper presents a joint encryption and compression framework in which video data are scrambled efficiently in the frequency domain by employing selective bit scrambling, block shuffling and block rotation of the transform coefficients and motion vectors. The new approach is very simple to implement, yet provides considerable level of security, has minimum adverse impact on the compression efficiency, and allows transparency, transcodability, and other content processing functionalities without accessing the cryptographic key.
Wenjun Zeng 0001, Shawmin Lei
ACM Multimedia (1)1
1999 Geometric-structure-based error concealment with novel applications in block-based low-bit-rate coding
abstract
This paper first proposes a computationally efficient spatial directional interpolation scheme, which makes use of the local geometric information extracted from the surrounding blocks. The proposed error-concealment scheme produces results that are superior to those of other approaches, in terms of both peak signal-to-noise ratio and visual quality. Then a novel approach that incorporates this directional spatial interpolation at the receiver is proposed for block-based low-bit-rate coding. The key observation is that the directional spatial interpolation at the receiver can reconstruct faithfully a large percentage of the blocks that are intentionally not sent. A rate-distortion optimal way to drop the blocks is shown. The new approach can be made compatible with standard JPEG and MPEG decoders. The block-dropping approach also has an important application for dynamic rate shaping in transmitting precompressed videos over channels of dynamic bandwidth. Experimental results show that the proposed coding and rate-shaping systems can provide significant subjective and objective gains over conventional approaches.
Wenjun Zeng 0001, Bede Liu
IEEE Trans. Circuits Syst. Video Technol.1
1999 A statistical watermark detection technique without using original images for resolving rightful ownerships of digital images
abstract
Digital watermarking has been proposed as the means for copyright protection of multimedia data. Many of existing watermarking schemes focused on the robust means to mark an image invisibly without really addressing the ends of these schemes. This paper first discusses some scenarios in which many current watermarking schemes fail to resolve the rightful ownership of an image. The key problems are then identified, and some crucial requirements for a valid invisible watermark detection are discussed. In particular, we show that, for the particular application of resolving rightful ownership using invisible watermarks, it might be crucial to require that the original image not be directly involved in the watermark detection process. A general framework for validly detecting the invisible watermarks is then proposed. Some requirements on the claimed signature/watermarks to be used for detection are discussed to prevent the existence of any counterfeit scheme. The optimal detection strategy within the framework is derived. We show the effectiveness of this technique based on some visual-model-based watermark encoding schemes.
Wenjun Zeng 0001, Bede Liu
IEEE Trans. Image Process.1
1998 Adaptive Wavelet Transforms with Spatially Varying Filters for Scalable Image Coding
abstract
An adaptive wavelet transform algorithm for scalable image coding is proposed. A quadtree segmentation scheme based on an entropy criterion is proposed to segment the image into several regions to allow for adaptive filtering using different types of prototype wavelet filters. A joint bit allocation and coding method is presented to efficiently code the segmented image in an embedding fashion. Experimental results show that the proposed scheme generally provides higher peak signal-to-noise ratios (PSNRs) at higher bit rates, and significantly better visual quality at lower bit rates, compared to fixed-filter wavelet transform based coding schemes.
Wenjun Zeng 0001, Shawmin Lei
ICIP (1)1
1998 Image-adaptive watermarking using visual models
abstract
The huge success of the Internet allows for the transmission, wide distribution, and access of electronic data in an effortless manner. Content providers are faced with the challenge of how to protect their electronic data. This problem has generated a flurry of research activity in the area of digital watermarking of electronic content for copyright protection. The challenge here is to introduce a digital watermark that does not alter the perceived quality of the electronic content, while being extremely robust to attack. For instance, in the case of image data, editing the picture or illegal tampering should not destroy or transform the watermark into another valid signature. Equally important, the watermark should not alter the perceived visual quality of the image. From a signal processing perspective, the two basic requirements for an effective watermarking scheme, robustness and transparency, conflict with each other. We propose two watermarking techniques for digital images that are based on utilizing visual models which have been developed in the context of image compression. Specifically, we propose watermarking schemes where visual models are used to determine image dependent upper bounds on watermark insertion. This allows us to provide the maximum strength transparent watermark which, in turn, is extremely robust to common image processing and editing such as JPEG compression, rescaling, and cropping. We propose perceptually based watermarking schemes in two frameworks: the block-based discrete cosine transform and multiresolution wavelet framework and discuss the merits of each one. Our schemes are shown to provide very good results both in terms of image transparency and robustness.
Christine Podilchuk, Wenjun Zeng 0001
IEEE J. Sel. Areas Commun.2
1997 On Resolving Rightful Ownership's of Digital Images by Invisible Watermarks
abstract
Digital watermarking has been proposed as a means for copyright protection of multimedia data. This paper first discusses some scenarios in which many current watermarking schemes fail to resolve the rightful ownership of an image. We then identify the key problems and discuss some crucial requirements for a valid invisible watermark detection. In particular, we argue that, for the particular application of resolving rightful ownership using invisible watermarks, it might be crucial to require that the original image not be directly involved in the watermark detection process. A general framework for validly detecting the invisible watermarks is then proposed. We show the effectiveness of this technique based on some visual-model-based watermarking schemes.
Wenjun Zeng 0001, Bede Liu
ICIP (1)1
1997 Feature-Oriented Rate Shaping of Pre-Compressed Image/Video
abstract
Rate shaping plays an important role in transmitting compressed video on a network with dynamic bandwidth. We present a scheme for rate shaping JPEG/MPEG-pre-compressed image/video with feature analysis in scale space. The proposed scheme draws its strengths from a technique for identifying perceptually important features based on geometry-driven diffusion, an interpolation scheme allowing many intentionally dropped blocks to be recovered faithfully from neighboring transmitted blocks, and some careful observations about the close interactions between block dropping and feature analysis. Compared with a successful rate shaping scheme based on joint block dropping and uniform coefficient truncation, the new scheme not only produces images of significantly higher visual quality, but also better PSNR at low bit rates.
Wenjun Zeng 0001, Baining Guo, Bede Liu
ICIP (2)1
1997 Perceptual watermarking of still images
abstract
Content providers on the Internet are faced with the problem of how to secure electronic data. This problem has generated research activity in the area of digital watermarking of electronic content. The challenge is to introduce a digital watermark that is both transparent and highly robust to common signal processing and possible attacks. The two basic requirements for an effective watermarking scheme, robustness and transparency, conflict with each other. We propose a watermarking technique for digital images that is based on utilizing visual models which have been developed in the context of image compression. The visual models give us a direct way to determine the maximum amount of watermark signal that each portion of an image can tolerate without affecting the visual quality of the image. This allows us to provide the maximum strength watermark which in turn, is extremely robust to common image processing and editing such as JPEG compression, rescaling, and cropping. Our watermarking scheme is based on a DCT framework which allows for the possibility of directly watermarking the JPEG bitstream. Our scheme is shown to provide very good results both in terms of image transparency and robustness.
Christine Podilchuk, Wenjun Zeng 0001
MMSP2
1996 Directional spatial interpolation for DCT-based low bit rate coding
abstract
A novel approach which combines directional spatial interpolation at the receiver and block-based coding is proposed for low bit rate coding. The directional spatial interpolation at the receiver can reconstruct faithfully a large percentage of the blocks that are intentionally not sent. Experimental results show that the new proposed coding system can provide significant subjective and objective gains over conventional approaches for still image compression, and has the potential to be extended for video coding.
Wenjun Zeng 0001, Bede Liu
ICASSP1
1996 Rate Shaping by Block Dropping for Transmission of MPEG-Precoded Video over Channels of Dynamic Bandwidth
abstract
In the transmission of a precoded, stored video such as that in video-on-demand systems , it is often necessary to reduce the bit rate of the compressed video in cases that the network capacity is reduced. This paper proposes a novel block-dropping approach for rate shaping of MPEG-precompressed video. A directional spatial interpolation scheme is incorporated at the receiver to reconstruct faithfully a large percentage of the blocks that are intelligently dropped to meet the constraint of reduced bandwidth. The blocks are dropped in a way that is optimal in the rate-distortion sense. The new approach is conceptually quite different from conventional approaches in which coefficients (usually high frequency coefficients) are eliminated. Experimental results show that the proposed scheme can provide significantly higher peak signal-to-noise ratio (PSNR) and much better visual quality than conventional approaches. We also show that by jointly dropping blocks and coefficients, the bit rate...
Wenjun Zeng 0001, Bede Liu
ACM Multimedia1