Wenyu Liu 0001

dblp:42/4110-1 · DBLP profile ↗
← Back
271ranked-venue papers
5as first author
97since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 154 · 3 first-author · 77 since 2021Graphics, computer vision, multimedia, augmented reality and games · 130 · 54 since 2021Databases, data management, data science and information retrieval · 19Computer networks · 18 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 18 · 1 first-author · 5 since 2021Systems, architecture and hardware · 13Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Gait Recognition via Collaborating Discriminative and Generative Diffusion Models
abstract
Gait recognition offers a non-intrusive biometric solution by identifying individuals through their walking patterns. Although discriminative models have achieved notable success in this domain, the full potential of generative models remains largely unexplored. In this paper, we introduce CoD², a novel framework that combines the data distribution modeling capabilities of diffusion models with the semantic representation learning strengths of discriminative models to extract robust gait features. We propose a Multi-level Conditional Control strategy that integrates both high-level identity-aware semantic conditions and low-level visual details. Specifically, the high-level condition, extracted by the discriminative extractor, guides the generation of identity-consistent gait sequences, while low-level visual details, such as appearance and motion, are preserved to enhance consistency. Moreover, the generated sequences facilitate the discriminative extractor's learning, enabling it to capture more comprehensive high-level semantic features. Extensive experiments on four datasets (SUSTech1K, CCPG, GREW, and Gait3D) demonstrate that CoD² achieves state-of-the-art performance and can be seamlessly integrated with existing discriminative methods, yielding consistent improvements.
Haijun Xiong, Bin Feng 0001, Bang Wang 0001, Xinggang Wang, Wenyu Liu 0001
AAAI5
2026 MolSight: Optical Chemical Structure Recognition with SMILES Pretraining, Multi-Granularity Learning and Reinforcement Learning
abstract
Optical Chemical Structure Recognition (OCSR) plays a pivotal role in modern chemical informatics, enabling the automated conversion of chemical structure images from scientific literature, patents, and educational materials into machine-readable molecular representations. This capability is essential for large-scale chemical data mining, drug discovery pipelines, and Large Language Model (LLM) applications in related domains. However, existing OCSR systems face significant challenges in accurately recognizing stereochemical information due to the subtle visual cues that distinguish stereoisomers, such as wedge and dash bonds, ring conformations, and spatial arrangements. To address these challenges, we propose MolSight, a comprehensive learning framework for OCSR that employs a three-stage training paradigm. In the first stage, we conduct pre-training on large-scale but noisy datasets to endow the model with fundamental perception capabilities for chemical structure images. In the second stage, we perform multi-granularity fine-tuning using datasets with richer supervisory signals, systematically exploring how auxiliary tasks—specifically chemical bond classification and atom localization—contribute to molecular formula recognition. Finally, we employ reinforcement learning for post-training optimization and introduce a novel stereochemical structure dataset. Remarkably, we find that even with MolSight's relatively compact parameter size, the Group Relative Policy Optimization (GRPO) algorithm can further enhance the model's performance on stereomolecular. Through extensive experiments across diverse datasets, our results demonstrate that MolSight achieves state-of-the-art performance in (stereo)chemical optical structure recognition.
Xinggang Wang, Bin Feng 0001, Wenyu Liu 0001
AAAI4
2026 LENS: Learning to Segment Anything with Unified Reinforced Reasoning
abstract
Text-prompted image segmentation enables fine-grained visual understanding and is critical for applications such as human-computer interaction and robotics. However, existing supervised fine-tuning methods typically ignore explicit chain-of-thought (CoT) reasoning at test time, which limits their ability to generalize to unseen prompts and domains. To address this issue, we introduce LENS, a scalable reinforcement-learning framework that jointly optimizes the reasoning process and segmentation in an end-to-end manner. We propose unified reinforcement-learning rewards that span sentence-, box-, and segment-level cues, encouraging the model to generate informative CoT rationales while refining mask quality. Using a publicly available 3-billion-parameter vision–language model, i.e., Qwen2.5-VL-3B-Instruct, LENS achieves an average cIoU of 81.2% on the RefCOCO, RefCOCO+, and RefCOCOg benchmarks, outperforming the strong fine-tuned method, i.e., GLaMM, by up to 5.6%. These results demonstrate that RL-driven CoT reasoning significantly enhances text-prompted segmentation and offers a practical path toward more generalizable Segment Anything models (SAM).
Lianghui Zhu, Bin Ouyang, Tianheng Cheng, Haocheng Shen, Longjin Ran, Xiaoxin Chen 0001, Li Yu 0003, Wenyu Liu 0001, Xinggang Wang
AAAI10
2026 Turbo-VAED: Fast and Stable Transfer of Video-VAEs to Mobile Devices
abstract
There is a growing demand for deploying large generative AI models on mobile devices. For recent popular video generative models, however, the Variational AutoEncoder (VAE) represents one of the major computational bottlenecks. Both large parameter sizes and mismatched kernels cause out-of-memory errors or extremely slow inference on mobile devices. To address this, we propose a low-cost solution that efficiently transfers widely used video VAEs to mobile devices. (1) We analyze redundancy in existing VAE architectures and get empirical design insights. By integrating 3D depthwise separable convolutions into our model, we significantly reduce the number of parameters. (2) We observe that the upsampling techniques in mainstream video VAEs are poorly suited to mobile hardware and form the main bottleneck. In response, we propose a decoupled 3D pixel shuffle scheme that slashes end-to-end delay. Building upon these, we develop a universal mobile-oriented VAE decoder, Turbo-VAED. (3) We propose an efficient VAE decoder training method. Since only the decoder is used during deployment, we distill it to Turbo-VAED instead of retraining the full VAE, enabling fast mobile adaptation with minimal performance loss. To our knowledge, our method enables real-time 720p video VAE decoding on mobile devices for the first time. This approach is widely applicable to most video VAEs. When integrated into four representative models, with training cost as low as $95, it accelerates original VAEs by up to 84.5× at 720p resolution on GPUs, uses as low as 17.5% of original parameter count, and retains 96.9% of the original reconstruction quality. Compared to mobile-optimized VAEs, Turbo-VAED achieves a 2.9× speedup in FPS and better reconstruction quality on the iPhone 16 Pro.
Ya Zou, Jingfeng Yao, Shuai Zhang 0050, Wenyu Liu 0001, Xinggang Wang
AAAI5
2026 Better early detector for high-performance detection transformer
Bin Hu 0020, Bencheng Liao, Jiyang Qi, Shusheng Yang, Wenyu Liu 0001
Image Vis. Comput.5
2026 EVF-SAM: Early Vision-Language Fusion for text-prompted Segment Anything Model
Tianheng Cheng, Lianghui Zhu, Lei Liu 0049, Longjin Ran, Xiaoxin Chen 0001, Wenyu Liu 0001, Xinggang Wang
Image Vis. Comput.8
2026 Fourier-based adaptive counterfactual intervention for object re-identification
Haijun Xiong, Bin Feng 0001, Bang Wang 0001, Xinggang Wang, Wenyu Liu 0001
Neural Networks5
2026 SceneVTG++: Controllable Multilingual Visual Text Generation in the Wild
abstract
Generating visual text in natural scene images is a challenging task with many unsolved problems. Different from generating text on artificially designed images (such as posters, covers, and cartoons), existing methods for natural scene visual text generation still have significant deficiencies: methods based on rendering engines rely on manually crafted rules, which struggle to adapt to diverse backgrounds and leave obvious artificial traces, while their text layouts may be placed in unreasonable areas (e.g., sky or ground) and text content is semantically disconnected from the scene; diffusion model-based methods, on the other hand, face difficulties in generating small characters, depend on manually designed prompts to ensure reasonable layout and content, fail to generate text at precise locations, and cannot effectively control text attributes (e.g., font and color). In this paper, we propose a two-stage method named SceneVTG++ to address these issues. SceneVTG++ comprises two core components: a Text Layout and Content Generator (TLCG) and a Controllable Local Text Diffusion (CLTD). The former leverages the world knowledge and visual reasoning capabilities of multimodal large language models to identify reasonable text areas and recommend scene-relevant text content based on natural scene background images; the latter generates controllable multilingual text using a diffusion model, ensuring alignment with the outputs of TLCG. Through extensive experiments, we verified the effectiveness of both TLCG and CLTD, and demonstrated that SceneVTG++ achieves state-of-the-art performance in natural scene visual text generation. Additionally, the images generated by SceneVTG++ exhibit superior utility for training natural scene optical character recognition (OCR) tasks, including text detection and text recognition. Codes and datasets will be made publicly available.
Jiawei Liu 0006, Feiyu Gao, Zhibo Yang 0003, Peng Wang 0028, Junyang Lin, Xinggang Wang, Wenyu Liu 0001
IEEE Trans. Circuits Syst. Video Technol.8
2026 WeakTr: Exploring Plain Vision Transformer for Weakly-Supervised Semantic Segmentation
abstract
Transformer has been very successful in various computer vision tasks and understanding the working mechanism of transformer is important. As touchstones, weakly-supervised semantic segmentation (WSSS) and class activation map (CAM) are useful tasks for analyzing vision transformers (ViT). Based on the plain ViT pre-trained with ImageNet classification, we find that multi-layer, multi-head self-attention maps can provide rich and diverse information for weakly-supervised semantic segmentation and CAM generation, e.g., different attention heads of ViT focus on different image areas and object categories. Thus we propose a novel method to end-to-end estimate the importance of attention heads, where the self-attention maps are adaptively fused for high-quality CAM results that tend to have more complete objects. Besides, we propose a ViT-based gradient clipping decoder for online retraining with the CAM results efficiently and effectively. Furthermore, the gradient clipping decoder can make good use of the knowledge in large-scale pre-trained ViT and has a scalable ability. The proposed plain Transformer-based Weakly-supervised learning method (WeakTr) obtains the superior WSSS performance on standard benchmarks, i.e., 78.5% mIoU on the $val$ set of PASCAL VOC 2012 and 51.1% mIoU on the $val$ set of COCO 2014. Source code and checkpoints are available at https://github.com/hustvl/WeakTr.
Lianghui Zhu, Yingyue Li, Jiemin Fang, Yan Liu 0069, Xin Hao, Wenyu Liu 0001, Xinggang Wang
IEEE Trans. Image Process.6
2025 GaraMoSt: Parallel Multi-Granularity Motion and Structural Modeling for Efficient Multi-Frame Interpolation in DSA Images
abstract
The rapid and accurate direct multi-frame interpolation method for Digital Subtraction Angiography (DSA) images is crucial for reducing radiation and providing real-time assistance to physicians for precise diagnostics and treatment. DSA images contain complex vascular structures and various motions. Applying natural scene Video Frame Interpolation (VFI) methods results in motion artifacts, structural dissipation, and blurriness. Recently, MoSt-DSA has specifically addressed these issues for the first time and achieved SOTA results. However, MoSt-DSA's focus on real-time performance leads to insufficient suppression of high-frequency noise and incomplete filtering of low-frequency noise in the generated images. To address these issues within the same computational time scale, we propose GaraMoSt. Specifically, we optimize the network pipeline with a parallel design and propose a module named MG-MSFE. MG-MSFE extracts frame-relative motion and structural features at various granularities in a fully convolutional parallel manner and supports independent, flexible adjustment of context-aware granularity at different scales, thus enhancing computational efficiency and accuracy. Extensive experiments demonstrate that GaraMoSt achieves the SOTA performance in accuracy, robustness, visual effects, and noise suppression, comprehensively surpassing MoSt-DSA and other natural scene VFI methods.
Huangxuan Zhao, Wenyu Liu 0001, Xinggang Wang
AAAI3
2025 GaussTR: Foundation Model-Aligned Gaussian Transformer for Self-Supervised 3D Spatial Understanding
abstract
3D Semantic Occupancy Prediction is fundamental for spatial understanding, yet existing approaches face challenges in scalability and generalization due to their reliance on extensive labeled data and computationally intensive voxel-wise representations. In this paper, we introduce GaussTR, a novel Gaussian-based TRansformer framework that unifies sparse 3D modeling with foundation model alignment through Gaussian representations to advance 3D spatial understanding. GaussTR predicts sparse sets of Gaussians in a feed-forward manner to represent 3D scenes. By splatting the Gaussians into 2D views and aligning the rendered features with foundation models, GaussTR facilitates self-supervised 3D representation learning and enables open-vocabulary semantic occupancy prediction without requiring explicit annotations. Empirical experiments on the Occ3D-nuScenes dataset demonstrate GaussTR’s state-of-the-art zero-shot performance of 12.27 mIoU, along with a 40% reduction in training time. These results highlight the efficacy of GaussTR for scalable and holistic 3D spatial understanding, with promising implications in autonomous driving and embodied agents. The code is available at https://github.com/hustvl/GaussTR.
Haoyi Jiang, Tianheng Cheng, Zhizhong Su, Wenyu Liu 0001, Xinggang Wang
CVPR7
2025 Mask-Adapter: The Devil is in the Masks for Open-Vocabulary Segmentation
abstract
Recent open-vocabulary segmentation methods adopt mask generators to predict segmentation masks and leverage pretrained vision-language models, e.g., CLIP, to classify these masks via mask pooling. Although these approaches show promising results, it is counterintuitive that accurate masks often fail to yield accurate classification results through pooling CLIP image embeddings within the mask regions. In this paper, we reveal the performance limitations of mask pooling and introduce Mask-Adapter, a simple yet effective method to address these challenges in open-vocabulary segmentation. Compared to directly using proposal masks, our proposed Mask-Adapter extracts semantic activation maps from proposal masks, providing richer contextual information and ensuring alignment between masks and CLIP. Additionally, we propose a mask consistency loss that encourages proposal masks with similar IoUs to obtain similar CLIP embeddings to enhance models’ robustness to varying predicted masks. Mask-Adapter integrates seamlessly into open-vocabulary segmentation methods based on mask pooling in a plug-and-play manner, delivering more accurate classification results. Extensive experiments across several zero-shot benchmarks demonstrate significant performance gains for the proposed Mask-Adapter on several well-established methods. Notably, Mask-Adapter also extends effectively to SAM and achieves impressive results on several open-vocabulary segmentation datasets. Code and models are available at https://github.com/hustvl/MaskAdapter.
Yongkang Li 0005, Tianheng Cheng, Bin Feng 0001, Wenyu Liu 0001, Xinggang Wang
CVPR4
2025 GroundingSuite: Measuring Complex Multi-Granular Pixel Grounding
abstract
Pixel grounding, encompassing tasks such as Referring Expression Segmentation (RES), has garnered considerable attention due to its immense potential for bridging the gap between vision and language modalities. However, advancements in this domain are currently constrained by limitations inherent in existing datasets, including limited object categories, insufficient textual diversity, and a scarcity of high-quality annotations. To mitigate these limitations, we introduce GroundingSuite, which comprises: (1) an automated data annotation framework leveraging multiple Vision-Language Model (VLM) agents; (2) a large-scale training dataset encompassing 9.56 million diverse referring expressions and their corresponding segmentations; and (3) a meticulously curated evaluation benchmark consisting of 3,800 images. The GroundingSuite training dataset facilitates substantial performance improvements, enabling models trained on it to achieve state-of-the-art results. Specifically, a cIoU of 68.9 on gRefCOCO and a gIoU of 55.3 on RefCOCOm. Moreover, the GroundingSuite annotation framework demonstrates superior efficiency compared to the current leading data annotation method, i.e., $4.5 \times$ faster than GLaMM.
Lianghui Zhu, Tianheng Cheng, Lei Liu 0049, Longjin Ran, Xiaoxin Chen 0001, Wenyu Liu 0001, Xinggang Wang
ICCV9
2025 MaTVLM: Hybrid Mamba-Transformer for Efficient Vision-Language Modeling
Yingyue Li, Bencheng Liao, Wenyu Liu 0001, Xinggang Wang
ICCV3
2025 ControlAR: Controllable Image Generation with Autoregressive Models
abstract
Autoregressive (AR) models have reformulated image generation as next-token prediction, demonstrating remarkable potential and emerging as strong competitors to diffusion models. However, control-to-image generation, akin to ControlNet, remains largely unexplored within AR models. Although a natural approach, inspired by advancements in Large Language Models, is to tokenize control images into tokens and prefill them into the autoregressive model before decoding image tokens, it still falls short in generation quality compared to ControlNet and suffers from inefficiency. To this end, we introduce ControlAR, an efficient and effective framework for integrating spatial controls into autoregressive image generation models. Firstly, we explore control encoding for AR models and propose a lightweight control encoder to transform spatial inputs (e.g., canny edges or depth maps) into control tokens. Then ControlAR exploits the conditional decoding method to generate the next image token conditioned on the per-token fusion between control and image tokens, similar to positional encodings. Compared to prefilling tokens, using conditional decoding significantly strengthens the control capability of AR models but also maintains the model efficiency. Furthermore, the proposed ControlAR surprisingly empowers AR models with arbitrary-resolution image generation via conditional decoding and specific controls. Extensive experiments can demonstrate the controllability of the proposed ControlAR for the autoregressive control-to-image generation across diverse inputs, including edges, depths, and segmentation masks. Furthermore, both quantitative and qualitative results indicate that ControlAR surpasses previous state-of-the-art controllable diffusion models, e.g., ControlNet++.
Zongming Li, Tianheng Cheng, Shoufa Chen, Peize Sun, Haocheng Shen, Longjin Ran, Xiaoxin Chen 0001, Wenyu Liu 0001, Xinggang Wang
ICLR8
2025 STP4D: Spatio-Temporal-Prompt Consistent Modeling for Text-To-4D Gaussian Splatting
abstract
Text-To-4D generation is rapidly developing and widely applied in various scenarios. However, existing methods often fail to incorporate adequate spatio-temporal modeling and prompt alignment within a unified framework, resulting in temporal inconsistencies, geometric distortions, or low-quality 4D content that deviates from the provided texts. Therefore, we propose STP4D, a novel approach that aims to integrate comprehensive spatio-temporal-prompt consistency modeling for high-quality text-to-4D generation. Specifically, STP4D employs three carefully designed modules: Time-Varying Prompt Embedding, Geometric Information Enhancement, and Temporal Extension Deformation, which collaborate to accomplish this goal. Furthermore, STP4D is among the first methods to exploit the Diffusion model to generate 4D Gaussians, combining the fine-grained modeling capabilities and the real-time rendering process of 4DGS with the rapid inference speed of the Diffusion model. Extensive experiments demonstrate that STP4D excels in generating high-fidelity 4D content with exceptional efficiency (approximately 4.6s per asset), surpassing existing methods in both quality and speed.
Yunze Deng, Haijun Xiong, Bin Feng 0001, Xinggang Wang, Wenyu Liu 0001
ICME5
2025 Dynamic 2D Gaussians: Geometrically Accurate Radiance Fields for Dynamic Objects
abstract
Reconstructing objects and extracting high-quality surfaces play a vital role in the real world. Current 4D representations show the ability to render high-quality novel views for dynamic objects, but cannot reconstruct high-quality meshes due to their implicit or geometrically inaccurate representations. In this paper, we propose a novel representation that can reconstruct accurate meshes from sparse image input, named Dynamic 2D Gaussians (D-2DGS). We adopt 2D Gaussians for basic geometry representation and use sparse-controlled points to capture the 2D Gaussian's deformation. By extracting the object mask from the rendered high-quality image and masking the rendered depth map, we remove floaters that are prone to occur during reconstruction and can extract high-quality dynamic mesh sequences of dynamic objects. Experiments demonstrate that our D-2DGS is outstanding in reconstructing detailed and smooth high-quality meshes from sparse inputs. The code is available at https://github.com/hustvl/Dynamic-2DGS.
Shuai Zhang 0050, Guanjun Wu, Zhoufeng Xie, Xinggang Wang, Bin Feng 0001, Wenyu Liu 0001
ACM Multimedia6
2025 RAD: Training an End-to-End Driving Policy via Large-Scale 3DGS-based Reinforcement Learning
abstract
Existing end-to-end autonomous driving (AD) algorithms typically follow the Imitation Learning (IL) paradigm, which faces challenges such as causal confusion and an open-loop gap. In this work, we propose RAD, a 3DGS-based closed-loop Reinforcement Learning (RL) framework for end-to-end Autonomous Driving. By leveraging 3DGS techniques, we construct a photorealistic digital replica of the real physical world, enabling the AD policy to extensively explore the state space and learn to handle out-of-distribution scenarios through large-scale trial and error. To enhance safety, we design specialized rewards to guide the policy in effectively responding to safety-critical events and understanding real-world causal relationships. To better align with human driving behavior, we incorporate IL into RL training as a regularization term. We introduce a closed-loop evaluation benchmark consisting of diverse, previously unseen 3DGS environments. Compared to IL-based methods, RAD achieves stronger performance in most closed-loop metrics, particularly exhibiting a 3× lower collision rate. Abundant closed-loop results are presented in the supplementary material. Code is available at https://github.com/hustvl/RAD for facilitating future research.
Shaoyu Chen, Bo Jiang 0011, Bencheng Liao, Yiang Shi, Yuechuan Pu, Xinbang Zhang, Wenyu Liu 0001, Qian Zhang 0001, Xinggang Wang
NeurIPS12
2025 Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency
abstract
We present Genesis, a unified world model for joint generation of multi-view driving videos and LiDAR sequences with spatio-temporal and cross-modal consistency. Genesis employs a two-stage architecture that integrates a DiT-based video diffusion model with 3D-VAE encoding, and a BEV-represented LiDAR generator with NeRF-based rendering and adaptive sampling. Both modalities are directly coupled through a shared condition input, enabling coherent evolution across visual and geometric domains. To guide the generation with structured semantics, we introduce DataCrafter, a captioning module built on vision-language models that provides scene-level and instance-level captions. Extensive experiments on the nuScenes benchmark demonstrate that Genesis achieves state-of-the-art performance across video and LiDAR metrics (FVD 16.95, FID 4.24, Chamfer 0.611), and benefits downstream tasks including segmentation and 3D detection, validating the semantic fidelity and practical utility of the synthetic data.
Zhanqian Wu, Kaixin Xiong, Gangwei Xu, Shaoqing Xu, Hangjun Ye, Wenyu Liu 0001, Xinggang Wang
NeurIPS12
2025 MapTRv2: An End-to-End Framework for Online Vectorized HD Map Construction
Bencheng Liao, Shaoyu Chen, Yunchi Zhang, Bo Jiang 0011, Qian Zhang 0009, Wenyu Liu 0001, Chang Huang, Xinggang Wang
Int. J. Comput. Vis.6
2025 Skim then Focus: Integrating Contextual and Fine-grained Views for Repetitive Action Counting
Zhengqi Zhao, Xiaohu Huang, Hao Zhou 0039, Errui Ding, Jingdong Wang 0001, Xinggang Wang, Wenyu Liu 0001, Bin Feng 0001
Int. J. Comput. Vis.8
2025 MIM4D: Masked Modeling with Multi-View Video for Autonomous Driving Representation Learning
Jialv Zou, Bencheng Liao, Wenyu Liu 0001, Xinggang Wang
Int. J. Comput. Vis.4
2025 Warped convolutional neural networks for large homography transformation with psl(3) algebra
Xinrui Zhan, Wenyu Liu 0001, Risheng Yu, Jianke Zhu, Yang Li 0041
Neurocomputing2
2025 PolarDETR: Polar Parametrization for vision-based surround-view 3D detection
Shaoyu Chen, Xinggang Wang, Tianheng Cheng, Qian Zhang 0009, Chang Huang, Wenyu Liu 0001
Image Vis. Comput.6
2025 Dynamic feature extraction and histopathology domain shift alignment for mitosis detection
Jiangxiao Han, Shikang Wang, Lianjun Wu, Wenyu Liu 0001
Image Vis. Comput.4
2025 MambaGait: Gait recognition approach combining explicit representation and implicit state space model
Haijun Xiong, Bin Feng 0001, Bang Wang 0001, Xinggang Wang, Wenyu Liu 0001
Image Vis. Comput.5
2025 Cross-layer attentive feature upsampling for low-latency semantic segmentation
Tianheng Cheng, Xinggang Wang, Junchao Liao, Wenyu Liu 0001
Mach. Vis. Appl.4
2025 Personvit: large-scale self-supervised vision transformer for person re-identification
Bin Hu 0020, Xinggang Wang, Wenyu Liu 0001
Mach. Vis. Appl.3
2025 Partial Scene Text Retrieval
abstract
The task of partial scene text retrieval involves localizing and searching for text instances that are the same or similar to a given query text from an image gallery. However, existing methods can only handle text-line instances, leaving the problem of searching for partial patches within these text-line instances unsolved due to a lack of patch annotations in the training data. To address this issue, we propose a network that can simultaneously retrieve both text-line instances and their partial patches. Our method embeds the two types of data (query text and scene text instances) into a shared feature space and measures their cross-modal similarities. To handle partial patches, our proposed approach adopts a Multiple Instance Learning (MIL) approach to learn their similarities with query text, without requiring extra annotations. However, constructing bags, which is a standard step of conventional MIL approaches, can introduce numerous noisy samples for training, and lower inference speed. To address this issue, we propose a Ranking MIL (RankMIL) approach to adaptively filter those noisy samples. Additionally, we present a Dynamic Partial Match Algorithm (DPMA) that can directly search for the target partial patch from a text-line instance during the inference stage, without requiring bags. This greatly improves the search efficiency and the performance of retrieving partial patches. We evaluate the proposed method on both English and Chinese datasets in two tasks: retrieving text-line instances and partial patches. For English text retrieval, our method outperforms state-of-the-art approaches by 8.04% mAP and 12.71% mAP on average, respectively, among three datasets for the two tasks. For Chinese text retrieval, our approach surpasses state-of-the-art approaches by 24.45% mAP and 38.06% mAP on average, respectively, among three datasets for the two tasks. The source code and dataset are available at https://github.com/lanfeng4659/PSTR.
Hao Wang 0207, Minghui Liao, Zhouyi Xie, Wenyu Liu 0001, Xiang Bai
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 SparseTrack: Multi-Object Tracking by Performing Scene Decomposition Based on Pseudo-Depth
abstract
Exploring robust and efficient association methods has always been an important issue in multi-object tracking (MOT). Although existing tracking methods have achieved impressive performance, congestion and frequent occlusions still pose challenging problems in multi-object tracking. We reveal that performing sparse decomposition on dense scenes is a crucial step to enhance the performance of associating occluded targets. To this end, we propose a pseudo-depth estimation method for obtaining the relative depth of targets from 2D images. Secondly, we design a depth cascading matching (DCM) algorithm, which can use the obtained depth information to convert a dense target set into multiple sparse target subsets and perform data association on these sparse target subsets in order from near to far. By integrating the pseudo-depth method and the DCM strategy into the data association process, we propose a new tracker, called SparseTrack. SparseTrack provides a new perspective for solving the challenging crowded scene MOT problem. Only using IoU matching, SparseTrack achieves comparable performance with the state-of-the-art (SOTA) methods on the MOT17 and MOT20 benchmarks. Code and models are publicly available at https://github.com/hustvl/SparseTrack.
Xinggang Wang, Cheng Wang 0048, Wenyu Liu 0001, Xiang Bai
IEEE Trans. Circuits Syst. Video Technol.4
2025 TOGS: Gaussian Splatting With Temporal Opacity Offset for Real-Time 4D DSA Rendering
abstract
Four-dimensional Digital Subtraction Angiography (4D DSA) is a medical imaging technique that provides a series of 2D images captured at different stages and angles during the process of contrast agent filling blood vessels. It plays a significant role in the diagnosis of cerebrovascular diseases. Improving the rendering quality and speed under sparse sampling is important for observing the status and location of lesions. The current methods exhibit inadequate rendering quality in sparse views and suffer from slow rendering speed. To overcome these limitations, we propose TOGS, a Gaussian splatting method with opacity offset over time, which can effectively improve the rendering quality and speed of 4D DSA. We introduce an opacity offset table for each Gaussian to model the opacity offsets of the Gaussian, using these opacity-varying Gaussians to model the temporal variations in the radiance of the contrast agent. By interpolating the opacity offset table, the opacity variation of the Gaussian at different time points can be determined. This enables us to render the 2D DSA image at that specific moment. Additionally, we introduced a Smooth loss term in the loss function to mitigate overfitting issues that may arise in the model when dealing with sparse view scenarios. During the training phase, we randomly prune Gaussians, thereby reducing the storage overhead of the model. The experimental results demonstrate that compared to previous methods, this model achieves state-of-the-art render quality under the same number of training views. Additionally, it enables real-time rendering while maintaining low storage overhead.
Shuai Zhang 0050, Huangxuan Zhao, Zhenghong Zhou, Guanjun Wu, Chuansheng Zheng, Xinggang Wang, Wenyu Liu 0001
IEEE J. Biomed. Health Informatics7
2024 Fast High Dynamic Range Radiance Fields for Dynamic Scenes
abstract
Neural Radiances Fields (NeRF) and their extensions have shown great success in representing 3D scenes and synthesizing novel-view images. However, most NeRF methods take in low-dynamic-range (LDR) images, which may lose details, especially with nonuniform illumination. Some previous NeRF methods attempt to introduce high-dynamic-range (HDR) techniques but mainly target static scenes. To extend HDR NeRF methods to wider applications, we propose a dynamic HDR NeRF framework, named HDR-HexPlane, which can learn 3D scenes from dynamic 2D images captured with various exposures. A learnable exposure mapping function is constructed to obtain adaptive exposure values for each image. Based on the monotonically increasing prior, a camera response function is designed for stable learning. With the proposed model, high- quality novel-view images at any time point can be rendered with any desired exposure. We further construct a dataset containing multiple dynamic scenes captured with diverse exposures for evaluation. All the datasets and code are available at https://guanjunwu.github.io/HDR-HexPlane/.
Guanjun Wu, Taoran Yi, Jiemin Fang, Wenyu Liu 0001, Xinggang Wang
3DV4
2024 MobileInst: Video Instance Segmentation on the Mobile
abstract
Video instance segmentation on mobile devices is an important yet very challenging edge AI problem. It mainly suffers from (1) heavy computation and memory costs for frame-by-frame pixel-level instance perception and (2) complicated heuristics for tracking objects. To address these issues, we present MobileInst, a lightweight and mobile-friendly framework for video instance segmentation on mobile devices. Firstly, MobileInst adopts a mobile vision transformer to extract multi-level semantic features and presents an efficient query-based dual-transformer instance decoder for mask kernels and a semantic-enhanced mask decoder to generate instance segmentation per frame. Secondly, MobileInst exploits simple yet effective kernel reuse and kernel association to track objects for video instance segmentation. Further, we propose temporal query passing to enhance the tracking ability for kernels. We conduct experiments on COCO and YouTube-VIS datasets to demonstrate the superiority of MobileInst and evaluate the inference latency on one single CPU core of the Snapdragon 778G Mobile Platform, without other methods of acceleration. On the COCO dataset, MobileInst achieves 31.2 mask AP and 433 ms on the mobile CPU, which reduces the latency by 50% compared to the previous SOTA. For video instance segmentation, MobileInst achieves 35.0 AP and 30.1 AP on YouTube-VIS 2019 & 2021.
Renhong Zhang, Tianheng Cheng, Shusheng Yang, Haoyi Jiang, Shuai Zhang 0050, Jiancheng Lyu, Xin Li 0034, Xiaowen Ying, Dashan Gao 0001, Wenyu Liu 0001, Xinggang Wang
AAAI10
2024 YOLO-World: Real-Time Open-Vocabulary Object Detection
abstract
The You Only Look Once (YOLO) series of detectors have established themselves as efficient and practical tools. However, their reliance on predefined and trained object categories limits their applicability in open scenarios. Addressing this limitation, we introduce YOLO-World, an innovative approach that enhances YOLO with open-vocabulary detection capabilities through vision-language modeling and pre-training on large-scale datasets. Specifically, we propose a new Re-parameterizable Vision-Language Path Aggregation Network (RepVL-PAN) and region-text contrastive loss to facilitate the interaction between visual and linguistic information. Our method excels in detecting a wide range of objects in a zero-shot manner with high efficiency. On the challenging LVIS dataset, YOLO-World achieves 35.4 AP with 52.0 FPS on V100, which outperforms many state-of-the-art methods in terms of both accuracy and speed. Furthermore, the finetuned YOLO-World achieves remarkable performance on several downstream tasks, including object detection and open-vocabulary instance segmentation. Code and models are available at: https://github.com/AILab-eve/YOLO-World.
Tianheng Cheng, Lin Song 0002, Yixiao Ge, Wenyu Liu 0001, Xinggang Wang, Ying Shan
CVPR4
2024 Symphonize 3D Semantic Scene Completion with Contextual Instance Queries
abstract
3D Semantic Scene Completion (SSC) has emerged as a nascent and pivotal undertaking in autonomous driving, aiming to predict the voxel occupancy within volumetric scenes. However, prevailing methodologies primarily focus on voxel-wise feature aggregation, while neglecting instance semantics and scene context. In this paper, we present a novel paradigm termed Symphonies (Scene-from-Insts), that delves into the integration of instance queries to orchestrate 2D-to-3D reconstruction and 3D scene modeling. Leveraging our proposed Serial Instance-Propagated Attentions, Symphonies dynamically encodes instance-centric semantics, facilitating intricate interactions between the image and volumetric domains. Simultaneously, Symphonies fosters holistic scene comprehension by capturing context through the efficient fusion of instance queries, alleviating geometric ambiguities such as occlusion and perspective errors through contextual scene reasoning. Experimental results demonstrate that Symphonies achieves state-of-the-art performance on the chal-lenging SemanticKITTI and SSCBench-KITTI-360 benchmarks, yielding remarkable mIoU scores of 15.04 and 18.58, respectively. These results showcase the promising advancements of our paradigm. The code for our method is available at https://github.com/hustvl/Symphonies.
Haoyi Jiang, Tianheng Cheng, Naiyu Gao, Wenyu Liu 0001, Xinggang Wang
CVPR6
2024 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering
abstract
Representing and rendering dynamic scenes has been an important but challenging task. Especially, to accurately model complex motions, high efficiency is usually hard to guarantee. To achieve real-time dynamic scene rendering while also enjoying high training and storage efficiency, we propose 4D Gaussian Splatting (4D-GS) as a holistic representation for dynamic scenes rather than applying 3D-GS for each individual frame. In 4D-GS, a novel explicit representation containing both 3D Gaussians and 4D neural voxels is proposed. A decomposed neural voxel encoding algorithm inspired by HexPlane is proposed to efficiently build Gaussian features from 4D neural voxels and then a lightweight MLP is applied to predict Gaussian deformations at novel timestamps. Our 4D-GS method achieves real-time rendering under high resolutions, 82 FPS at an 800x800 resolution on an RTX 3090 GPU while maintaining comparable or better quality than previous state- of-the-art methods. More demos and code are available at https://guanjunwu.github.io/4dgs/.
Guanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie, Xiaopeng Zhang 0008, Wei Wei 0002, Wenyu Liu 0001, Qi Tian 0001, Xinggang Wang
CVPR7
2024 GaussianDreamer: Fast Generation from Text to 3D Gaussians by Bridging 2D and 3D Diffusion Models
abstract
In recent times, the generation of 3D assets from text prompts has shown impressive results. Both 2D and 3D diffusion models can help generate decent 3D objects based on prompts. 3D diffusion models have good 3D consistency, but their quality and generalization are limited as trainable 3D data is expensive and hard to obtain. 2D diffusion models enjoy strong abilities of generalization and fine generation, but 3D consistency is hard to guarantee. This paper attempts to bridge the power from the two types of diffusion models via the recent explicit and efficient 3D Gaussian splatting representation. A fast 3D object gener-ation framework, named as GaussianDreamer, is proposed, where the 3D diffusion model provides priors for initial-ization and the 2D diffusion model enriches the geometry and appearance. Operations of noisy point growing and color perturbation are introduced to enhance the initialized Gaussians. Our GaussianDreamer can generate a high-quality 3D instance or 3D avatar within 15 minutes on one GPU, much faster than previous methods, while the generated instances can be directly rendered in real time. Demos and code are available at https://taoranyi.com/gaussiandreamer/.
Taoran Yi, Jiemin Fang, Junjie Wang 0012, Guanjun Wu, Lingxi Xie, Xiaopeng Zhang 0008, Wenyu Liu 0001, Qi Tian 0001, Xinggang Wang
CVPR7
2024 MoSt-DSA: Modeling Motion and Structural Interactions for Direct Multi-Frame Interpolation in DSA Images
abstract
Artificial intelligence has become a crucial tool for medical image analysis. As an advanced cerebral angiography technique, Digital Subtraction Angiography (DSA) poses a challenge where the radiation dose to humans is proportional to the image count. By reducing images and using AI interpolation instead, the radiation can be cut significantly. However, DSA images present more complex motion and structural features than natural scenes, making interpolation more challenging. We propose MoSt-DSA, the first work that uses deep learning for DSA frame interpolation. Unlike natural scene Video Frame Interpolation (VFI) methods that extract unclear or coarse-grained features, we devise a general module that models motion and structural context interactions between frames in an efficient full convolution manner by adjusting optimal context range and transforming contexts into linear functions. Benefiting from this, MoSt-DSA is also the first method that directly achieves any number of interpolations at any time steps with just one forward pass during both training and testing. We conduct extensive comparisons with 7 representative VFI models for interpolating 1 to 3 frames, MoSt-DSA demonstrates robust results across 470 DSA image sequences (each typically 152 images), with average SSIM over 0.93, average PSNR over 38 (standard deviations of less than 0.030 and 3.6, respectively), comprehensively achieving state-of-the-art performance in accuracy, speed, visual effect, and memory usage. Our code is available at https://github.com/ZyoungXu/MoSt-DSA.
Huangxuan Zhao, Ziwei Cui, Wenyu Liu 0001, Chuansheng Zheng, Xinggang Wang
ECAI4
2024 Lane Graph as Path: Continuity-Preserving Path-Wise Modeling for Online Lane Graph Construction
Bencheng Liao, Shaoyu Chen, Bo Jiang 0011, Tianheng Cheng, Qian Zhang 0009, Wenyu Liu 0001, Chang Huang, Xinggang Wang
ECCV (44)6
2024 Occupancy as Set of Points
Yiang Shi, Tianheng Cheng, Qian Zhang 0009, Wenyu Liu 0001, Xinggang Wang
ECCV (61)4
2024 Causality-Inspired Discriminative Feature Learning in Triple Domains for Gait Recognition
Haijun Xiong, Bin Feng 0001, Xinggang Wang, Wenyu Liu 0001
ECCV (54)4
2024 Visual Text Generation in the Wild
Jiawei Liu 0006, Feiyu Gao, Wenyu Liu 0001, Xinggang Wang, Peng Wang 0028, Fei Huang 0002, Cong Yao, Zhibo Yang 0003
ECCV (53)4
2024 Gaitgs: Temporal Feature Learning in Granularity And Span Dimension for Gait Recognition
abstract
Gait recognition, a growing field in biological recognition technology, utilizes distinct walking patterns for accurate individual identification. However, existing methods lack the incorporation of temporal information. To reach the full potential of gait recognition, we advocate for the consideration of temporal features at varying granularities and spans. This paper introduces a novel framework, GaitGS, which aggregates temporal features simultaneously in both granularity and span dimensions. Specifically, the Multi-Granularity Feature Extractor (MGFE) is designed to capture micro-motion and macro-motion information at fine and coarse levels respectively, while the Multi-Span Feature Extractor (MSFE) generates local and global temporal representations. Through extensive experiments on two datasets, our method demonstrates state-of-the-art performance, achieving Rank-1 accuracy of 98.2%, 96.5%, and 89.7% on CASIA-B under different conditions, and 97.6% on OU-MVLP. The source code will be available at https://github.com/Haijun-Xiong/GaitGS.
Haijun Xiong, Yunze Deng, Bin Feng 0001, Xinggang Wang, Wenyu Liu 0001
ICIP5
2024 Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
abstract
Recently the state space models (SSMs) with efficient hardware-aware designs, i.e., the Mamba deep learning model, have shown great potential for long sequence modeling. Meanwhile building efficient and generic vision backbones purely upon SSMs is an appealing direction. However, representing visual data is challenging for SSMs due to the position-sensitivity of visual data and the requirement of global context for visual understanding. In this paper, we show that the reliance on self-attention for visual representation learning is not necessary and propose a new generic vision backbone with bidirectional Mamba blocks (Vim), which marks the image sequences with position embeddings and compresses the visual representation with bidirectional state space models. On ImageNet classification, COCO object detection, and ADE20k semantic segmentation tasks, Vim achieves higher performance compared to well-established vision transformers like DeiT, while also demonstrating significantly improved computation & memory efficiency. For example, Vim is 2.8x faster than DeiT and saves 86.8% GPU memory when performing batch inference to extract features on images with a resolution of 1248x1248. The results demonstrate that Vim is capable of overcoming the computation & memory constraints on performing Transformer-style understanding for high-resolution images and it has great potential to be the next-generation backbone for vision foundation models.
Lianghui Zhu, Bencheng Liao, Qian Zhang 0009, Wenyu Liu 0001, Xinggang Wang
ICML5
2024 WeakSAM: Segment Anything Meets Weakly-supervised Instance-level Recognition
abstract
Weakly-supervised visual recognition using inexact supervision is a critical yet challenging learning problem. It significantly reduces human labeling costs and traditionally relies on multi-instance learning and pseudo-labeling. This paper introduces WeakSAM and solves the weakly-supervised object detection (WSOD) and segmentation by utilizing the pre-learned world knowledge contained in a vision foundation model, i.e., the Segment Anything Model (SAM). WeakSAM addresses two critical limitations in traditional WSOD retraining, i.e., pseudo ground truth (PGT) incompleteness and noisy PGT instances, through adaptive PGT generation and Region of Interest (RoI) drop regularization. It also addresses the SAM's shortcomings of requiring human prompts and category unawareness in object detection and segmentation. Our results indicate that WeakSAM significantly surpasses previous state-of-the-art methods in WSOD and WSIS benchmarks with large margins, i.e. average improvements of 7.4% and 8.5%, respectively.
Lianghui Zhu, Junwei Zhou 0003, Yan Liu 0069, Xin Hao, Wenyu Liu 0001, Xinggang Wang
ACM Multimedia5
2024 FasterDiT: Towards Faster Diffusion Transformers Training without Architecture Modification
abstract
Diffusion Transformers (DiT) have attracted significant attention in research. However, they suffer from a slow convergence rate. In this paper, we aim to accelerate DiT training without any architectural modification. We identify the following issues in the training process: firstly, certain training strategies do not consistently perform well across different data. Secondly, the effectiveness of supervision at specific timesteps is limited. In response, we propose the following contributions: (1) We introduce a new perspective for interpreting the failure of the strategies. Specifically, we slightly extend the definition of Signal-to-Noise Ratio (SNR) and suggest observing the Probability Density Function (PDF) of SNR to understand the essence of the data robustness of the strategy. (2) We conduct numerous experiments and report over one hundred experimental results to empirically summarize a unified accelerating strategy from the perspective of PDF. (3) We develop a new supervision method that further accelerates the training process of DiT. Based on them, we propose FasterDiT, an exceedingly simple and practicable design strategy. With few lines of code modifications, it achieves 2.30 FID on ImageNet at 256x256 resolution with 1000 iterations, which is comparable to DiT (2.27 FID) but 7 times faster in training.
Jingfeng Yao, Cheng Wang 0048, Wenyu Liu 0001, Xinggang Wang
NeurIPS3
2024 Stabilized activation scale estimation for precise Post-Training Quantization
Zhenyang Hao, Xinggang Wang, Jiawei Liu 0006, Zhihang Yuan, Wenyu Liu 0001
Neurocomputing6
2024 Eliminating and mining strategies for open-world object proposal
Cheng Wang 0048, Guoli Wang 0004, Qian Zhang 0009, Peng Guo 0001, Wenyu Liu 0001, Xinggang Wang
Neurocomputing5
2024 Video text tracking with transformer-based local search
Xingsheng Zhou, Cheng Wang 0048, Xinggang Wang, Wenyu Liu 0001
Neurocomputing4
2024 Learning accurate monocular 3D voxel representation via bilateral voxel transformer
Tianheng Cheng, Haoyi Jiang, Shaoyu Chen, Bencheng Liao, Qian Zhang 0009, Wenyu Liu 0001, Xinggang Wang
Image Vis. Comput.6
2024 High-performance mitosis detection using single-level feature and hybrid label assignment
Jiangxiao Han, Shikang Wang, Xianbo Deng, Wenyu Liu 0001
Image Vis. Comput.4
2024 Matte anything: Interactive natural image matting with segment anything model
Jingfeng Yao, Xinggang Wang, Lang Ye, Wenyu Liu 0001
Image Vis. Comput.4
2024 Transgaze: exploring plain vision transformers for gaze estimation
Lang Ye, Xinggang Wang, Jingfeng Yao, Wenyu Liu 0001
Mach. Vis. Appl.4
2024 OpenInst: A simple query-based method for open-world instance segmentation
Cheng Wang 0048, Guoli Wang 0004, Qian Zhang 0009, Peng Guo 0001, Wenyu Liu 0001, Xinggang Wang
Pattern Recognit.5
2024 Efficient Task-Specific Feature Re-Fusion for More Accurate Object Detection and Instance Segmentation
abstract
Feature pyramid representations have been widely adopted in the object detection literature for better handling of variations in scale, which provide abundant information from various spatial levels for classification and localization sub-tasks. We find that inter sub-task feature disentanglement and intra sub-task feature re-fusion are crucial for final prediction performance, but are hard to be achieved simultaneously considering the computational efficiency. We find this issue can be addressed by delicate module design. In this paper, we propose an Efficient Task-specific Feature Re-fusion (ETFR) module to mitigate the dilemma. ETFR disentangles inter sub-task features, reduces the output channels of multi-scale features based on their importance and re-fuses intra sub-task features via concatenation operation. As a plug-and-play module, ETFR can remarkably and consistently improve the well-established and highly-optimized object detection and instance segmentation methods, such as RetinaNet, FCOS, BlendMask and CondInst, with neglectable extra computation cost. Extensive experiments demonstrate that ETFR has good generalization ability on various changeling datasets, including COCO, LVIS and Cityscapes.
Cheng Wang 0048, Jiemin Fang, Peng Guo 0001, Rui Wu 0018, Xinggang Wang, Chang Huang, Wenyu Liu 0001
IEEE Trans. Circuits Syst. Video Technol.9
2023 BoxTeacher: Exploring High-Quality Pseudo Labels for Weakly Supervised Instance Segmentation
abstract
Labeling objects with pixel-wise segmentation requires a huge amount of human labor compared to bounding boxes. Most existing methods for weakly supervised instance segmentation focus on designing heuristic losses with priors from bounding boxes. While, we find that box-supervised methods can produce some fine segmentation masks and we wonder whether the detectors could learn from these fine masks while ignoring low-quality masks. To answer this question, we present BoxTeacher, an efficient and end-to-end training framework for high-performance weakly supervised instance segmentation, which leverages a sophisticated teacher to generate high-quality masks as pseudo labels. Considering the massive noisy masks hurt the training, we present a mask-aware confidence score to estimate the quality of pseudo masks, and propose the noiseaware pixel loss and noise-reduced affinity loss to adaptively optimize the student with pseudo masks. Extensive experiments can demonstrate effectiveness of the proposed BoxTeacher. Without bells and whistles, BoxTeacher remarkably achieves 35.0 mask AP and 36.5 mask AP with ResNet-50 and ResNet-101 respectively on the challenging COCO dataset, which outperforms the previous state-of-the-art methods by a significant margin and bridges the gap between box-supervised and mask-supervised methods.
Tianheng Cheng, Xinggang Wang, Shaoyu Chen, Qian Zhang 0009, Wenyu Liu 0001
CVPR5
2023 PD-Quant: Post-Training Quantization Based on Prediction Difference Metric
abstract
Post-training quantization (PTQ) is a neural network compression technique that converts a full-precision model into a quantized model using lower-precision data types. Although it can help reduce the size and computational cost of deep neural networks, it can also introduce quantization noise and reduce prediction accuracy, especially in extremely low-bit settings. How to determine the appropriate quantization parameters (e.g., scaling factors and rounding of weights) is the main problem facing now. Existing methods attempt to determine these parameters by minimize the distance between features before and after quantization, but such an approach only considers local information and may not result in the most optimal quantization parameters. We analyze this issue and propose PD-Quant, a method that addresses this limitation by considering global information. It determines the quantization parameters by using the information of differences between network prediction before and after quantization. In addition, PD-Quant can alleviate the overfitting problem in PTQ caused by the small number of calibration sets by adjusting the distribution of activations. Experiments show that PD-Quant leads to better quantization parameters and improves the prediction accuracy of quantized models, especially in low-bit settings. For example, PD-Quant pushes the accuracy of ResNet-18 up to 53.14% and RegNetX-600MF up to 40.67% in weight 2-bit activation 2-bit. The code is released at https://github.com/hustv1/PD-Quant.
Jiawei Liu 0006, Lin Niu, Zhihang Yuan, Xinggang Wang, Wenyu Liu 0001
CVPR6
2023 VAD: Vectorized Scene Representation for Efficient Autonomous Driving
abstract
Autonomous driving requires a comprehensive understanding of the surrounding environment for reliable trajectory planning. Previous works rely on dense rasterized scene representation (e.g., agent occupancy and semantic map) to perform planning, which is computationally intensive and misses the instance-level structure information. In this paper, we propose VAD, an end-to-end vectorized paradigm for autonomous driving, which models the driving scene as a fully vectorized representation. The proposed vectorized paradigm has two significant advantages. On one hand, VAD exploits the vectorized agent motion and map elements as explicit instance-level planning constraints which effectively improves planning safety. On the other hand, VAD runs much faster than previous end-to-end planning methods by getting rid of computation-intensive rasterized representation and hand-designed post-processing steps. VAD achieves state-of-the-art end-to-end planning performance on the nuScenes dataset, outperforming the previous best method by a large margin. Our base model, VAD-Base, greatly reduces the average collision rate by 29.0% and runs 2.5× faster. Besides, a lightweight variant, VAD-Tiny, greatly improves the inference speed (up to 9.3×) while achieving comparable planning performance. We believe the excellent performance and the high efficiency of VAD are critical for the real-world deployment of an autonomous driving system. Code and models are available at https://github.com/hustvl/VAD for facilitating future research.
Bo Jiang 0011, Shaoyu Chen, Bencheng Liao, Helong Zhou, Qian Zhang 0009, Wenyu Liu 0001, Chang Huang, Xinggang Wang
ICCV8
2023 Graph Contrastive Learning for Skeleton-based Action Recognition
Xiaohu Huang, Hao Zhou 0039, Jian Wang 0066, Haocheng Feng, Junyu Han, Errui Ding, Jingdong Wang 0001, Xinggang Wang, Wenyu Liu 0001, Bin Feng 0001
ICLR9
2023 MapTR: Structured Modeling and Learning for Online Vectorized HD Map Construction
Bencheng Liao, Shaoyu Chen, Xinggang Wang, Tianheng Cheng, Qian Zhang 0009, Wenyu Liu 0001, Chang Huang
ICLR6
2023 Circuit as Set of Points
abstract
As the size of circuit designs continues to grow rapidly, artificial intelligence technologies are being extensively used in Electronic Design Automation (EDA) to assist with circuit design. Placement and routing are the most time-consuming parts of the physical design process, and how to quickly evaluate the placement has become a hot research topic. Prior works either transformed circuit designs into images using hand-crafted methods and then used Convolutional Neural Networks (CNN) to extract features, which are limited by the quality of the hand-crafted methods and could not achieve end-to-end training, or treated the circuit design as a graph structure and used Graph Neural Networks (GNN) to extract features, which require time-consuming preprocessing. In our work, we propose a novel perspective for circuit design by treating circuit components as point clouds and using Transformer-based point cloud perception methods to extract features from the circuit. This approach enables direct feature extraction from raw data without any preprocessing, allows for end-to-end training, and results in high performance. Experimental results show that our method achieves state-of-the-art performance in congestion prediction tasks on both the CircuitNet and ISPD2015 datasets, as well as in design rule check (DRC) violation prediction tasks on the CircuitNet dataset. Our method establishes a bridge between the relatively mature point cloud perception methods and the fast-developing EDA algorithms, enabling us to leverage more collective intelligence to solve this task. To facilitate the research of open EDA design, source codes and pre-trained models are released at https://github.com/hustvl/circuitformer.
Jialv Zou, Xinggang Wang, Wenyu Liu 0001, Qian Zhang 0009, Chang Huang
NeurIPS4
2023 TinyDet: accurately detecting small objects within 1 GFLOPs
Shaoyu Chen, Tianheng Cheng, Jiemin Fang, Qian Zhang 0009, Wenyu Liu 0001, Xinggang Wang
Sci. China Inf. Sci.6
2023 What Makes for Hierarchical Vision Transformer?
abstract
Recent studies indicate that hierarchical Vision Transformer (ViT) with a macro architecture of interleaved non-overlapped window-based self-attention & shifted-window operation can achieve state-of-the-art performance in various visual recognition tasks, and challenges the ubiquitous convolutional neural networks (CNNs) using densely slid kernels. In most recently proposed hierarchical ViTs, self-attention is the de-facto standard for spatial information aggregation. In this paper, we question whether self-attention is the only choice for hierarchical ViT to attain strong performance, and study the effects of different kinds of cross-window communication methods. To this end, we replace self-attention layers with embarrassingly simple linear mapping layers, and the resulting proof-of-concept architecture termed TransLinear can achieve very strong performance in ImageNet-[Formula: see text] image recognition. Moreover, we find that TransLinear is able to leverage the ImageNet pre-trained weights and demonstrates competitive transfer learning properties on downstream dense prediction tasks such as object detection and instance segmentation. We also experiment with other alternatives to self-attention for content aggregation inside each non-overlapped window under different cross-window communication approaches. Our results reveal that the macro architecture, other than specific aggregation layers or cross-window communication mechanisms, is more responsible for hierarchical ViT's strong performance and is the real challenger to the ubiquitous CNN's dense sliding window paradigm.
Xinggang Wang, Rui Wu 0018, Wenyu Liu 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 CCNet: Criss-Cross Attention for Semantic Segmentation
abstract
Contextual information is vital in visual understanding problems, such as semantic segmentation and object detection. We propose a criss-cross network (CCNet) for obtaining full-image contextual information in a very effective and efficient way. Concretely, for each pixel, a novel criss-cross attention module harvests the contextual information of all the pixels on its criss-cross path. By taking a further recurrent operation, each pixel can finally capture the full-image dependencies. Besides, a category consistent loss is proposed to enforce the criss-cross attention module to produce more discriminative features. Overall, CCNet is with the following merits: 1) GPU memory friendly. Compared with the non-local block, the proposed recurrent criss-cross attention module requires 11× less GPU memory usage. 2) High computational efficiency. The recurrent criss-cross attention significantly reduces FLOPs by about 85 percent of the non-local block. 3) The state-of-the-art performance. We conduct extensive experiments on semantic segmentation benchmarks including Cityscapes, ADE20K, human parsing benchmark LIP, instance segmentation benchmark COCO, video segmentation benchmark CamVid. In particular, our CCNet achieves the mIoU scores of 81.9, 45.76 and 55.47 percent on the Cityscapes test set, the ADE20K validation set and the LIP validation set respectively, which are the new state-of-the-art results. The source codes are available at https://github.com/speedinghzl/CCNethttps://github.com/speedinghzl/CCNet.
Xinggang Wang, Yunchao Wei, Lichao Huang, Humphrey Shi, Wenyu Liu 0001, Thomas S. Huang
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 VoxelTrack: Multi-Person 3D Human Pose Estimation and Tracking in the Wild
abstract
We present VoxelTrack for multi-person 3D pose estimation and tracking from a few cameras which are separated by wide baselines. It employs a multi-branch network to jointly estimate 3D poses and re-identification (Re-ID) features for all people in the environment. In contrast to previous efforts which require to establish cross-view correspondence based on noisy 2D pose estimates, it directly estimates and tracks 3D poses from a 3D voxel-based representation constructed from multi-view images. We first discretize the 3D space by regular voxels and compute a feature vector for each voxel by averaging the body joint heatmaps that are inversely projected from all views. We estimate 3D poses from the voxel representation by predicting whether each voxel contains a particular body joint. Similarly, a Re-ID feature is computed for each voxel which is used to track the estimated 3D poses over time. The main advantage of the approach is that it avoids making any hard decisions based on individual images. The approach can robustly estimate and track 3D poses even when people are severely occluded in some cameras. It outperforms the state-of-the-art methods by a large margin on four public datasets including Shelf, Campus, Human3.6 M and CMU Panoptic.
Chunyu Wang 0001, Xinggang Wang, Wenyu Liu 0001, Wenjun Zeng 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Weakly-supervised semantic segmentation via online pseudo-mask correcting
Jiapei Feng, Xinggang Wang, Shanshan Ji, Wenyu Liu 0001
Pattern Recognit. Lett.5
2023 Condition-Adaptive Graph Convolution Learning for Skeleton-Based Gait Recognition
abstract
Graph convolutional networks have been widely applied in skeleton-based gait recognition. A key challenge in this task is to distinguish the individual walking styles of different subjects across various views. Existing state-of-the-art methods employ uniform convolutions to extract features from diverse sequences and ignore the effects of viewpoint changes. To overcome these limitations, we propose a condition-adaptive graph (CAG) convolution network that can dynamically adapt to the specific attributes of each skeleton sequence and the corresponding view angle. In contrast to using fixed weights for all joints and sequences, we introduce a joint-specific filter learning (JSFL) module in the CAG method, which produces sequence-adaptive filters at the joint level. The adaptive filters capture fine-grained patterns that are unique to each joint, enabling the extraction of diverse spatial-temporal information about body parts. Additionally, we design a view-adaptive topology learning (VATL) module that generates adaptive graph topologies. These graph topologies are used to correlate the joints adaptively according to the specific view conditions. Thus, CAG can simultaneously adjust to various walking styles and viewpoints. Experiments on the two most widely used datasets (i.e., CASIA-B and OU-MVLP) show that CAG surpasses all previous skeleton-based methods. Moreover, the recognition performance can be enhanced by simply combining CAG with appearance-based methods, demonstrating the ability of CAG to provide useful complementary information.
Xiaohu Huang, Xinggang Wang, Zhidianqiu Jin, Botao He, Bin Feng 0001, Wenyu Liu 0001
IEEE Trans. Image Process.7
2022 AziNorm: Exploiting the Radial Symmetry of Point Cloud for Azimuth-Normalized 3D Perception
abstract
Studying the inherent symmetry of data is of great importance in machine learning. Point cloud, the most important data format for 3D environmental perception, is naturally endowed with strong radial symmetry. In this work, we exploit this radial symmetry via a divide-and-conquer strategy to boost 3D perception performance and ease optimization. We propose Azimuth Normalization (AziNorm), which normalizes the point clouds along the radial direction and eliminates the variability brought by the difference of azimuth. AziNorm can be flexibly incorporated into most LiDAR-based perception methods. To validate its effectiveness and generalization ability, we apply AziNorm in both object detection and semantic segmentation. For detection, we integrate AziNorm into two representative detection methods, the one-stage SECOND detector and the state-of-the-art two-stage PV-RCNN detector. Experiments on Waymo Open Dataset demonstrate that AziNorm improves SECOND and PV-RCNN by 7.03 mAPH and 3.01 mAPH respectively. For segmentation, we integrate AziNorm into KPConv. On SemanticKitti dataset, AziNorm improves KPConv by 1.6/1.1 mIoU on val/test set. Besides, AziNorm remarkably improves data efficiency and accelerates convergence, reducing the requirement of data amounts or training epochs by an order of magnitude. SECOND w/ AziNorm can significantly outperform fully trained vanilla SECOND, even trained with only 10% data or 10% epochs. Code and models are available at https://github.com/hustvl/AziNorm.
Shaoyu Chen, Xinggang Wang, Tianheng Cheng, Qian Zhang 0009, Chang Huang, Wenyu Liu 0001
CVPR7
2022 Sparse Instance Activation for Real-Time Instance Segmentation
abstract
In this paper, we propose a conceptually novel, efficient, and fully convolutional framework for real-time instance segmentation. Previously, most instance segmentation methods heavily rely on object detection and perform mask prediction based on bounding boxes or dense centers. In contrast, we propose a sparse set of instance activation maps, as a new object representation, to high-light informative regions for each foreground object. Then instance-level features are obtained by aggregating features according to the highlighted regions for recognition and segmentation. Moreover, based on bipartite matching, the instance activation maps can predict objects in a one-to-one style, thus avoiding non-maximum suppression (NMS) in post-processing. Owing to the simple yet effective designs with instance activation maps, SparseInst has extremely fast inference speed and achieves 40 FPS and 37.9 AP on the COCO benchmark, which significantly out-performs the counterparts in terms of speed and accuracy. Code and models are available at https://github.com/hustvl/SparseInst.
Tianheng Cheng, Xinggang Wang, Shaoyu Chen, Qian Zhang 0009, Chang Huang, Zhaoxiang Zhang 0001, Wenyu Liu 0001
CVPR8
2022 MSG-Transformer: Exchanging Local Spatial Information by Manipulating Messenger Tokens
abstract
Transformers have offered a new methodology of designing neural networks for visual recognition. Compared to convolutional networks, Transformers enjoy the ability of referring to global features at each stage, yet the attention module brings higher computational overhead that obstructs the application of Transformers to process highresolution visual data. This paper aims to alleviate the conflict between efficiency and flexibility, for which we propose a specialized token for each region that serves as a messenger (MSG). Hence, by manipulating these MSG tokens, one can flexibly exchange visual information across regions and the computational complexity is reduced. We then integrate the MSG token into a multi-scale architecture named MSG-Transformer. In standard image classification and object detection, MSG-Transformer achieves competitive performance and the inference on both GPU and CPU is accelerated. Code is available at https://github.com/hustvl/MSG-Transformer.
Jiemin Fang, Lingxi Xie, Xinggang Wang, Xiaopeng Zhang 0008, Wenyu Liu 0001, Qi Tian 0001
CVPR5
2022 Knowledge Mining with Scene Text for Fine-Grained Recognition
abstract
Recently, the semantics of scene text has been proven to be essential in fine-grained image classification. However, the existing methods mainly exploit the literal meaning of scene text for fine-grained recognition, which might be irrelevant when it is not significantly related to objects/scenes. We propose an end-to-end trainable network that mines implicit contextual knowledge behind scene text image and enhance the semantics and correlation to fine-tune the image representation. Unlike the existing methods, our model integrates three modalities: visual feature extraction, text semantics extraction, and correlating background knowledge to fine-grained image classification. Specifically, we employ KnowBert to retrieve relevant knowledge for semantic representation and combine it with image features for fine-grained classification. Experiments on two benchmark datasets, Con-Text, and Drink Bottle, show that our method outperforms the state-of-the-art by 3.72% mAP and 5.39% mAp, respectively. To further validate the effectiveness of the proposed method, we create a new dataset on crowd activity recognition for the evaluation. The source code and new dataset of this work are available at this repository11https://github.com/lanfeng4659/KnowledgeMiningWithSceneText.
Hao Wang 0207, Junchao Liao, Tianheng Cheng, Zewen Gao, Hao Liu 0003, Bo Ren 0002, Xiang Bai, Wenyu Liu 0001
CVPR8
2022 Temporally Efficient Vision Transformer for Video Instance Segmentation
abstract
Recently vision transformer has achieved tremendous success on image-level visual recognition tasks. To effectively and efficiently model the crucial temporal information within a video clip, we propose a Temporally Efficient Vision Transformer (TeViT) for video instance segmentation (VIS). Different from previous transformer-based VIS methods, TeViT is nearly convolution-free, which contains a transformer backbone and a query-based video instance segmentation head. In the backbone stage, we propose a nearly parameter-free messenger shift mechanism for early temporal context fusion. In the head stages, we propose a parameter-shared spatiotemporal query interaction mechanism to build the one-to-one correspondence between video instances and queries. Thus, TeViT fully utilizes both frame-level and instance-level temporal context information and obtains strong temporal modeling capacity with negligible extra computational cost. On three widely adopted VIS benchmarks, i.e., YouTube-VIS-2019, YouTube-VIS-2021, and OVIS, TeViT obtains state-of-the-art results and maintains high inference speed, e.g., 46.6 AP with 68.9 FPS on YouTube-VIS-2019. Code is available at https://github.com/hustvl/TeViT.
Shusheng Yang, Xinggang Wang, Yu Li 0003, Jiemin Fang, Wenyu Liu 0001, Ying Shan
CVPR6
2022 TopFormer: Token Pyramid Transformer for Mobile Semantic Segmentation
abstract
Although vision transformers (ViTs) have achieved great success in computer vision, the heavy computational cost hampers their applications to dense prediction tasks such as semantic segmentation on mobile devices. In this paper, we present a mobile-friendly architecture named Token Pyramid Vision Transformer (TopFormer). The proposed TopFormer takes Tokens from various scales as input to produce scale-aware semantic features, which are then in-Jected into the corresponding tokens to augment the representation. Experimental results demonstrate that our method significantly outperforms CNN- and ViT-based networks across several semantic segmentation datasets and achieves a good trade-off between accuracy and latency. On the ADE20K dataset, TopFormer achieves 5% higher accuracy in mIoU than MobileNetV3 with lower latency on an ARM-based mobile device. Furthermore, the tiny version of TopFormer achieves real-time inference on an ARM-based mobile device with competitive results. The code and models are available at: https://github.com/hustvl/TopFormer.
Guozhong Luo, Tao Chen 0003, Xinggang Wang, Wenyu Liu 0001, Gang Yu 0002, Chunhua Shen
CVPR6
2022 Box-Supervised Instance Segmentation with Level Set Evolution
Wentong Li 0001, Wenyu Liu 0001, Jianke Zhu, Miaomiao Cui, Xian-Sheng Hua 0001, Lei Zhang 0006
ECCV (29)2
2022 When Counting Meets HMER: Counting-Aware Network for Handwritten Mathematical Expression Recognition
Bohan Li 0010, Dingkang Liang, Xiao Liu 0040, Zhilong Ji, Jinfeng Bai, Wenyu Liu 0001, Xiang Bai
ECCV (28)7
2022 ByteTrack: Multi-object Tracking by Associating Every Detection Box
Peize Sun, Yi Jiang 0009, Dongdong Yu, Fucheng Weng, Zehuan Yuan, Ping Luo 0002, Wenyu Liu 0001, Xinggang Wang
ECCV (22)8
2022 Robust Multi-object Tracking by Marginal Inference
Chunyu Wang 0001, Xinggang Wang, Wenjun Zeng 0001, Wenyu Liu 0001
ECCV (22)5
2022 Local Point Matching Network for Stabilized Crowd Counting and Localization
Lin Niu, Xinggang Wang, Chen Duan, Qiongxia Shen, Wenyu Liu 0001
PRCV (1)5
2022 Fast Dynamic Radiance Fields with Time-Aware Neural Voxels
abstract
Neural radiance fields (NeRF) have shown great success in modeling 3D scenes and synthesizing novel-view images. However, most previous NeRF methods take much time to optimize one single scene. Explicit data structures, e.g. voxel features, show great potential to accelerate the training process. However, voxel features face two big challenges to be applied to dynamic scenes, i.e. modeling temporal information and capturing different scales of point motions. We propose a radiance field framework by representing scenes with time-aware voxel features, named as TiNeuVox. A tiny coordinate deformation network is introduced to model coarse motion trajectories and temporal information is further enhanced in the radiance network. A multi-distance interpolation method is proposed and applied on voxel features to model both small and large motions. Our framework significantly accelerates the optimization of dynamic radiance fields while maintaining high rendering quality. Empirical evaluation is performed on both synthetic and real scenes. Our TiNeuVox completes training with only 8 minutes and 8-MB storage cost while showing similar or even better rendering performance than previous dynamic NeRF methods. Code is available at https://github.com/hustvl/TiNeuVox.
Jiemin Fang, Taoran Yi, Xinggang Wang, Lingxi Xie, Xiaopeng Zhang 0008, Wenyu Liu 0001, Matthias Nießner, Qi Tian 0001
SIGGRAPH Asia6
2022 AlignSeg: Feature-Aligned Segmentation Networks
abstract
Aggregating features in terms of different convolutional blocks or contextual embeddings has been proven to be an effective way to strengthen feature representations for semantic segmentation. However, most of the current popular network architectures tend to ignore the misalignment issues during the feature aggregation process caused by step-by-step downsampling operations and indiscriminate contextual information fusion. In this paper, we explore the principles in addressing such feature misalignment issues and inventively propose Feature-Aligned Segmentation Networks (AlignSeg). AlignSeg consists of two primary modules, i.e., the Aligned Feature Aggregation (AlignFA) module and the Aligned Context Modeling (AlignCM) module. First, AlignFA adopts a simple learnable interpolation strategy to learn transformation offsets of pixels, which can effectively relieve the feature misalignment issue caused by multi-resolution feature aggregation. Second, with the contextual embeddings in hand, AlignCM enables each pixel to choose private custom contextual information adaptively, making the contextual embeddings be better aligned. We validate the effectiveness of our AlignSeg network with extensive experiments on Cityscapes and ADE20K, achieving new state-of-the-art mIoU scores of 82.6 and 45.95 percent, respectively. Our source code is available at https://github.com/speedinghzl/AlignSeg.
Yunchao Wei, Xinggang Wang, Wenyu Liu 0001, Thomas S. Huang, Humphrey Shi
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Boundary TextSpotter: Toward Arbitrary-Shaped Scene Text Spotting
abstract
Reading arbitrary-shaped text in an end-to-end fashion has received particularly growing interested in computer vision. In this paper, we study the problem of scene text spotting, which aims to detect and recognize text from cluttered images simultaneously and propose an end-to-end trainable neural network named Boundary TextSpotter. Different from existing methods that describe the shape of text instance with bounding box or shape mask, Boundary TextSpotter formulates it as a set of boundary points. Besides, the representation of such boundary points provides the order of reading text. Benefiting from the representation on both detection and recognition, Boundary TextSpotter can easily deal with the text of arbitrary shapes. Further, to efficiently detect the boundary points of the text, a single-stage text detector is proposed, which can almost perform at a real-time speed. Experiments on three challenging datasets, including ICDAR2015, Total-Text and CTW1500 demonstrate that the proposed method achieves state-of-the-art or competitive results, meanwhile significantly improving the inference speed.
Pu Lu, Hao Wang 0207, Shenggao Zhu, Jing Wang 0221, Xiang Bai, Wenyu Liu 0001
IEEE Trans. Image Process.6
2021 Dynamic Class Queue for Large Scale Face Recognition in the Wild
abstract
Learning discriminative representation using large-scale face datasets in the wild is crucial for real-world applications, yet it remains challenging. The difficulties lie in many aspects and this work focus on computing resource constraint and long-tailed class distribution. Recently, classification-based representation learning with deep neural networks and well-designed losses have demonstrated good recognition performance. However, the computing and memory cost linearly scales up to the number of identities (classes) in the training set, and the learning process suffers from unbalanced classes. In this work, we propose a dynamic class queue (DCQ) to tackle these two problems. Specifically, for each iteration during training, a subset of classes for recognition are dynamically selected and their class weights are dynamically generated on-the-fly which are stored in a queue. Since only a subset of classes is selected for each iteration, the computing requirement is reduced. By using a single server without model parallel, we empirically verify in large-scale datasets that 10% of classes are sufficient to achieve similar performance as using all classes. Moreover, the class weights are dynamically generated in a few-shot manner and therefore suitable for tail classes with only a few instances. We show clear improvement over a strong baseline in the largest public dataset Megaface Challenge2 (MF2) which has 672K identities and over 88% of them have less than 10 instances. Code is available at https://github.com/bilylee/DCQ
Bi Li 0005, Teng Xi, Haocheng Feng, Junyu Han, Jingtuo Liu, Errui Ding, Wenyu Liu 0001
CVPR8
2021 Scene Text Retrieval via Joint Text Detection and Similarity Learning
abstract
Scene text retrieval aims to localize and search all text instances from an image gallery, which are the same or similar with a given query text. Such a task is usually realized by matching a query text to the recognized words, outputted by an end-to-end scene text spotter. In this paper, we address this problem by directly learning a cross-modal similarity between a query text and each text instance from natural images. Specifically, we establish an end-to-end trainable network, jointly optimizing the procedures of scene text detection and cross-modal similarity learning. In this way, scene text retrieval can be simply performed by ranking the detected text instances with the learned similarity. Experiments on three benchmark datasets demonstrate our method consistently outperforms the state-of-the-art scene text spotting/retrieval approaches. In particular, the proposed framework of joint detection and similarity learning achieves significantly better performance than separated methods. Code is available at: https://github.com/lanfeng4659/STR-TDSL.
Hao Wang 0207, Xiang Bai, Shenggao Zhu, Jing Wang 0221, Wenyu Liu 0001
CVPR6
2021 Weakly-Supervised Instance Segmentation via Class-Agnostic Learning With Salient Images
abstract
Humans have a strong class-agnostic object segmentation ability and can outline boundaries of unknown objects precisely, which motivates us to propose a box-supervised class-agnostic object segmentation (BoxCaseg) based solution for weakly-supervised instance segmentation. The BoxCaseg model is jointly trained using box-supervised images and salient images in a multi-task learning manner. The fine-annotated salient images provide class-agnostic and precise object localization guidance for box-supervised images. The object masks predicted by a pretrained BoxCaseg model are refined via a novel merged and dropped strategy as proxy ground truth to train a Mask R-CNN for weakly-supervised instance segmentation. Only using 7991 salient images, the weakly-supervised Mask R-CNN is on par with fully-supervised Mask R-CNN on PASCAL VOC and significantly outperforms previous state-of-the-art box-supervised instance segmentation methods on COCO. The source code, pretrained models and datasets are available at https://github.com/hustvl/BoxCaseg.
Xinggang Wang, Jiapei Feng, Bin Hu 0020, Longjin Ran, Xiaoxin Chen 0001, Wenyu Liu 0001
CVPR7
2021 Hierarchical Aggregation for 3D Instance Segmentation
abstract
Instance segmentation on point clouds is a fundamental task in 3D scene perception. In this work, we propose a concise clustering-based framework named HAIS, which makes full use of spatial relation of points and point sets. Considering clustering-based methods may result in over-segmentation or under-segmentation, we introduce the hierarchical aggregation to progressively generate instance proposals, i.e., point aggregation for preliminarily clustering points to sets and set aggregation for generating complete instances from sets. Once the complete 3D instances are obtained, a sub-network of intra-instance prediction is adopted for noisy points filtering and mask quality scoring. HAIS is fast (only 410ms per frame on Titan X)) and does not require non-maximum suppression. It ranks 1st on the ScanNet v2 benchmark1, achieving the highest 69.9% AP50and surpassing previous state-of-the-art (SOTA) methods by a large margin. Besides, the SOTA results on the S3DIS dataset validate the good generalization ability. Code is available at https://github.com/hustvl/HAIS.
Shaoyu Chen, Jiemin Fang, Qian Zhang 0009, Wenyu Liu 0001, Xinggang Wang
ICCV4
2021 Instances as Queries
abstract
We present QueryInst, a new perspective for instance segmentation. QueryInst is a multi-stage end-to-end system that treats instances of interest as learnable queries, enabling query based object detectors, e.g., Sparse RCNN, to have strong instance segmentation performance. The attributes of instances such as categories, bounding boxes, instance masks, and instance association embeddings are represented by queries in a unified manner. In QueryInst, a query is shared by both detection and segmentation via dynamic convolutions and driven by parallellysupervised multi-stage learning. We conduct extensive experiments on three challenging benchmarks, i.e., COCO, CityScapes, and YouTube-VIS to evaluate the effectiveness of QueryInst in object detection, instance segmentation, and video instance segmentation tasks. For the first time, we demonstrate that a simple end-to-end query based framework can achieve the state-of-the-art performance in various instance-level recognition tasks. Code is available at https://github.com/hustvl/QueryInst.
Shusheng Yang, Xinggang Wang, Yu Li 0003, Ying Shan, Bin Feng 0001, Wenyu Liu 0001
ICCV8
2021 Context-Sensitive Temporal Feature Learning for Gait Recognition
abstract
Although gait recognition has drawn increasing research attention recently, it remains challenging to learn discriminative temporal representation since the silhouette differences are quite subtle in spatial domain. Inspired by the observation that humans can distinguish gaits of different subjects by adaptively focusing on temporal sequences with different time scales, we propose a context-sensitive temporal feature learning (CSTL) network in this paper, which aggregates temporal features in three scales to obtain motion representation according to the temporal contextual information. Specifically, CSTL introduces relation modeling among multi-scale features to evaluate feature importances, based on which network adaptively enhances more important scale and suppresses less important scale. Besides that, we propose a salient spatial feature learning (SSFL) module to tackle the misalignment problem caused by temporal operation, e.g., temporal convolution. SSFL recombines a frame of salient spatial features by extracting the most discriminative parts across the whole sequence. In this way, we achieve adaptive temporal learning and salient spatial mining simultaneously. Extensive experiments conducted on two datasets demonstrate the state-of-the-art performance. On CASIA-B dataset, we achieve rank-1 accuracies of 98.0%, 95.4% and 87.0% under normal walking, bag-carrying and coat-wearing conditions. On OU-MVLP dataset, we achieve rank-1 accuracy of 90.2%. The source code will be published at https://github.com/OliverHxh/CSTL.
Xiaohu Huang, Duowang Zhu, Hao Wang 0207, Xinggang Wang, Botao He, Wenyu Liu 0001, Bin Feng 0001
ICCV7
2021 Crossover Learning for Fast Online Video Instance Segmentation
abstract
Modeling temporal visual context across frames is critical for video instance segmentation (VIS) and other video understanding tasks. In this paper, we propose a fast on-line VIS model termed CrossVIS. For temporal information modeling in VIS, we present a novel crossover learning scheme that uses the instance feature in the current frame to pixel-wisely localize the same instance in other frames. Different from previous schemes, crossover learning does not require any additional network parameters for feature enhancement. By integrating with the instance segmentation loss, crossover learning enables efficient cross-frame instance-to-pixel relation learning and brings cost-free improvement during inference. Besides, a global balanced instance embedding branch is proposed for better and more stable online instance association. We conduct extensive experiments on three challenging VIS benchmarks, i.e., YouTube-VIS-2019, OVIS, and YouTube-VIS-2021 to evaluate our methods. CrossVIS achieves state-of-the-art online VIS performance and shows a decent trade-off between latency and accuracy. Code is available at https://github.com/hustvl/CrossVIS.
Shusheng Yang, Xinggang Wang, Yu Li 0003, Ying Shan, Bin Feng 0001, Wenyu Liu 0001
ICCV8
2021 You Only Look at One Sequence: Rethinking Transformer in Vision through Object Detection
abstract
Can Transformer perform $2\mathrm{D}$ object- and region-level recognition from a pure sequence-to-sequence perspective with minimal knowledge about the $2\mathrm{D}$ spatial structure? To answer this question, we present You Only Look at One Sequence (YOLOS), a series of object detection models based on the vanilla Vision Transformer with the fewest possible modifications, region priors, as well as inductive biases of the target task. We find that YOLOS pre-trained on the mid-sized ImageNet-$1k$ dataset only can already achieve quite competitive performance on the challenging COCO object detection benchmark, e.g., YOLOS-Base directly adopted from BERT-Base architecture can obtain $42.0$ box AP on COCO val. We also discuss the impacts as well as limitations of current pre-train schemes and model scaling strategies for Transformer in vision through YOLOS. Code and pre-trained models are available at https://github.com/hustvl/YOLOS.
Bencheng Liao, Xinggang Wang, Jiemin Fang, Jiyang Qi, Rui Wu 0018, Jianwei Niu 0004, Wenyu Liu 0001
NeurIPS8
2021 Pyramid Self-attention for Semantic Segmentation
Jiyang Qi, Xinggang Wang, Yao Hu 0002, Xu Tang 0007, Wenyu Liu 0001
PRCV (1)5
2021 Learning to focus: cascaded feature matching network for few-shot image recognition
Xinggang Wang, Yifeng Geng, Wenyu Liu 0001
Sci. China Inf. Sci.5
2021 EAT-NAS: elastic architecture transfer for accelerating large-scale neural architecture search
Jiemin Fang, Yukang Chen, Xinbang Zhang, Qian Zhang 0009, Chang Huang, Gaofeng Meng, Wenyu Liu 0001, Xinggang Wang
Sci. China Inf. Sci.7
2021 Deep graph cut network for weakly-supervised semantic segmentation
Jiapei Feng, Xinggang Wang, Wenyu Liu 0001
Sci. China Inf. Sci.3
2021 EfficientPose: Efficient human pose estimation with neural architecture search
abstract
Human pose estimation from image and video is a key task in many multimedia applications. Previous methods achieve great performance but rarely take efficiency into consideration, which makes it difficult to implement the networks on lightweight devices. Nowadays, real-time multimedia applications call for more efficient models for better interaction. Moreover, most deep neural networks for pose estimation directly reuse networks designed for image classification as the backbone, which are not optimized for the pose estimation task. In this paper, we propose an efficient framework for human pose estimation with two parts, an efficient backbone and an efficient head. By implementing a differentiable neural architecture search method, we customize the backbone network design for pose estimation, and reduce computational cost with negligible accuracy degradation. For the efficient head, we slim the transposed convolutions and propose a spatial information correction module to promote the performance of the final prediction. In experiments, we evaluate our networks on the MPII and COCO datasets. Our smallest model requires only 0.65 GFLOPs with 88.1% [email protected] on MPII and our large model needs only 2 GFLOPs while its accuracy is competitive with the state-of-the-art large model, HRNet, which takes 9.5 GFLOPs.
Jiemin Fang, Xinggang Wang, Wenyu Liu 0001
Comput. Vis. Media4
2021 FairMOT: On the Fairness of Detection and Re-identification in Multiple Object Tracking
Chunyu Wang 0001, Xinggang Wang, Wenjun Zeng 0001, Wenyu Liu 0001
Int. J. Comput. Vis.5
2021 Deep High-Resolution Representation Learning for Visual Recognition
abstract
High-resolution representations are essential for position-sensitive vision problems, such as human pose estimation, semantic segmentation, and object detection. Existing state-of-the-art frameworks first encode the input image as a low-resolution representation through a subnetwork that is formed by connecting high-to-low resolution convolutions in series (e.g., ResNet, VGGNet), and then recover the high-resolution representation from the encoded low-resolution representation. Instead, our proposed network, named as High-Resolution Network (HRNet), maintains high-resolution representations through the whole process. There are two key characteristics: (i) Connect the high-to-low resolution convolution streams in parallel and (ii) repeatedly exchange the information across resolutions. The benefit is that the resulting representation is semantically richer and spatially more precise. We show the superiority of the proposed HRNet in a wide range of applications, including human pose estimation, semantic segmentation, and object detection, suggesting that the HRNet is a stronger backbone for computer vision problems. All the codes are available at https://github.com/HRNet.
Jingdong Wang 0001, Ke Sun 0009, Tianheng Cheng, Borui Jiang, Chaorui Deng, Yang Zhao 0019, Dong Liu 0002, Yadong Mu, Mingkui Tan, Xinggang Wang, Wenyu Liu 0001, Bin Xiao 0004
IEEE Trans. Pattern Anal. Mach. Intell.11
2021 FNA++: Fast Network Adaptation via Parameter Remapping and Architecture Search
abstract
Deep neural networks achieve remarkable performance in many computer vision tasks. Most state-of-the-art (SOTA) semantic segmentation and object detection approaches reuse neural network architectures designed for image classification as the backbone, commonly pre-trained on ImageNet. However, performance gains can be achieved by designing network architectures specifically for detection and segmentation, as shown by recent neural architecture search (NAS) research for detection and segmentation. One major challenge though is that ImageNet pre-training of the search space representation (a.k.a. super network) or the searched networks incurs huge computational cost. In this paper, we propose a Fast Network Adaptation (FNA++) method, which can adapt both the architecture and parameters of a seed network (e.g., an ImageNet pre-trained network) to become a network with different depths, widths, or kernel sizes via a parameter remapping technique, making it possible to use NAS for segmentation and detection tasks a lot more efficiently. In our experiments, we apply FNA++ on MobileNetV2 to obtain new networks for semantic segmentation, object detection, and human pose estimation that clearly outperform existing networks designed both manually and by NAS. We also implement FNA++ on ResNets and NAS networks, which demonstrates a great generalization ability. The total computation cost of FNA++ is significantly less than SOTA segmentation and detection NAS approaches: 1737× less than DPC, 6.8× less than Auto-DeepLab, and 8.0× less than DetNAS. A series of ablation studies are performed to demonstrate the effectiveness, and detailed analysis is provided for more insights into the working mechanism. Codes are available at https://github.com/JaminFong/FNA.
Jiemin Fang, Yuzhu Sun, Qian Zhang 0009, Kangjian Peng, Wenyu Liu 0001, Xinggang Wang
IEEE Trans. Pattern Anal. Mach. Intell.6
2020 Diversity Transfer Network for Few-Shot Learning
abstract
Few-shot learning is a challenging task that aims at training a classifier for unseen classes with only a few training examples. The main difficulty of few-shot learning lies in the lack of intra-class diversity within insufficient training samples. To alleviate this problem, we propose a novel generative framework, Diversity Transfer Network (DTN), that learns to transfer latent diversities from known categories and composite them with support features to generate diverse samples for novel categories in feature space. The learning problem of the sample generation (i.e., diversity transfer) is solved via minimizing an effective meta-classification loss in a single-stage network, instead of the generative loss in previous works. Besides, an organized auxiliary task co-training over known categories is proposed to stabilize the meta-training process of DTN. We perform extensive experiments and ablation studies on three datasets, i.e., miniImageNet, CIFAR100 and CUB. The results show that DTN, with single-stage training and faster convergence speed, obtains the state-of-the-art results among the feature generation based few-shot learning methods. Code and supplementary material are available at: https://github.com/Yuxin-CV/DTN.
Xinggang Wang, Yifeng Geng, Chang Huang, Wenyu Liu 0001, Bo Wang 0044
AAAI8
2020 All You Need Is Boundary: Toward Arbitrary-Shaped Text Spotting
abstract
Recently, end-to-end text spotting that aims to detect and recognize text from cluttered images simultaneously has received particularly growing interest in computer vision. Different from the existing approaches that formulate text detection as bounding box extraction or instance segmentation, we localize a set of points on the boundary of each text instance. With the representation of such boundary points, we establish a simple yet effective scheme for end-to-end text spotting, which can read the text of arbitrary shapes. Experiments on three challenging datasets, including ICDAR2015, TotalText and COCO-Text demonstrate that the proposed method consistently surpasses the state-of-the-art in both scene text detection and end-to-end text recognition tasks.
Hao Wang 0207, Pu Lu, Hui Zhang 0085, Xiang Bai, Yongchao Xu, Mengchao He, Yongpan Wang, Wenyu Liu 0001
AAAI9
2020 Densely Connected Search Space for More Flexible Neural Architecture Search
abstract
Neural architecture search (NAS) has dramatically advanced the development of neural network design. We revisit the search space design in most previous NAS methods and find the number and widths of blocks are set manually. However, block counts and block widths determine the network scale (depth and width) and make a great influence on both the accuracy and the model cost (FLOPs/latency). In this paper, we propose to search block counts and block widths by designing a densely connected search space, i.e., DenseNAS. The new search space is represented as a dense super network, which is built upon our designed routing blocks. In the super network, routing blocks are densely connected and we search for the best path between them to derive the final architecture. We further propose a chained cost estimation algorithm to approximate the model cost during the search. Both the accuracy and model cost are optimized in DenseNAS. For experiments on the MobileNetV2-based search space, DenseNAS achieves 75.3% top-1 accuracy on ImageNet with only 361MB FLOPs and 17.9ms latency on a single TITAN-XP. The larger model searched by DenseNAS achieves 76.1% accuracy with only 479M FLOPs. DenseNAS further promotes the ImageNet classification accuracies of ResNet-18, -34 and -50-B by 1.5%, 0.5% and 0.3% with 200M, 600M and 680M FLOPs reduction respectively. The related code is available at https://github.com/JaminFong/DenseNAS.
Jiemin Fang, Yuzhu Sun, Qian Zhang 0009, Wenyu Liu 0001, Xinggang Wang
CVPR5
2020 Maximum Entropy Regularization and Chinese Text Recognition
Changxu Cheng, Wuheng Xu, Xiang Bai, Bin Feng 0001, Wenyu Liu 0001
DAS5
2020 Boundary-Preserving Mask R-CNN
Tianheng Cheng, Xinggang Wang, Lichao Huang, Wenyu Liu 0001
ECCV (14)4
2020 Fast Neural Network Adaptation via Parameter Remapping and Architecture Search
Jiemin Fang, Yuzhu Sun, Kangjian Peng, Qian Zhang 0009, Wenyu Liu 0001, Xinggang Wang
ICLR6
2020 Learning Global Structure Consistency for Robust Object Tracking
abstract
Fast appearance variations and the distractions of similar objects are two of the most challenging problems in visual object tracking. Unlike many existing trackers that focus on modeling only the target, in this work, we consider the transient variations of the whole scene. The key insight is that the object correspondence and spatial layout of the whole scene are consistent (i.e., global structure consistency) in consecutive frames which helps to disambiguate the target from distractors. Moreover, modeling transient variations enables to localize the target under fast variations. Specifically, we propose an effective and efficient short-term model that learns to exploit the global structure consistency in a short time and thus can handle fast variations and distractors. Since short-term modeling falls short of handling occlusion and out of the views, we adopt the long-short term paradigm and use a long-term model that corrects the short-term model when it drifts away from the target or the target is not present. These two components are carefully combined to achieve the balance of stability and plasticity during tracking. We empirically verify that the proposed tracker can tackle the two challenging scenarios and validate it on large scale benchmarks. Remarkably, our tracker improves state-of-the-art-performance on VOT2018 from 0.440 to 0.460, GOT-10k from 0.611 to 0.640, and NFS from 0.619 to 0.629.
Bi Li 0005, Chengquan Zhang, Zhibin Hong, Xu Tang 0007, Jingtuo Liu, Junyu Han, Errui Ding, Wenyu Liu 0001
ACM Multimedia8
2020 EEG responses to emotional videos can quantitatively predict big-five personality traits
Wenyu Liu 0001, Xuefei Long, Lilu Tang, Fei Wang 0032, Dan Zhang 0014
Neurocomputing1
2020 Deep multi-metric learning for text-independent speaker verification
Xinggang Wang, Bin Feng 0001, Wenyu Liu 0001
Neurocomputing4
2020 PCL: Proposal Cluster Learning for Weakly Supervised Object Detection
abstract
Weakly Supervised Object Detection (WSOD), using only image-level annotations to train object detectors, is of growing importance in object recognition. In this paper, we propose a novel deep network for WSOD. Unlike previous networks that transfer the object detection problem to an image classification problem using Multiple Instance Learning (MIL), our strategy generates proposal clusters to learn refined instance classifiers by an iterative process. The proposals in the same cluster are spatially adjacent and associated with the same object. This prevents the network from concentrating too much on parts of objects instead of whole objects. We first show that instances can be assigned object or background labels directly based on proposal clusters for instance classifier refinement, and then show that treating each cluster as a small new bag yields fewer ambiguities than the directly assigning label method. The iterative instance classifier refinement is implemented online using multiple streams in convolutional neural networks, where the first is an MIL network and the others are for instance classifier refinement supervised by the preceding one. Experiments are conducted on the PASCAL VOC, ImageNet detection, and MS-COCO benchmarks for WSOD. Results show that our method outperforms the previous state of the art significantly.
Peng Tang 0005, Xinggang Wang, Song Bai 0001, Wei Shen 0002, Xiang Bai, Wenyu Liu 0001, Alan L. Yuille
IEEE Trans. Pattern Anal. Mach. Intell.6
2020 Object Detection in Videos by High Quality Object Linking
abstract
Compared with object detection in static images, object detection in videos is more challenging due to degraded image qualities. An effective way to address this problem is to exploit temporal contexts by linking the same object across video to form tubelets and aggregating classification scores in the tubelets. In this paper, we focus on obtaining high quality object linking results for better classification. Unlike previous methods that link objects by checking boxes between neighboring frames, we propose to link in the same frame. To achieve this goal, we extend prior methods in following aspects: (1) a cuboid proposal network that extracts spatio-temporal candidate cuboids which bound the movement of objects; (2) a short tubelet detection network that detects short tubelets in short video segments; (3) a short tubelet linking algorithm that links temporally-overlapping short tubelets to form long tubelets. Experiments on the ImageNet VID dataset show that our method outperforms both the static image detector and the previous state of the art. In particular, our method improves results by 8.8 percent over the static image detector for fast moving objects.
Peng Tang 0005, Chunyu Wang 0001, Xinggang Wang, Wenyu Liu 0001, Wenjun Zeng 0001, Jingdong Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2020 Pose Anchor: A Single-Stage Hand Keypoint Detection Network
abstract
This paper presents an effective single network for hand keypoint detection, instead of relying on the frequently-used two-stage pipeline consisting of localizing the hand and detecting the key points. Our method trains a fully convolutional neural network in an end-to-end manner, based on a novelly proposed pose anchor network, which can be deemed as an extension of the region proposal network (RPN) in Faster Region-based convolutional network. Moreover, we generate our pose anchor in a data-driven way, i.e., a K-means cluster algorithm based on object keypoint similarity (OKS), instead of manually design. In this way, we can obtain multiple representative pose anchors with various gestures, angles, and scales. By introducing the pose anchor, we are capable of utilizing the prior knowledge of the hand structure, mitigating the problem of occlusion to some extent. We demonstrate the feasibility and effectiveness of our method with extensive experiments on the challenging large-scale multiview 3D hand pose dataset (LSM-HPD) and New Zealand Sign Language Dataset (NZSL).
Xinggang Wang, Wenyu Liu 0001, Bin Feng 0001
IEEE Trans. Circuits Syst. Video Technol.3
2020 Cascaded Boundary Network for High-Quality Temporal Action Proposal Generation
abstract
Creating high-quality temporal action proposals is fundamental yet challenging for accurate action detection in untrimmed videos due to the complexity of the background and variation in actions' durations and magnitudes. In this paper, we propose a cascaded boundary network (CBN) to predict the action boundaries by considering the importance of precise boundary information to develop accurate action proposals. Specifically, the first stage of CBN locates the temporal boundaries by predicting the probability that each frame corresponds to an action, the start position, and the end position. A temporal convolutional network is used in this stage to capture short-term context information. Next, the predicted probabilities are forwarded to the second stage, in which a long short-term memory (LSTM) network is utilized for further refinement by exploiting the correlation between the predicted probabilities to capture long-term context information. Finally, we combine the results from both stages to produce a long- and short-term information fusion. The experiments on THUMOS14 and ActivityNet-1.3 show that CBN achieves state-of-the-art recall performance. The performance improvement is especially remarkable for a small average number (AN) of retrieved proposals; e.g., the average recall at AN=50 on THUMOS14 is improved from 37.46% to 43.06%. Further experiments are performed by introducing proposals generated by CBN into an existing action detection framework. CBN also achieves state-of-the-art average mAP@tIoU on the THUMOS14 detection benchmark.
Xinggang Wang, Wenyu Liu 0001, Bin Feng 0001
IEEE Trans. Circuits Syst. Video Technol.3
2020 Semantic Image Segmentation by Scale-Adaptive Networks
abstract
Semantic image segmentation is an important yet unsolved problem. One of the major challenges is the large variability of the object scales. To tackle this scale problem, we propose a Scale-Adaptive Network (SAN) which consists of multiple branches with each one taking charge of the segmentation of the objects of a certain range of scales. Given an image, SAN first computes a dense scale map indicating the scale of each pixel which is automatically determined by the size of the enclosing object. Then the features of different branches are fused according to the scale map to generate the final segmentation map. To ensure that each branch indeed learns the features for a certain scale, we propose a scale-induced ground-truth map and enforce a scale-aware segmentation loss for the corresponding branch in addition to the final loss. Extensive experiments over the PASCAL-Person-Part, the PASCAL VOC 2012, and the Look into Person datasets demonstrate that our SAN can handle the large variability of the object scales and outperforms the state-of-the-art semantic segmentation methods.
Chunyu Wang 0001, Xinggang Wang, Wenyu Liu 0001, Jingdong Wang 0001
IEEE Trans. Image Process.4
2020 A Weakly-Supervised Framework for COVID-19 Classification and Lesion Localization From Chest CT
abstract
Accurate and rapid diagnosis of COVID-19 suspected cases plays a crucial role in timely quarantine and medical treatment. Developing a deep learning-based model for automatic COVID-19 diagnosis on chest CT is helpful to counter the outbreak of SARS-CoV-2. A weakly-supervised deep learning framework was developed using 3D CT volumes for COVID-19 classification and lesion localization. For each patient, the lung region was segmented using a pre-trained UNet; then the segmented 3D lung region was fed into a 3D deep neural network to predict the probability of COVID-19 infectious; the COVID-19 lesions are localized by combining the activation regions in the classification network and the unsupervised connected components. 499 CT volumes were used for training and 131 CT volumes were used for testing. Our algorithm obtained 0.959 ROC AUC and 0.976 PR AUC. When using a probability threshold of 0.5 to classify COVID-positive and COVID-negative, the algorithm obtained an accuracy of 0.901, a positive predictive value of 0.840 and a very high negative predictive value of 0.982. The algorithm took only 1.93 seconds to process a single patient's CT volume using a dedicated GPU. Our weakly-supervised deep learning model can accurately predict the COVID-19 infectious probability and discover lesion regions in chest CT without the need for annotating the lesions for training. The easily-trained and high-performance deep learning algorithm provides a fast way to identify COVID-19 patients, which is beneficial to control the outbreak of SARS-CoV-2. The developed deep learning software is available at https://github.com/sydney0zq/covid-19-detection.
Xinggang Wang, Xianbo Deng, Qing Fu, Jiapei Feng, Wenyu Liu 0001, Chuansheng Zheng
IEEE Trans. Medical Imaging7
2019 CCNet: Criss-Cross Attention for Semantic Segmentation
abstract
Full-image dependencies provide useful contextual information to benefit visual understanding problems. In this work, we propose a Criss-Cross Network (CCNet) for obtaining such contextual information in a more effective and efficient way. Concretely, for each pixel, a novel criss-cross attention module in CCNet harvests the contextual information of all the pixels on its criss-cross path. By taking a further recurrent operation, each pixel can finally capture the full-image dependencies from all pixels. Overall, CCNet is with the following merits: 1) GPU memory friendly. Compared with the non-local block, the proposed recurrent criss-cross attention module requires 11x less GPU memory usage. 2) High computational efficiency. The recurrent criss-cross attention significantly reduces FLOPs by about 85% of the non-local block in computing full-image dependencies. 3) The state-of-the-art performance. We conduct extensive experiments on popular semantic segmentation benchmarks including Cityscapes, ADE20K, and instance segmentation benchmark COCO. In particular, our CCNet achieves the mIoU score of 81.4 and 45.22 on Cityscapes test set and ADE20K validation set, respectively, which are the new state-of-the-art results. The source code is available at https://github.com/speedinghzl/CCNet.
Xinggang Wang, Lichao Huang, Chang Huang, Yunchao Wei, Wenyu Liu 0001
ICCV6
2019 Patch Aggregator for Scene Text Script Identification
abstract
Script identification in the wild is of great importance in a multi-lingual robust-reading system. The scripts deriving from the same language family share a large set of characters, which makes script identification a fine-grained classification problem. Most existing methods make efforts to learn a single representation that combines the local features by making a weighted average or other clustering methods, which may reduce the discriminatory power of some important parts in each script for the interference of redundant features. In this paper, we present a novel module named Patch Aggregator (PA), which learns a more discriminative representation for script identification by taking into account the prediction scores of local patches. Specifically, we design a CNN-based method consisting of a standard CNN classifier and a PA module. Experiments demonstrate that the proposed PA module brings significant performance improvements over the baseline CNN model, achieving the state-of-the-art results on three benchmark datasets for script identification: SIW-13, CVSI 2015 and RRC-MLT 2017.
Changxu Cheng, Qiuhui Huang, Xiang Bai, Bin Feng 0001, Wenyu Liu 0001
ICDAR5
2019 Multiple Comparative Attention Network for Offline Handwritten Chinese Character Recognition
abstract
Recent advances in deep learning have made great progress in offline Handwritten Chinese Character Recognition (HCCR). However, most existing CNN-based methods only utilize global image features as contextual guidance to classify characters, while neglecting the local discriminative features which is very important for HCCR. To overcome this limitation, in this paper, we present a convolutional neural network with multiple comparative attention (MCANet) in order to produce separable local attention regions with discriminative feature across different categories. Concretely, our MCANet takes the last convolutional feature map as input and outputs multiple attention maps, a contrastive loss is used to restrict different attention selectively focus on different sub-regions. Moreover, we apply a region-level center loss to pull the features that learned from the same class and different regions closer to further obtain robust features invariant to large intra-class variance. Combining with classification loss, our method can learn which parts of images are relevant for recognizing characters and adaptively integrates information from different regions to make the final prediction. We conduct experiments on ICDAR2013 offline HCCR competition dataset with our proposed approach and achieves an accuracy of 97.66%, outperforming all single-network methods trained only on handwritten data.
Qingquan Xu, Xiang Bai, Wenyu Liu 0001
ICDAR3
2019 Bag similarity network for deep multi-instance learning
Xinggang Wang, Yongluan Yan, Peng Tang 0005, Wenyu Liu 0001, Xiaojie Guo 0001
Inf. Sci.4
2019 Weakly- and Semi-Supervised Fast Region-Based CNN for Object Detection
Xinggang Wang, Jiasi Wang, Peng Tang 0005, Wenyu Liu 0001
J. Comput. Sci. Technol.4
2019 Weakly supervised mitosis detection in breast histopathology images using concentric loss
Chao Li 0007, Xinggang Wang, Wenyu Liu 0001, Longin Jan Latecki, Bo Wang 0044, Junzhou Huang
Medical Image Anal.3
2019 Learning to Update for Object Tracking With Recurrent Meta-Learner
abstract
Model update lies at the heart of object tracking. Generally, model update is formulated as an online learning problem where a target model is learned over the online training set. Our key innovation is to formulate the model update problem in the meta-learning framework and learn the online learning algorithm itself using large numbers of offline videos, i.e., learning to update. The learned updater takes as input the online training set and outputs an updated target model. As a first attempt, we design the learned updater based on recurrent neural networks (RNNs) and demonstrate its application in a template-based tracker and a correlation filter-based tracker. Our learned updater consistently improves the base trackers and runs faster than realtime on GPU while requiring small memory footprint during testing. Experiments on standard benchmarks demonstrate that our learned updater outperforms commonly used update baselines including the efficient exponential moving average (EMA)-based update and the well-designed stochastic gradient descent (SGD)-based update. Equipped with our learned updater, the template-based tracker achieves state-of-the-art performance among realtime trackers on GPU.
Bi Li 0005, Wenxuan Xie, Wenjun Zeng 0001, Wenyu Liu 0001
IEEE Trans. Image Process.4
2019 Deep FisherNet for Image Classification
abstract
Despite the great success of convolutional neural networks (CNNs) for the image classification task on data sets such as Cifar and ImageNet, CNN's representation power is still somewhat limited in dealing with images that have a large variation in size and clutter, where Fisher vector (FV) has shown to be an effective encoding strategy. FV encodes an image by aggregating local descriptors with a universal generative Gaussian mixture model (GMM). FV, however, has limited learning capability and its parameters are mostly fixed after constructing the codebook. To combine together the best of the two worlds, we propose in this brief a neural network structure with FV layer being part of an end-to-end trainable system that is differentiable; we name our network FisherNet that is learnable using back propagation. Our proposed FisherNet combines CNN training and FV encoding in a single end-to-end structure. We observe a clear advantage of FisherNet over plain CNN and standard FV in terms of both classification accuracy and computational efficiency on the challenging PASCAL visual object classes object classification and emotion image classification tasks.
Peng Tang 0005, Xinggang Wang, Baoguang Shi, Xiang Bai, Wenyu Liu 0001, Zhuowen Tu
IEEE Trans. Neural Networks Learn. Syst.5
2018 Deep Multi-instance Learning with Dynamic Pooling
abstract
End-to-end optimization of multi-instance learning (MIL) using neural networks is an important problem with many applications, in which a core issue is how to design a permutation-invariant pooling function without losing much instance-level information. Inspired by the dynamic routing in recent capsule networks, we propose a novel dynamic pooling function for MIL. It is an adaptive scheme for both key instance selection and modeling the contextual information among instances in a bag. The dynamic pooling iteratively updates the instance contribution to its bag. It is permutation-invariant and can interpret instance-to-bag relationship. The proposed dynamic pooling based multi-instance neural network has been validated on many MIL tasks and outperforms other MIL methods.
Yongluan Yan, Xinggang Wang, Xiaojie Guo 0001, Jiemin Fang, Wenyu Liu 0001, Junzhou Huang
ACML5
2018 Weakly-Supervised Semantic Segmentation Network With Deep Seeded Region Growing
abstract
This paper studies the problem of learning image semantic segmentation networks only using image-level labels as supervision, which is important since it can significantly reduce human annotation efforts. Recent state-of-the-art methods on this problem first infer the sparse and discriminative regions for each object class using a deep classification network, then train semantic a segmentation network using the discriminative regions as supervision. Inspired by the traditional image segmentation methods of seeded region growing, we propose to train a semantic segmentation network starting from the discriminative regions and progressively increase the pixel-level supervision using by seeded region growing. The seeded region growing module is integrated in a deep segmentation network and can benefit from deep features. Different from conventional deep networks which have fixed/static labels, the proposed weakly-supervised network generates new labels using the contextual information within an image. The proposed method significantly outperforms the weakly-supervised semantic segmentation methods using static labels, and obtains the state-of-the-art performance, which are 63.2% mIoU score on the PASCAL VOC 2012 test set and 26.0% mIoU score on the COCO dataset.
Xinggang Wang, Jiasi Wang, Wenyu Liu 0001, Jingdong Wang 0001
CVPR4
2018 Weakly Supervised Region Proposal Network and Object Detection
Peng Tang 0005, Xinggang Wang, Angtian Wang, Yongluan Yan, Wenyu Liu 0001, Junzhou Huang, Alan L. Yuille
ECCV (11)5
2018 Mancs: A Multi-task Attentional Network with Curriculum Sampling for Person Re-Identification
Cheng Wang 0048, Qian Zhang 0009, Chang Huang, Wenyu Liu 0001, Xinggang Wang
ECCV (4)4
2018 Weakly- and Semi-supervised Faster R-CNN with Curriculum Learning
abstract
Object detection is a core problem in computer vision and pattern recognition. In this paper, we study the problem of learning an effective object detector using weakly-annotated images (i.e., only the image level annotation is given) and a small proportion of fully-annotated images (i.e., bounding box level annotation is given) with curriculum learning. Our method is built upon Faster R-CNN. Different from previous weakly-supervised object detectors which rely on hand-craft object proposals, the proposed method learns a region proposal network using weakly- and semi-supervised training data. And the weakly-labeled images are fed into the deep network in a meaningful order which illustrates from easy to gradually more complex examples with curriculum learning. We name the Faster R-CNN trained using Weakly- And Semi-Supervised data with Curriculum Learning as WASSCL R-CNN. The WASSCL R-CNN is validated on the PASCAL VOC 2007 benchmark, and obtains 90% of a fully-supervised Faster R-CNN's performance (measured using mAP) with only 15% of fully-supervised annotations together with weak supervision. The results show that the proposed learning framework can significantly reduce the labeling efforts for obtaining reliable object detectors.
Jiasi Wang, Xinggang Wang, Wenyu Liu 0001
ICPR3
2018 Monocular Camera Based Real-Time Dense Mapping Using Generative Adversarial Network
abstract
Monocular simultaneous localization and mapping (SLAM) is a key enabling technique for many computer vision and robotics applications. However, existing methods either can obtain only sparse or semi-dense maps in highly-textured image areas or fail to achieve a satisfactory reconstruction accuracy. In this paper, we present a new method based on a generative adversarial network,named DM-GAN, for real-time dense mapping based on a monocular camera. Specifcally, our depth generator network takes a semidense map obtained from motion stereo matching as a guidance to supervise dense depth prediction of a single RGB image. The depth generator is trained based on a combination of two loss functions, i.e. an adversarial loss for enforcing the generated depth maps to reside on the manifold of the true depth maps and a pixel-wise mean square error (MSE) for ensuring the correct absolute depth values. Extensive experiments on three public datasets demonstrate that our DM-GAN signifcantly outperforms the state-of-the-art methods in terms of greater reconstruction accuracy and higher depth completeness.
Xin Yang 0008, Zhiwei Wang 0002, Qiaozhe Zhang, Wenyu Liu 0001, Chunyuan Liao, Kwang-Ting Cheng
ACM Multimedia5
2018 Structured random forest for label distribution learning
Xinggang Wang, Bin Feng 0001, Wenyu Liu 0001
Neurocomputing4
2018 Deep attention network for joint hand gesture localization and recognition using static RGB-D images
Xinggang Wang, Wenyu Liu 0001, Bin Feng 0001
Inf. Sci.3
2018 DeepMitosis: Mitosis detection via deep detection, verification and segmentation networks
Chao Li 0007, Xinggang Wang, Wenyu Liu 0001, Longin Jan Latecki
Medical Image Anal.3
2018 Revisiting multiple instance neural networks
Xinggang Wang, Yongluan Yan, Peng Tang 0005, Xiang Bai, Wenyu Liu 0001
Pattern Recognit.5
2018 Cascaded Segmentation-Detection Networks for Text-Based Traffic Sign Detection
abstract
In this paper, we propose a novel text-based traffic sign detection framework with two deep learning components. More precisely, we apply a fully convolutional network to segment candidate traffic sign areas providing candidate regions of interest (RoI), followed by a fast neural network to detect texts on the extracted RoI. The proposed method makes full use of the characteristics of traffic signs to improve the efficiency and accuracy of text detection. On one hand, the proposed two-stage detection method reduces the search area of text detection and removes texts outside traffic signs. On the other hand, it solves the problem of multi-scales for the text detection part to a large extent. Extensive experimental results show that the proposed method achieves the state-of-the-art results on the publicly available traffic sign data set: Traffic Guide Panel data set. In addition, we collect a data set of text-based traffic signs including Chinese and English traffic signs. Our method also performs well on this data set, which demonstrates that the proposed method is general in detecting traffic signs of different languages.
Yingying Zhu 0005, Minghui Liao, Wenyu Liu 0001
IEEE Trans. Intell. Transp. Syst.4
2018 Face Alignment With Deep Regression
abstract
In this paper, we present a deep regression approach for face alignment. The deep regressor is a neural network that consists of a global layer and multistage local layers. The global layer estimates the initial face shape from the whole image, while the following local layers iteratively update the shape with local image observations. Combining standard derivations and numerical approximations, we make all layers able to backpropagate error differentials, so that we can apply the standard backpropagation to jointly learn the parameters from all layers. We show that the resulting deep regressor gradually and evenly approaches the true facial landmarks stage by stage, avoiding the tendency that often occurs in the cascaded regression methods and deteriorates the overall performance: yielding early stage regressors with high alignment accuracy gains but later stage regressors with low alignment accuracy gains. Experimental results on standard benchmarks demonstrate that our approach brings significant improvements over previous cascaded regression algorithms.
Baoguang Shi, Xiang Bai, Wenyu Liu 0001, Jingdong Wang 0001
IEEE Trans. Neural Networks Learn. Syst.3
2017 TextBoxes: A Fast Text Detector with a Single Deep Neural Network
abstract
This paper presents an end-to-end trainable fast scene text detector, named TextBoxes, which detects scene text with both high accuracy and efficiency in a single network forward pass, involving no post-process except for a standard non-maximum suppression. TextBoxes outperforms competing methods in terms of text localization accuracy and is much faster, taking only 0.09s per image in a fast implementation. Furthermore, combined with a text recognizer, TextBoxes significantly outperforms state-of-the-art approaches on word spotting and end-to-end text recognition tasks.
Minghui Liao, Baoguang Shi, Xiang Bai, Xinggang Wang, Wenyu Liu 0001
AAAI5
2017 Multiple Instance Detection Network with Online Instance Classifier Refinement
abstract
Of late, weakly supervised object detection is with great importance in object recognition. Based on deep learning, weakly supervised detectors have achieved many promising results. However, compared with fully supervised detection, it is more challenging to train deep network based detectors in a weakly supervised manner. Here we formulate weakly supervised detection as a Multiple Instance Learning (MIL) problem, where instance classifiers (object detectors) are put into the network as hidden nodes. We propose a novel online instance classifier refinement algorithm to integrate MIL and the instance classifier refinement procedure into a single deep network, and train the network end-to-end with only image-level supervision, i.e., without object location information. More precisely, instance labels inferred from weak supervision are propagated to their spatially overlapped instances to refine instance classifier online. The iterative instance classifier refinement procedure is implemented using multiple streams in deep network, where each stream supervises its latter stream. Weakly supervised object detection experiments are carried out on the challenging PASCAL VOC 2007 and 2012 benchmarks. We obtain 47% mAP on VOC 2007 that significantly outperforms the previous state-of-the-art.
Peng Tang 0005, Xinggang Wang, Xiang Bai, Wenyu Liu 0001
CVPR4
2017 Auto-Encoder Guided GAN for Chinese Calligraphy Synthesis
abstract
In this paper, we investigate the Chinese calligraphy synthesis problem: synthesizing Chinese calligraphy images with specified style from standard font(eg. Hei font) images (Fig. 1(a)). Recent works mostly follow the stroke extraction and assemble pipeline which is complex in the process and limited by the effect of stroke extraction. In this work we treat the calligraphy synthesis problem as an image-to-image translation problem and propose a deep neural network based model which can generate calligraphy images from standard font images directly. Besides, we also construct a large scale benchmark that contains various styles for Chinese calligraphy synthesis. We evaluate our method as well as some baseline methods on the proposed dataset, and the experimental results demonstrate the effectiveness of our proposed model.
Pengyuan Lv, Xiang Bai, Cong Yao, Zhen Zhu 0006, Tengteng Huang, Wenyu Liu 0001
ICDAR6
2017 Joint Classification Loss and Histogram Loss for Sketch-Based Image Retrieval
Yongluan Yan, Xinggang Wang, Xin Yang 0008, Xiang Bai, Wenyu Liu 0001
ICIG (1)5
2017 Barrier Coverage Lifetime Maximization in a Randomly Deployed Bistatic Radar Network
Bang Wang 0001, Wenyu Liu 0001
MSN3
2017 Neural features for pedestrian detection
Chao Li 0007, Xinggang Wang, Wenyu Liu 0001
Neurocomputing3
2017 Learning extremely shared middle-level image representation for scene classification
Peng Tang 0005, Xinggang Wang, Bin Feng 0001, Fabio Roli, Wenyu Liu 0001
Knowl. Inf. Syst.6
2017 Multiscale salient region-based visual tracking
Sihua Yi, Wenyu Liu 0001
Mach. Vis. Appl.2
2017 Deep patch learning for weakly supervised object classification and discovery
Peng Tang 0005, Xinggang Wang, Xiang Bai, Wenyu Liu 0001
Pattern Recognit.5
2017 Depth-Projection-Map-Based Bag of Contour Fragments for Robust Hand Gesture Recognition
abstract
This paper presents a novel and robust descriptor, depth-projection-map-based bag of contour fragments, which is applied to extraction of hand shape and structure information from depth maps. Our method projects depth maps onto three orthogonal planes to generate the depth projection maps. Then, the bag of contour fragment descriptors are extracted from the three depth projection maps and concatenated as a final shape representation of the original depth data. A support vector machine with a linear kernel is used as a shape classifier. The proposed description method is evaluated on three public datasets, as well as a new and more challenging dataset for hand gesture recognition. Results demonstrate that the proposed method significantly outperforms the previous methods on all tested datasets for both static digit recognition and letter gesture recognition. For the challenging HUST-ASL dataset, in particular, the proposed method improves on the previous state-of-the-art methods from 40.1% to 64.6%.
Bin Feng 0001, Fangzi He, Xinggang Wang, Yongjiang Wu, Hao Wang 0207, Sihua Yi, Wenyu Liu 0001
IEEE Trans. Hum. Mach. Syst.7
2017 Learning Multi-Instance Deep Discriminative Patterns for Image Classification
abstract
Finding an effective and efficient representation is very important for image classification. The most common approach is to extract a set of local descriptors, and then aggregate them into a high-dimensional, more semantic feature vector, like unsupervised bag-of-features and weakly supervised part-based models. The latter one is usually more discriminative than the former due to the use of information from image labels. In this paper, we propose a weakly supervised strategy that using multi-instance learning (MIL) to learn discriminative patterns for image representation. Specially, we extend traditional multi-instance methods to explicitly learn more than one patterns in positive class, and find the "most positive" instance for each pattern. Furthermore, as the positiveness of instance is treated as a continuous variable, we can use stochastic gradient decent to maximize the margin between different patterns meanwhile considering MIL constraints. To make the learned patterns more discriminative, local descriptors extracted by deep convolutional neural networks are chosen instead of hand-crafted descriptors. Some experimental results are reported on several widely used benchmarks (Action 40, Caltech 101, Scene 15, MIT-indoor, SUN 397), showing that our method can achieve very remarkable performance.
Peng Tang 0005, Xinggang Wang, Bin Feng 0001, Wenyu Liu 0001
IEEE Trans. Image Process.4
2016 Multi-oriented Text Detection with Fully Convolutional Networks
abstract
In this paper, we propose a novel approach for text detection in natural images. Both local and global cues are taken into account for localizing text lines in a coarse-to-fine procedure. First, a Fully Convolutional Network (FCN) model is trained to predict the salient map of text regions in a holistic manner. Then, text line hypotheses are estimated by combining the salient map and character components. Finally, another FCN classifier is used to predict the centroid of each character, in order to remove the false hypotheses. The framework is general for handling text in multiple orientations, languages and fonts. The proposed method consistently achieves the state-of-the-art performance on three text detection benchmarks: MSRA-TD500, ICDAR2015 and ICDAR2013.
Zheng Zhang 0022, Chengquan Zhang, Wei Shen 0002, Cong Yao, Wenyu Liu 0001, Xiang Bai
CVPR5
2016 Location-Aware Image Classification
Xinggang Wang, Xin Yang 0008, Wenyu Liu 0001, Chen Duan, Longin Jan Latecki
MMM (1)3
2016 Similarity Fusion for Visual Tracking
Yu Zhou 0016, Xiang Bai, Wenyu Liu 0001, Longin Jan Latecki
Int. J. Comput. Vis.3
2016 Individual adaptive metric learning for visual tracking
Sihua Yi, Nan Jiang 0016, Xinggang Wang, Wenyu Liu 0001
Neurocomputing4
2016 Traffic sign detection and recognition using fully convolutional network guided proposals
Yingying Zhu 0005, Chengquan Zhang, Duoyou Zhou, Xinggang Wang, Xiang Bai, Wenyu Liu 0001
Neurocomputing6
2016 Representing conditional preference by boosted regression trees for recommendation
Chenhong Sui, Dewei Deng, Bin Feng 0001, Wenyu Liu 0001, Caihua Wu
Inf. Sci.6
2016 Unsupervised local deep feature for image recognition
Xinggang Wang, Wenyu Liu 0001
Inf. Sci.3
2016 Online similarity learning for visual tracking
Sihua Yi, Nan Jiang 0016, Bin Feng 0001, Xinggang Wang, Wenyu Liu 0001
Inf. Sci.5
2016 Recurrent neural network based recommendation for time heterogeneous feedback
Caihua Wu, Wenyu Liu 0001
Knowl. Based Syst.4
2016 Renal compartment segmentation in DCE-MRI images
Xin Yang 0008, Hung Le Minh, Kwang-Ting Cheng, Kyung Hyun Sung, Wenyu Liu 0001
Medical Image Anal.5
2016 Minimum Cost Placement of Bistatic Radar Sensors for Belt Barrier Coverage
abstract
How to construct barrier coverage efficiently is a critical problem for many wireless sensor network applications, such as boundary surveillance and intrusion detection. In this paper, we study the belt barrier coverage in bistatic radar sensor networks. Much different from the disk and sector coverage, the coverage area of bistatic radar is dependent on the distance between a pair of radar transmitter and receiver. To improve coverage quality, we require to construct a belt barrier with the breadth not smaller than a predefined threshold. Furthermore, the unit cost of a radar transmitter may be different from a receiver. The bistatic radar placement problem is to construct a belt barrier with the minimum total placement cost. To solve the minimum cost placement problem, we propose a line-based equipartition placement strategy such that all radars placed on a deployment line can form a barrier with some breadth and one or more such placement lines can form a belt barrier with the required breadth. We first study the barrier property of different placement patterns on one deployment line, and prove the structure property of the optimal placement sequence on one deployment line. When multiple deployment lines are needed for belt barrier construction, we propose algorithms to find out the number of deployment lines and the number of receivers in the optimal placement pattern on each deployment line to minimize the total placement cost. The efficiency of the proposed algorithm is also validated by our simulation results.
Bang Wang 0001, Wenyu Liu 0001, Laurence T. Yang
IEEE Trans. Computers3
2016 Strokelets: A Learned Multi-Scale Mid-Level Representation for Scene Text Recognition
abstract
In this paper, we are concerned with the problem of automatic scene text recognition, which involves localizing and reading characters in natural images. We investigate this problem from the perspective of representation and propose a novel multi-scale representation, which leads to accurate, robust character identification and recognition. This representation consists of a set of mid-level primitives, termed strokelets, which capture the underlying substructures of characters at different granularities. The Strokelets possess four distinctive advantages: 1) usability: automatically learned from character level annotations; 2) robustness: insensitive to interference factors; 3) generality: applicable to variant languages; and 4) expressivity: effective at describing characters. Extensive experiments on standard benchmarks verify the advantages of the strokelets and demonstrate the effectiveness of the text recognition algorithm built upon the strokelets. Moreover, we show the method to incorporate the strokelets to improve the performance of scene text detection.
Xiang Bai, Cong Yao, Wenyu Liu 0001
IEEE Trans. Image Process.3
2016 Multiple Stage Residual Model for Image Classification and Vector Compression
abstract
Feature coding is a fundamental issue with many vision tasks, such as image classification, image retrieval and image segmentation, etc. There is no doubt that the encoding procedure leads to information loss, due to the existence of quantization error. The residual vector, defined as the difference between the feature and its corresponding visual word, is the chief culprit to be responsible for the quantization error. Many previous algorithms consider it as a coding issue, and focus on reducing the quantization error by reconstructing the feature with more than one visual word, or by the so-called soft-assignment strategy. In this paper, we consider the problem from a different point of view, and propose an effective and efficient model called multiple stage residual model (MSRM). It makes full use of the residual vector to generate a multiple stage code. MSRM is a hierarchical structure, with the bottom stage producing the coarsest quantization, and the top stage producing the finest quantization. Moreover, our proposed model is a generic framework, which can be built upon many coding algorithms. The interplay of such a coarse-to-fine quantization procedure with a discriminative classifier (e.g., SVM) can improve the classification accuracy of the baseline algorithms significantly. As a special case of MSRM, multiple stage vector quantization (MSVQ) can be directly used for vector compression and approximate nearest neighbor search, and achieves competitive performances with high efficiency.
Song Bai 0001, Xiang Bai, Wenyu Liu 0001
IEEE Trans. Multim.3
2016 Minimizing Content Reorganization and Tolerating Imperfect Workload Prediction for Cloud-Based Video-on-Demand Services
abstract
Video-on-demand (VoD) services historically rely on commercial content distribution networks (CDNs) for on-demand capacity provisioning. Content providers gradually prefer a self-managed content infrastructure because of its full control and customization. However, such a dedicated physical infrastructure could be costly in initial capital investment, and complex in management. It has become a promising alternative to host VoD services on pay-as-you-go cloud platforms, on which using dynamic server provisioning to reduce server rental cost is the key objective of content providers. In this paper we address two major challenges to reducing cost: to minimize content reorganization and to tolerate imperfect workload prediction. We first present a practical VoD servicing system design based on a pay-as-you-go cloud. We prove that previous works, focusing exclusively on cost savings, cause significant content reorganization and are vulnerable to imperfect workload prediction. To address such issues, we propose a novel idea called workload absorber, and design a provisioning algorithm called Absorb Window based on the idea. Workload absorbers eliminate the bandwidth wastage and significantly reduce content reorganization. We conduct extensive evaluations with real VoD access traces, and demonstrate the superior scalability of the proposed algorithm by producing highly optimized provisioning in seconds for thousands of servers.
Chen Tian 0001, Yi Wang 0049, Yan Luo 0001, Hongbo Jiang 0001, Wenyu Liu 0001, Jie Wu 0003
IEEE Trans. Serv. Comput.5
2015 Automatic Segmentation of Renal Compartments in DCE-MRI Images
Xin Yang 0008, Hung Le Minh, Kwang-Ting Cheng, Kyung Hyun Sung, Wenyu Liu 0001
MICCAI (1)5
2015 Conditional preference in recommender systems
Wenyu Liu 0001, Caihua Wu, Bin Feng 0001
Expert Syst. Appl.1
2015 B-spline-based shape coding with accurate distortion measurement using analytical model
Zhongyuan Lai, Zhen Zuo, Zhijun Yao, Wenyu Liu 0001
Neurocomputing5
2015 Constructing perimeter barrier coverage with bistatic radar sensors
Bang Wang 0001, Wenyu Liu 0001
J. Netw. Comput. Appl.3
2015 Neural shape codes for 3D model retrieval
Song Bai 0001, Xiang Bai, Wenyu Liu 0001, Fabio Roli
Pattern Recognit. Lett.3
2015 The Optimal Node Placement for Long Belt Coverage in Wireless Networks
abstract
The optimal node placement for a very large plane without boundary effect has been proven to be the regular triangular-lattice pattern in 1939. However, the regular triangular-lattice placement may not be optimal in a long belt with an upper and lower boundary. This paper proposes an optimal node deployment pattern to minimize the number of nodes for completely covering a long belt. The optimal pattern uses shifted node strips for belt coverage, and we compute the best node distance, strip offset, and strip distance for different belt heights. Mathematical analysis are provided to prove its optimality in terms of the minimum node density for belt coverage. Numerical computations are used to show its superiority, compared with other well-known placement patterns and our previously proposed equipartition placement.
Bang Wang 0001, Han Xu 0003, Wenyu Liu 0001, Laurence T. Yang
IEEE Trans. Computers3
2015 Sensor Scheduling for Multi-Modal Confident Information Coverage in Sensor Networks
abstract
Network lifetime maximization with guaranteed coverage is an important issue in wireless sensor networks. Based on our recently proposed confident information coverage (CIC) model, this paper studies the multi-modal confident information coverage (M2CIC) problem. Assuming that each node is equipped with different types of sensors, the objective is to schedule the multi-modal sensors' activity, such that the confident information coverage for each sensing modality can be guaranteed while the network lifetime can be maximized. We model the M2CIC problem as a multi-modal set cover problem (M2SC) and prove its NP-completeness. For solving the M2SC problem, we design two energy-efficient heuristics including a centralized one and a distributed one. In the proposed algorithms, different modal sensors are organized into a family of set covers, each of which can provide confident information coverage for all the monitored physical phenomena. Simulation results show that both the proposed algorithms can efficiently prolong the network lifetime and outperform two classical peer algorithms in terms of the extended network lifetime.
Xianjun Deng, Bang Wang 0001, Wenyu Liu 0001, Laurence T. Yang
IEEE Trans. Parallel Distributed Syst.3
2015 An Approximate Convex Decomposition Protocol for Wireless Sensor Network Localization in Arbitrary-Shaped Fields
abstract
Accurate localization in wireless sensor networks is the foundation for many applications, such as geographic routing and position-aware data processing. In this paper, we develop a new localization protocol based on approximate convex decomposition (ACDL), with reliance on network connectivity information only. ACDL can calculate the node virtual locations for a large-scale sensor network with a complex shape. We first examine one representative localization algorithm and study the influential factors on the localization accuracy, including the sharpness of the angle at the concave point and the depth of the concave valley. We show that after decomposition, the depth of the concave valley becomes irrelevant. We thus define the concavity according to the angle at a concave point, which reflects the localization error. We then propose ACDL protocol for network localization. It consists of four main steps. First, convex and concave nodes are recognized and network boundaries are segmented. As the sensor network is discrete, we show that it is acceptable to approximately identify the concave nodes to control the localization error. Second, an approximate convex decomposition is conducted. Our convex decomposition requires only local information and we show that it has low message overhead. Third, for each convex section of the network, an improved MDS algorithm is proposed to compute a relative location map. Fourth, a fast and low complexity merging algorithm is developed to construct the global location map. Besides, by slight modification on the third step, we propose a variant of ACDL, denoted by ACDL-Tri, which is fully distributed and scalable while the localization accuracy is still comparable. We finally show the efficiency of ACDL by extensive simulations.
Wenping Liu 0001, Dan Wang 0002, Hongbo Jiang 0001, Wenyu Liu 0001, Chonggang Wang
IEEE Trans. Parallel Distributed Syst.4
2014 Strokelets: A Learned Multi-scale Representation for Scene Text Recognition
abstract
Driven by the wide range of applications, scene text detection and recognition have become active research topics in computer vision. Though extensively studied, localizing and reading text in uncontrolled environments remain extremely challenging, due to various interference factors. In this paper, we propose a novel multi-scale representation for scene text recognition. This representation consists of a set of detectable primitives, termed as strokelets, which capture the essential substructures of characters at different granularities. Strokelets possess four distinctive advantages: (1) Usability: automatically learned from bounding box labels, (2) Robustness: insensitive to interference factors, (3) Generality: applicable to variant languages, and (4) Expressivity: effective at describing characters. Extensive experiments on standard benchmarks verify the advantages of strokelets and demonstrate the effectiveness of the proposed algorithm for text recognition.
Cong Yao, Xiang Bai, Baoguang Shi, Wenyu Liu 0001
CVPR4
2014 Adaptive Edge Encoding Schemes for the Rate-Distortion Optimal Polygon-Based Shape Coding
abstract
In this paper, we present two adaptive edge encoding schemes for the operational rate-distortion optimal polygon-based shape coding. The encoding edge is represented by an octant number, a major component, and a minor component, where the ranges of the two components are determined at two levels. For the object-level, these ranges are either determined by users or adaptive to the contour characteristics and the predefined admissible distortions using the discrete contour evolution method. For the edge-level, the range of the minor component is further adaptive to the magnitude of the major component. The appropriate code tables are selected for the two components according to their ranges. Experiments on MPEG-4 test sequences showed that our schemes outperform existing schemes in terms of bit-rate at the same distortion level.
Junhuan Zhu, Zhongyuan Lai, Wenyu Liu 0001, Jiebo Luo 0001
DCC3
2014 Human Detection Using Learned Part Alphabet and Pose Dictionary
Cong Yao, Xiang Bai, Wenyu Liu 0001, Longin Jan Latecki
ECCV (5)3
2014 Online Multiple targets Detection and Tracking from Mobile robot in Cluttered indoor Environments with Depth Camera
abstract
Indoor environment is a common scene in our everyday life, and detecting and tracking multiple targets in this environment is a key component for many applications. However, this task still remains challenging due to limited space, intrinsic target appearance variation, e.g. full or partial occlusion, large pose deformation, and scale change. In the proposed approach, we give a novel framework for detection and tracking in indoor environments, and extend it to robot navigation. One of the key components of our approach is a virtual top view created from an RGB-D camera, which is named ground plane projection (GPP). The key advantage of using GPP is the fact that the intrinsic target appearance variation and extrinsic noise is far less likely to appear in GPP than in a regular side-view image. Moreover, it is a very simple task to determine free space in GPP without any appearance learning even from a moving camera. Hence GPP is very different from the top-view image obtained from a ceiling mounted camera. We perform both object detection and tracking in GPP. Two kinds of GPP images are utilized: gray GPP, which represents the maximal height of 3D points projecting to each pixel, and binary GPP, which is obtained by thresholding the gray GPP. For detection, a simple connected component labeling is used to detect footprints of targets in binary GPP. For tracking, a novel Pixel Level Association (PLA) strategy is proposed to link the same target in consecutive frames in gray GPP. It utilizes optical flow in gray GPP, which to our best knowledge has never been done before. Then we "back project" the detected and tracked objects in GPP to original, side-view (RGB) images. Hence we are able to detect and track objects in the side-view (RGB) images. Our system is able to robustly detect and track multiple moving targets in real time. The detection process does not rely on any target model, which means we do not need any training process. Moreover, tracking does not require any manual initialization, since all entering objects are robustly detected. We also extend the novel framework to robot navigation by tracking. As our experimental results demonstrate, our approach can achieve near prefect detection and tracking results. The performance gain in comparison to state-of-the-art trackers is most significant in the presence of occlusion and background clutter.
Yu Zhou 0016, Yinfei Yang, Meng Yi, Xiang Bai, Wenyu Liu 0001, Longin Jan Latecki
Int. J. Pattern Recognit. Artif. Intell.5
2014 List-wise probabilistic matrix factorization for recommendation
Caihua Wu, Wenyu Liu 0001
Inf. Sci.4
2014 Robust Subspace Discovery via Relaxed Rank Minimization
abstract
This letter examines the problem of robust subspace discovery from input data samples (instances) in the presence of overwhelming outliers and corruptions. A typical example is the case where we are given a set of images; each image contains, for example, a face at an unknown location of an unknown size; our goal is to identify or detect the face in the image and simultaneously learn its model. We employ a simple generative subspace model and propose a new formulation to simultaneously infer the label information and learn the model using low-rank optimization. Solving this problem enables us to simultaneously identify the ownership of instances to the subspace and learn the corresponding subspace model. We give an efficient and effective algorithm based on the alternating direction method of multipliers and provide extensive simulations and experiments to verify the effectiveness of our method. The proposed scheme can also be used to tackle many related high-dimensional combinatorial selection problems.
Xinggang Wang, Zhengdong Zhang 0001, Yi Ma 0001, Xiang Bai, Wenyu Liu 0001, Zhuowen Tu
Neural Comput.5
2014 Bag of contour fragments for robust shape classification
Xinggang Wang, Bin Feng 0001, Xiang Bai, Wenyu Liu 0001, Longin Jan Latecki
Pattern Recognit.4
2014 Subarea Localization Performance of the Divide-and-Cover Node Deployment in a Long-Bounded Belt Scenario
abstract
Subarea localization has been recently proposed to locate a mobile device within a certain subarea delimitated by the overlapping ranges of monitoring nodes. In our previous work, we have proposed a divide-and-cover node placement for complete coverage of a long belt, but the subarea localization has not been considered in the design of such a placement. This paper studies the subarea localization performance for the divide-and-cover deployment. We first obtain a formula to compute the mean subarea localization error of the whole belt and then analyze the optimal placement parameter to minimize the subarea localization error. Theoretical analysis show that compared with other node placement including the well-known regular triangular-lattice placement, the proposed scheme can achieve a lower mean localization error. Field experiments validate the effectiveness of the proposed scheme.
Han Xu 0003, Wenyu Liu 0001, Bang Wang 0001
IEEE Trans. Computers2
2014 Data-Driven Spatially-Adaptive Metric Adjustment for Visual Tracking
abstract
Matching visual appearances of the target over consecutive video frames is a fundamental yet challenging task in visual tracking. Its performance largely depends on the distance metric that determines the quality of visual matching. Rather than using fixed and predefined metric, recent attempts of integrating metric learning-based trackers have shown more robust and promising results, as the learned metric can be more discriminative. In general, these global metric adjustment methods are computationally demanding in real-time visual tracking tasks, and they tend to underfit the data when the target exhibits dynamic appearance variation. This paper presents a nonparametric data-driven local metric adjustment method. The proposed method finds a spatially adaptive metric that exhibits different properties at different locations in the feature space, due to the differences of the data distribution in a local neighborhood. It minimizes the deviation of the empirical misclassification probability to obtain the optimal metric such that the asymptotic error as if using an infinite set of training samples can be approximated. Moreover, by taking the data local distribution into consideration, it is spatially adaptive. Integrating this new local metric learning method into target tracking leads to efficient and robust tracking performance. Extensive experiments have demonstrated the superiority and effectiveness of the proposed tracking method in various tracking scenarios.
Nan Jiang 0016, Wenyu Liu 0001
IEEE Trans. Image Process.2
2014 A Unified Framework for Multioriented Text Detection and Recognition
abstract
High level semantics embodied in scene texts are both rich and clear and thus can serve as important cues for a wide range of vision applications, for instance, image understanding, image indexing, video search, geolocation, and automatic navigation. In this paper, we present a unified framework for text detection and recognition in natural images. The contributions of this paper are threefold: 1) text detection and recognition are accomplished concurrently using exactly the same features and classification scheme; 2) in contrast to methods in the literature, which mainly focus on horizontal or near-horizontal texts, the proposed system is capable of localizing and reading texts of varying orientations; and 3) a new dictionary search method is proposed, to correct the recognition errors usually caused by confusions among similar yet different characters. As an additional contribution, a novel image database with texts of different scales, colors, fonts, and orientations in diverse real-world scenarios, is generated and released. Extensive experiments on standard benchmarks as well as the proposed database demonstrate that the proposed system achieves highly competitive performance, especially on multioriented texts.
Cong Yao, Xiang Bai, Wenyu Liu 0001
IEEE Trans. Image Process.3
2014 Vehicle Color Recognition on Urban Road by Feature Context
abstract
Vehicle information recognition is a key component of intelligent transportation systems. Color plays an important role in vehicle identification. As a vehicle has its inner structure, the main challenge of vehicle color recognition is to select the region of interest (ROI) for recognizing its dominant color. In this paper, we propose a method to implicitly select the ROI for color recognition. Preprocessing is performed to overcome the influence of image quality degradation. Then, the ROI in vehicle images is selected by assigning the subregions with different weights that are learned by a classifier trained on the vehicle images. We train the classifier by linear support vector machine for its efficiency and high precision. The experiments are extensively validated on both images and videos, which are collected on urban roads. The proposed method outperforms other competing color recognition methods.
Pan Chen 0004, Xiang Bai, Wenyu Liu 0001
IEEE Trans. Intell. Transp. Syst.3
2014 Learning Conditional Preference Networks from Inconsistent Examples
abstract
The problem of learning conditional preference networks (CP-nets) from a set of examples has received great attention recently. However, because of the randomicity of the users' behaviors and the observation errors, there is always some noise making the examples inconsistent, namely, there exists at least one outcome preferred over itself (by transferring) in examples. Existing CP-nets learning methods cannot handle inconsistent examples. In this work, we introduce the model of learning consistent CP-nets from inconsistent examples and present a method to solve this model. We do not learn the CP-nets directly. Instead, we first learn a preference graph from the inconsistent examples, because dominance testing and consistency testing in preference graphs are easier than those in CP-nets. The problem of learning preference graphs is translated into a 0-1 programming and is solved by the branch-and-bound search. Then, the obtained preference graph is transformed into a CP-net equivalently, which can entail a subset of examples with maximal sum of weight. Examples are given to show that our method can obtain consistent CP-nets over both binary and multivalued variables from inconsistent examples. The proposed method is verified on both simulated data and real data, and it is also compared with existing methods.
Caihua Wu, Zhijun Yao, Wenyu Liu 0001
IEEE Trans. Knowl. Data Eng.5
2013 Learning context sensitive similarity measure on pair fusion graph
abstract
In this paper, we present a new approach for shape/image retrieval by efficiently fusing different shape similarities, called Pair-Graph Diffusion. Different from other algorithms which linearly integrate different similarity measures, our algorithm adopts Tensor Product Graph(TPG) to combine two shape similarities by fusing two single-graphs into a multi-graph for fusion process. In such way, we gain more shape information in a higher order, and the multigraph is able to better reveal the intrinsic relation between shapes especially when the two input similarities are very complementary. We perform the experiments on two popular image datasets: MPEG-7 shape dataset and Nistér and Stewénius (N-S) dataset, and achieve state-of-arts retrieval rates: 98.87% on MPEG-7 dataset and 3.69 on N-S dataset. The results demonstrate that the proposed method can effectively fuse two similarities. In addition, Multi-graph Diffusion is a general similarity learning algorithm, and it can be easily applied other tasks for ranking/retrieval.
Cheng Wang 0048, Yingying Zhu 0005, Wenyu Liu 0001
ICIP4
2013 On contrast combinations for visual saliency detection
abstract
Saliency detection is an important task in computer vision and image processing. The most influential factor in bottom-up visual saliency is contrast operation. In this paper, we propose a unified model to combine widely used contrast measurements, namely, center-surround, corner-surround and global contrast to detect visual saliency. The proposed model benefits from the advantages of each individual contrast operation, and thus produces more robust and accurate saliency maps. Extensive experimental results on natural images show the effectiveness of the proposed model for visual saliency detection task, and demonstrate the combination is superior than individual subcomponent.
Quan Zhou 0004, Shiwei Ren, Yu Zhou 0016, Jun Chen 0019, Wenyu Liu 0001
ICIP6
2013 Max-Margin Multiple-Instance Dictionary Learning
abstract
Dictionary learning has became an increasingly important task in machine learning, as it is fundamental to the representation problem. A number of emerging techniques specifically include a codebook learning step, in which a critical knowledge abstraction process is carried out. Existing approaches in dictionary (codebook) learning are either generative (unsupervised e.g. k-means) or discriminative (supervised e.g. extremely randomized forests). In this paper, we propose a multiple instance learning (MIL) strategy (along the line of weakly supervised learning) for dictionary learning. Each code is represented by a classifier, such as a linear SVM, which naturally performs metric fusion for multi-channel features. We design a formulation to simultaneously learn mixtures of codes by maximizing classification margins in MIL. State-of-the-art results are observed in image classification benchmarks based on the learned codebooks, which observe both compactness and effectiveness.
Xinggang Wang, Baoyuan Wang, Xiang Bai, Wenyu Liu 0001, Zhuowen Tu
ICML (3)4
2013 Improvement of soil moisture monitoring using EVI as a key parameter based on TVDI in the north China plain
abstract
Soil moisture is a key component of land surface parameterization. The triangle/trapezoid feature space can be used to monitor soil moisture effectively. This research aims to use enhanced vegetation index (EVI) as an alternative for the normalized difference vegetation index (NDVI) in estimation of temperature vegetation dryness index (TVDI) to improve its ability of retrieving soil moisture and to discuss the flexibility of EVI in retrieval of soil moisture. The result shows that LST-EVI has higher R2in linear regression and more significant in validation at depth of 0–10cm than LST-NDVI. Influence caused by altitude for LST-EVI in monitoring soil moisture will be discussed in this study. In comparison with precipitation, LST-EVI and soil moisture shows a better mirror symmetry in dry season than in rainy season.
Yue Shan, Adu Gong, Yongrong Su, Wenyu Liu 0001, Jing Li 0018, Weiguo Jiang
IGARSS4
2013 Sensor scheduling for confident information coverage in wireless sensor networks
abstract
Many applications in wireless sensor networks have strict coverage accuracy requirements and need to operate as long as possible. In this paper, based on the new confident information coverage model proposed in our previous study (Wang et al., 2012), we design a novel sensor scheduling algorithm to prolong the network lifetime. This algorithm organizes the sensors into a maximal number of set covers, each capable of providing required coverage. The task of reconstructing the physical phenomena will be accomplished only by the sensors from one working set cover, while all the other sensors are in sleep mode. And we rotate the working set cover to prolong the network lifetime. Our simulation results show that the proposed algorithm outperforms two typical peer algorithms in terms of longer network lifetime.
Xianjun Deng, Bang Wang 0001, Nuoya Wang, Wenyu Liu 0001, Yijun Mo
WCNC4
2013 Mending barrier gaps via mobile sensor nodes with adjustable sensing ranges
abstract
Barrier coverage is an important topic in wireless sensor networks. When sensors are randomly deployed, barrier gaps may occur if the number of deployed sensors is not large enough or some sensors start malfunctioning or run out of energy. How to efficiently mend these barrier gaps is an important research issue. In this paper, we study the gap mending problem in a hybrid sensor network which consists of both stationary and mobile sensors with adjustable sensing ranges. We propose two gap mending schemes: the min-max scheme and the max-lifetime scheme. The first is to minimize the maximal energy consumption to move sensors, and the second is to maximize the lifetime of barrier coverage after mending all gaps. Simulation results show that the min-max scheme can achieve a lower maximal moving distance and the max-lifetime scheme can efficiently extend the barrier lifetime.
Xianjun Deng, Bang Wang 0001, Han Xu 0003, Wenyu Liu 0001
WCNC5
2013 Bayesian Probabilistic Matrix Factorization with Social Relations and Item Contents for recommendation
Caihua Wu, Wenyu Liu 0001
Decis. Support Syst.3
2013 Extracting robust distribution using adaptive Gaussian Mixture Model and online feature selection
Zhijun Yao, Wenyu Liu 0001
Neurocomputing2
2013 Perceptually friendly shape decomposition by resolving segmentation points with minimum cost
Wenyu Liu 0001, Zhongyuan Lai
J. Vis. Commun. Image Represent.2
2013 Learning conditional preference network from noisy samples using hypothesis testing
Zhijun Yao, Wenyu Liu 0001, Caihua Wu
Knowl. Based Syst.4
2013 A Sink-Oriented Layered Clustering Protocol for Wireless Sensor Networks
Yijun Mo, Bang Wang 0001, Wenyu Liu 0001, Laurence T. Yang
Mob. Networks Appl.3
2013 Minimum Near-Convex Shape Decomposition
abstract
Shape decomposition is a fundamental problem for part-based shape representation. We propose the minimum near-convex decomposition (MNCD) to decompose arbitrary shapes into minimum number of "near-convex" parts. The near-convex shape decomposition is formulated as a discrete optimization problem by minimizing the number of nonintersecting cuts. Two perception rules are imposed as constraints into our objective function to improve the visual naturalness of the decomposition. With the degree of near-convexity a user-specified parameter, our decomposition is robust to local distortions and shape deformation. The optimization can be efficiently solved via binary integer linear programming. Both theoretical analysis and experiment results show that our approach outperforms the state-of-the-art results without introducing redundant parts and thus leads to robust shape representation.
Zhou Ren, Junsong Yuan 0001, Wenyu Liu 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2013 A Novel Node Placement for Long Belt Coverage in Wireless Networks
abstract
Coverage is an important issue in many wireless networks. In this paper, we address the problem of node placement for ensuring complete coverage in a long belt scenario and propose a novel placement approach to minimize the number of nodes needed. In our work, each node is assumed to be able to cover a disk area centered at itself with a fixed radius, then a divide-and-cover node placement method is proposed. In the proposed method, a long belt is divided into some sub-belts (if necessary), and then a string of nodes are placed parallel to the long side of each sub-belt to completely cover the sub-belt. We then determine the optimal distance between two adjacent nodes in a string and the number of such strings to minimize the number of nodes for complete belt coverage. Theoretical proofs and analysis show that compared with other node placement including the well-known regular triangular-lattice placement, the proposed method can achieve lower node density in some cases when the belt height is not very large. A combination of the proposed method and the triangular-lattice placement is then proposed, and the optimal ranges of the belt height for their respective applications to achieve the lowest node density are computed.
Bang Wang 0001, Han Xu 0003, Wenyu Liu 0001
IEEE Trans. Computers3
2013 Learning Dynamic Hybrid Markov Random Field for Image Labeling
abstract
Using shape information has gained increasing concerns in the task of image labeling. In this paper, we present a dynamic hybrid Markov random field (DHMRF), which explicitly captures middle-level object shape and low-level visual appearance (e.g., texture and color) for image labeling. Each node in DHMRF is described by either a deformable template or an appearance model as visual prototype. On the other hand, the edges encode two types of intersections: co-occurrence and spatial layered context, with respect to the labels and prototypes of connected nodes. To learn the DHMRF model, an iterative algorithm is designed to automatically select the most informative features and estimate model parameters. The algorithm achieves high computational efficiency since a branch-and-bound schema is introduced to estimate model parameters. Compared with previous methods, which usually employ implicit shape cues, our DHMRF model seamlessly integrates color, texture, and shape cues to inference labeling output, and thus produces more accurate and reliable results. Extensive experiments validate its superiority over other state-of-the-art methods in terms of recognition accuracy and implementation efficiency on: 1) the MSRC 21-class dataset, and 2) the lotus hill institute 15-class dataset.
Quan Zhou 0004, Wenyu Liu 0001
IEEE Trans. Image Process.3
2013 Distance Transform-Based Skeleton Extraction and Its Applications in Sensor Networks
abstract
We study the problem of skeleton extraction for large-scale sensor networks with reliance purely on connectivity information. Existing efforts in this line highly depend on the boundary detection algorithms, which are used to extract accurate boundary nodes. One challenge is that in practical this could limit the applicability of the boundary detection algorithms. For instance, in low node density networks where boundary detection algorithms do not work well, the extracted boundary nodes are often incomplete. This paper brings a new view to skeleton extraction from a distance transform perspective, bridging the distance transform of the network and the incomplete boundaries. As such, we propose a distributed and scalable algorithm for skeleton extraction, called DIST, based on DIStance Transform, while incurring low communication overhead. The proposed algorithm does not require that the boundaries are complete or accurate, which makes the proposed algorithm more practical in applications. First, we compute the distance transform of the network. Specifically, the distance (hop count) of each node to the boundaries of a sensor network is estimated. The node map consisting of the distance values is considered as the distance transform (the distance map). The distance map is then used to identify skeleton nodes. Next, skeleton arcs are generated by controlled flooding within the identified skeleton nodes, thereby connecting these skeleton arcs, to extract a coarse skeleton. Finally, we refine the coarse skeleton by building shortest path trees followed by a prune phase. The obtained skeleton is robust to boundary noise or shape variations. Besides, we present two specific applications that benefit from the extracted skeleton: identifying complete boundaries and shape segmentation. First, with the extracted skeleton using DIST, we propose to identify more boundary nodes to form a meaningful boundary curve. Second, the utilization of the derived skeleton to segment the network into approximately convex pieces has been shown to be effective.
Wenping Liu 0001, Hongbo Jiang 0001, Xiang Bai, Guang Tan, Chonggang Wang, Wenyu Liu 0001, Kechao Cai
IEEE Trans. Parallel Distributed Syst.6
2012 One-Class Multiple Instance Learning via Robust PCA for Common Object Discovery
Xinggang Wang, Zhengdong Zhang 0001, Yi Ma 0001, Xiang Bai, Wenyu Liu 0001, Zhuowen Tu
ACCV (1)5
2012 Order determination and sparsity-regularized metric learning adaptive visual tracking
abstract
Recent attempts of integrating metric learning in visual tracking have produced encouraging results. Instead of using fixed and pre-specified metric in visual appearance matching, these methods are able to learn and adjust the metric adaptively by finding the best projection of the feature space. Such learned metric is by design the best to discriminate the target of interest and its distracters from the background. However, an important issue remained unaddressed is how we can determine the optimal dimensionality of the projection to achieve best discrimination. Using inappropriate dimensions for the projection is likely to result in larger classification error, or higher computational costs and over-fitting. This paper presents a novel solution to this structural order determination problem, by introducing sparsity regularization for metric learning (or SRML). This regularization leads to the lowest possible dimensionality of the projection and thus determining the best order. This can actually be viewed as the minimum description length regularization in metric learning. The experiments validate this new approach on standard benchmark datasets, and demonstrate its effectiveness in visual tracking applications.
Nan Jiang 0016, Wenyu Liu 0001, Ying Wu 0001
CVPR2
2012 Fan Shape Model for object detection
abstract
We propose a novel shape model for object detection called Fan Shape Model (FSM). We model contour sample points as rays of final length emanating for a reference point. As in folding fan, its slats, which we call rays, are very flexible. This flexibility allows FSM to tolerate large shape variance. However, the order and the adjacency relation of the slats stay invariant during fan deformation, since the slats are connected with a thin fabric. In analogy, we enforce the order and adjacency relation of the rays to stay invariant during the deformation. Therefore, FSM preserves discriminative power while allowing for a substantial shape deformation. FSM allows also for precise scale estimation during object detection. Thus, there is not need to scale the shape model or image in order to perform object detection. Another advantage of FSM is the fact that it can be applied directly to edge images, since it does not require any linking of edge pixels to edge fragments (contours).
Xinggang Wang, Xiang Bai, Tianyang Ma, Wenyu Liu 0001, Longin Jan Latecki
CVPR4
2012 Detecting texts of arbitrary orientations in natural images
abstract
With the increasing popularity of practical vision systems and smart phones, text detection in natural scenes becomes a critical yet challenging task. Most existing methods have focused on detecting horizontal or near-horizontal texts. In this paper, we propose a system which detects texts of arbitrary orientations in natural images. Our algorithm is equipped with a two-level classification scheme and two sets of features specially designed for capturing both the intrinsic characteristics of texts. To better evaluate our algorithm and compare it with other competing algorithms, we generate a new dataset, which includes various texts in diverse real-world scenarios; we also propose a protocol for performance evaluation. Experiments on benchmark datasets and the proposed dataset demonstrate that our algorithm compares favorably with the state-of-the-art algorithms when handling horizontal texts and achieves significantly enhanced performance on texts of arbitrary orientations in complex natural scenes.
Cong Yao, Xiang Bai, Wenyu Liu 0001, Yi Ma 0001, Zhuowen Tu
CVPR3
2012 Skeleton Extraction from Incomplete Boundaries in Sensor Networks Based on Distance Transform
abstract
This paper proposes a novel approach, named DIST, to skeleton extraction from incomplete boundaries using the idea of {\em distance transform}, a concept in the computer graphics area. The main contribution is a distributed and low-cost algorithm that produces accurate network skeletons without requiring that the boundaries be complete or tight. The algorithm first establishes the network's distance transform -- the hop distance of each node to the network's boundaries. Based on this, some {\em critical skeleton nodes} are identified. Next, a set of {\em skeleton arcs} are generated by controlled flooding, connecting these skeleton arcs then gives us a coarse skeleton. The algorithm finally refines the coarse skeleton by building shortest path trees, followed by a prune phase. The obtained skeletons are robust to boundary noise and shape variations.
Wenping Liu 0001, Hongbo Jiang 0001, Xiang Bai, Guang Tan, Chonggang Wang, Wenyu Liu 0001, Kechao Cai
ICDCS6
2012 Connectivity-based and Boundary-Free Skeleton Extraction in Sensor Networks
abstract
In sensor networks, skeleton (also known as medial axis) extraction is recognized as an appealing approach to support many applications such as load-balanced routing and location free segmentation. Existing solutions in the literature rely heavily on the identified boundaries, which puts limitations on the applicability of the skeleton extraction algorithm. In this paper, we conduct the first work of a connectivity-based and boundary free skeleton extraction scheme, in sensor networks. In detail, we propose a simple, distributed and scalable algorithm that correctly identifies a few skeleton nodes and connects them into a meaningful representation of the network, without reliance on any constraint on communication radio model or boundary information. The key idea of our algorithm is to exploit the necessary (but not sufficient) condition of skeleton points: the intersection area of the disk centered at a skeleton point x should be the largest one as compared to other points on the chord generated by x, where the chord is referred to as the line segment connecting x and the tangent point in the boundary. To that end, we present the concept of \epsilon-centrality of a point, quantitatively measuring how "central" a point is. Accordingly, a skeleton point should have the largest value of \epsilon-centrality as compared to other points on the chord generated by this point. Our simulation results show that the proposed algorithm works well even for networks with low node density or skewed nodal distribution, etc. In addition, we obtain two by-products, the boundaries and the segmentation result of the network.
Wenping Liu 0001, Hongbo Jiang 0001, Chonggang Wang, Yang Yang 0060, Wenyu Liu 0001, Bo Li 0001
ICDCS6
2012 Online Random Ferns for robust visual tracking
Cong Rao, Cong Yao, Xiang Bai, Weichao Qiu, Wenyu Liu 0001
ICPR5
2012 Adjacent coding for image classification
Xinggang Wang, Shaojun Zhu, Xiang Bai, Wenyu Liu 0001
ICPR5
2012 A fast and effective appearance model-based particle filtering object tracking algorithm
Zhijun Yao, Wenyu Liu 0001
ICPR4
2012 Corner-surround Contrast for saliency detection
Quan Zhou 0004, Nianyi Li, Pan Chen 0004, Wenyu Liu 0001
ICPR5
2012 Retrieval of land surface temperature (LST) based on Support Vector Machine (SVM) from HJ-1B data with single-channel
abstract
Land surface temperature (LST) is a very key variable for land surface process research. However, the retrieval of LST is still underdetermined issue because of the fact that the unknowns are always more than the measurements even the atmospheric condition can be acquired completely. Currently, Support Vector Machine (SVM) as an effective machine learning tool has been used widely in the domain of quantitative remote sensing because its optimization and generalization. This paper used SVM to retrieve LST based on only one thermal band in HJ-1B satellite launched by China. The radiance and water vapor content were selected as the independent variables. The validation result indicates that the errors of the SVM-MOD07 are lower than the Qin's-MOD07. Additionally, the sensitivity analysis indicates that when the errors of the water vapor content increase, the errors for the SVM model change insignificantly. In the end the SVM model was applied in Beijing area.
Adu Gong, Wenyu Liu 0001, Yue Shan, Jianwei Yue
IGARSS2
2012 Approximate convex decomposition based localization in wireless sensor networks
abstract
Accurate localization in wireless sensor networks is the foundation for many applications, such as geographic routing and position-aware data processing. An important research direction for localization is to develop schemes using connectivity information only. These schemes primary apply hop counts to distance estimation. Not surprisingly, they work well only when the network topology has a convex shape. In this paper, we develop a new Localization protocol based on Approximate Convex Decomposition (ACDL). It can calculate the node virtual locations for a large-scale sensor network with arbitrary shapes. The basic idea is to decompose the network into convex subregions. It is not straight-forward, however. We first examine the influential factors on the localization accuracy when the network is concave such as the sharpness of concave angle and the depth of the concave valley. We show that after decomposition, the depth of the concave valley becomes irrelevant. We thus define concavity according to the angle at a concave point, which can reflect the localization error. We then propose ACDL protocol for network localization. It consists of four main steps. First, convex and concave nodes are recognized and network boundaries are segmented. As the sensor network is discrete, we show that it is acceptable to approximately identify the concave nodes to control the localization error. Second, an approximate convex decomposition is conducted. Our convex decomposition requires only local information and we show that it has low message overhead. Third, for each convex subsection of the network, an improved Multi-Dimensional Scaling (MDS) algorithm is proposed to compute a relative location map. Fourth, a fast and low complexity merging algorithm is developed to construct the global location map. Our simulation on several representative networks demonstrated that ACDL has localization error that is 60%-90% smaller as compared with the typical MDS-MAP algorithm and 20%-30% smaller as compared to a recent state-of-the-art localization algorithm CATL.
Wenping Liu 0001, Dan Wang 0002, Hongbo Jiang 0001, Wenyu Liu 0001, Chonggang Wang
INFOCOM4
2012 Fusion with Diffusion for Robust Visual Tracking
abstract
A weighted graph is used as an underlying structure of many algorithms like semi-supervised learning and spectral clustering. The edge weights are usually deter-mined by a single similarity measure, but it often hard if not impossible to capture all relevant aspects of similarity when using a single similarity measure. In par-ticular, in the case of visual object matching it is beneficial to integrate different similarity measures that focus on different visual representations. In this paper, a novel approach to integrate multiple similarity measures is pro-posed. First pairs of similarity measures are combined with a diffusion process on their tensor product graph (TPG). Hence the diffused similarity of each pair of ob-jects becomes a function of joint diffusion of the two original similarities, which in turn depends on the neighborhood structure of the TPG. We call this process Fusion with Diffusion (FD). However, a higher order graph like the TPG usually means significant increase in time complexity. This is not the case in the proposed approach. A key feature of our approach is that the time complexity of the dif-fusion on the TPG is the same as the diffusion process on each of the original graphs, Moreover, it is not necessary to explicitly construct the TPG in our frame-work. Finally all diffused pairs of similarity measures are combined as a weighted sum. We demonstrate the advantages of the proposed approach on the task of visual tracking, where different aspects of the appearance similarity between the target object in frame t and target object candidates in frame t+1 are integrated. The obtained method is tested on several challenge video sequences and the experimental results show that it outperforms state-of-the-art tracking methods.
Yu Zhou 0016, Xiang Bai, Wenyu Liu 0001, Longin Jan Latecki
NIPS3
2012 Energy-Efficient Barrier Coverage in WSNs with Adjustable Sensing Ranges
abstract
Energy-efficient barrier coverage is an important issue in wireless sensor networks. In this paper we study the problem of how to maximize the lifetime of a barrier, where sensors have adjustable sensing ranges. In our approach, each node is divided into some virtual sub-nodes according to the available sensing ranges. For small-scale sensor networks, we construct a barrier coverage graph where a link exists between two sub-nodes in two different nodes, if their sensing ranges overlap. We propose to use a linear programming optimization method based on the exhaustive search of all possible barriers in the constructed graph to find the optimal barriers and their respective operation times. For large-scale sensor networks, we propose two distributed heuristics: one is to randomly select, from its neighboring sub-nodes, a next sub-node to construct barriers; another is to greedily select a next sub- node to best match the lifetime of the barrier constructed before choosing this sub-node. Simulation results show that compared with the randomized one, the greedy scheme can achieve longer lifetime and lower message overhead.
Bang Wang 0001, Han Xu 0003, Wenyu Liu 0001
VTC Spring4
2012 Shape matching and classification using height functions
Xiang Bai, Xinge You, Wenyu Liu 0001, Longin Jan Latecki
Pattern Recognit. Lett.4
2012 Co-Transduction for Shape Retrieval
abstract
In this paper, we propose a new shape/object retrieval algorithm, namely, co-transduction. The performance of a retrieval system is critically decided by the accuracy of adopted similarity measures (distances or metrics). In shape/object retrieval, ideally, intraclass objects should have smaller distances than interclass objects. However, it is a difficult task to design an ideal metric to account for the large intraclass variation. Different types of measures may focus on different aspects of the objects: for example, measures computed based on contours and skeletons are often complementary to each other. Our goal is to develop an algorithm to fuse different similarity measures for robust shape retrieval through a semisupervised learning framework. We name our method co-transduction, which is inspired by the co-training algorithm. Given two similarity measures and a query shape, the algorithm iteratively retrieves the most similar shapes using one measure and assigns them to a pool for the other measure to do a re-ranking, and vice versa. Using co-transduction, we achieved an improved result of 97.72% (bull's-eye measure) on the MPEG-7 data set over the state-of-the-art performance. We also present an algorithm called tri-transduction to fuse multiple-input similarities, and it achieved 99.06% on the MPEG-7 data set. Our algorithm is general, and it can be directly applied on input similarity measures/metrics; it is not limited to object shape retrieval and can be applied to other tasks for ranking/retrieval.
Xiang Bai, Bo Wang 0044, Cong Yao, Wenyu Liu 0001, Zhuowen Tu
IEEE Trans. Image Process.4
2012 Discriminative Metric Preservation for Tracking Low-Resolution Targets
abstract
Tracking low-resolution (LR) targets is a practical yet quite challenging problem in real video analysis applications. Lack of discriminative details in the visual appearance of the LR target leads to the matching ambiguity, which confronts most existing tracking methods. Although artificially enhancing the video resolution by superresolution (SR) techniques before analyzing might be an option, the high demand of computational cost can hardly meet the requirements of the tracking scenario. This paper presents a novel solution to track LR targets without explicitly performing SR. This new approach is based on discriminative metric preservation that preserves the data affinity structure in the high-resolution (HR) feature space for effective and efficient matching of LR images. In addition, we substantialize this new approach in a solid case study of differential tracking under metric preservation and derive a closed-form solution to motion estimation for LR video. In addition, this paper extends the basic linear metric preservation method to a more powerful nonlinear kernel metric preservation method. Such a solution to LR target tracking is discriminative, robust, and efficient. Extensive experiments validate the entrustments and effectiveness of the proposed approach and demonstrate the improved performance of the proposed method in tracking LR targets.
Nan Jiang 0016, Heng Su, Wenyu Liu 0001, Ying Wu 0001
IEEE Trans. Image Process.3
2012 Revisiting Dynamic Query Protocols in Unstructured Peer-to-Peer Networks
abstract
In unstructured peer-to-peer networks, the average response latency and traffic cost of a query are two main performance metrics. Controlled-flooding resource query algorithms are widely used in unstructured networks such as peer-to-peer networks. In this paper, we propose a novel algorithm named Selective Dynamic Query (SDQ). Based on mathematical programming, SDQ calculates the optimal combination of an integer TTL value and a set of neighbors to control the scope of the next query. Our results demonstrate that SDQ provides finer grained control than other algorithms: its response latency is close to the well-known minimum one via Expanding Ring; in the mean time, its traffic cost is also close to the minimum. To our best knowledge, this is the first work capable of achieving a best trade-off between response latency and traffic cost.
Chen Tian 0001, Hongbo Jiang 0001, Xue (Steve) Liu, Wenyu Liu 0001
IEEE Trans. Parallel Distributed Syst.4
2011 Tracking low resolution objects by metric preservation
abstract
Tracking low resolution (LR) targets is a practical yet quite challenging problem in real applications. The loss of discriminative details in the visual appearance of the L-R targets confronts most existing visual tracking methods. Although the resolution of the LR video inputs may be enhanced by super resolution (SR) techniques, the large computational cost for high-quality SR does not make it an attractive option. This paper presents a novel solution to track LR targets without performing explicit SR. This new approach is based on discriminative metric preservation that preserves the structure in the high resolution feature space for LR matching. In addition, we integrate metric preservation with differential tracking to derive a closed-form solution to motion estimation for LR video. Extensive experiments have demonstrated the effectiveness and efficiency of the proposed approach.
Nan Jiang 0016, Wenyu Liu 0001, Heng Su, Ying Wu 0001
CVPR2
2011 Adaptive and discriminative metric differential tracking
abstract
Matching the visual appearances of the target over consecutive image frames is the most critical issue in video-based object tracking. Choosing an appropriate distance metric for matching determines its accuracy and robustness, and significantly influences the tracking performance. This paper presents a new tracking approach that incorporates adaptive metric into differential tracking method. This new approach automatically learns an optimal distance metric for more accurate matching, and obtains a closed-form analytical solution to motion estimation and differential tracking. Extensive experiments validate the effectiveness of adaptive metric, and demonstrate the improved performance of the proposed new tracking method.
Nan Jiang 0016, Wenyu Liu 0001, Ying Wu 0001
CVPR2
2011 Feature context for image classification and object detection
abstract
In this paper, we presents a new method to encode the spatial information of local image features, which is a natural extension of Shape Context (SC), so we call it Feature Context (FC). Given a position in a image, SC computes histogram of other points belonging to the target binary shape based on their distances and angles to the position. The value of each histogram bin of SC is the number of the shape points in the region assigned to the bin. Thus, SC requires knowing the location of the points of the target shape. In other words, an image point can have only two labels, it belongs to the shape or not. In contrast, FC can be applied to the whole image without knowing the location of the target shape in the image. Each image point can have multiple labels depending on its local features. The value of each histogram bin of FC is a histogram of various features assigned to points in the bin region. We also introduce an efficient coding method to encode the local image features, call Radial Basis Coding (RBC). Combining RBC and FC together, and using a linear SVM classifier, our method is suitable for both image classification and object detection.
Xinggang Wang, Xiang Bai, Wenyu Liu 0001, Longin Jan Latecki
CVPR3
2011 A Hybrid Admissible Distortion Checking Algorithm for the B-Spline-Based Operational Rate-Distortion Optimal Shape Coding
abstract
Admissible distortion checking algorithm plays a very important role in both rate distortion performance and computational efficiency of B-spline-based operational rate-distortion optimal shape coding framework under the minimum-maximum criterion. Existing distortion measurement using chord-length parameterization (DMCLP) is fast but results in extra bit-rate problem. In contrast, the up to date accurate distortion measurement using analytical model (ADMAM) can achieve the smallest bit-rate but is very time consuming. It motivates us to develop a hybrid admissible checking algorithm that can take full use of each advantage. Recalling the definitions of both DMCLP and ADMAM for each associated contour point, the authors show that DMCLP is the distance from the parameterized B-spline point while ADMAM is the shortest distance from the approximating B-splines.
Zhongyuan Lai, Zhen Zuo, Wenyu Liu 0001
DCC4
2011 Accurate Distortion Measurement Using Analytical Model for the B-Spline-Based Shape Coding
abstract
Summary form only given. Existing distortion measurements for the B-spline-based shape coding include ap proximation, quantization, or parameterization process, so they are approximate techniques. They may inaccurately predict the actual distortion value, which motivates us to construct a model that can accurately measure the actual distortion. It was reported that the actual distortion for reconstruction quality assessment is the minimal Euclidean distance between each associated contour point and the reconstruction contour.
Zhongyuan Lai, Zhen Zuo, Wenyu Liu 0001
DCC4
2011 SHARP: A Scalable Framework for Dynamic Joint Replica Placement and Request Routing Scheduling
abstract
This paper presents SHARP: a scalable framework for Dynamic Joint Replica Placement and Request Routing (DJRPRR) scheduling in content delivery networks. After grouping similar proxies and modeling them by a single section, we propose a hierarchical scheduling framework to greatly reduce the dimensions of the mathematical formulation. In every phase the obtained shaped formulation has an easy-solvable form and the complete optimization process is highly scalable. To verify the scalability and effectiveness of our approach, SHARP is evaluated by comprehensive experiment settings which are derived from realistic data/topology of an operational commercial CDN.
Yi Wang 0049, Chen Tian 0001, Hongbo Jiang 0001, Xue (Steve) Liu, Wenyu Liu 0001
GLOBECOM6
2011 Optimal Joint Multi-Path Routing and Sampling Rates Assignment for Real-Time Wireless Sensor Networks
abstract
Maximizing the aggregate network performance over constrained computing and communication resources has been an active research area. Real-time wireless sensor network (RTWSN) is an important application of real-time and networked embedded systems. Due to the severe resources constraints and associated real-time requirements, new challenges arise in RTWSN. In this paper, we study an integrated scheme to optimize the total system performance of real-time flows over an RTWSN by exploiting multi-path routing and dynamic sampling rate assignment. We formally model the problem using nonlinear optimization and design an online distributed algorithm to obtain the optimal rate assignments on multiple paths. Extensive simulation studies demonstrates the significant performance improvements over existing proposals.
Lei Rao, Xue (Steve) Liu, Kyoung-Don Kang, Wenyu Liu 0001, Liang Liu 0010, Ying Chen 0004
ICC4
2011 Energy Efficient Broadcasting Using Network Coding Aware Protocol in Wireless Ad Hoc Network
abstract
Energy efficient broadcasting is of paramount importance for many broadcast applications in wireless ad hoc networks. With respects network coding, it has been proved that the energy gain is upper bounded by 3. However, the coding opportunity is often highly dependent on the established routing paths, resulting in that a lot of coding opportunities could be lost in practice. By combining network coding with the Connected Dominating Set (CDS)-based broadcasting, we take full use of network coding. The intuition behind our algorithm is to intersect information flows at nodes in CDS to increase the coding opportunities. We propose a novel scheme named NCAB, a Network Coding Aware based Broadcast routing mechanism, integrating the network coding and the dynamic implementation of connected dominating set. Our experimental results show that NCAB provides up to 169% gains compared to flooding, and 41% gains compared to CDS-based broadcasting.
Shuai Wang 0008, Athanasios V. Vasilakos, Hongbo Jiang 0001, Xiaoqiang Ma, Wenyu Liu 0001, Kai Peng 0001, Bo Liu 0104, Yan Dong 0001
ICC5
2011 Minimum-Latency Aggregation Scheduling in Underwater Wireless Sensor Networks
abstract
Abstract-Underwater Wireless Sensor Networks (UWSNs) can enable a broad range of applications; data aggregation is a fundamental task in such multi-hop wireless sensor networks. To the best of our knowledge, none of existing research works have addressed the interference-free data aggregation scheduling problem in UWSNs. In this paper, we formally define the data aggregation model in UWSNs. We propose a realistic aggregation scheduling scheme together with its theoretical latency bound Rh(C(Δ - 1) + D), where Rhand Δ are the hop radius and the max degree of the network respectively while C and D is a constant. Specifically, we introduce the concept of Virtual Slot to efficiently exploit multiplexing opportunities of time domain. Compared with naively adapted terrestrial algorithms, the evaluation results show that our proposed algorithm achieve far better performance especially when the packet size is small or the node density is high.
Zuodong Wu, Chen Tian 0001, Hongbo Jiang 0001, Wenyu Liu 0001
ICC4
2011 Measurements and Analysis of an Unconstrained User Generated Content System
abstract
User-Generated Content (UGC) is overwhelming the Internet with its interactivity and various contents. However, traditional UGC still have constrains on videos' length and size, which block out a wide variety of potential popular contents. In this paper, we present the first experimental measurements and analysis of an Unconstrained User-Generated Content (UUGC) system - a test site (so-called "T" site in this paper) of a leading VOD service provider in China. This test site is a video-sharing portal just like traditional UGC, while its contents are not constrained by either duration or size. As an UUGC system, its most distinguishing characteristics are the various types of contents uploaded (movie, TV episode, TV show, music, documentary, sports, etc.) and the wide range of uploaders, which make it an interesting case study. By matching relative key words in video's index, we classify the contents into several basic types and analyze the statistics of three major types - movie, TV episode and TV show (labeled MVI, TV-E and TV-S). For further study of various contents, we demonstrate the patterns of flash crowd triggering of MVI, TV-E and TV-S with several typical cases. To find out the viewers' consumption pattern, we investigate daily & weekly cycles, as well as grouping the videos by age and exhibiting the popularity evolution. By means of curve fitting with multiple known distributions to video view traces, we show that power law with exponential cutoff best fits the videos' popularity distribution for this UUGC system.
Tianlong Yu, Chen Tian 0001, Hongbo Jiang 0001, Wenyu Liu 0001
ICC4
2011 Minimum near-convex decomposition for robust shape representation
abstract
Shape decomposition is a fundamental problem for part-based shape representation. We propose a novel shape decomposition method called Minimum Near-Convex Decomposition (MNCD), which decomposes 2D and 3D arbitrary shapes into minimum number of “near-convex” parts. With the degree of near-convexity a user specified parameter, our decomposition is robust to large local distortions and shape deformation. The shape decomposition is formulated as a combinatorial optimization problem by minimizing the number of non-intersection cuts. Two major perception rules are also imposed into our scheme to improve the visual naturalness of the decomposition. The global optimal solution of this challenging discrete optimization problem is obtained by a dynamic subgradient-based branch-and-bound search. Both theoretical analysis and experiment results show that our approach outperforms the state-of-the-art results without introducing redundant parts. Finally we also show the superiority of our method in the application of hand gesture recognition.
Zhou Ren, Junsong Yuan 0001, Chunyuan Li, Wenyu Liu 0001
ICCV4
2011 Shape Matching Using Points Co-occurrence Pattern
abstract
Shape matching is a very critical problem in computer vision, and many smart features have been designed in recent literature for improving the similarity measure between pairs of shapes, and most of them consider either distribution of the sample contour points, or convexity/concavity property of the contour. In this paper, we design a novel shape feature to capture the Co-Occurrence Pattern (COP) of the points sampled from any given shape contour, and each pattern is described by \textbf{Self-Similarity} which investigates the spatial co-occurrence relation among all the sample points. We test our feature on three famous shape databases: MPEG-7 CE-Shape-1 part B, Tari1000, and Kimia99 data set for shape matching and retrieval. The experimental results show that the proposed descriptor achieves higher computational efficiency with no significant performance loss.
Yu Zhou 0016, Quan Zhou 0004, Xiang Bai, Wenyu Liu 0001
ICIG5
2011 Accurate distortion measurement for B-spline-based shape coding
abstract
In this paper, we present a new contour point distortion measurement, called accurate distortion measurement for B-spline-based shape coding (ADMBSC). Different from existing distortion measurements containing approximation, quantization or parameterization, our distortion is defined as the shortest distance from the original B-spline to the associated contour point. This is in line with the subjective-based objective quality metric. Geometric relationships are introduced to simplify computation, followed by a hybrid admissible distortion checking algorithm to reduce execution time. Theoretical analysis and experimental results demonstrate that when the operational rate-distortion optimal shape coding framework under the minimum-maximum criterion is applied, the ADMBSC can lead to the smallest bit-rate among all the distortion measurements that can guarantee the admissible distortion. Moreover, if the original contour has NCpoints, it takes only O(NC) time for segment distortion measuring paradigms, whose computational complexity is the same as the lowest one among the existing distortion measurements.
Zhongyuan Lai, Zhen Zuo, Zhijun Yao, Wenyu Liu 0001
ICIP5
2011 A symmetric KL divergence based spatiogram similarity measure
abstract
Spatiogram is a generalization of histogram to capture higher-order spatial moments information. To apply spatiogram to object tracking, suitable similarity measure is critical. Although there is a series of work on introducing improved distance measure over the original method, their performance in object tracking is very limited due to insufficient discriminative power. In this paper, we present a symmetric KL divergence based spatiogram similarity measure and show both theoretically and experimentally that, the proposed measure gives superior discriminative power than existing methods, and achieved promising performance in tracking object from single or sequence of images.
Zhijun Yao, Zhongyuan Lai, Wenyu Liu 0001
ICIP3
2011 Image labeling by multiple segmentation
abstract
In this paper, we provide a method for image labeling by combining the local features and contextual cues in a multiple segmentation framework. Our main insight is to weight the classification results of each image region in different levels, which are obtained by a series of learned discriminative models based on bag of features. The contextual cues are implicitly embedded as feature selection in learning process. Multiple segmentation framework provides robust representation, allowing a wide variety of cues to contribute to the confidence in each semantic label. Our algorithm has been applied on the lotus hill institute(LHI) 15-class dataset and outperforms other state-of-the-art methods.
Quan Zhou 0004, Canxiang Yan, Yingying Zhu 0005, Xiang Bai, Wenyu Liu 0001
ICIP5
2011 Estimation of the land surface instantaneous net radiation and its diurnal cycle integrating multi-source remote sensing data under clear sky
abstract
Land surface net radiation is one of critical factors controlling variable land surface processes. Because of the contradiction between spatial resolution and temporal resolution, it is difficult to estimate the hourly or half-hourly surface net radiation with high spatial resolution (<;100m) using remote observation. A simple approach is proposed to retrieve the surface net radiation with high spatial resolution of 30m using multi-source remote sensing data (MODIS atmospheric products and ASTER surface products) under clear sky. The sinusoidal diurnal cycle function for land surface temperature is applied to the net radiation to estimate the hourly net radiation. The estimation results of instantaneous and hourly net radiation show good agreement with field meteorological datum.
Wenyu Liu 0001, Adu Gong, Ji Zhou 0001, Wenfeng Zhan
IGARSS1
2011 Maximal Cliques that Satisfy Hard Constraints with Application to Deformable Object Model Learning
abstract
We propose a novel inference framework for finding maximal cliques in a weighted graph that satisfy hard constraints. The constraints specify the graph nodes that must belong to the solution as well as mutual exclusions of graph nodes, i.e., sets of nodes that cannot belong to the same solution. The proposed inference is based on a novel particle filter algorithm with state permeations. We apply the inference framework to a challenging problem of learning part-based, deformable object models. Two core problems in the learning framework, matching of image patches and finding salient parts, are formulated as two instances of the problem of finding maximal cliques with hard constraints. Our learning framework yields discriminative part based object models that achieve very good detection rate, and outperform other methods on object classes with large deformation.
Xinggang Wang, Xiang Bai, Xingwei Yang, Wenyu Liu 0001, Longin Jan Latecki
NIPS4
2011 Shape Matching and Recognition Using Group-Wised Points
Yu Zhou 0016, Xiang Bai, Wenyu Liu 0001
PSIVT (2)4
2011 Sharpening Thermal Imageries: A Generalized Theoretical Framework From an Assimilation Perspective
abstract
Land surface temperature (LST) plays an important role in many fields. However, thermal bands in prevailing sensors that are onboard satellites have limited spatial resolutions, which seriously impede their potential applications. Many approaches that aim to downscale thermal imageries to finer spatial resolution levels have been developed in recent years. This paper managed to construct a Generalized Theoretical Framework from an Assimilation Perspective for them with semiempirical regression and modulation integration techniques. Based on three hierarchical sharpening levels, which include digital number, radiance, and surface temperature, many of them can be brought into such a unified framework as derivatives. Two typical land cover patterns were chosen as case study areas to evaluate the capabilities of various kernels to represent the LST distribution. The results demonstrate that there are great discrepancies among those kernels. The single-band kernels are dependent on different land cover types, while the band-derivative kernels perform better in most circumstances when portraying the LST variations. In addition, the simulated imageries that were resampled by scaling up the original thermal bands with an aggregation technique were utilized to validate a localization approach of temperature vegetation dryness index (TVDI). The results indicate that the TVDI has satisfactory effects when depicting slight LST variations due to soil anomalies. More intercomparisons between the approach presented here and other different methods, including artificial neural network and Gram-Schmidt techniques, were made thoroughly, coupling with the Moderate Resolution Imaging Spectroradiometer and Advanced Spaceborne Thermal Emission Reflection Radiometer data. Consequently, the generalized framework opens up the foreground for sharpening thermal images with high efficiency over a solid theoretical foundation.
Wenfeng Zhan, Ji Zhou 0001, Jing Li 0018, Wenyu Liu 0001
IEEE Trans. Geosci. Remote. Sens.5
2011 Learning Adaptive Metric for Robust Visual Tracking
abstract
Matching the visual appearances of the target over consecutive image frames is the most critical issue in video-based object tracking. Choosing an appropriate distance metric for matching determines its accuracy and robustness, and thus significantly influences the tracking performance. Most existing tracking methods employ fixed pre-specified distance metrics. However, this simple treatment is problematic and limited in practice, because a pre-specified metric does not likely to guarantee the closest match to be the true target of interest. This paper presents a new tracking approach that incorporates adaptive metric learning into the framework of visual object tracking. Collecting a set of supervised training samples on-the-fly in the observed video, this new approach automatically learns the optimal distance metric for more accurate matching. The design of the learned metric ensures that the closest match is very likely to be the true target of interest based on the supervised training. Such a learned metric is discriminative and adaptive. This paper substantializes this new approach in a solid case study of adaptive-metric differential tracking, and obtains a closed-form analytical solution to motion estimation and visual tracking. Moreover, this paper extends the basic linear distance metric learning method to a more powerful nonlinear kernel metric learning method. Extensive experiments validate the effectiveness of the proposed approach, and demonstrate the improved performance of the proposed new tracking method.
Nan Jiang 0016, Wenyu Liu 0001, Ying Wu 0001
IEEE Trans. Image Process.2
2011 Spread Spectrum Image Watermarking Based on Perceptual Quality Metric
abstract
Efficient image watermarking calls for full exploitation of the perceptual distortion constraint. Second-order statistics of visual stimuli are regarded as critical features for perception. This paper proposes a second-order statistics (SOS)-based image quality metric, which considers the texture masking effect and the contrast sensitivity in Karhunen-Loève transform domain. Compared with the state-of-the-art metrics, the quality prediction by SOS better correlates with several subjectively rated image databases, in which the images are impaired by the typical coding and watermarking artifacts. With the explicit metric definition, spread spectrum watermarking is posed as an optimization problem: we search for a watermark to minimize the distortion of the watermarked image and to maximize the correlation between the watermark pattern and the spread spectrum carrier. The simple metric guarantees the optimal watermark a closed-form solution and a fast implementation. The experiments show that the proposed watermarking scheme can take full advantage of the distortion constraint and improve the robustness in return.
Fan Zhang 0093, Wenyu Liu 0001, Weisi Lin, King Ngi Ngan
IEEE Trans. Image Process.2
2011 Improving Application Placement for Cluster-Based Web Applications
abstract
Dynamic application placement for clustered web applications heavily influences system performance and quality of user experience. Existing approaches claim that they strive to maximize the throughput, keep resource utilization balanced across servers, and minimize the start/stop cost of application instances. However, they fail to minimize the worst case of server utilization; the load balancing performance is not optimal. What's more, some applications need to communicate with each other, which we called dependent applications; the network cost of them also should be taken into consideration. In this paper, we investigate how to minimize the resource utilization of servers in the worst case, aiming at improving load balancing among clustered servers. Our contribution is two-fold. First we propose and define a new optimization objectives: limiting the worst case of each individual server's utilization, formulated by a min-max problem. A novel framework based on binary search is proposed to detect an optimal load balancing solution. Second, we define system cost as the weighted combination of both placement change and inter-application communication cost. By maximizing the number of instances of dependent applications that reside in the same set of servers, the basic load-shifting and placement-change procedures are enhanced to minimize whole system cost. Extensive experiments have been conducted and effectively demonstrate that: 1) the proposed framework achieves a good allocation for clustered web applications. In other words, requests are evenly allocated among servers, and throughput is still maximized; 2) the total system cost maintains at a low level; 3) our algorithm has the capacity of approximating an optimal solution within polynomial time and is promising for practical implementation in real deployments.
Chen Tian 0001, Hongbo Jiang 0001, Arun Iyengar, Xue (Steve) Liu, Zuodong Wu, Wenyu Liu 0001, Chonggang Wang
IEEE Trans. Netw. Serv. Manag.7
2010 Inference Scene Labeling by Incorporating Object Detection with Explicit Shape Model
Quan Zhou 0004, Wenyu Liu 0001
ACCV (3)2
2010 Convex shape decomposition
abstract
In this paper, we propose a new shape decomposition method, called convex shape decomposition. We formalize the convex decomposition problem as an integer linear programming problem, and obtain approximate optimal solution by minimizing the total cost of decomposition under some concavity constraints. Our method is based on Morse theory and combines information from multiple Morse functions. The obtained decomposition provides a compact representation, both geometrical and topological, of original object. Our experiments show that such representation is very useful in many applications.
Hairong Liu, Wenyu Liu 0001, Longin Jan Latecki
CVPR2
2010 Two-Step Coding for High Definition Video Compression
abstract
High definition (HD) video has come into people’s life from movie theaters to HDTV. However, the compression of HD videos is a challenging problem due to flicker noise, caused by film grain. The flicker noise significantly limits the applicability of motion estimation (ME), which is a key factor of the efficient video compression in block-based coding standards. Due to the flicker noise, it is difficult to obtain a perfect match between a current block and a reference block. In block-based video coding standards including H.264 a given block is either encoded by inter-frame or intra-frame prediction. We propose a new coding scheme called Two-Step Coding (TSC) that utilizes both for each block. TSC first reduces the resolution of each frame by replacing each block with its DC coefficient of the DCT to the original color values. The flicker noise is greatly reduced in the obtained lower resolution frame, which we call DC frame. The key benefit is that ME becomes very efficient on DC frames, and consequently, the DC frame can be efficiently inter-frame coded. The difference between the original frame and DC frame is actually described by the AC coefficients of the DCT of the original frame. We utilize the existing H.264 tools to combine the intra-frame and inter-frame coded parts of blocks both on the encoder and decoder sides. The key benefit of the proposed TSC in comparison to the most popular standards, in particular, in comparison to H.264 lies in better utilization of inter-frame coding.. Due to flicker noise, H.264 mostly employs intra block coding on HD videos. However, it is well-known that inter-frame coding significantly outperforms intra coding in video compression rate if the temporal correllation is correctly utilized. By reducing each frame to DC frame, TSC makes it possible to apply inter-frame coding. We provide experimental data and analysis to illustrate this fact.
Wenfei Jiang, Wenyu Liu 0001, Longin Jan Latecki, Bing Feng
DCC2
2010 Arbitrary Directional Edge Encoding Schemes for the Operational Rate-Distortion Optimal Shape Coding Framework
abstract
We present two edge encoding schemes, namely 8-sector scheme and 16-sector scheme, for the operational rate-distortion (ORD) optimal shape coding framework. Different from the traditional 8-direction scheme that can only encode edges with angles being an integer multiple of π/4, our proposals can encode edges with arbitrary angles. We partition the digital coordinate plane into 8 and 16 sectors, and design the corresponding differential schemes to encode the short and the long component of each vertex. Experiment results demonstrate that our two proposals can reduce a large number of encoding vertices and therefore reduce 10%~20% bits for the basic ORD optimal algorithms and 10%~30% bits for all the ORD optimal algorithms under the same distortion thresholds, respectively. Moreover, the reconstruction contours are more compact compared with those using the traditional 8-direction edge encoding scheme.
Zhongyuan Lai, Junhuan Zhu, Zhou Ren, Wenyu Liu 0001, Baolan Yan
DCC4
2010 Co-transduction for Shape Retrieval
Xiang Bai, Bo Wang 0044, Xinggang Wang, Wenyu Liu 0001, Zhuowen Tu
ECCV (3)4
2010 Object Recognition Using Junctions
Bo Wang 0044, Xiang Bai, Xinggang Wang, Wenyu Liu 0001, Zhuowen Tu
ECCV (5)4
2010 Efficient Data Collection with Sampling in WSNs: Making Use of Matrix Completion Techniques
abstract
Data collection is of paramount importance in many applications of wireless sensor networks (WSNs). Especially, to accommodate ever increasing demands of signal source coding applications, the capacity of processing multi-user data query is crucial in WSNs where the efficiency is one key consideration. To that end, this paper presents EDCA: an Efficient Data Collection Approach for data query in WSNs, which exploits recent matrix completion techniques. Specifically, for the efficiency of energy consumption, we randomly select a part of nodes from the sensor network to sample at each time instance and directly forward the data to the sink. Then, to recover the data precisely, we shift the rank minimization problem, which is NP-hard, to a convex optimization one. Compared with the centralized scheme, energy consumption using EDCA is significantly reduced due to lower sampling rate and fewer packets to transmit. The experimental results demonstrate that EDCA significantly outperforms the existing naive method in terms of energy consumption and the introduced errors are quite trivial.
Jie Cheng 0003, Hongbo Jiang 0001, Xiaoqiang Ma, Lanchao Liu, Lijun Qian, Chen Tian 0001, Wenyu Liu 0001
GLOBECOM7
2010 Shape Classification Using Tree -Unions
abstract
In this paper, we proposed a novel approach to shape classification. A new shape tree based on junction nodes can represent the global structure in a simple way. The statistic distribution of junctions can be learned by merging the shape trees. In the process of learning, context of a junction node is obtained to improve the rate of classification. We illustrate the utility of the proposed method on the problem of 2D shape classification using the new shape tree representation.
Bo Wang 0044, Wei Shen 0002, Wenyu Liu 0001, Xinge You, Xiang Bai
ICPR3
2010 Minimizing Electricity Cost: Optimization of Distributed Internet Data Centers in a Multi-Electricity-Market Environment
abstract
The study of Cyber-Physical System (CPS) has been an active area of research. Internet Data Center (IDC) is an important emerging Cyber-Physical System. As the demand on Internet services drastically increases in recent years, the power used by IDCs has been skyrocketing. While most existing research focuses on reducing power consumptions of IDCs, the power management problem for minimizing the total electricity cost has been overlooked. This is an important problem faced by service providers, especially in the current multi-electricity market, where the price of electricity may exhibit time and location diversities. Further, for these service providers, guaranteeing quality of service (i.e. service level objectives-SLO) such as service delay guarantees to the end users is of paramount importance. This paper studies the problem of minimizing the total electricity cost under multiple electricity markets environment while guaranteeing quality of service geared to the location diversity and time diversity of electricity price. We model the problem as a constrained mixed-integer programming and propose an efficient solution method. Extensive evaluations based on real-life electricity price data for multiple IDC locations illustrate the efficiency and efficacy of our approach.
Lei Rao, Xue (Steve) Liu, Le Xie 0001, Wenyu Liu 0001
INFOCOM4
2010 Learning Context-Sensitive Shape Similarity by Graph Transduction
abstract
Shape similarity and shape retrieval are very important topics in computer vision. The recent progress in this domain has been mostly driven by designing smart shape descriptors for providing better similarity measure between pairs of shapes. In this paper, we provide a new perspective to this problem by considering the existing shapes as a group, and study their similarity measures to the query shape in a graph structure. Our method is general and can be built on top of any existing shape similarity measure. For a given similarity measure, a new similarity is learned through graph transduction. The new similarity is learned iteratively so that the neighbors of a given shape influence its final similarity to the query. The basic idea here is related to PageRank ranking, which forms a foundation of Google Web search. The presented experimental results demonstrate that the proposed approach yields significant improvements over the state-of-art shape matching algorithms. We obtained a retrieval rate of 91.61 percent on the MPEG-7 data set, which is the highest ever reported in the literature. Moreover, the learned similarity by the proposed method also achieves promising improvements on both shape classification and shape clustering.
Xiang Bai, Xingwei Yang, Longin Jan Latecki, Wenyu Liu 0001, Zhuowen Tu
IEEE Trans. Pattern Anal. Mach. Intell.4
2010 Connectivity-Based Skeleton Extraction in Wireless Sensor Networks
abstract
Many sensor network applications are tightly coupled with the geometric environment where the sensor nodes are deployed. The topological skeleton extraction for the topology has shown great impact on the performance of such services as location, routing, and path planning in wireless sensor networks. Nonetheless, current studies focus on using skeleton extraction for various applications in wireless sensor networks. How to achieve a better skeleton extraction has not been thoroughly investigated. There are studies on skeleton extraction from the computer vision community; their centralized algorithms for continuous space, however, are not immediately applicable for the discrete and distributed wireless sensor networks. In this paper, we present a novel Connectivity-bAsed Skeleton Extraction (CASE) algorithm to compute skeleton graph that is robust to noise, and accurate in preservation of the original topology. In addition, CASE is distributed as no centralized operation is required, and is scalable as both its time complexity and its message complexity are linearly proportional to the network size. The skeleton graph is extracted by partitioning the boundary of the sensor network to identify the skeleton points, then generating the skeleton arcs, connecting these arcs, and finally refining the coarse skeleton graph. We believe that CASE has broad applications and present a skeleton-assisted segmentation algorithm as an example. Our evaluation shows that CASE is able to extract a well-connected skeleton graph in the presence of significant noise and shape variations, and outperforms the state-of-the-art algorithms.
Hongbo Jiang 0001, Wenping Liu 0001, Dan Wang 0002, Chen Tian 0001, Xiang Bai, Xue (Steve) Liu, Ying Wu 0001, Wenyu Liu 0001
IEEE Trans. Parallel Distributed Syst.8
2009 Skeleton Graph Matching Based on Critical Points Using Path Similarity
Bo Wang 0044, Wenyu Liu 0001, Xiang Bai
ACCV (3)3
2009 Shape band: A deformable object detection approach
abstract
In this paper, we focus on the problem of detecting/matching a query object in a given image. We propose a new algorithm, shape band, which models an object within a bandwidth of its sketch/contour. The features associated with each point on the sketch are the gradients within the bandwidth. In the detection stage, the algorithm simply scans an input image at various locations and scales for good candidates. We then perform fine scale shape matching to locate the precise object boundaries, also by taking advantage of the information from the shape band. The overall algorithm is very easy to implement, and our experimental results show that it can outperform stat-of-the-art contour based object detection algorithms.
Xiang Bai, Quannan Li, Longin Jan Latecki, Wenyu Liu 0001, Zhuowen Tu
CVPR4
2009 Perceptual Relevance Measure for Generic Shape Coding
abstract
Approximation metric is a significant factor for subjective quality improvement of polygonal vertex-based shape codec. Usually this metric is defined by absolute distance measure (ADM). ADM only considers the shortest absolute distance from candidate vertices on the original object contour segment to corresponding approximating line segment and ignores other visual characteristics of the contour segment. Thus it cannot describe the original contour appropriately in visual aspects. As a result, it may severely degrade the subjective reconstruction quality, especially for contours with sharp salience. We propose perceptual relevance measure (PRM) to address this problem. We first describe the original relative position relationship between candidate vertex C and approximating line segment AB by three parameters having definite visual meanings, namely turn angle of and lengths of two adjacent line segments a and b. And then we provide the following three visual properties to find the exactly expression of PRM.
Zhongyuan Lai, Wenyu Liu 0001
DCC2
2009 Tri-Message: A Lightweight Time Synchronization Protocol for High Latency and Resource-Constrained Networks
abstract
Existing terrestrial synchronization protocols including RBS, FTSP, TPSN, LTS and TSHL have already achieved high precision in radio networks, but none of them perform well in high latency networks like acoustic sensor networks. In this paper, we present tri-message: a lightweight time synchronization protocol for high latency and resource-constrained networks. As its name suggests, only three message exchanges are required in one synchronization process. Meanwhile, tri-message utilizes very simple mathematical operations to calculate the clock skew and offset. Specially, tri-message is feasible for many extremely long latency applications such as space exploration because it has an increasing synchronization precision with the increasement of distance.
Chen Tian 0001, Hongbo Jiang 0001, Xue (Steve) Liu, Xinbing Wang, Wenyu Liu 0001, Yi Wang 0049
ICC5
2009 Active skeleton for non-rigid object detection
abstract
We present a shape-based algorithm for detecting and recognizing non-rigid objects from natural images. The existing literature in this domain often cannot model the objects very well. In this paper, we use the skeleton (medial axis) information to capture the main structure of an object, which has the particular advantage in modeling articulation and non-rigid deformation. Given a set of training samples, a tree-union structure is learned on the extracted skeletons to model the variation in configuration. Each branch on the skeleton is associated with a few part-based templates, modeling the object boundary information. We then apply sum-and-max algorithm to perform rapid object detection by matching the skeleton-based active template to the edge map extracted from a test image. The algorithm reports the detection result by a composition of the local maximum responses. Compared with the alternatives on this topic, our algorithm requires less training samples. It is simple, yet efficient and effective. We show encouraging results on two widely used benchmark image sets: the Weizmann horse dataset [7] and the ETHZ dataset [16].
Xiang Bai, Xinggang Wang, Longin Jan Latecki, Wenyu Liu 0001, Zhuowen Tu
ICCV4
2009 CASE: Connectivity-Based Skeleton Extraction in Wireless Sensor Networks
abstract
Many sensor network applications are tightly coupled with the geometric environment where the sensor nodes are deployed. The topological skeleton extraction has shown great impact on the performance of such services as location, routing, and path planning in sensor networks. Nonetheless, current studies focus on using skeleton extraction for various applications in sensor networks. How to achieve a better skeleton extraction has not been thoroughly investigated. There are studies on skeleton extraction from the computer vision community; their centralized algorithms for continuous space, however, is not immediately applicable for the discrete and distributed sensor networks. In this paper we present CASE: a novel connectivity-based skeleton extraction algorithm to compute skeleton graph that is robust to noise, and accurate in preservation of the original topology. In addition, no centralized operation is required. The skeleton graph is extracted by partitioning the boundary of the sensor network to identify the skeleton points, then generating the skeleton arcs, connecting these arcs, and finally refining the coarse skeleton graph. Our evaluation shows that CASE is able to extract a well-connected skeleton graph in the presence of significant noise and shape variations, and outperforms state-of-the-art algorithms.
Hongbo Jiang 0001, Wenping Liu 0001, Dan Wang 0002, Chen Tian 0001, Xiang Bai, Xue (Steve) Liu, Ying Wu 0001, Wenyu Liu 0001
INFOCOM8
2009 Contour Grouping with Partial Shape Similarity
Chengqian Wu, Xiang Bai, Quannan Li, Xingwei Yang, Wenyu Liu 0001
PSIVT5
2009 Transmission distortion estimation for real-time video delivery over hybrid channels with bit errors and packet erasures
Wenyu Liu 0001
Comput. Commun.1
2009 A Video Coding Scheme Based on Joint Spatiotemporal and Adaptive Prediction
abstract
We propose a video coding scheme that departs from traditional Motion Estimation/DCT frameworks and instead uses Karhunen-Loeve Transform (KLT)/Joint Spatiotemporal Prediction framework. In particular, a novel approach that performs joint spatial and temporal prediction simultaneously is introduced. It bypasses the complex H.26x interframe techniques and it is less computationally intensive. Because of the advantage of the effective joint prediction and the image-dependent color space transformation (KLT), the proposed approach is demonstrated experimentally to consistently lead to improved video quality, and in many cases to better compression rates and improved computational speed.
Wenfei Jiang, Longin Jan Latecki, Wenyu Liu 0001, Ken Gorman
IEEE Trans. Image Process.3
2008 Mining user similarity based on location history
abstract
The pervasiveness of location-acquisition technologies (GPS, GSM networks, etc.) enable people to conveniently log the location histories they visited with spatio-temporal data. The increasing availability of large amounts of spatio-temporal data pertaining to an individual's trajectories has given rise to a variety of geographic information systems, and also brings us opportunities and challenges to automatically discover valuable knowledge from these trajectories. In this paper, we move towards this direction and aim to geographically mine the similarity between users based on their location histories. Such user similarity is significant to individuals, communities and businesses by helping them effectively retrieve the information with high relevance. A framework, referred to as hierarchical-graph-based similarity measurement (HGSM), is proposed for geographic information systems to consistently model each individual's location history and effectively measure the similarity among users. In this framework, we take into account both the sequence property of people's movement behaviors and the hierarchy property of geographic spaces. We evaluate this framework using the GPS data collected by 65 volunteers over a period of 6 months in the real world. As a result, HGSM outperforms related similarity measures, such as the cosine similarity and Pearson similarity measures.
Quannan Li, Yu Zheng 0004, Xing Xie 0001, Wenyu Liu 0001, Wei-Ying Ma
GIS5
2008 Improving BitTorrent Traffic Performance by Exploiting Geographic Locality
abstract
Current implementations of BitTorrent-like P2P applications ignore the underlying Internet topology hence incur a large amount of traffic both inside an Internet service provider (ISP)' national backbone networks and over cross-ISP Internet working links. These traffics not only occupy costly bandwidth, but also increase user perceived response latency. ISP-biased neighbor selection proposes to exploit peers' topological locality by biased neighbor selection, in which a peer chooses the majority of its neighbors from peers within the same ISP. In this paper, we propose to further exploit peers' geographic locality. First we improved ISP-biased neighbor selection (ISP-Biased+) to take into consideration network locality (or, city locations) within the same ISP. When required neighbor number is relatively much less than seeds available, ISP-biased neighbor selection+ performs much better than original approach, proved by simulations. Next, we propose that a peer could also choose its neighbors from peers of different ISPs within the same city with priority: assist by a well-know Chinese operator's unique ISP-internetworking content distribution network (CDN), these local cross-ISP traffics can be routed through local CDN cite. Using simulations, we show that cross-ISP traffic burden can be completely shifted to CDN local links and backbone traffic. At the same time, user perceived delay can be significantly reduced.
Chen Tian 0001, Xue (Steve) Liu, Hongbo Jiang 0001, Wenyu Liu 0001, Yi Wang 0049
GLOBECOM4
2008 Skeletonization of gray-scale image from incomplete boundaries
abstract
Skeletonization of gray-scale images is a challenging problem in computer vision due to the difficulty of segmenting grayscale images to get the complete contour. Compared with previous skeletonization algorithms which use computational methods to avoid segmentation, this paper reveals that it is applicable to skeletonize gray-scale images from boundaries directly. We start from boundaries of gray-scale images and perform Euclidean Distance Transform on boundaries. Then we compute the gradient magnitude of the distance transform and perform isotropic vector diffusion. After diffusion, the Skeleton Strength Map (SSM) is computed and skeleton can be extracted from SSM. The experiments show that this method can obtain good performance from boundaries so long as major boundary segments are preserved.
Quannan Li, Xiang Bai, Wenyu Liu 0001
ICIP3
2008 Towards Minimum Traffic Cost and Minimum Response Latency: A Novel Dynamic Query Protocol in Unstructured P2P Networks
abstract
Controlled-flooding algorithms are widely used in unstructured networks. Expanding ring (ER) achieves low response delay, while its traffic cost is huge; dynamic querying (DQ) is known for its desirable behavior in traffic control, but it achieves lower search cost at the price of an undesirable latency performance; Enhanced dynamic querying (DQ+) can reduce the search latency too, while it is hard to determine a general optimum parameters set. In this paper, a novel algorithm named selective dynamic query (SDQ) is proposed. Unlike previous works that awkwardly processing floating TTL values, SDQ properly select an integer TTL value and a set of neighbors to narrow the scope of next query. Our experiments demonstrate that SDQ provides finer-grained control than other algorithms: its latency is close to the well-known minimum one via ER; in the mean time its traffic cost also close to the minimum. To our best knowledge, this is the first work capable of achieving best performance in terms of both response latency and traffic cost. In addition, our experiments also demonstrate that SDQ works well in various network topologies.
Chen Tian 0001, Hongbo Jiang 0001, Xue (Steve) Liu, Wenyu Liu 0001, Yi Wang 0049
ICPP4
2008 Computing Stable Skeletons with Particle Filters
Xiang Bai, Xingwei Yang, Longin Jan Latecki, Yanbo Xu, Wenyu Liu 0001
PRICAI5
2008 A Uniform Host Protocol Framework Planning to Change
abstract
As a substantial foundation of the Internet, traditional layered host protocol stack is found to be a framework "not planning to change". This framework is hard to accommodate new layers and new protocols, hence hard to provide new services including security and mobility. What's more, it is also hard to share information among layers, which is critical for cross-layer optimization in MANET implementation. A novel protocol stack framework, altitude based architecture (ABA), is presented in this paper. The proposed framework is cater to framework extensibility, and by negotiation mechanisms, a universal communication model is also presented for backward-compatibility. ABA, which is planning to change, aims to build a uniform host protocol framework for future Internet and MANET. Analyses and simulation illustrations are also given in the paper.
Wenyu Liu 0001, Wang Yi 0004, Huaien Luo
VTC Spring2
2008 A Unified Curvature Definition for Regular, Polygonal, and Digital Planar Curves
Hairong Liu, Longin Jan Latecki, Wenyu Liu 0001
Int. J. Comput. Vis.3
2008 Image Decomposition and Texture Segmentation via Sparse Representation
abstract
Decomposing an image into a texture part and a non-texture (cartoon) part, as well as grouping the texture part into several homogeneous subparts, is studied in this letter. The so-called texture part is composed of both the self-similar structure and the oscillatory noise. The self-similar structure of each homogenous subtexture is captured in its principal subspace. Both the segmentation and the decomposition are essentially related to sparse representation and are united to a framework.
Fan Zhang 0093, Xiaoqiong Ye, Wenyu Liu 0001
IEEE Signal Process. Lett.3
2007 High Capacity Watermarking in Nonedge Texture Under Statistical Distortion Constraint
Fan Zhang 0019, Wenyu Liu 0001
ACCV (1)2
2007 Objects Similarity Measurement Based on Skeleton Tree Descriptor Matching
abstract
In this paper, we proposed a framework to address the problem of binary object (2D or 3D) recognition. In our method, a binary object is represented as a Skeleton Tree (ST), transformed from its skeleton (or centerline). Both topological and geometrical features are embedded in the ST and this allows comparisons between different objects by tree matching algorithms. Tree descriptor is used to represent the topological features of the ST, and the maximal isomorphic subtrees (MIST) are obtained by searching for the longest matching substrings in the tree descriptors. A novel method of ST matching based on tree descriptor is also presented. The problems with cyclic skeleton and noise on the skeleton are discussed too. Experiments on a variety of objects get satisfying results, which show the potential of our method in the presence of rotation, scaling, translation and reflection. The time complexity of the algorithm is o(n3), where n is the number of the skeleton branches in ST.
Wenyu Liu 0001, Caihua Wu
CAD/Graphics2
2007 Visual Curvature
abstract
In this paper, we propose a new definition of curvature, called visual curvature. It is based on statistics of the extreme points of the height functions computed over all directions. By gradually ignoring relatively small heights, a single parameter multi-scale curvature is obtained. It does not modify the original contour and the scale parameter has an obvious geometric meaning. The theoretical properties and the experiments presented demonstrate that multi-scale visual curvature is stable, even in the presence of significant noise. In particular, it can deal with contours with significant gaps. We also show a relation between multi-scale visual curvature and convexity of simple closed curves. To our best knowledge, the proposed definition of visual curvature is the first ever that applies to regular curves as defined in differential geometry as well as to turn angles of polygonal curves. Moreover, it yields stable curvature estimates of curves in digital images even under sever distortions.
Hairong Liu, Longin Jan Latecki, Wenyu Liu 0001, Xiang Bai
CVPR3
2007 Skeletonization using SSM of the Distance Transform
abstract
This paper proposes a new approach for skeletonization based on the skeleton strength map (SSM) caculated by Euclidean distance transform of a binary image. After the distance transform and gradient are computed, isotropic diffusion is performed on the gradient vector field and the skeleton strength map is computed from the diffused vector field. A critical point set is then selected from local maxima of the SSM. The critical points are located on significant visual parts of the object. The skeleton is obtained by connecting the critical points with geodesic paths. This approach overcomes intrinsic drawbacks of distance transform based skeletons, since it yields stable and connected skeletons without losing significant visual parts.
Longin Jan Latecki, Quannan Li, Xiang Bai, Wenyu Liu 0001
ICIP (5)4
2007 Localization and Synchronization for 3D Underwater Acoustic Sensor Networks
Chen Tian 0001, Wenyu Liu 0001, Jiang Jin, Yi Wang 0049, Yijun Mo
UIC2
2007 Skeleton Pruning by Contour Partitioning with Discrete Curve Evolution
abstract
In this paper, we introduce a new skeleton pruning method based on contour partitioning. Any contour partition can be used, but the partitions obtained by Discrete Curve Evolution (DCE) yield excellent results. The theoretical properties and the experiments presented demonstrate that obtained skeletons are in accord with human visual perception and stable, even in the presence of significant noise and shape variations, and have the same topology as the original skeletons. In particular, we have proven that the proposed approach never produces spurious branches, which are common when using the known skeleton pruning methods. Moreover, the proposed pruning method does not displace the skeleton points. Consequently, all skeleton points are centers of maximal disks. Again, many existing methods displace skeleton points in order to produces pruned skeletons.
Xiang Bai, Longin Jan Latecki, Wenyu Liu 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2006 Fast adaptive inter mode decision method for P slices in H.264
abstract
H.264 defines 7 different coding modes for macroblocks (MBs) in P slices. In order to achieve a coding performance as high as possible, the H.264 encoder calculates rate distortion costs of all possible modes to determine the best mode of a MB. The computation complexity is so large that make it difficult to be used in practical applications especially in real time environment. In this paper, a fast adaptive inter modes decision method is proposed to reduce the complexity of H.264 encoder. The candidate inter modes used in rate distortion optimization can be limited in a small mode group (MG) by using the characteristics of the motion compensated residual image. The overlapped mode groups and dynamic adjusted thresholds adopted in the proposed method can make the best mode lies within the chosen MG with great possibility which leads to extensively computation reduction without any perceivable loss in quality. The experimental results show that the proposed method can save the encoding time up to 52% on average with -0.05dB performance degradation that is negligible. Keywords-H.264; inter modes; mode group; adaptive; overlapped
Bin Feng 0001, Guangxi Zhu, Wenyu Liu 0001
CCNC3
2006 Fast adaptive inter-prediction mode decision method for H.264 based on spatial correlation
abstract
H.264 defines 7 different coding modes for macroblocks (MBs) in P slices. In order to achieve a coding performance as high as possible, the H.264 encoder calculates rate distortion costs of all possible modes to determine the best mode of a MB. The computation complexity is so large that make it difficult to be used in practical applications especially in real time environment. In this paper, a fast adaptive intermodes decision method is proposed to reduce the complexity of H.264 encoder. Firstly the candidate inter modes used in rate distortion optimization can be limited in a small mode group (MG) by using the characteristics of the motion compensated residual image. Then the two most probable modes of the chosen MG are obtained on the basis of the modes of the up MB and the left MB. By calculating and comparing the rate distortion cost of the two modes, the optimum mode of the MB is determined. The overlapped mode groups and dynamic adjusted thresholds adopted in the proposed method can make the best mode lie within the chosen MG with great possibility which leads to extensive computation reduction with acceptable loss in quality. The experimental results show that the proposed method can save the encoding time up to 65% on average with -0.24dB performance degradation.
Bin Feng 0001, Guangxi Zhu, Wenyu Liu 0001
ISCAS3
2005 Fast Adaptive Inter Mode Decision Method in H.264 Based on Spatial Correlation
abstract
The computation complexity of H.264 is so large that make it difficult to be used in practical applications especially in real time environment. In this paper, a fast adaptive inter modes decision method is proposed to reduce the complexity of H.264 encoder. Firstly the candidate inter modes can be limited in a small mode group (MG) by using the characteristics of the motion compensated residual image. Then the most probable mode (MPM) of the MB is predicted on the basis of the modes of the neighboring macroblocks. The overlapped mode groups and dynamic adjusted thresholds adopted in the proposed method can make the best mode lies within the chosen MG with great possibility which leads to extensively computation reduction with acceptable loss in quality. The experimental results show that the proposed method can save the encoding time up to 64% on average with -0.45dB performance degradation.
Bin Feng 0001, Guangxi Zhu, Wenyu Liu 0001
ISM3
2004 Blind source separation of more sources than mixtures using generalized exponential mixture models
Huanwen Tang, Wenyu Liu 0001, Yiyuan Tang
Neurocomputing3
2001 A fast algorithm for corner detection using the morphologic skeleton
Wenyu Liu 0001, Li Hua, Guangxi Zhu
Pattern Recognit. Lett.1