Wenxi Liu

dblp:68/8221 · DBLP profile ↗
← Back
82ranked-venue papers
12as first author
57since 2021 · last 2026
0000-0002-3630-6322ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 47 · 6 first-author · 29 since 2021Artificial intelligence and machine learning · 43 · 4 first-author · 33 since 2021Systems, architecture and hardware · 4 · 2 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Seeing in Double: Dual-Granularity BEV Segmentation via Mamba-Driven Alignment and Polar-Decoupled Experts
abstract
Bird's Eye View (BEV) representation has become pivotal for autonomous driving, yet existing polar coordinate-based approaches face two critical limitations: (1) distant semantic misprojection caused by radial resolution decay, and (2) region-specific geometric distortions from non-uniform polar discretization. To address these issues, we propose a novel framework addressing these challenges through three key innovations. First, we present a bilateral heterogeneous network constructs multi-granularity BEV spaces, efficiently exploiting dual-resolution visual information for distant detail preservation. Second, we employ an align-fusion strategy for multi-granularity feature aggregation. Specifically, the Mamba-Based Cross-Resolution Alignment module establishes semantic consistency for perspective features through shared state-space optimization. In the later stage, the Adaptive BEV Space Selector dynamically aggregates multi-granularity BEV features. Third, we introduce a Mixture of Radial-Angular Decoupled Experts, which employs polar-aware expert routing to disentangle radial compression and angular shear distortions through specialized geometric refinement. Comprehensive experiments on nuScenes and Lyft L5 demonstrate the state-of-the-art performance of our model across various resolution settings, visibility filtering, and perception ranges.
Jingze Su, Qi Li 0038, Wenjie Yang 0005, Yuanlong Yu 0001, Wenxi Liu
AAAI7
2026 AIPO: Adaptive Information Guided Token-Level Reinforcement Learning for Large Language Model Reasoning
abstract
Reinforcement Learning with Verifiable Rewards (RLVR) improves the reasoning capability of Large Language Models (LLMs). Current RLVR trains LLMs on all generated tokens, rather than exploring which tokens actually contribute to reasoning. We propose AIPO(Adaptive–Information Policy Optimization), which focuses updates on those decisive tokens discovered on the fly. AIPO estimates each hidden state’s mutual information to score tokens. Policy gradients are then computed only on these critical tokens, using an advantage that blends information gain and verifiable correctness. To improve the efficiency of mutual-information estimation, AIPO adopts a Random–Fourier approximation of the Hilbert–Schmidt Independence Criterion. Across five math and science benchmarks, AIPO yields up to +20% accuracy over strong RLVR baselines while updating merely 10% of tokens, demonstrating superior efficiency and effectiveness. Our findings highlight the importance of information–driven token selection for efficient and effective reinforcement learning of LLM reasoning.
Bin Chen 0006, Hongfei Ye, Wenxi Liu, Yu Zhang 0296, Furui Liu
ACL (1)4
2026 ParaSuite: Boosting LLM Reasoning via Paradox Resolution
abstract
Logical reasoning is a key capability of large language models, yet current benchmarks focus almost entirely on tasks that just check basic logical consistency and overlook the reflective reasoning required for paradox detection and resolution.To fill the gap, we present ParaSuite, the first pipeline dedicated to paradox research that automates data synthesis, evaluation, and training.We introduce PARADOX, a synthetic, high-quality data spanning two difficulty tiers and three academic domains, accompanied by specialized evaluation metrics and solving algorithms.We propose ParadoxBreaker-7B, trained with Mutual-Information Guided Fine-Tuning and reinforcement learning step verify paradox reward(PAPO).Experiments demonstrate significant improvements in both paradoxical and general STEM reasoning.
Bin Chen 0006, Yu Zhang 0296, Hongfei Ye, Wenxi Liu, Hongyang Chen 0001
ACL (1)5
2026 Assessing Color Vision Test in Large Vision-language Models
abstract
With the widespread adoption of large vision-language models, the capacity for color vision in these models is crucial. However, the color vision abilities of large visual-language models have not yet been thoroughly explored. To address this gap, we define a color vision testing task for large vision-language models and construct a dataset that covers multiple categories of test questions and tasks of varying difficulty levels. Furthermore, we analyze the types of errors made by large vision-language models and propose a chain-of-thought prompting strategy to enhance their performance in color vision tests.
Hongfei Ye, Bin Chen 0006, Wenxi Liu, Yu Zhang 0296, Zhao Li 0007, Dandan Ni, Hongyang Chen 0001
ICMR3
2026 ITGO: A general framework for text-guided image outpainting
Bin Chen 0006, Yuanbo Zhou, Xinlin Zhang, Yuanbin Chen, Qinquan Gao, Wenxi Liu, Tong Tong 0001
Expert Syst. Appl.8
2026 Perception-Aware Offloading With Collaborative Ground-Space Beamforming for Resilient SAGIN Communications
abstract
The integration of space, air, and ground segments into unified Space-Air-Ground Integrated Networks (SAGINs) enables low-latency, ubiquitous, and scalable computing. However, such systems face critical challenges: ground terminals suffer from weak satellite links, UAV-based edge nodes have limited resources, and highly dynamic environments make it difficult to make efficient offloading and resource allocation decisions. Prior approaches often optimize either communication or computation in isolation and lack adaptability to real-time environmental feedback. This paper presents a novel perception-aware hybrid-action deep reinforcement learning (DRL) framework for joint optimization of task offloading, beamforming, and resource allocation in SAGINs. To improve tractability, the original non-convex problem is first decomposed using Block Coordinate Descent (BCD) and approximated with Successive Convex Approximation (SCA), generating a structured feasible action space. A Soft Actor-Critic (SAC) agent then learns policies over this space, informed by real-time UAV perception via mmWave radar and vision sensors that detect user density, link quality, and environmental blockages. The DRL agent operates over a hybrid action space, combining discrete offloading decisions with continuous controls such as beamforming weights, CPU frequency, and transmission power. We employ a constraint-aware action masking mechanism that prunes infeasible hybrid actions violating delay, power, or SNR limits, thereby accelerating learning while respecting SAGIN-specific constraints. Extensive simulations show that the proposed framework significantly outperforms greedy, no-perception DRL, and state-of-the-art DRL offloading algorithms in reducing latency and energy consumption, while improving offloading success and resource stability. These results highlight the effectiveness of combining analytical optimization structure with adaptive perception-driven learning for robust and scalable control in future SAGINs.
Syed Muhammad Waqas, Anhui Liang, Xingsi Xue, Wenxi Liu, Jia Hu 0001, Mu-En Wu, Salman Raza, Fakhar Abbas
IEEE Internet Things J.5
2026 Distributed Edge Intelligence Framework for Secure and Efficient Data Sharing in 6G-IoV
abstract
This paper introduces a distributed edge intelligence framework for secure and efficient data sharing in 6G-Internet-of-Vehicles (6G-IoV) environments. The approach integrates a hierarchical blockchain architecture with an innovative asynchronous federated learning algorithm to address the challenges of privacy preservation, communication efficiency, and scalability in 6G-IoV scenarios. The proposed framework employs a multi-layer structure comprising vehicles, roadside units, and cloud servers, each playing a distinct role in the distributed learning process. We present a comprehensive latency model that accounts for computation and communication aspects, considering the unique characteristics of vehicular networks. The asynchronous federated learning algorithm incorporates genetic algorithm-based resource optimization to adaptively allocate communication resources, significantly enhancing efficiency in heterogeneous 6G-IoV environments. Furthermore, we introduce the hierarchical edge intelligence consensus mechanism, a lightweight consensus protocol tailored for edge intelligence scenarios, which accelerates blockchain consensus while maintaining security. Simulations using MNIST and SVHN datasets demonstrate that the proposed framework outperforms state-of-the-art methods regarding model accuracy, communication efficiency, and privacy preservation.
Hai Zhu 0001, Wenji Zhu, Quanzhen Huang, Hengzhou Xu, Zhongyang Yu, Wenxi Liu, Xingsi Xue
IEEE Internet Things J.7
2026 Adapting SAM to nuclei instance segmentation and classification via Cooperative Fine-Grained Refinement
Jingze Su, Tianle Zhu, Qi Li 0038, Tong Tong 0001, Wenxi Liu
Medical Image Anal.9
2026 ViFIT-assisted histopathology: From H&E style standardization to virtual fiber image transformation
Xingfu Wang, Chenyong Lv, Xiahui Han, Xiong Lin, Deyong Kang, Ruolan Lin, Liwen Hu 0008, Haohua Tu, Wenxi Liu
Medical Image Anal.12
2026 Integrating perceptual cues with mixture-of-experts for low-light image restoration
Yuezhou Li, Yuzhen Niu, Huangbiao Xu, Rui Xu 0028, Hui Da, Wenxi Liu, Lifang Wei
Neural Networks7
2026 Political Response Analysis of Twitter/X Users Using Topic-Based Sentiment Analysis
abstract
The heavy use of social media platforms is generating a high volume of affective data over the internet. This data is being used by researchers in various domains for prediction, qualitative, and quantitative analytical problems such as stock market prediction, opinion mining of online reviews on products, events, and many more. This article leverages X data for the political response analysis of users towards the 2019 Indian General election. In this article, a methodology is proposed that analyses X data to know what topics were mostly discussed during the election time under the #LoksabhaElection2019 hashtag. Also, we have tried to find out the sentiments of people towards different political terms (words) in the topics inferred. For this task, the study has used topic modeling and sentiment analysis of Tweets. This research may be useful for political parties or newsgroups to mine main topics and analyze the sentiments of people towards different entities.
Xingsi Xue, Priyavrat Chauhan, Sachin Kumar 0002, Himanshu Dhumras, Zhe Liu 0041, Wenxi Liu, G. Thippa Reddy
IEEE Trans. Comput. Soc. Syst.6
2026 Multimodal Image Representation Learning With Limited Visual-Tactile Data
abstract
Previous multimodal visual-tactile image representation learning (VTL) methods have achieved significant success in object understanding through large-scale training data. However, obtaining sufficient training data is often infeasible, and the above methods struggle to effectively focus on discriminative visual and tactile features with limited data, resulting in degraded performance. To solve the above issue, we introduce a new task called visual-tactile image representation learning with limited data (VTL-L), which better facilitates real-world applications. To address the challenges of limited data and modality discrepancy in the VTL-L task, we propose a novel multi-order feature enhancement-based, alignment-free fusion network (MOA-Net). First, we introduce a multi-order feature enhancement (MFE) module to hierarchically strengthen the detailed and structural representation by aggregating the low- and high-order topological information. This approach can effectively reduce the attention noise and obtain discriminative features with limited data. Then, we propose the alignment-free visual-tactile fusion (AVTF) module to achieve representative spatial and channel features and perform the cross-modality fusion without alignment, which efficiently mitigates the modality discrepancy. Finally, we develop a dual counterfactual intervention (DCI) loss to jointly optimize fused visual-tactile feature and probability distributions, thereby improving the performance of the MOA-Net in the VTL-L task. Extensive experiments demonstrate the superiority of the proposed method across three types of tasks on four datasets under diverse limited-data settings (source code available at: https://github.com/liuxiangqiu007/MOA-Net).
Liuxiang Qiu, Hui Da, Wenxi Liu, Yuzhen Niu, Hanli Wang, Tiesong Zhao
IEEE Trans. Image Process.3
2026 Dynamic Resource Allocation for RIS-Assisted Full-Duplex ISAC via Hybrid Lagrangian-DRL Approach
Syed Muhammad Waqas, Fakhar Abbas, Salman Raza, Wenxi Liu, Xingwang Li 0001, Xingsi Xue
IEEE Trans. Wirel. Commun.5
2025 URWKV: Unified RWKV Model with Multi-state Perspective for Low-light Image Restoration
abstract
Existing low-light image enhancement (LLIE) and joint LLIE and deblurring (LLIE-deblur) models have made strides in addressing predefined degradations, yet they are often constrained by dynamically coupled degradations. To address these challenges, we introduce a Unified Receptance Weighted Key Value (URWKV) model with multi-state perspective, enabling flexible and effective degradation restoration for low-light images. Specifically, we customize the core URWKV block to perceive and analyze complex degradations by leveraging multiple intra- and inter-stage states. First, inspired by the pupil mechanism in the human visual system, we propose Luminance-adaptive Normalization (LAN) that adjusts normalization parameters based on rich inter-stage states, allowing for adaptive, scene-aware luminance modulation. Second, we aggregate multiple intra-stage states through exponential moving average approach, effectively capturing subtle variations while mitigating information loss inherent in the single-state mechanism. To reduce the degradation effects commonly associated with conventional skip connections, we propose the State-aware Selective Fusion (SSF) module, which dynamically aligns and integrates multi-state features across encoder stages, selectively fusing contextual information. In comparison to state-of-the-art models, our URWKV model achieves superior performance on various benchmarks, while requiring significantly fewer parameters and computational resources. Code is available at: https://github.com/FZU-N/URWKV.
Rui Xu 0028, Yuzhen Niu, Yuezhou Li, Huangbiao Xu, Wenxi Liu, Yuzhong Chen 0001
CVPR5
2025 Keep the Balance: A Parameter-Efficient Symmetrical Framework for RGB+X Semantic Segmentation
abstract
Multimodal semantic segmentation is a critical challenge in computer vision, with early methods suffering from high computational costs and limited transferability due to full fine-tuning of RGB-based pre-trained parameters. Recent studies, while leveraging additional modalities as supplementary prompts to RGB, still predominantly rely on RGB, which restricts the full potential of other modalities. To address these issues, we propose a novel symmetric parameter-efficient fine-tuning framework for multimodal segmentation, featuring with a modality-aware prompting and adaptation scheme, to simultaneously adapt the capabilities of a powerful pre-trained model to both RGB and X modalities. Furthermore, prevalent approaches use the global cross-modality correlations of attention mechanism for modality fusion, which inadvertently introduces noise across modalities. To mitigate this noise, we propose a dynamic sparse cross-modality fusion module to facilitate effective and efficient cross-modality fusion. To further strengthen the above two modules, we propose a training strategy that leverages accurately predicted dual-modality results to self-teach the single-modality outcomes. In comprehensive experiments, we demonstrate that our method outperforms previous state-of-the-art approaches across six multimodal segmentation scenarios with minimal computation cost.
Jingze Su, Qi Li 0038, Wenjie Yang 0005, Tiesong Zhao, Shengfeng He, Wenxi Liu
CVPR8
2025 Stable Score Distillation
abstract
Text-guided image and 3D editing have advanced with diffusion-based models, yet methods like Delta Denoising Score often struggle with stability, spatial control, and editing strength. These limitations stem from reliance on complex auxiliary structures, which introduce conflicting optimization signals and restrict precise, localized edits. We introduce Stable Score Distillation (SSD), a streamlined framework that enhances stability and alignment in the editing process by anchoring a single classifier to the source prompt. Specifically, SSD utilizes Classifier-Free Guidance (CFG) equation to achieves cross-prompt alignment, and introduces a constant term null-text branch to stabilize the optimization process. This approach preserves the original content's structure and ensures that editing trajectories are closely aligned with the source prompt, enabling smooth, prompt-specific modifications while maintaining coherence in surrounding regions. Additionally, SSD incorporates a prompt enhancement branch to boost editing strength, particularly for style transformations. Our method achieves state-of-the-art results in 2D and 3D editing tasks, including NeRF and text-driven style edits, with faster convergence and reduced complexity, providing a robust and efficient solution for text-guided editing.
Haiming Zhu, Yangyang Xu 0003, Chenshu Xu, Tingrui Shen, Wenxi Liu, Yong Du 0003, Jun Yu 0002, Shengfeng He
ICCV5
2025 IPCMoE: Integrating Perceptual Cues with Mixture-of-Experts for Joint Low-Light Image Enhancement and Deblurring
abstract
Visual perception of nighttime images is often compromised by co-existing low-light and blur degradations. While recent methods have made progress in jointly solving these degradations, the diversity of patterns and intensities in degradation has not been properly considered, leading to inconsistent illumination and unintended artifacts. In response, we propose to integrate perceptual cues with mixture-of-experts (IPCMoE) to achieve flexible processing for low-light blurry images. By exploiting the perceptual cues, we strategically combine dedicated experts with the selective collaboration approach for feature enlightening and texture restoration. To this end, we develop perceptual-integrated MoEs by designing customized routers and task-depended experts. Specifically, the texture memorial MoE is developed to preserve valuable features to restore high-fidelity details, and the enhancement MoE that adaptively integrates enlightening cues and texture cues is designed to formulate the relationship between feature enlightening and texture restoration, thereby achieving dynamic image processing. Extensive experiments show that our method achieves state-of-the-art performance on LOL-Blur and Real-LOL-Blur datasets.
Yuezhou Li, Yuzhen Niu, Huangbiao Xu, Hui Da, Rui Xu 0028, Wenxi Liu
ACM Multimedia6
2025 Multi-UNet: An effective Multi-U convolutional networks for semantic segmentation
Qiangwei Zhao, Jingjing Cao, Junjie Ge, Wenxi Liu
Knowl. Based Syst.6
2025 Regional crowd flow estimation from aerial view
Huibin Wei, Qi Li 0038, Xindai Lin, Shengfeng He, Antoni B. Chan, Wenxi Liu
Neural Networks8
2025 Category-Contrastive Fine-Grained Crowd Counting and Beyond
abstract
Crowd counting has drawn increasing attention across various fields. However, existing crowd counting tasks primarily focus on estimating the overall population, ignoring the behavioral and semantic information of different social groups within the crowd. In this paper, we aim to address a newly proposed research problem, namely fine-grained crowd counting, which involves identifying different categories of individuals and accurately counting them in static images. In order to fully leverage the categorical information in static crowd images, we propose a two-tier salient feature propagation module designed to sequentially extract semantic information from both the crowd and its surrounding environment. Additionally, we introduce a category difference loss to refine the feature representation by highlighting the differences between various crowd categories. Moreover, our proposed framework can adapt to a novel problem setup called few-example fine-grained crowd counting. This setup, unlike the original fine-grained crowd counting, requires only a few exemplar point annotations instead of dense annotations from predefined categories, making it applicable in a wider range of scenarios. The baseline model for this task can be established by substituting the loss function in our proposed model with a novel hybrid loss function that integrates point-oriented cross-entropy loss and category contrastive loss. Through comprehensive experiments, we present results in both the formulation and application of fine-grained crowd counting.
Meijing Zhang, Mengxue Chen, Qi Li 0038, Yanchen Chen, Xiaolian Li, Shengfeng He, Wenxi Liu
IEEE Trans. Multim.8
2025 Another Perspective of Over-Smoothing: Alleviating Semantic Over-Smoothing in Deep GNNs
abstract
Graph neural networks (GNNs) are widely used for analyzing graph-structural data and solving graph-related tasks due to their powerful expressiveness. However, existing off-the-shelf GNN-based models usually consist of no more than three layers. Deeper GNNs usually suffer from severe performance degradation due to several issues including the infamous "over-smoothing" issue, which restricts the further development of GNNs. In this article, we investigate the over-smoothing issue in deep GNNs. We discover that over-smoothing not only results in indistinguishable embeddings of graph nodes, but also alters and even corrupts their semantic structures, dubbed semantic over-smoothing. Existing techniques, e.g., graph normalization, aim at handling the former concern, but neglect the importance of preserving the semantic structures in the spatial domain, which hinders the further improvement of model performance. To alleviate the concern, we propose a cluster-keeping sparse aggregation strategy to preserve the semantic structure of embeddings in deep GNNs (especially for spatial GNNs). Particularly, our strategy heuristically redistributes the extent of aggregations for all the nodes from layers, instead of aggregating them equally, so that it enables aggregate concise yet meaningful information for deep layers. Without any bells and whistles, it can be easily implemented as a plug-and-play structure of GNNs via weighted residual connections. Last, we analyze the over-smoothing issue on the GNNs with weighted residual structures and conduct experiments to demonstrate the performance comparable to the state-of-the-arts.
Jin Li 0032, Qirong Zhang, Wenxi Liu, Antoni B. Chan, Yanggeng Fu
IEEE Trans. Neural Networks Learn. Syst.3
2025 Upright-Net+: Enhanced Learning of Upright Orientation for 3D Point Clouds
abstract
Automatic 3D shape analysis is heavily influenced by the pose of input 3D models, as the continuous nature of pose space introduces complexities that usually exceed the encoding capacities of standard deep learning frameworks. To tackle this challenge, we present Upright-Net+, an enhancement of our previous model, Upright-Net, specifically developed for estimating upright orientation in 3D point clouds. Our approach is grounded in the design principle that "form ever follows function," treating the natural base of an object as a functional structure that stabilizes it in its typical pose, influenced by physical laws and geometric properties. We reformulate the continuous orientation problem into a discrete classification task, focusing on learning the points that constitute the natural base of a 3D model. The upright orientation is determined by aligning the normal orientation of this base towards the mass center. To mitigate over-smoothing in the global feature embeddings from stacked graph convolutional layers, we introduce a Global Positional Encoding Module using Relative Distance Histogram Statistics Embedding (GPE-RDHS), which reduces structural ambiguity and enhances orientation estimation. We also enhanced a weighted residual loss term to penalize false positive predictions, enhancing overall model performance. Our method demonstrates exceptional performance in upright orientation estimation and reveals that the learned orientation-aware features significantly benefit downstream tasks, particularly in classification.
Xufang Pang, Hongjie Zhuang, Ning Ding 0003, Xiaopin Zhong, Shengfeng He, Wenxi Liu
IEEE Trans. Vis. Comput. Graph.7
2025 Tooth Motion Monitoring in Orthodontic Treatment by Mobile Device-Based Multi-View Stereo
abstract
Nowadays, orthodontics has become an important part of modern personal life to assist one in improving mastication and raising self-esteem. However, the quality of orthodontic treatment still heavily relies on the empirical evaluation of experienced doctors, which lacks quantitative assessment and requires patients to visit clinics frequently for in-person examination. To resolve the aforementioned problem, we propose a novel and practical mobile device-based framework for precisely measuring tooth movement in treatment, so as to simplify and strengthen the traditional tooth monitoring process. To this end, we formulate the tooth movement monitoring task as a multi-view multi-object pose estimation problem via different views that capture multiple texture-less and severely occluded objects (i.e. teeth). Specifically, we exploit a pre-scanned 3D tooth model and a sparse set of multi-view tooth images as inputs for our proposed tooth monitoring framework. After extracting tooth contours and localizing the initial camera pose of each view from the initial configuration, we propose a joint pose estimation scheme to precisely estimate the 3D pose of each individual tooth, so as to infer their relative offsets during treatment. Furthermore, we introduce the metric of Relative Pose Bias to evaluate the individual tooth pose accuracy in a small scale. We demonstrate that our approach is capable of reaching high accuracy and efficiency as practical orthodontic treatment monitoring requires.
Jiaming Xie, Congyi Zhang 0001, Guangshun Wei, Peng Wang 0099, Guodong Wei, Wenxi Liu, Min Gu 0003, Ping Luo 0002, Wenping Wang 0001
IEEE Trans. Vis. Comput. Graph.6
2024 Efficient Semantic Segmentation for Compressed Video
abstract
Robots, constrained by limited onboard computing resources, often encounter situations wherein high-resolution and high-bit-rate videos captured by their cameras necessitate compression before further analysis. In this paper, we propose a novel video semantic segmentation paradigm for compressed video. Specifically, our framework draws the inspiration from the principle of Wavelet Transform, and thus we design the network structure, WTDecomNet, approximating the decomposition of high-resolution image into its low-resolution counterpart and axial details. The aim is to well preserve the image content through decomposition and maintain model efficiency by obtaining semantics from low-resolution image. To facilitate this purpose, we propose an efficient axial subband approximation module for extracting axial details and a lightweight temporal alignment module for associating keyframes and non-keyframes of compressed video. Through comprehensive experiments, we show that our model can achieve the state-of-the-art performance on public benchmarks. Especially on CamVid, comparing to baseline, our proposed model reduces the computational overhead by ∼70% while improving mIoU by ∼4%.
Qi Li 0038, Jia Pan 0001, Wenxi Liu
ICRA5
2024 Revisiting Network Perturbation for Semi-supervised Semantic Segmentation
Sien Li, Tao Wang 0047, Ruizhe Hu, Wenxi Liu
PRCV (12)4
2024 Glance to Count: Learning to Rank with Anchors for Weakly-supervised Crowd Counting
abstract
Crowd image is arguably one of the most laborious data to annotate. In this paper, we aim to reduce the massive demand for densely labeled crowd data, and propose a novel weakly-supervised setting, in which we leverage the binary ranking of two images with high-contrast crowd counts as training guidance. To enable training under this new setting, we convert the crowd count regression problem to a ranking potential prediction problem. In particular, we tailor a Siamese Ranking Network that predicts the potential scores of two images indicating the ordering of the counts. Hence, the ultimate goal is to assign appropriate potentials for all the crowd images to ensure their orderings obey the ranking labels. On the other hand, potentials reveal the relative crowd sizes but cannot yield an exact crowd count. We resolve this problem by introducing "anchors" during the inference stage. Concretely, anchors are a few images with count labels used for referencing the corresponding counts from potential scores by a simple linear mapping function. We conduct extensive experiments to study various combinations of supervision, and we show that our method outperforms existing weakly-supervised methods by a large margin without additional labeling effort. The code is available at https://github.com/pandaszzzzz/CCRanking.
Zheng Xiong, Liangyu Chai, Wenxi Liu, Yongtuo Liu, Sucheng Ren, Shengfeng He
WACV3
2024 Ultra-High Resolution Image Segmentation via Locality-Aware Context Fusion and Alternating Local Enhancement
Wenxi Liu, Qi Li 0038, Xindai Lin, Weixiang Yang, Shengfeng He, Yuanlong Yu 0001
Int. J. Comput. Vis.1
2024 Pose-aware video action segmentation
Meijing Zhang, Chenyang Liao, Qi Li 0038, Wenxi Liu
Neural Comput. Appl.5
2024 FE-Net: Feature enhancement segmentation network
Zhangyan Zhao, Jingjing Cao, Qiangwei Zhao, Wenxi Liu
Neural Networks5
2024 Monocular BEV Perception of Road Scenes via Front-to-Top View Projection
abstract
HD map reconstruction is crucial for autonomous driving. LiDAR-based methods are limited due to expensive sensors and time-consuming computation. Camera-based methods usually need to perform road segmentation and view transformation separately, which often causes distortion and missing content. To push the limits of the technology, we present a novel framework that reconstructs a local map formed by road layout and vehicle occupancy in the bird's-eye view given a front-view monocular image only. We propose a front-to-top view projection (FTVP) module, which takes the constraint of cycle consistency between views into account and makes full use of their correlation to strengthen the view transformation and scene understanding. In addition, we apply multi-scale FTVP modules to propagate the rich spatial information of low-level features to mitigate spatial deviation of the predicted object location. Experiments on public benchmarks show that our method achieves various tasks on road layout estimation, vehicle occupancy estimation, and multi-class semantic estimation, at a performance level comparable to the state-of-the-arts, while maintaining superior efficiency.
Wenxi Liu, Qi Li 0038, Weixiang Yang, Yuanlong Yu 0001, Yuexin Ma, Shengfeng He, Jia Pan 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Semi-supervised domain generalization with evolving intermediate domain
Luojun Lin, Zhishu Sun, Weijie Chen 0006, Wenxi Liu, Yuanlong Yu 0001, Lei Zhang 0038
Pattern Recognit.5
2024 Prototype learning based generic multiple object tracking via point-to-box supervision
Wenxi Liu, Qi Li 0038, Yinhua She, Yuanlong Yu 0001, Jia Pan 0001, Jason Gu
Pattern Recognit.1
2024 STD-Net: Spatio-Temporal Decomposition Network for Video Demoiréing With Sparse Transformers
abstract
The problem of video demoiréing is a new challenge in video restoration. Unlike image demoiréing, which involves removing static and uniform patterns, video demoiréing requires tackling dynamic and varied moiré patterns while maintaining video details, colors, and temporal consistency. It is particularly challenging to model moiré patterns for videos with camera or object motions, where separating moiré from the original video content across frames is extremely difficult. Nonetheless, we observe that the spatial distribution of moiré patterns is often sparse on each frame, and their long-range temporal correlation is not significant. To fully leverage this phenomenon, a sparsity-constrained spatial self-attention scheme is proposed to concentrate on removing sparse moiré efficiently for each frame without being distracted by dynamic video content. The frame-wise spatial features are then correlated and aggregated via the local temporal cross-frame-attention module to produce temporal-consistent high-quality moiré-free videos. The above decoupled spatial and temporal transformers constitute the Spatio-Temporal Decomposition Network, dubbed STD-Net. For evaluation, we present a large-scale video demoiréing benchmark featuring various real-life scenes, camera motions, and object motions. We demonstrate that our proposed model can effectively and efficiently achieve superior performance on video demoiréing and single image demoiréing tasks.The proposed dataset will be released after the paper is accepted.
Yuzhen Niu, Rui Xu 0028, Zhihua Lin, Wenxi Liu
IEEE Trans. Circuits Syst. Video Technol.4
2024 Attentive and Contrastive Image Manipulation Localization With Boundary Guidance
abstract
In recent years, the rapid advancement of image generation techniques has resulted in the widespread abuse of manipulated images, leading to a crisis of trust and affecting social equity. Thus, the goal of our work is to detect and localize tampered regions in images. Many deep learning based approaches have been proposed to address this problem, but they can hardly handle the tampered regions that are manually fine-tuned to blend into image background. By observing that the boundaries of tempered regions are critical to separating tampered and non-tampered parts, we present a novel boundary-guided approach to image manipulation detection, which introduces an inherent bias towards exploiting the boundary information of tampered regions. Our model follows an encoder-decoder architecture, with multi-scale localization mask prediction, and is guided to utilize the prior boundary knowledge through an attention mechanism and contrastive learning. In particular, our model is unique in that 1) we propose a boundary-aware attention module in the network decoder, which predicts the boundary of tampered regions and thus uses it as crucial contextual cues to facilitate the localization; and 2) we propose a multi-scale contrastive learning scheme with a novel boundary-guided sampling strategy, leading to more discriminative localization features. Our state-of-art performance on several public benchmarks demonstrates the superiority of our model over prior works.
Wenxi Liu, Hao Zhang 0170, Xinyang Lin, Qi Li 0038, Xiaoxiang Liu, Ying Cao 0001
IEEE Trans. Inf. Forensics Secur.1
2024 Learning Nighttime Semantic Segmentation the Hard Way
abstract
Nighttime semantic segmentation is an important but challenging research problem for autonomous driving. The major challenges lie in the small objects or regions from the under-/over-exposed areas or suffer from motion blur caused by the camera deployed on moving vehicles. To resolve this, we propose a novel hard-class-aware module that bridges the main network for full-class segmentation and the hard-class network for segmenting aforementioned hard-class objects. In specific, it exploits the shared focus of hard-class objects from the dual-stream network, enabling the contextual information flow to guide the model to concentrate on the pixels that are hard to classify. In the end, the estimated hard-class segmentation results will be utilized to infer the final results via an adaptive probabilistic fusion refinement scheme. Moreover, to overcome over-smoothing and noise caused by extreme exposures, our model is modulated by a carefully crafted pretext task of constructing an exposure-aware semantic gradient map, which guides the model to faithfully perceive the structural and semantic information of hard-class objects while mitigating the negative impact of noises and uneven exposures. In experiments, we demonstrate that our unique network design leads to superior segmentation performance over existing methods, featuring the strong ability of perceiving hard-class objects under adverse conditions.
Wenxi Liu, Qi Li 0038, Chenyang Liao, Jingjing Cao, Shengfeng He, Yuanlong Yu 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2023 Frame-Event Alignment and Fusion Network for High Frame Rate Tracking
abstract
Most existing RGB-based trackers target low frame rate benchmarks of around 30 frames per second. This setting restricts the tracker's functionality in the real world, especially for fast motion. Event-based cameras as bioinspired sensors provide considerable potential for high frame rate tracking due to their high temporal resolution. However, event-based cameras cannot offer fine-grained texture information like conventional cameras. This unique complementarity motivates us to combine conventional frames and events for high frame rate object tracking under various challenging conditions. In this paper, we propose an end-to-end network consisting of multi-modality alignment and fusion modules to effectively combine meaningful information from both modalities at different measurement rates. The alignment module is responsible for cross-style and cross-frame-rate alignment between frame and event modalities under the guidance of the moving cues furnished by events. While the fusion module is accountable for emphasizing valuable features and suppressing noise information by the mutual complement between the two modalities. Extensive experiments show that the proposed approach outper-forms state-of-the-art trackers by a significant margin in high frame rate tracking. With the FE240Hz dataset, our approach achieves high frame rate tracking up to 240Hz.
Jiqing Zhang, Yuanchen Wang, Wenxi Liu, Meng Li 0072, Jinpeng Bai, Xin Yang 0011
CVPR3
2023 Diffuse3D: Wide-Angle 3D Photography via Bilateral Diffusion
abstract
This paper aims to resolve the challenging problem of wide-angle novel view synthesis from a single image, a.k.a. wide-angle 3D photography. Existing approaches rely on local context and treat them equally to inpaint occluded RGB and depth regions, which fail to deal with large-region occlusion (i.e., observing from an extreme angle) and foreground layers might blend into background inpainting. To address the above issues, we propose Diffuse3D which employs a pre-trained diffusion model for global synthesis, while amending the model to activate depth-aware inference. Our key insight is to alter the convolution mechanism in the denoising process. We inject depth information into the denoising convolution operation with bilateral kernels, i.e., a depth kernel and a spatial kernel, to consider layered correlations among pixels. In this way, foreground regions are overlooked in background inpainting and only pixels close in depth are leveraged. On the other hand, we propose a global-local balancing approach to maximize both contextual understandings. Extensive experiments demonstrate that our approach outperforms state-of-the-art methods in novel view synthesis, especially in wide-angle scenarios. More importantly, our method does not require any training and is a plug-and-play module that can be integrated with any diffusion model. Our code can be found at https://github.com/yutaojiang1/Diffuse3D.
Yutao Jiang, Yang Zhou 0038, Wenxi Liu, Jianbo Jiao, Yuhui Quan, Shengfeng He
ICCV4
2023 CIRI: Curricular Inactivation for Residue-aware One-shot Video Inpainting
abstract
Video inpainting aims at filling in missing regions of a video. However, when dealing with dynamic scenes with camera or object movements, annotating the inpainting target becomes laborious and impractical. In this paper, we resolve the one-shot video inpainting problem in which only one annotated first frame is provided. A naive solution is to propagate the initial target to the other frames with techniques like object tracking. In this context, the main obstacles are the unreliable propagation and the partially inpainted artifacts due to the inaccurate mask. For the former problem, we propose curricular inactivation to replace the hard masking mechanism for indicating the in-painting target, which is robust to erroneous predictions in long-term video inpainting. For the latter, we explore the properties of inpainting residue and present an online residue removal method in an iterative detect-and-refine manner. Extensive experiments on several real-world datasets demonstrate the quantitative and qualitative superiorities of our proposed method in one-shot video inpainting. More importantly, our method is extremely flexible that can be integrated with arbitrary traditional inpainting models, activating them to perform the reliable one-shot video inpainting task. Video demonstrations can be found in our supplement, and our code can be found at https://github.com/Arise-zwy/CIRI.
Weiying Zheng, Xuemiao Xu, Wenxi Liu, Shengfeng He
ICCV4
2023 Single-View View Synthesis with Self-rectified Pseudo-Stereo
Yang Zhou 0038, Hanjie Wu, Wenxi Liu, Zheng Xiong, Harry Qin, Shengfeng He
Int. J. Comput. Vis.3
2023 Monocular Depth Estimation for Glass Walls With Context: A New Dataset and Method
abstract
Traditional monocular depth estimation assumes that all objects are reliably visible in the RGB color domain. However, this is not always the case as more and more buildings are decorated with transparent glass walls. This problem has not been explored due to the difficulties in annotating the depth levels of glass walls, as commercial depth sensors cannot provide correct feedbacks on transparent objects. Furthermore, estimating depths from transparent glass walls requires the aids of surrounding context, which has not been considered in prior works. To cope with this problem, we introduce the first Glass Walls Depth Dataset (GW-Depth dataset). We annotate the depth levels of transparent glass walls by propagating the context depth values within neighboring flat areas, and the glass segmentation mask and instance level line segments of glass edges are also provided. On the other hand, a tailored monocular depth estimation method is proposed to fully activate the glass wall contextual understanding. First, we propose to exploit the glass structure context by incorporating the structural prior knowledge embedded in glass boundary line segment detections. Furthermore, to make our method adaptive to scenes without structure context where the glass boundary is either absent in the image or too narrow to be recognized, we propose to derive a reflection context by utilizing the depth reliable points sampled according to the variance between two depth estimations from different resolutions. High-resolution depth is thus estimated by the weighted summation of depths by those reliable points. Extensive experiments are conducted to evaluate the effectiveness of the proposed dual context design. Superior performances of our method is also demonstrated by comparing with state-of-the-art methods. We present the first feasible solution for monocular depth estimation in the presence of glass walls, which can be widely adopted in autonomous navigation.
Bailin Deng, Wenxi Liu, Harry Qin, Shengfeng He
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Monocular Camera-Based Complex Obstacle Avoidance via Efficient Deep Reinforcement Learning
abstract
Deep reinforcement learning has achieved great success in laser-based collision avoidance works because the laser can sense accurate depth information without too much redundant data, which can maintain the robustness of the algorithm when it is migrated from the simulation environment to the real world. However, high-cost laser devices are not only difficult to deploy for a large scale of robots but also demonstrate unsatisfactory robustness towards the complex obstacles, including irregular obstacles, e.g., tables, chairs, and shelves, as well as complex ground and special materials. In this paper, we propose a novel monocular camera-based complex obstacle avoidance framework. Particularly, we innovatively transform the captured RGB images to pseudo-laser measurements for efficient deep reinforcement learning. Compared to the traditional laser measurement captured at a certain height that only contains one-dimensional distance information away from the neighboring obstacles, our proposed pseudo-laser measurement fuses the depth and semantic information of the captured RGB image, which makes our method effective for complex obstacles. We also design a feature extraction guidance module to weight the input pseudo-laser measurement, and the agent has more reasonable attention for the current state, which is conducive to improving the accuracy and efficiency of the obstacle avoidance policy. Besides, we adaptively add the synthesized noise to the laser measurement during the training stage to decrease the sim-to-real gap and increase the robustness of our model in the real environment. Finally, the experimental results show that our framework achieves state-of-the-art performance in several virtual and real-world scenarios.
Jianchuan Ding, Lingping Gao, Wenxi Liu, Haiyin Piao, Jia Pan 0001, Zhenjun Du, Xin Yang 0011
IEEE Trans. Circuits Syst. Video Technol.3
2023 Comment-Guided Semantics-Aware Image Aesthetics Assessment
abstract
Existing image aesthetics assessment methods mainly rely on the visual features of images but ignore their rich semantics. Nowadays, with the widespread application of social media, the comments corresponding to images in the form of texts can be easily accessed and provide rich semantic information, which can be utilized to effectively complement image features. This paper proposes a comment-guided semantics-aware image aesthetics assessment method, which is built upon a multi-task learning framework for image aesthetics prediction and comment-guided semantics classification. To assist image aesthetics assessment, we first model the semantics of an image as the topic features of its corresponding comments using Latent Dirichlet Allocation. We then propose a two-stream multitask learning framework for both topic feature prediction and aesthetic score distribution prediction. Topic feature prediction task enables to infer the semantics from images, since the comments are usually unavailable during inference and comment-guided semantics can only serve as supervision during training. We further propose to deeply fuse aesthetics and semantic features using a layerwise feature fusion method. Experimental results demonstrate that the proposed method outperforms state-of-the-art image aesthetics assessment methods.
Yuzhen Niu, Bingrui Song, Wenxi Liu
IEEE Trans. Circuits Syst. Video Technol.5
2023 Progressive Moire Removal and Texture Complementation for Image Demoireing
abstract
Taking photos of digital screens often produces color-distorting moire patterns caused by inconsistency between the color filter array of cameras and the sub-pixel layout of screens, which severely degrades the quality of photos. Most existing demoireing methods employ multi-stream network architecture to simultaneously process the same moire image with different resolutions, but they neglect the complementarity among different resolutions. In this paper, we propose a novel moire removal model to address this issue. Unlike the existing multi-stream based approaches, in this model, we present a progressive texture complementation block to exploit the complementary information from different resolutions in order to progressively remove moire textures and restore image content. Additionally, we propose a residual moire removal block, in which the depthwise separable convolution is utilized to remove moire from image while reducing computation overhead. This block also includes a local color correction structure, which is used to correct color shifts presented in the moire images. Experimental results on two public datasets show that our method outperforms state-of-the-art methods. Besides, the quantity of parameters and FLOPs of our model are tens of times fewer than the off-the-shelf models. Furthermore, our network framework can adapt well to another low-level vision task, rain removal, in which our model also achieves state-of-the-art performance.
Yuzhen Niu, Zhihua Lin, Wenxi Liu, Wenzhong Guo
IEEE Trans. Circuits Syst. Video Technol.3
2023 Finite-Time Synchronization of Fractional-Order Fuzzy Time-Varying Coupled Neural Networks Subject to Reaction-Diffusion
abstract
In this article, finite-time synchronization is investigated for fractional-order fuzzy time-varying coupled neural networks subject to reaction–diffusion by establishing a new framework under fuzzy-based feedback control and fuzzy-based adaptive control. For the considered networks, we put forward an innovative graph-theory-based time-varying Lyapunov function. To overcome the difficulty of estimating the fractional derivative of this function, this article proposes a novel fractional derivative rule. Through graph theory and the Lyapunov method, several finite-time synchronous criteria are obtained for the considered networks, and the estimation of the settling time is derived. Finally, the numerical results are shown to demonstrate the practicability of the given results.
Wenxi Liu, Yongbao Wu, Wenxue Li 0001
IEEE Trans. Fuzzy Syst.2
2023 Distractor-Aware Event-Based Tracking
abstract
Event cameras, or dynamic vision sensors, have recently achieved success from fundamental vision tasks to high-level vision researches. Due to its ability to asynchronously capture light intensity changes, event camera has an inherent advantage to capture moving objects in challenging scenarios including objects under low light, high dynamic range, or fast moving objects. Thus event camera are natural for visual object tracking. However, the current event-based trackers derived from RGB trackers simply modify the input images to event frames and still follow conventional tracking pipeline that mainly focus on object texture for target distinction. As a result, the trackers may not be robust dealing with challenging scenarios such as moving cameras and cluttered foreground. In this paper, we propose a distractor-aware event-based tracker that introduces transformer modules into Siamese network architecture (named DANet). Specifically, our model is mainly composed of a motion-aware network and a target-aware network, which simultaneously exploits both motion cues and object contours from event data, so as to discover motion objects and identify the target object by removing dynamic distractors. Our DANet can be trained in an end-to-end manner without any post-processing and can run at over 80 FPS on a single V100. We conduct comprehensive experiments on two large event tracking datasets to validate the proposed model. We demonstrate that our tracker has superior performance against the state-of-the-art trackers in terms of both accuracy and efficiency.
Yingkai Fu, Meng Li 0072, Wenxi Liu, Yuanchen Wang, Jiqing Zhang, Xiaopeng Wei, Xin Yang 0011
IEEE Trans. Image Process.3
2022 End-to-End Trajectory Distribution Prediction Based on Occupancy Grid Maps
abstract
In this paper, we aim to forecast a future trajectory distribution of a moving agent in the real world, given the social scene images and historical trajectories. Yet, it is a challenging task because the ground-truth distribution is unknown and unobservable, while only one of its samples can be applied for supervising model learning, which is prone to bias. Most recent works focus on predicting diverse trajectories in order to cover all modes of the real distribution, but they may despise the precision and thus give too much credit to unrealistic predictions. To address the issue, we learn the distribution with symmetric cross-entropy using occupancy grid maps as an explicit and scene-compliant approximation to the ground-truth distribution, which can effectively penalize unlikely predictions. In specific, we present an inverse reinforcement learning based multi-modal trajectory distribution forecasting framework that learns to plan by an approximate value iteration network in an end-to-end manner. Besides, based on the predicted distribution, we generate a small set of representative trajectories through a differentiable Transformer-based network, whose attention mechanism helps to model the relations of trajectories. In experiments, our method achieves state-of-the-art performance on the Stanford Drone Dataset and Intersection Drone Dataset.
Wenxi Liu, Jia Pan 0001
CVPR2
2022 Perceptual-Aware and Restorable Real-Time Image Downscaling
abstract
Image downscaling has been a classical problem and has recently been linked to super-resolution (SR). In this paper, we aim to propose a learning-based image downscaling model, FastDownscaler, which can efficiently produce low-resolution (LR) images that not only preserve the rich details of the original high-resolution images but also be highly restorable for existing SR models. We first present two separate lightweight networks with different upsampling losses, the bilinear loss and the bicubic loss, which are better for SR restoration and LR downscaling, respectively. To produce versatile LR images, we then propose to distill bilinear loss guided network with bicubic loss guided one. To our best knowledge, we establish the first image downscaling quality assessment dataset to evaluate the downscaling performance. Experimental results demonstrate the superior performance of the proposed model on image downscaling and SR. Furthermore, our model can achieve over 600 FPS for downscaling a$1920\times 1280$image.
Yuzhen Niu, Luwei Zheng, Jianbin Wu, Wenxi Liu
ICME5
2022 Continuous Transformation Superposition for Visual Comfort Enhancement of Casual Stereoscopic Photography
abstract
Casual stereoscopic photography allows ordinary users to create a stereoscopic photo using two photos taken casually by a monocular camera. The visual comfort of a casual stereoscopic photo can greatly affect its visual experience. In this paper, we present a novel visual comfort enhancement method for casual stereoscopic photography via reinforcement learning based on continuous transformation superposition. We consider the transformation, in a continuous transformation space, to transform each view as superpositions of several basic continuous transformations, enabling more subtle and flexible image transformation operations to approach better solutions. To achieve the continuous transformation superposition, we prepare a collection of continuous transformation models for translation, rotation, and perspective transformations. Then we train a policy model to determine an optimal transformation chain to recurrently handle both the geometric constraints and disparity adjustment, and thereby enhance the visual comfort of casual stereoscopic images. We further propose an attention-based stereo feature fusion module that enhances and integrates the binocular information between the left and right views. Experimental results on three datasets demonstrate that our proposed method achieves superior performance to state-of-the-art methods.
Yuzhong Chen 0001, Qijin Shen, Yuzhen Niu, Wenxi Liu
VR4
2022 CrowdGAN: Identity-Free Interactive Crowd Video Generation and Beyond
abstract
In this paper, we introduce a novel yet challenging research problem, interactive crowd video generation, committed to producing diverse and continuous crowd video, and relieving the difficulty of insufficient annotated real-world datasets in crowd analysis. Our goal is to recursively generate realistic future crowd video frames given few context frames, under the user-specified guidance, namely individual positions of the crowd. To this end, we propose a deep network architecture specifically designed for crowd video generation that is composed of two complementary modules, each of which combats the problems of crowd dynamic synthesis and appearance preservation respectively. Particularly, a spatio-temporal transfer module is proposed to infer the crowd position and structure from guidance and temporal information, and a point-aware flow prediction module is presented to preserve appearance consistency by flow-based warping. Then, the outputs of the two modules are integrated by a self-selective fusion unit to produce an identity-preserved and continuous video. Unlike previous works, we generate continuous crowd behaviors beyond identity annotations or matching. Extensive experiments show that our method is effective for crowd video generation. More importantly, we demonstrate the generated video can produce diverse crowd behaviors and be used for augmenting different crowd analysis tasks, i.e., crowd counting, anomaly detection, crowd video prediction. Code is available at https://github.com/Icep2020/CrowdGAN.
Liangyu Chai, Yongtuo Liu, Wenxi Liu, Guoqiang Han 0002, Shengfeng He
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 3-D Deconvolutional Networks for the Unsupervised Representation Learning of Human Motions
abstract
Data representation learning is one of the most important problems in machine learning. Unsupervised representation learning becomes meritorious as it has no necessity of label information with observed data. Due to the highly time-consuming learning of deep-learning models, there are many machine-learning models directly adapting well-trained deep models that are obtained in a supervised and end-to-end manner as feature abstractors to distinct problems. However, it is obvious that different machine-learning tasks require disparate representation of original input data. Taking human action recognition as an example, it is well known that human actions in a video sequence are 3-D signals containing both visual appearance and motion dynamics of humans and objects. Therefore, the data representation approaches with the capabilities to capture both spatial and temporal correlations in videos are meaningful. Most of the existing human motion recognition models build classifiers based on deep-learning structures such as deep convolutional networks. These models require a large quantity of training videos with annotations. Meanwhile, these supervised models cannot recognize samples from the distinct dataset without retraining. In this article, we propose a new 3-D deconvolutional network (3DDN) for representation learning of high-dimensional video data, in which the high-level features are obtained through the optimization approach. The proposed 3DDN decomposes the video frames into spatiotemporal features under a sparse constraint in an unsupervised way. In addition, it also can be regarded as a building block to develop deep architectures by stacking. The high-level representation of input sequential data can be used in multiple downstream machine-learning tasks, we evaluate the proposed 3DDN and its deep models in human action recognition. The experimental results from three datasets: 1) KTH data; 2) HMDB-51; and 3) UCF-101, demonstrate that the proposed 3DDN is an alternative approach to feedforward convolutional neural networks (CNNs), that attains comparable results.
Chun-Yang Zhang, Yong-Yi Xiao, Jin-Cheng Lin, C. L. Philip Chen, Wenxi Liu, Yu-Hong Tong
IEEE Trans. Cybern.5
2021 Reciprocal Transformations for Unsupervised Video Object Segmentation
abstract
Unsupervised video object segmentation (UVOS) aims at segmenting the primary objects in videos without any human intervention. Due to the lack of prior knowledge about the primary objects, identifying them from videos is the major challenge of UVOS. Previous methods often regard the moving objects as primary ones and rely on optical flow to capture the motion cues in videos, but the flow information alone is insufficient to distinguish the primary objects from the background objects that move together. This is because, when the noisy motion features are combined with the appearance features, the localization of the primary objects is misguided. To address this problem, we propose a novel reciprocal transformation network to discover primary objects by correlating three key factors: the intra-frame contrast, the motion cues, and temporal coherence of recurring objects. Each corresponds to a representative type of primary object, and our reciprocal mechanism enables an organic coordination of them to effectively remove ambiguous distractions from videos. Additionally, to exclude the information of the moving background objects from motion features, our transformation module enables to reciprocally transform the appearance features to enhance the motion features, so as to focus on the moving objects with salient appearance while removing the co-moving outliers. Experiments on the public benchmarks demonstrate that our model significantly outperforms the state-of-the-art methods. Code is available at https://github.com/OliverRensu/RTNet.
Sucheng Ren, Wenxi Liu, Yongtuo Liu, Haoxin Chen, Guoqiang Han 0002, Shengfeng He
CVPR2
2021 Projecting Your View Attentively: Monocular Road Scene Layout Estimation via Cross-View Transformation
abstract
HD map reconstruction is crucial for autonomous driving. LiDAR-based methods are limited due to the deployed expensive sensors and time-consuming computation. Camera-based methods usually need to separately perform road segmentation and view transformation, which often causes distortion and the absence of content. To push the limits of the technology, we present a novel framework that enables reconstructing a local map formed by road layout and vehicle occupancy in the bird’s-eye view given a front-view monocular image only. In particular, we propose a cross-view transformation module, which takes the constraint of cycle consistency between views into account and makes full use of their correlation to strengthen the view transformation and scene understanding. Considering the relationship between vehicles and roads, we also design a context-aware discriminator to further refine the results. Experiments on public benchmarks show that our method achieves the state-of-the-art performance in the tasks of road layout estimation and vehicle occupancy estimation. Especially for the latter task, our model outperforms all competitors by a large margin. Furthermore, our model runs at 35 FPS on a single GPU, which is efficient and applicable for real-time panorama HD map reconstruction.
Weixiang Yang, Qi Li 0038, Wenxi Liu, Yuanlong Yu 0001, Yuexin Ma, Shengfeng He, Jia Pan 0001
CVPR3
2021 Coarse-To-Fine Person Re-Identification With Auxiliary-Domain Classification and Second-Order Information Bottleneck
abstract
Person re-identification (Re-ID) is to retrieve a particular person captured by different cameras, which is of great significance for security surveillance and pedestrian behavior analysis. However, due to the large intra-class variation of a person across cameras, e.g., occlusions, illuminations, viewpoints, and poses, Re-ID is still a challenging task in the field of computer vision. In this paper, to attack the issues concerning with intra-class variation, we propose a coarse-to-fine Re-ID framework with the incorporation of auxiliary-domain classification (ADC) and second-order information bottleneck (2O-IB). In particular, as an auxiliary task, ADC is introduced to extract the coarse-grained essential features to distinguish a person from miscellaneous backgrounds, which leads to the effective coarse- and fine-grained feature representations for Re-ID. On the other hand, to cope with the redundancy, irrelevance, and noise contained in the Re-ID features caused by intra-class variations, we integrate 2O-IB into the network to compress and optimize the features, without increasing additional computation overhead during inference. Experimental results demonstrate that our proposed method significantly reduces the neural network output variance of intra-class person images and achieves the superior performance to state-of-the-art methods.
Anguo Zhang, Yueming Gao, Yuzhen Niu, Wenxi Liu, Yongcheng Zhou
CVPR4
2021 From Contexts to Locality: Ultra-high Resolution Image Segmentation via Locality-aware Contextual Correlation
abstract
Ultra-high resolution image segmentation has raised increasing interests in recent years due to its realistic applications. In this paper, we innovate the widely used high-resolution image segmentation pipeline, in which an ultrahigh resolution image is partitioned into regular patches for local segmentation and then the local results are merged into a high-resolution semantic mask. In particular, we introduce a novel locality-aware contextual correlation based segmentation model to process local patches, where the relevance between local patch and its various contexts are jointly and complementarily utilized to handle the semantic regions with large variations. Additionally, we present a contextual semantics refinement network that associates the local segmentation result with its contextual semantics, and thus is endowed with the ability of reducing boundary artifacts and refining mask contours during the generation of final high-resolution mask. Furthermore, in comprehensive experiments, we demonstrate that our model outperforms other state-of-the-art methods in public benchmarks. Our released codes are available at https://github.com/liqiokkk/FCtL.
Qi Li 0038, Weixiang Yang, Wenxi Liu, Yuanlong Yu 0001, Shengfeng He
ICCV3
2021 A Vision-based Irregular Obstacle Avoidance Framework via Deep Reinforcement Learning
abstract
Deep reinforcement learning has achieved great success in laser-based collision avoidance work because the laser can sense accurate depth information without too much redundant data, which can maintain the robustness of the algorithm when it is migrated from the simulation environment to the real world. However, high-cost laser devices are not only difficult to apply on a large scale but also have poor robustness to irregular objects, e.g., tables, chairs, shelves, etc. In this paper, we propose a vision-based collision avoidance framework to solve the challenging problem. Our method attempts to estimate the depth and incorporate the semantic information from RGB data to obtain a new form of data, pseudo-laser data, which combines the advantages of visual information and laser information. Compared to traditional laser data that only contains the one-dimensional distance information captured at a certain height, our proposed pseudo-laser data encodes the depth information and semantic information within the image, which makes our method more effective for irregular obstacles. Besides, we adaptively add noise to the laser data during the training stage to increase the robustness of our model in the real world, due to the estimated depth information is not accurate. Experimental results show that our framework achieves state-of-the-art performance in several unseen virtual and real-world scenarios.
Lingping Gao, Jianchuan Ding, Wenxi Liu, Haiyin Piao, Yuxin Wang 0001, Xin Yang 0011
IROS3
2021 EBRB cascade classifier for imbalanced data via rule weight updating
Yanggeng Fu, Hong-Yun Huang, Ying-Ming Wang 0001, Wenxi Liu, Weijie Fang
Knowl. Based Syst.5
2021 HDR-GAN: HDR Image Reconstruction From Multi-Exposed LDR Images With Large Motions
abstract
Synthesizing high dynamic range (HDR) images from multiple low-dynamic range (LDR) exposures in dynamic scenes is challenging. There are two major problems caused by the large motions of foreground objects. One is the severe misalignment among the LDR images. The other is the missing content due to the over-/under-saturated regions caused by the moving objects, which may not be easily compensated for by the multiple LDR exposures. Thus, it requires the HDR generation model to be able to properly fuse the LDR images and restore the missing details without introducing artifacts. To address these two problems, we propose in this paper a novel GAN-based model, HDR-GAN, for synthesizing HDR images from multi-exposed LDR images. To our best knowledge, this work is the first GAN-based approach for fusing multi-exposed LDR images for HDR reconstruction. By incorporating adversarial learning, our method is able to produce faithful information in the regions with missing content. In addition, we also propose a novel generator network, with a reference-based residual merging block for aligning large object motions in the feature domain, and a deep HDR supervision scheme for eliminating artifacts of the reconstructed HDR images. Experimental results demonstrate that our model achieves state-of-the-art reconstruction performance over the prior HDR methods on diverse scenes.
Yuzhen Niu, Jianbin Wu, Wenxi Liu, Wenzhong Guo, Rynson W. H. Lau
IEEE Trans. Image Process.3
2020 Mapping in a Cycle: Sinkhorn Regularized Unsupervised Learning for Point Cloud Shapes
Lei Yang 0048, Wenxi Liu, Zhiming Cui 0001, Nenglun Chen, Wenping Wang 0001
ECCV (10)2
2020 Learning Resilient Behaviors for Navigation Under Uncertainty
abstract
Deep reinforcement learning has great potential to acquire complex, adaptive behaviors for autonomous agents automatically. However, the underlying neural network polices have not been widely deployed in real-world applications, especially in these safety-critical tasks (e.g., autonomous driving). One of the reasons is that the learned policy cannot perform flexible and resilient behaviors as traditional methods to adapt to diverse environments. In this paper, we consider the problem that a mobile robot learns adaptive and resilient behaviors for navigating in unseen uncertain environments while avoiding collisions. We present a novel approach for uncertainty-aware navigation by introducing an uncertainty-aware predictor to model the environmental uncertainty, and we propose a novel uncertainty-aware navigation network to learn resilient behaviors in the prior unknown environments. To train the proposed uncertainty-aware network more stably and efficiently, we present the temperature decay training paradigm, which balances exploration and exploitation during the training process. Our experimental evaluation demonstrates that our approach can learn resilient behaviors in diverse environments and generate adaptive trajectories according to environmental uncertainties.
Tingxiang Fan, Pinxin Long, Wenxi Liu, Jia Pan 0001, Ruigang Yang, Dinesh Manocha
ICRA3
2020 Recurrent Enhancement of Visual Comfort for Casual Stereoscopic Photography
abstract
Creating stereoscopic 3D media content has wide applications in virtual reality. In this paper, we are interested in a challenging application, casual stereoscopic photography, that allows ordinary users to create a stereoscopic photo using two images captured by a hand-held monocular camera. To handle the geometric constraints and disparity adjustment for casually captured left and right images, we present a coarse-to-fine framework. In the coarse stage, we propose a unified reinforcement learning-based method, in which the produced stereo image is iteratively adjusted and evaluated in the term of visual comfort. In addition, to further enhance the visual comfort of the stereoscopic image produced in the coarse stage, we introduce another independent recurrent network to fine-tune its disparity range. Lastly, we perform comprehensive experiments to evaluate our method and demonstrate the applicability of our model for real images.
Yuzhen Niu, Qingyang Zheng, Wenxi Liu, Wenzhong Guo
VR3
2020 Generating Stereoscopic Images With Convergence Control Ability From a Light Field Image Pair
abstract
With the advances in commercial light field cameras, light field image processing has attracted considerable attention from researchers. In this paper, we propose a novel method for generating stereoscopic images from a light field image pair with flexible control over the convergence of the virtual stereo cameras. We have developed a light field image-capturing prototype that consists of two horizontally arranged light field cameras (i.e., with their optical axes being parallel to each other). When using our proposed device for image/video capture, stereo photographers can concentrate on how to capture the desired visual experience without being frequently disturbed by having to manipulate the stereo camera parameters, i.e., the convergence angle of a stereo camera. During postprocessing, our method estimates accurate disparity maps for the light field image pair and then generates the target stereoscopic images that satisfy the desired stereo camera convergence requirements by adopting a novel view synthesis method for light field images. We have conducted extensive experiments to demonstrate the effectiveness of our proposed method.
Tao Yan 0001, Yiming Mao 0004, Wenxi Liu, Xiaohua Qian, Rynson W. H. Lau
IEEE Trans. Circuits Syst. Video Technol.4
2020 Crowd Counting Via Cross-Stage Refinement Networks
abstract
Crowd counting is challenging due to unconstrained imaging factors, e.g., background clutters, non-uniform distribution of people, large scale and perspective variations. Dealing with these problems using deep neural networks requires rich prior knowledge and multi-scale contextual representations. In this paper, we propose a Cross-stage Refinement Network (CRNet) that can refine predicted density maps progressively based on hierarchical multi-level density priors. In particular, CRNet is composed of several fully convolutional networks. They are stacked together recursively with the previous output as the next input, and each of them serves to utilize previous density output to gradually correct prediction errors of crowd areas and refine the predicted density maps at different stages. Cross-stage multi-level density priors are further exploited in our recurrent framework by the cross-stage skip layers based on ConvLSTM. To cope with different challenges of unconstrained crowd scenes, we explore different crowd-specific data augmentation methods to mimic real-world scenarios and enrich crowd feature representations from different aspects. Extensive experiments show the proposed method achieves superior performances against state-of-the-art methods on four widely-used challenging benchmarks in terms of counting accuracy and density map quality. Code and models are available at this https://github.com/lytgftyf/Crowd-Counting-via-Cross-stage-Refinement-Networks.
Yongtuo Liu, Haoxin Chen, Wenxi Liu, Harry Qin, Guoqiang Han 0002, Shengfeng He
IEEE Trans. Image Process.4
2020 Boundary-Aware RGBD Salient Object Detection With Cross-Modal Feature Sampling
abstract
Mobile devices usually mount a depth sensor to resolve ill-posed problems, like salient object detection on cluttered background. The main barrier of exploring RGBD data is to handle the information from two different modalities. To cope with this problem, in this paper, we propose a boundary-aware cross-modal fusion network for RGBD salient object detection. In particular, to enhance the fusion of color and depth features, we present a cross-modal feature sampling module to balance the contribution of the RGB and depth features based on the statistics of their channel values. In addition, in our multi-scale dense fusion network architecture, we not only incorporate edge-sensitive losses to preserve the boundary of the detected salient region, but also refine its structure by merging the estimated saliency maps of different scales. We accomplish the multi-scale saliency map merging using two alternative methods which produce refined saliency maps via per-pixel weighted combination and an encoder-decoder network. Extensive experimental evaluations demonstrate that our proposed framework can achieve the state-of-the-art performance on several public RGBD-based datasets.
Yuzhen Niu, Guanchao Long, Wenxi Liu, Wenzhong Guo, Shengfeng He
IEEE Trans. Image Process.3
2020 Real-Time Hierarchical Supervoxel Segmentation via a Minimum Spanning Tree
abstract
Supervoxel segmentation algorithm has been applied as a preprocessing step for many vision tasks. However, existing supervoxel segmentation algorithms cannot generate hierarchical supervoxel segmentation well preserving the spatiotemporal boundaries in real time, which prevents the downstream applications from accurate and efficient processing. In this paper, we propose a real-time hierarchical supervoxel segmentation algorithm based on the minimum spanning tree (MST), which achieves state-of-the-art accuracy meanwhile at least 11× faster than existing methods. In particular, we present a dynamic graph updating operation into the iterative construction process of the MST, which can geometrically decrease the numbers of vertices and edges. In this way, the proposed method is able to generate arbitrary scales of supervoxels on the fly. We prove the efficiency of our algorithm that can produce hierarchical supervoxels in the time complexity of O(n) , where n denotes the number of voxels in the input video. Quantitative and qualitative evaluations on public benchmarks demonstrate that our proposed algorithm significantly outperforms the state-of-the-art algorithms in terms of supervoxel segmentation accuracy and computational efficiency. Furthermore, we demonstrate the effectiveness of the proposed method on a downstream application of video object segmentation.
Bo Wang 0057, Yiliang Chen, Wenxi Liu, Harry Qin, Yong Du 0003, Guoqiang Han 0002, Shengfeng He
IEEE Trans. Image Process.3
2020 Learning Long-Term Structural Dependencies for Video Salient Object Detection
abstract
Existing video salient object detection (VSOD) methods focus on exploring either short-term or long-term temporal information. However, temporal information is exploited in a global frame-level or regular grid structure, neglecting interframe structural dependencies. In this paper, we propose to learn long-term structural dependencies with a structure-evolving graph convolutional network (GCN). Particularly, we construct a graph for the entire video using a fast supervoxel segmentation method, in which each node is connected according to spatio-temporal structural similarity. We infer the inter-frame structural dependencies of salient object using convolutional operations on the graph. To prune redundant connections in the graph and better adapt to the moving salient object, we present an adaptive graph pooling to evolve the structure of the graph by dynamically merging similar nodes, learning better hierarchical representations of the graph. Experiments on six public datasets show that our method outperforms all other state-of-the-art methods. Furthermore, We also demonstrate that our proposed adaptive graph pooling can effectively improve the supervoxel algorithm in the term of segmentation accuracy.
Bo Wang 0057, Wenxi Liu, Guoqiang Han 0002, Shengfeng He
IEEE Trans. Image Process.2
2020 Stereoscopic Image Generation From Light Field With Disparity Scaling and Super-Resolution
abstract
In this paper, we propose a novel method to generate stereoscopic images from light-field images with the intended depth range and simultaneously perform image super-resolution. Subject to the small baseline of neighboring subaperture views and low spatial resolution of light-field images captured using compact commercial light-field cameras, the disparity range of any two subaperture views is usually very small. We propose a method to control the disparity range of the target stereoscopic images with linear or nonlinear disparity scaling and properly resolve the disocclusion problem with the aid of a smooth energy term previously used for texture synthesis. The left and right views of the target stereoscopic image are simultaneously generated by a unified optimization framework, which preserves content coherence between the left and right views by a coherence energy term. The disparity range of the target stereoscopic image can be larger than that of the input light field image. This benefits many light field image-based applications, e.g., displaying light field images on various stereo display devices and generating stereoscopic panoramic images from a light field image montage. An extensive experimental evaluation demonstrates the effectiveness of our method.
Tao Yan 0001, Jianbo Jiao, Wenxi Liu, Rynson W. H. Lau
IEEE Trans. Image Process.3
2019 Context-Aware Spatio-Recurrent Curvilinear Structure Segmentation
abstract
Curvilinear structures are frequently observed in various images in different forms, such as blood vessels or neuronal boundaries in biomedical images. In this paper, we propose a novel curvilinear structure segmentation approach using context-aware spatio-recurrent networks. Instead of directly segmenting the whole image or densely segmenting fixed-sized local patches, our method recurrently samples patches with varied scales from the target image with learned policy and processes them locally, which is similar to the behavior of changing retinal fixations in the human visual system and it is beneficial for capturing the multi-scale or hierarchical modality of the complex curvilinear structures. In specific, the policy of choosing local patches is attentively learned based on the contextual information of the image and the historical sampling experience. In this way, with more patches sampled and refined, the segmentation of the whole image can be progressively improved. To validate our approach, comparison experiments on different types of image data are conducted and the sampling procedures for exemplar images are illustrated. We demonstrate that our method achieves the state-of-the-art performance in public datasets.
Feigege Wang, Wenxi Liu, Yuanlong Yu 0001, Shengfeng He, Jia Pan 0001
CVPR3
2019 Single Image Reflection Removal Beyond Linearity
abstract
Due to the lack of paired data, the training of image reflection removal relies heavily on synthesizing reflection images. However, existing methods model reflection as a linear combination model, which cannot fully simulate the real-world scenarios. In this paper, we inject non-linearity into reflection removal from two aspects. First, instead of synthesizing reflection with a fixed combination factor or kernel, we propose to synthesize reflection images by predicting a non-linear alpha blending mask. This enables a free combination of different blurry kernels, leading to a controllable and diverse reflection synthesis. Second, we design a cascaded network for reflection removal with three tasks: predicting the transmission layer, reflection layer, and the non-linear alpha blending mask. The former two tasks are the fundamental outputs, while the latter one being the side output of the network. This side output, on the other hand, making the training a closed loop, so that the separated transmission and reflection layers can be recombined together for training with a reconstruction loss. Extensive quantitative and qualitative experiments demonstrate the proposed synthesis and removal approaches outperforms state-of-the-art methods on two standard benchmarks, as well as in real-world scenarios.
Yinjie Tan, Harry Qin, Wenxi Liu, Guoqiang Han 0002, Shengfeng He
CVPR4
2019 Visualizing the Invisible: Occluded Vehicle Segmentation and Recovery
abstract
In this paper, we propose a novel iterative multi-task framework to complete the segmentation mask of an occluded vehicle and recover the appearance of its invisible parts. In particular, firstly, to improve the quality of the segmentation completion, we present two coupled discriminators that introduce an auxiliary 3D model pool for sampling authentic silhouettes as adversarial samples. In addition, we propose a two-path structure with a shared network to enhance the appearance recovery capability. By iteratively performing the segmentation completion and the appearance recovery, the results will be progressively refined. To evaluate our method, we present a dataset, Occluded Vehicle dataset, containing synthetic and real-world occluded vehicle images. Based on this dataset, we conduct comparison experiments and demonstrate that our model outperforms the state-of-the-arts in both tasks of recovering segmentation mask and appearance for occluded vehicles. Moreover, we also demonstrate that our appearance recovery approach can benefit the occluded vehicle tracking in real-world videos.
Xiaosheng Yan, Yuanlong Yu 0001, Feigege Wang, Wenxi Liu, Shengfeng He, Jia Pan 0001
ICCV4
2019 Deep compression of probabilistic graphical networks
Chun-Yang Zhang, C. L. Philip Chen, Wenxi Liu
Pattern Recognit.4
2019 Extraversion Measure for Crowd Trajectories
abstract
In this paper, we propose an approach to estimate and quantify the degree of extraversion for crowd motion based on individual trajectories. Extraversion is a typical personality that is often observed in human behaviors. We present a composite motion descriptor, which integrates the basic motion information and social metrics, to describe the extraversion of each individual in a crowd. In order to train a universal scoring function that can measure the degrees of extraversion, we incorporate the active learning technique with the relative attribute approach based on the social grouping behavior in crowd motions. In addition, we demonstrate the performance of the proposed method by measuring the degree of extraversion for real individual trajectories in a crowd and analyzing crowd scenes from a real-world dataset.
Wenxi Liu, Chun-Yang Zhang, Genggeng Liu, Yaru Su, Naixue Xiong
IEEE Trans. Ind. Informatics1
2019 Deformable Object Tracking With Gated Fusion
abstract
The tracking-by-detection framework receives growing attention through the integration with the convolutional neural networks (CNNs). Existing tracking-by-detection-based methods, however, fail to track objects with severe appearance variations. This is because the traditional convolutional operation is performed on fixed grids, and thus may not be able to find the correct response while the object is changing pose or under varying environmental conditions. In this paper, we propose a deformable convolution layer to enrich the target appearance representations in the tracking-by-detection framework. We aim to capture the target appearance variations via deformable convolution, which adaptively enhances its original features. In addition, we also propose a gated fusion scheme to control how the variations captured by the deformable convolution affect the original appearance. The enriched feature representation through deformable convolution facilitates the discrimination of the CNN classifier on the target object and background. The extensive experiments on the standard benchmarks show that the proposed tracker performs favorably against the state-of-the-art methods.
Wenxi Liu, Yibing Song, Dengsheng Chen, Shengfeng He, Yuanlong Yu 0001, Tao Yan 0001, Gerhard P. Hancke 0002, Rynson W. H. Lau
IEEE Trans. Image Process.1
2018 Towards Optimally Decentralized Multi-Robot Collision Avoidance via Deep Reinforcement Learning
abstract
Developing a safe and efficient collision avoidance policy for multiple robots is challenging in the decentralized scenarios where each robot generates its paths without observing other robots' states and intents. While other distributed multi-robot collision avoidance systems exist, they often require extracting agent-level features to plan a local collision-free action, which can be computationally prohibitive and not robust. More importantly, in practice the performance of these methods are much lower than their centralized counterparts. We present a decentralized sensor-level collision avoidance policy for multi-robot systems, which directly maps raw sensor measurements to an agent's steering commands in terms of movement velocity. As a first step toward reducing the performance gap between decentralized and centralized methods, we present a multi-scenario multi-stage training framework to learn an optimal policy. The policy is trained over a large number of robots on rich, complex environments simultaneously using a policy gradient based reinforcement learning algorithm. We validate the learned sensor-level collision avoidance policy in a variety of simulated scenarios with thorough performance evaluations and show that the final learned policy is able to find time efficient, collision-free paths for a large-scale robot system. We also demonstrate that the learned policy can be well generalized to new scenarios that do not appear in the entire training period, including navigating a heterogeneous group of robots and a large-scale scenario with 100 robots. Videos are available at https://sites.google.com/view/drlmaca.
Pinxin Long, Tingxiang Fan, Xinyi Liao, Wenxi Liu, Hao Zhang 0170, Jia Pan 0001
ICRA4
2016 Robust individual and holistic features for crowd scene classification
Wenxi Liu, Rynson W. H. Lau, Dinesh Manocha
Pattern Recognit.1
2016 Exemplar-AMMs: Recognizing Crowd Movements From Pedestrian Trajectories
abstract
In this paper, we present a novel method to recognize the types of crowd movement from crowd trajectories using agent-based motion models (AMMs). Our idea is to apply a number of AMMs, referred to as exemplar-AMMs, to describe the crowd movement. Specifically, we propose an optimization framework that filters out the unknown noise in the crowd trajectories and measures their similarity to the exemplar-AMMs to produce a crowd motion feature. We then address our real-world crowd movement recognition problem as a multilabel classification problem. Our experiments show that the proposed feature outperforms the state-of-the-art methods in recognizing both simulated and real-world crowd movements from their trajectories. Finally, we have created a synthetic dataset, SynCrowd, which contains two-dimensional (2D) crowd trajectories in various scenarios, generated by various crowd simulators. This dataset can serve as a training set or benchmark for crowd analysis work.
Wenxi Liu, Rynson W. H. Lau, Xiaogang Wang 0001, Dinesh Manocha
IEEE Trans. Multim.1
2015 SuperCNN: A Superpixelwise Convolutional Neural Network for Salient Object Detection
Shengfeng He, Rynson W. H. Lau, Wenxi Liu, Zhe Huang 0004, Qingxiong Yang
Int. J. Comput. Vis.3
2015 Leveraging Long-Term Predictions and Online Learning in Agent-Based Multiple Person Tracking
abstract
We present a multiple-person tracking algorithm, based on combining particle filters (PFs) and reciprocal velocity obstacle (RVO), an agent-based crowd model that infers collision-free velocities so as to predict a pedestrian's motion. In addition to position and velocity, our tracking algorithm can estimate the internal goals (desired destination or desired velocity) of the tracked pedestrian in an online manner, thus removing the need to specify this information beforehand. Furthermore, we leverage the longer term predictions of RVO by deriving a higher order PF, which aggregates multiple predictions from different prior time steps. This yields a tracker that can recover from short-term occlusions and spurious noise in the appearance model. Experimental results show that our tracking algorithm is suitable for predicting pedestrians' behaviors online without needing scene priors or hand-annotated goal information, and improves tracking in real-world crowded scenes under low frame rates.
Wenxi Liu, Antoni B. Chan, Rynson W. H. Lau, Dinesh Manocha
IEEE Trans. Circuits Syst. Video Technol.1
2014 Data-driven sequential goal selection model for multi-agent simulation
abstract
With recent advances in distributed virtual worlds, online users have access to larger and more immersive virtual environments. Sometimes the number of users in virtual worlds is not large enough to make the virtual world realistic. In our paper, we present a crowd simulation algorithm that allows a large number of virtual agents to navigate around the virtual world autonomously by sequentially selecting the goals. Our approach is based on our sequential goal selection model (SGS) which can learn goal-selection patterns from synthetic sequences. We demonstrate our algorithm's simulation results in complex scenarios containing more than 20 goals.
Wenxi Liu, Zhe Huang 0004, Rynson W. H. Lau, Dinesh Manocha
VRST1
2012 Crowd simulation using Discrete Choice Model
abstract
We present a new algorithm to simulate a variety of crowd behaviors using the Discrete Choice Model (DCM). DCM has been widely studied in econometrics to examine and predict customers' or households' choices. Our DCM formulation can simulate virtual agents' goal selection and we highlight our algorithm by simulating heterogeneous crowd behaviors: evacuation, shopping, and rioting scenarios.
Wenxi Liu, Rynson W. H. Lau, Dinesh Manocha
VR1
2012 Predicting Pedestrian Trajectories Using Velocity-Space Reasoning
Sujeong Kim, Stephen J. Guy, Wenxi Liu, Rynson W. H. Lau, Ming C. Lin, Dinesh Manocha
WAFR3
2012 A statistical similarity measure for aggregate crowd dynamics
abstract
We present an information-theoretic method to measure the similarity between a given set of observed, real-world data and visual simulation technique for aggregate crowd motions of a complex system consisting of many individual agents. This metric uses a two-step process to quantify a simulator's ability to reproduce the collective behaviors of the whole system, as observed in the recorded real-world data. First, Bayesian inference is used to estimate the simulation states which best correspond to the observed data, then a maximum likelihood estimator is used to approximate the prediction errors. This process is iterated using the EM-algorithm to produce a robust, statistical estimate of the magnitude of the prediction error as measured by its entropy (smaller is better). This metric serves as a simulator-to-data similarity measurement. We evaluated the metric in terms of robustness to sensor noise, consistency across different datasets and simulation methods, and correlation to perceptual metrics.
Stephen J. Guy, Jur P. van den Berg, Wenxi Liu, Rynson W. H. Lau, Ming C. Lin, Dinesh Manocha
ACM Trans. Graph.3
2010 Combined X-ray and facial videos for phoneme-level articulator dynamics
Hui Chen 0020, Wenxi Liu, Pheng-Ann Heng
Vis. Comput.3