EDBT 2026 Demo / reviewers in the wild / expert
Ping Wei 0001
dblp:49/6362-1
· DBLP profile ↗
72ranked-venue papers
10as first author
47since 2021 · last 2026
0000-0002-8535-9527ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 54 · 8 first-author · 33 since 2021Graphics, computer vision, multimedia, augmented reality and games · 44 · 8 first-author · 27 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hierarchical motion control framework based on SLM-guided and learning-enhanced NMPC for autonomous underwater vehicles
Zhiteng Zhang, Meiqin Liu 0001, Ronghao Zheng, Ping Wei 0001 |
Expert Syst. Appl. | 5 |
| 2026 | IntentQA: Intent Question Answering in Videos by Cognitive Context ReasoningabstractVideo understanding requires intelligent agents to transcend mere recognition of visual facts and comprehend the underlying intents behind human actions-often termed the "dark matter" of social intelligence. To bridge the gap between visual observation and intent reasoning, we introduce a novel task, IntentQA, and contribute a large-scale VideoQA dataset specifically tailored for this purpose. However, recognizing that standard metrics may overestimate capabilities due to dataset biases, we go beyond simple accuracy to rigorously evaluate model robustness. We augment the benchmark by generating five distinct contrast sets via Large Language Models (LLMs) and introducing a "Contrast Performance Decline" metric. We propose the X-CaVIR(eXplainable Context-aware Video Intent Reasoning) framework, which leverages three types of "Cognitive Context" to enhance video analysis: i) Situational Context via a cross-modal Video Query Language (VQL) module, ii) Contrastive Context via a Contrastive Learning module, and iii) Commonsense Context via a Commonsense Reasoning module. Crucially, to overcome the lack of transparency in traditional models, we refine the integration of LLMs within X-CaVIR by employing a transparent pipeline that synergizes video captions with VQA model outputs. This approach not only improves performance by effectively utilizing rich commonsense knowledge but also renders the reasoning process explicitly interpretable. Extensive experiments demonstrate the effectiveness of our components, the superiority of X-CaVIR over state-of-the-art baselines, and its stability against perturbations on the contrast sets. Jiapeng Li 0003, Ping Wei 0001, Wenjuan Han, Song-Chun Zhu, Lifeng Fan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Learning unified patterns of multimodalities for video temporal grounding
Ping Wei 0001 |
Pattern Recognit. | 2 |
| 2025 | OTIAS: OcTree Implicit Adaptive Sampling for Multispectral and Hyperspectral Image FusionabstractImplicit Neural Representation (INR) methods have demonstrated great potential in arbitrary-scale super-resolution tasks. This success is primarily due to their ability to continuously represent images using coordinates. In the task of remote sensing image fusion, INR methods have also shown promising applications. However, the previous INR methods neglect channel-wise modeling, while sharing a single kernel across all channels at each position, resulting in a lack of sensitivity to data specificity. To address these issues, we propose the OcTree Implicit Adaptive Sampling (OTIAS) method, which innovatively applies the octree structure to restore data from both horizontal and vertical directions, effectively incorporating spatial and spectral information from hyperspectral data. Additionally, we introduce a novel method to adaptively generate interpolation kernels based on coordinates. This approach efficiently produces customized interpolation kernel parameters for octree nodes, tailored to different spectral information. Overall, our method achieves state-of-the-art performance on the CAVE and Harvard datasets with 4× and 8× scaling factors, outperforming existing approaches. Shangqi Deng, Liang-Jian Deng, Ping Wei 0001 |
AAAI | 4 |
| 2025 | Unveiling Multi-View Anomaly Detection: Intra-view Decoupling and Inter-view FusionabstractAnomaly detection has garnered significant attention for its extensive industrial application value. Most existing methods focus on single-view scenarios and fail to detect anomalies hidden in blind spots, leaving a gap in addressing the demands of multi-view detection in practical applications. Ensemble of multiple single-view models is a typical way to tackle the multi-view situation, but it overlooks the correlations between different views. In this paper, we propose a novel multi-view anomaly detection framework, Intra-view Decoupling and Inter-view Fusion (IDIF), to explore correlations among views. Our method contains three key components: 1) a proposed Consistency Bottleneck module extracting the common features of different views through information compression and mutual information maximization; 2) an Implicit Voxel Construction module fusing features of different views with prior knowledge represented in the form of voxels; and 3) a View-wise Dropout training strategy enabling the model to learn how to cope with missing views during test. The proposed IDIF achieves state-of-the-art performance on three datasets. Extensive ablation studies also demonstrate the superiority of our methods. Yiyang Lian, Meiqin Liu 0001, Nanning Zheng 0001, Ping Wei 0001 |
AAAI | 6 |
| 2025 | STCOcc: Sparse Spatial-Temporal Cascade Renovation for 3D Occupancy and Scene Flow Predictionabstract3D occupancy and scene flow offer a detailed and dynamic representation of 3D scene. Recognizing the sparsity and complexity of 3D space, previous vision-centric methods have employed implicit learning-based approaches to model spatial and temporal information. However, these approaches struggle to capture local details and diminish the model’s spatial discriminative ability. To address these challenges, we propose a novel explicit state-based modeling method designed to leverage the occupied state to renovate the 3D features. Specifically, we propose a sparse occlusion-aware attention mechanism, integrated with a cascade refinement strategy, which accurately renovates 3D features with the guidance of occupied state information. Additionally, we introduce a novel method for modeling long-term dynamic interactions, which reduces computational costs and preserves spatial information. Compared to the previous state-of-the-art methods, our efficient explicit renovation strategy not only delivers superior performance in terms of RayloU and mAVE for occupancy and scene flow prediction but also markedly reduces GPU memory usage during training, bringing it down to 8.7GB. Our code is available on https://github.com/lzzzzzm/STCOcc Zhimin Liao, Ping Wei 0001, Shuaijia Chen, Ziyang Ren |
CVPR | 2 |
| 2025 | Beyond Single-Modal Boundary: Cross-Modal Anomaly Detection through Visual Prototype and HarmonizationabstractAnomaly detection is a significant task for its application and research value. While existing methods have made impressive progress within the same modality, cross-modal anomaly detection remains an open and challenging problem. In this paper, we propose a cross-modal anomaly detection model that is trained using data from a variety of existing modalities and can be generalized well to unseen modalities. The model consists of three major components: 1) the Transferable Visual Prototype directly learns normal/abnormal semantics in visual space; 2) the Prototype Harmonization strategy adaptively utilizes the Transferable Visual Prototypes from various modalities for inference on the unknown modality; 3) the Visual Discrepancy Inference under the few-shot setting enhances performance. In the zero-shot setting, the proposed method achieves AUROC improvements of 4.1%, 6.1%, 7.6%, and 6.8% over the best competing methods in the RGB, 3D, MRI/CT, and Thermal modalities, respectively. In the few-shot setting, our model also achieves the highest AUROC/AP on ten datasets in four modalities, substantially outperforming existing methods. Codes are available at https://github.com/Kerio99/CMAD. Ping Wei 0001, Yiyang Lian, Nanning Zheng 0001 |
CVPR | 2 |
| 2025 | Stochastic-Aware Mamba Diffusion for Pedestrian Trajectory PredictionabstractPedestrian trajectory prediction plays a crucial role in understanding human behavior and intentions. Due to the inherent randomness in human movement, current research constructs trajectories in stochastic space and uses diffusion models to reverse the denoising process. The commonly used denoising model, Transformer, is affected by uneven noise sampling, leading to confusion between temporal consistency and randomness across multiple time points. To address this issue, we propose a novel framework named Stochastic-Aware Mamba Diffusion (SAMD), which combines Stochastic-Aware Mamba with the Diffusion model to predict motion noise and motion states in the stochastic space. It utilizes a temporal aggregator to extract temporal consistency features. We construct a motion-selective state space model, which includes the adaptive transition between consistency motion states and stochastic states for balancing stability and diversity. The motion gating unit activates pertinent information within motion states to extract motion noise for trajectory prediction. Our approach achieves state-of-the-art results on the ETH-UCY and SDD datasets while significantly reducing the computational cost. Ziyang Ren, Ping Wei 0001, Haowen Tang, Jialu Qin |
ICASSP | 2 |
| 2025 | $I^{\mathbf{2}}$-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting
Zhimin Liao, Ping Wei 0001, Shuaijia Chen, Ziyang Ren |
ICCV | 2 |
| 2025 | TOTP: Transferable Online Pedestrian Trajectory Prediction with Temporal-Adaptive Mamba Latent Diffusion
Ziyang Ren, Ping Wei 0001, Shangqi Deng, Haowen Tang, Jiapeng Li 0003 |
ICCV | 2 |
| 2025 | Alchemy: Amplifying Theorem-Proving Capability Through Symbolic MutationabstractFormal proofs are challenging to write even for experienced experts. Recent progress in Neural Theorem Proving (NTP) shows promise in expediting this process. However, the formal corpora available on the Internet are limited compared to the general text, posing a significant data scarcity challenge for NTP. To address this issue, this work proposes Alchemy, a general framework for data synthesis that constructs formal theorems through symbolic mutation. Specifically, for each candidate theorem in Mathlib, we identify all invocable theorems that can be used to rewrite or apply to it. Subsequently, we mutate the candidate theorem by replacing the corresponding term in the statement with its equivalent form or antecedent. As a result, our method increases the number of theorems in Mathlib by an order of magnitude, from 110k to 6M. Furthermore, we perform continual pretraining and supervised finetuning on this augmented corpus for large language models. Experimental results demonstrate the effectiveness of our approach, achieving a 4.70% absolute performance improvement on Leandojo benchmark. Additionally, our approach achieves a 2.47% absolute performance gain on the out-of-distribution miniF2F benchmark based on the synthetic data. To provide further insights, we conduct a comprehensive analysis of synthetic data composition and the training paradigm, offering valuable guidance for developing a strong theorem prover. Shaonan Wu, Yeyun Gong, Nan Duan 0001, Ping Wei 0001 |
ICLR | 5 |
| 2025 | Achieving Seamless Camouflage: Attention Fusion Diffusion Model for Image SynthesisabstractCamouflage image generation plays a vital role in various research fields. Current methods typically rely on manually selecting and blending objects with backgrounds, which often produce incongruous combinations where the object does not seamlessly integrate with the background, leading to unrealistic and unnatural visual outcomes. To address these challenges, we introduce a novel Attention Fusion Diffusion Model (AFDM) designed to generate realistic camouflage images from a single input image containing an object and its surrounding background. The AFDM framework is comprised of two essential components: an Attention Fusion Module, which adeptly integrates object features with surrounding background information to produce convincingly camouflaged objects, and a content guidance strategy designed to mitigate content drift during the fusion process, thereby ensuring that the camouflaged image remains faithfully aligned with the original content. In addition, we build the outdoor Solidier Dataset(OSD) for advancing camouflage target recognization and in-depth research on this topic. Extensive experiments and user studies demonstrate the performance of our method in camouflage image generation and its potential to enhance image segmentation-related fields. Our code and dataset will be available at https://github.com/xhxhzhz/AFDM. Hao Xi, Meiqin Liu 0001, Zechen Yang, Ping Wei 0001 |
ICME | 4 |
| 2025 | Decentralized but Not Compromised: Modular Architecture with Refined Observation for Multi-Agent Model-Based Reinforcement LearningabstractMulti-agent adversarial tasks such as swarm robotics and autonomous vehicle coordination, demand efficient decentralized collaboration under partial observability. While model-free multi-agent RL (MF-MARL) methods suffer from necessitating extensive environment interactions, most existing multi-agent model-based RL (MA-MBRL) methods fail to align with the Centralized Training with Decentralized Execution (CTDE) paradigm, which limits system flexibility. This paper proposes a novel modular architecture with refined observations (MARO) to achieve the CTDE paradigm by decoupling agents from the world model. Key innovations include: 1) an enhanced world model with weighted loss and history-augmented rollout for high-quality data generation; 2) a dual-stream semantic decomposition network (DSDN) that performs fine-grained decomposition of observations to refine action mapping and mitigate performance degradation from information loss. Extensive experiments on the StarCraft Multi-Agent Challenge (SMAC) demonstrate superior performance over opponents, validating the effectiveness and advancement of MARO. Meiqin Liu 0001, Ronghao Zheng, Shanling Dong, Ping Wei 0001 |
IROS | 5 |
| 2025 | Multi-scale Frequency-Space Fusion Camouflaged Object Detection
Linyu Zhang, Ping Wei 0001, Shuaijia Chen |
PRCV (17) | 2 |
| 2025 | Foreground natural anomaly synthesis for attention guided anomaly detection
Xinyuan Xiang, Meiqin Liu 0001, Senlin Zhang, Ping Wei 0001, Badong Chen |
Neurocomputing | 4 |
| 2025 | Semantic Consistency Reasoning for 3-D Object Detection in Point CloudsabstractPoint cloud-based 3-D object detection is a significant and critical issue in numerous applications. While most existing methods attempt to capitalize on the geometric characteristics of point clouds, they neglect the internal semantic properties of point and the consistency between the semantic and geometric clues. We introduce a semantic consistency (SC) mechanism for 3-D object detection in this article, by reasoning about the semantic relations between 3-D object boxes and its internal points. This mechanism is based on a natural principle: the semantic category of a 3-D bounding box should be consistent with the categories of all points within the box. Driven by the SC mechanism, we propose a novel SC network (SCNet) to detect 3-D objects from point clouds. Specifically, the SCNet is composed of a feature extraction module, a detection decision module, and a semantic segmentation module. In inference, the feature extraction and the detection decision modules are used to detect 3-D objects. In training, the semantic segmentation module is jointly trained with the other two modules to produce more robust and applicable model parameters. The performance is greatly boosted through reasoning about the relations between the output 3-D object boxes and segmented points. The proposed SC mechanism is model-agnostic and can be integrated into other base 3-D object detection models. We test the proposed model on three challenging indoor and outdoor benchmark datasets: ScanNetV2, SUN RGB-D, and KITTI. Furthermore, to validate the universality of the SC mechanism, we implement it in three different 3-D object detectors. The experiments show that the performance is impressively improved and the extensive ablation studies also demonstrate the effectiveness of the proposed model. Wenwen Wei, Ping Wei 0001, Zhimin Liao, Jialu Qin, Xiang Cheng 0001, Meiqin Liu 0001, Nanning Zheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Learning Scene-Goal-Aware Motion Representation for Trajectory Prediction
Ziyang Ren, Ping Wei 0001, Haowen Tang |
BMVC | 2 |
| 2024 | Salience DETR: Enhancing Detection Transformer with Hierarchical Salience Filtering RefinementabstractDETR-like methods have significantly increased detection performance in an end-to-end manner. The main-stream two-stage frameworks of them perform dense self-attention and select a fraction of queries for sparse cross-attention, which is proven effective for improving performance but also introduces a heavy computational burden and high dependence on stable query selection. This paper demonstrates that suboptimal two-stage selection strategies result in scale bias and redundancy due to the mismatch between selected queries and objects in two-stage initial-ization. To address these issues, we propose hierarchical salience filtering refinement, which performs transformer encoding only on filtered discriminative queries, for a bet-ter trade-off between computational efficiency and precision. The filtering process overcomes scale bias through a novel scale-independent salience supervision. To com-pensate for the semantic misalignment among queries, we introduce elaborate query refinement modules for stable two-stage initialization. Based on above improvements, the proposed Salience DETR achieves significant improvements of +4.0% AP, +0.2% AP, +4.4% AP on three challenging task-specific detection datasets, as well as 49.2% AP on COCO 2017 with less FLOPs. The code is available at https://github.com/xiuqhou/Salience-DETR. Xiuquan Hou, Meiqin Liu 0001, Senlin Zhang, Ping Wei 0001, Badong Chen |
CVPR | 4 |
| 2024 | Task-Driven Exploration: Decoupling and Inter-Task Feedback for Joint Moment Retrieval and Highlight DetectionabstractVideo moment retrieval and highlight detection are two highly valuable tasks in video understanding, but until recently they have been jointly studied. Although existing studies have made impressive advancement recently, they predominantly follow the data-driven bottom-up paradigm. Such paradigm overlooks task-specific and inter-task effects, resulting in poor model performance. In this paper, we propose a novel task-driven top-down framework TaskWeave for joint moment retrieval and highlight detection. The framework introduces a task-decoupled unit to capture task-specific and common representations. To investigate the interplay between the two tasks, we propose an inter-task feedback mechanism, which transforms the results of one task as guiding masks to assist the other task. Different from existing methods, we present a task-dependent joint loss function to optimize the model. Comprehensive experiments and in-depth ablation studies on QVHighlights, TVSum, and Charades-STA datasets corroborate the effectiveness and flexibility of the proposed framework. Codes are available at github.com/EdenGabriel/TaskWeave. Ping Wei 0001, Ziyang Ren |
CVPR | 2 |
| 2024 | Relation DETR: Exploring Explicit Position Relation Prior for Object Detection
Xiuquan Hou, Meiqin Liu 0001, Senlin Zhang, Ping Wei 0001, Badong Chen, Xuguang Lan |
ECCV (50) | 4 |
| 2024 | Local-to-Global Perception Network for Point Cloud SegmentationabstractLiDAR-based point cloud segmentation is a significant and challenging task for 3D scene understanding. Recent voxel-based methods are often built on submanifold sparse residual calculation with small kernel size, which limits the local feature interactions and neglects the global contextual information. In this paper, we propose a local-to-global perception LiDAR-based point cloud segmentation network LGPSeg. From the local perspective, we design the Dynamic Spatial Aggregation Convolution to expand the receptive field range while avoiding a large increase in the model parameters. From the global perspective, we propose the BEV-Voxel Fusion to aggregate the global contextual information on the BEV feature maps through advanced 2D operators. By combining the local and global features, the 3D object and background information can be better captured. Our method achieves state-of-the-art results on two large datasets, SemanticKITTI and nuScenes, and even outperformed multimodal-based methods. Ping Wei 0001, Shuaijia Chen, Zhimin Liao, Jialu Qin |
ICME | 2 |
| 2024 | TransOSV: Offline Signature Verification with Transformers
Ping Wei 0001, Zeyu Ma 0005, Changkai Li, Nanning Zheng 0001 |
Pattern Recognit. | 2 |
| 2024 | Cross Time-Frequency Transformer for Temporal Action LocalizationabstractMost modern approaches in temporal action localization (TAL) mainly focus on time domain information, while neglecting the advantages of information from other domains. How to effectively utilize information from different domains and their interactions in a reasonable manner has been an attractive yet challenging issue in TAL. In this paper, we propose a novel cross time-frequency Transformer model (TFFormer) for TAL. A dual-branch network architecture is designed to capture the time and frequency features at multiple scales, using the multi-scale transformer in the time branch and the DB1 Discrete Wavelet Transform (DWT) in the frequency branch. To fuse these features from different domains, we propose a cross time-frequency attention mechanism that includes a time pathway and a frequency pathway, enhancing the interaction between the temporal and frequency features. Furthermore, a gated control mechanism is designed to aggregate features from different scales, characterizing the respective contributions of features at different scales. We also design a new regression loss function for locating the time boundaries. Extensive experiments were carried out on four challenging benchmark datasets, including two third-person datasets and two first-person datasets. The proposed method achieves impressive results on these datasets. Specifically, TFFormer achieves an average mAP of 23.2% on Ego4D and 25.6% on EPIC-Kitchens 100, which outperform previous state-of-the-arts by a large margin. It also obtains competitive results on ActivityNet v1.3 and THUMOS14, with an average mAP of 36.2% and 67.8%. We also conducted extensive ablation studies to validate the effectiveness of each component in the proposed method. Ping Wei 0001, Nanning Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Hierarchical Heterogeneous Multi-Agent Cross-Domain Search Method Based on Deep Reinforcement LearningabstractMarine target searching is a complex task due to large search areas, unique signal propagation characteristics, and limited visibility, posing significant challenges for single-agent or homogeneous multi-agent systems. In response, we propose a novel hierarchical heterogeneous multi-agent (HHMA) framework designed for underwater search scenarios. This framework integrates three types of vehicles moving in different domains—unmanned aerial, surface, and underwater vehicles, effectively overcoming the limitations of single or double-agent configurations. We begin by elucidating the advantages of the HHMA system in target searching, providing the kinematic modeling, while also transforming sonar detecting data and defining the search problem. The mission is decomposed to three human-comprehensible subtasks that are adaptive to both environmental conditions and equipment capabilities: moving, target estimating and trajectory planning. The target estimating subtask is effectively modeled as a Markov Decision Process, retaining its memory capability. Additionally, we extend multi-agent reinforcement learning to multi-policy reinforcement learning, facilitating the training of interdependent policies. The efficacy of our approach is demonstrated through simulations, comparing it with rule-based methods. Simulation results underscore the significance of the HHMA system and validate the proposed training methodology. Shangqun Dong, Meiqin Liu 0001, Shanling Dong, Ronghao Zheng, Ping Wei 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | 3D Scene Graph Generation From Point CloudsabstractScene graph generation is a significant and challenging task for scene understanding. Most existing methods are confined to the 2D space (i.e. images) or additional use of segmentation information, while neglecting the richer spatial and geometric information of 3D space. In this paper, we propose a novel method to generate scene graphs from 3D point clouds. Specifically, our model consists of three parts: a point feature extraction backbone, a box head, and a relation head. The feature extraction backbone extracts base features directly from raw point clouds, and the box head produces detected 3D bounding boxes. Final 3D scene graphs are obtained from the relation head which takes the extracted features and 3D boxes as inputs. We also design a point RoI module which sequentially processes points inside 3D boxes with a bidirectional LSTM. To further leverage the geometric characteristics of point clouds, we propose a location attention module which learns the influence of relative locations between objects. We introduce the RelationScanNet dataset with densely annotated semantic and geometric relationships, which extends one of the most widely used dataset ScanNetV2 in 3D indoor scene understanding. We test the proposed method on the RelationScanNet dataset and 3DSSG dataset. The results prove the strength of our method. Wenwen Wei, Ping Wei 0001, Jialu Qin, Zhimin Liao, Shuaijie Wang, Xiang Cheng 0001, Meiqin Liu 0001, Nanning Zheng 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Gated Multi-Scale Transformer for Temporal Action LocalizationabstractTemporal action localization (TAL) is a critical task in video understanding. Effectively utilizing multi-scale information and handling interactions across various scales have consistently posed challenging issues within the realm of TAL. In this paper, we propose a novel gated multi-scale Transformer model (TransGMC) for temporal action localization. A gated control mechanism is designed to filter and aggregate the information at different scales, by which the contributions of contexts at different temporal scales are well characterized. To enhance the feature representation at each temporal scale, the rich global-local contexts are extracted at each temporal scale. A cascade attention module that contains two seamlessly integrated channel attention and moment attention is proposed for capturing global temporal contexts. We utilize a new regression loss function for locating the time boundaries. We conducted experiments on four challenging benchmark datasets, including two third-person view datasets and two first-person view datasets. Our method achieves an average mAP of 67.5% on THUMOS14, 36.1% on ActivityNet v1.3, 24.9% on EPIC-Kitchens 100, and 23.2% on Ego4D, which all outperform the previous state-of-the-arts methods. Extensive ablation studies also validate the effectiveness of the proposed method. Code is available at https://github.com/EdenGabriel/TransGMC. Ping Wei 0001, Ziyang Ren, Nanning Zheng 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Cross-Graph Transformer Network for Temporal Sentence Grounding
Jiahui Shang, Ping Wei 0001, Nanning Zheng 0001 |
ICANN (6) | 2 |
| 2023 | Temporal Deformable Transformer for Action Localization
Haoying Wang, Ping Wei 0001, Meiqin Liu 0001, Nanning Zheng 0001 |
ICANN (6) | 2 |
| 2023 | IntentQA: Context-aware Video Intent ReasoningabstractIn this paper, we propose a novel task IntentQA, a special VideoQA task focusing on video intent reasoning, which has become increasingly important for AI with its advantages in equipping AI agents with the capability of reasoning beyond mere recognition in daily tasks. We also contribute a large-scale VideoQA dataset for this task. We propose a Context-aware Video Intent Reasoning model (CaVIR) consisting of i) Video Query Language (VQL) for better cross-modal representation of the situational context, ii) Contrastive Learning module for utilizing the contrastive context, and iii) Commonsense Reasoning module for incorporating the commonsense context. Comprehensive experiments on this challenging task demonstrate the effectiveness of each model component, the superiority of our full model over other baselines, and the generalizability of our model to a new VideoQA task. The dataset and codes are open-sourced at: https://github.com/JoseponLee/IntentQA.git. Jiapeng Li 0003, Ping Wei 0001, Wenjuan Han, Lifeng Fan |
ICCV | 2 |
| 2023 | Inverse Compositional Learning for Weakly-supervised Relation GroundingabstractVideo relation grounding (VRG) is a significant and challenging problem in the domains of cross-modal learning and video understanding. In this study, we introduce a novel approach called inverse compositional learning (ICL) for weakly-supervised video relation grounding. Our approach represents relations at both the holistic and partial levels, formulating VRG as a joint optimization problem that encompasses reasoning at both levels. For holistic-level reasoning, we propose an inverse attention mechanism and a compositional encoder to generate compositional relevance features. Additionally, we introduce an inverse loss to evaluate and learn the relevance between visual features and relation features. At the partial-level reasoning, we introduce a grounding by classification scheme. By leveraging the learned holistic-level features and partial-level features, we train the entire model in an end-to-end manner. We conduct evaluations on two challenging datasets and demonstrate the substantial superiority of our proposed method over state-of-the-art methods. Extensive ablation studies confirm the effectiveness of our approach. Ping Wei 0001, Zeyu Ma 0005, Nanning Zheng 0001 |
ICCV | 2 |
| 2023 | Multi-scale interaction transformer for temporal action proposal generation
Jiahui Shang, Ping Wei 0001, Nanning Zheng 0001 |
Image Vis. Comput. | 2 |
| 2023 | A graph-based two-stage classification network for mobile screen defect inspectionabstractDefect inspection, also known as defect detection, is significant in mobile screen quality control. There are some challenging issues brought by the characteristics of screen defects, including the following: (1) the problem of interclass similarity and intraclass variation, (2) the difficulty in distinguishing low contrast, tiny-sized, or incomplete defects, and (3) the modeling of category dependencies for multi-label images. To solve these problems, a graph reasoning module, stacked on a classification module, is proposed to expand the feature dimension and improve low-quality image features by exploiting category-wise dependency, image-wise relations, and interactions between them. To further improve the classification performance, the classifier of the classification module is redesigned as a cosine similarity function. With the help of contrastive learning, the classification module can better initialize the category-wise graph of the reasoning module. Experiments on the mobile screen defect dataset show that our two-stage network achieves the following best performances: 97.7% accuracy and 97.3% F -measure. This proves that the proposed approach is effective in industrial applications. Chaofan Zhou, Meiqin Liu 0001, Senlin Zhang, Ping Wei 0001, Badong Chen |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2023 | CANet: Contextual Information and Spatial Attention Based Network for Detecting Small Defects in Manufacturing Industry
Xiuquan Hou, Meiqin Liu 0001, Senlin Zhang, Ping Wei 0001, Badong Chen |
Pattern Recognit. | 4 |
| 2023 | Multi-scale attention and dilation network for small defect detection
Xinyuan Xiang, Meiqin Liu 0001, Senlin Zhang, Ping Wei 0001, Badong Chen |
Pattern Recognit. Lett. | 4 |
| 2022 | Asymmetric Relation Consistency Reasoning for Video Relation Grounding
Ping Wei 0001, Jiapeng Li 0003, Zeyu Ma 0005, Jiahui Shang, Nanning Zheng 0001 |
ECCV (35) | 2 |
| 2022 | Offline Signature Verification with TransformersabstractSignature verification is a frequently-used forensics technology. Although the previous convolution neural network (CNN) based methods have made a great progress, the limitation of local neighborhood operation of CNN impedes reasoning about the relation of global signature strokes. To overcome this weakness, in this paper, we propose a novel holistic-part unified model named TransOSV based on the transformer framework. Signature images are encoded into patch sequences by the proposed holistic encoder to learn global representation. Considering the subtle local difference between the genuine signature and forged signature, we design a contrast based part decoder that is utilized to learn discriminative part features. To reduce the influence of sample imbalance, we formulate a new focal contrast loss function. Extensive experimental results and ablation studies prove the potential of the proposed model. Ping Wei 0001, Zeyu Ma 0005, Changkai Li, Nanning Zheng 0001 |
ICME | 2 |
| 2022 | HOIG: End-to-End Human-Object Interactions Grounding with TransformersabstractVisual grounding is a crucial and challenging problem in many applications. While it has been extensively investigated over the past years, human-centric grounding with multiple instances is still an open problem. In this paper, we introduce a new task of Human-Object Interactions (HOI) Grounding to localize all the referring human-object pair instances in an image with a given ⟨human, interaction, object⟩ phrase. We design an encoder-decoder architecture to model the task as a set prediction problem based on transformers. A vision-language alignment module and a grounding decoder are designed to learn accurate cross-modal contexts and interactions. Our model accomplishes alignment and prediction in an end-to-end manner without pre-trained detectors or post-processing. Experiments on two challenging datasets prove the strength of our model. Zeyu Ma 0005, Ping Wei 0001, Nanning Zheng 0001 |
ICME | 2 |
| 2022 | Relation Reasoning for Video Pedestrian Trajectory PredictionabstractPedestrian trajectory prediction is a challenging and important task in many applications, which aims to predict future pedestrians' trajectory coordinates from the input historical data. The existing methods usually use ready-made trajectory coordinates as inputs, which is, however, unavailable in video-based scenarios. In this paper, we propose a relation reasoning hypergraph (RRH) model to directly predict multiple pedestrian trajectories from raw videos. It is a challenging issue for the input and output are in different modalities and a video may contain multiple pedestrians. Our model integrates historical trajectory tracking, pedestrian relation reasoning, and future trajectory prediction into one framework. For capturing the subtle social relationships among pedestrians, we design a relation reasoning hypergraph network. We tested the proposed method on two public pedestrians datasets and the performance demonstrates the power of the model. Haowen Tang, Ping Wei 0001, Jiapeng Li 0003, Nanning Zheng 0001 |
ICME | 2 |
| 2022 | EvoSTGAT: Evolving spatiotemporal graph attention networks for pedestrian trajectory prediction
Haowen Tang, Ping Wei 0001, Jiapeng Li 0003, Nanning Zheng 0001 |
Neurocomputing | 2 |
| 2022 | AVN: An Adversarial Variation Network Model for Handwritten Signature VerificationabstractHandwritten signature verification is a crucial yet challenging problem. While previous studies have made great progress in this problem, they learn signature features passively from given existing data. In this paper, we propose a novel adversarial variation network (AVN) model for handwritten signature verification which mines effective features by actively varying existing data and generating new data. Powered by a proposed novel variation consistency mechanism, the AVN contains three different types of modules unified under one end-to-end framework: the extractor seeks to extract deep discriminative features of handwritten signatures, the discriminator aims to make verification decisions based on the extracted features, and the variator is designed to actively generate signature variants for constructing a more discriminative model. The proposed model is trained in an adversarial way with a min-max loss function, by which the three modules cooperate and compete to enhance the entire model’s ability and therefore the signature verification performance is improved. We test the proposed method on four challenging signature datasets of different languages: CEDAR, BHSig-Hindi, BHSig-Bengali, and GPDS Synthetic Signature. Extensive experiments with in-depth discussions validate the effectiveness of the proposed method. Ping Wei 0001 |
IEEE Trans. Multim. | 2 |
| 2021 | Static-Dynamic Interaction Networks for Offline Signature VerificationabstractOffline signature verification is a challenging issue that is widely used in various fields. Previous approaches model this task as a static feature matching or distance metric problem of two images. In this paper, we propose a novel Static-Dynamic Interaction Network (SDINet) model which introduces sequential representation into static signature images. A static signature image is converted to sequences by assuming pseudo dynamic processes in the static image. A static representation extracting deep features from signature images describes the global information of signatures. A dynamic representation extracting sequential features with LSTM networks characterizes the local information of signatures. A dynamic-to-static attention is learned from the sequences to refine the static features. Through the static-to-dynamic conversion and the dynamic-to-static attention, the static representation and dynamic representation are unified into a compact framework. The proposed method was evaluated on four popular datasets of different languages. The extensive experimental results manifest the strength of our model. Ping Wei 0001 |
AAAI | 2 |
| 2021 | Semantic Consistency Networks for 3D Object DetectionabstractDetecting 3D objects from point clouds is a significant yet challenging issue in many applications. While most existing approaches seek to leverage geometric information of point clouds, few studies accommodate the inherent semantic characteristics of each point and the consistency between the geometric and semantic cues. In this work, we propose a novel semantic consistency network (SCNet) driven by a natural principle: the class of a predicted 3D bounding box should be consistent with the classes of all the points inside this box. Specifically, our SCNet consists of a feature extraction structure, a detection decision structure, and a semantic segmentation structure. In inference, the feature extraction and the detection decision structures are used to detect 3D objects. In training, the semantic segmentation structure is jointly trained with the other two structures to produce more robust and applicative model parameters. A novel semantic consistency loss is proposed to regulate the output 3D object boxes and the segmented points to boost the performance. Our model is evaluated on two challenging datasets and achieves comparable results to the state-of-the-art methods. Wenwen Wei, Ping Wei 0001, Nanning Zheng 0001 |
AAAI | 2 |
| 2021 | Nesting spatiotemporal attention networks for action recognition
Jiapeng Li 0003, Ping Wei 0001, Nanning Zheng 0001 |
Neurocomputing | 2 |
| 2021 | A multilevel fusion network for 3D object detection
Chunlong Xia, Ping Wei 0001, Wenwen Wei, Nanning Zheng 0001 |
Neurocomputing | 2 |
| 2021 | A multi-cue guidance network for depth completion
Yongchi Zhang, Ping Wei 0001, Nanning Zheng 0001 |
Neurocomputing | 2 |
| 2021 | A Generalized Earley Parser for Human Activity Parsing and PredictionabstractDetection, parsing, and future predictions on sequence data (e.g., videos) require the algorithms to capture non-Markovian and compositional properties of high-level semantics. Context-free grammars are natural choices to capture such properties, but traditional grammar parsers (e.g., Earley parser) only take symbolic sentences as inputs. In this paper, we generalize the Earley parser to parse sequence data which is neither segmented nor labeled. Given the output of an arbitrary probabilistic classifier, this generalized Earley parser finds the optimal segmentation and labels in the language defined by the input grammar. Based on the parsing results, it makes top-down future predictions. The proposed method is generic, principled, and widely applicable. Experiment results clearly show the benefit of our method for both human activity parsing and prediction on three video datasets. Siyuan Qi, Baoxiong Jia, Siyuan Huang 0001, Ping Wei 0001, Song-Chun Zhu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Predicting Task-Driven Attention via Integrating Bottom-Up Stimulus and Top-Down GuidanceabstractTask-free attention has gained intensive interest in the computer vision community while relatively few works focus on task-driven attention (TDAttention). Thus this paper handles the problem of TDAttention prediction in daily scenarios where a human is doing a task. Motivated by the cognition mechanism that human attention allocation is jointly controlled by the top-down guidance and bottom-up stimulus, this paper proposes a cognitively-explanatory deep neural network model to predict TDAttention. Given an image sequence, bottom-up features, such as human pose and motion, are firstly extracted. At the same time, the coarse-grained task information and fine-grained task information are embedded as a top-down feature. The bottom-up features are then fused with the top-down feature to guide the model to predict TDAttention. Two public datasets are re-annotated to make them qualified for TDAttention prediction, and our model is widely compared with other models on the two datasets. In addition, some ablation studies are conducted to evaluate the individual modules in our model. Experiment results demonstrate the effectiveness of our model. Zhixiong Nan, Jingjing Jiang, Xiaofeng Gao 0002, Sanping Zhou, Weiliang Zuo, Ping Wei 0001, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 6 |
| 2020 | Local-Global Interactive Network For Face Age TransformationabstractFace age transformation aims to generate a face image in the past or future and has been receiving increasing attention for its significance recently. This paper proposes a local-global interactive framework for long-span face age transformation. We divide a face image into five local parts and design a local generative network for each of them to learn the local changes of a face image. Meanwhile a global generative network is utilized to learn the global changes. We introduce an interactive network and an age classification network, which are respectively used to integrate the local and global features and maintain the corresponding age features in different age groups. Given a face image at a certain age, our network can produce a realistic image of face aging or rejuvenation. We test and evaluate the model on complex datasets. Extensive qualitative comparison experiments prove the effectiveness and potential of our proposed method. Ping Wei 0001, Yongchi Zhang, Nanning Zheng 0001 |
ICPR | 2 |
| 2020 | Inferring Tasks and Fluents in Videos by Learning Causal RelationsabstractRecognizing time-varying object states in complex tasks is an important and challenging issue. In this paper, we propose a novel model to jointly infer object fluents and complex tasks in videos. A task is a complex human activity with specific goals and a fluent is defined as a time-varying object state. A hierarchical graph represents a task as a human action stream and multiple concurrent object fluents which vary as the human performs the actions. In this process, the human actions serve as the causes of object state changes which conversely reflect the effects of human actions. For a given input video, a causal sampling search algorithm is proposed to jointly infer the task category and the states of objects in each video frame. For model learning, a structural SVM framework is adopted to jointly train the task, fluent, cause, and effect parameters. We test the proposed method on a task and fluent dataset. Experimental results demonstrate the effectiveness of the proposed method. Haowen Tang, Ping Wei 0001, Nanning Zheng 0001 |
ICPR | 2 |
| 2020 | A Duplex Spatiotemporal Filtering Network for Video-based Person Re-identificationabstractVideo-based person re-identification plays important roles in surveillance video analysis. This paper proposes a duplex spatiotemporal filtering network (DSFN) to re-identify persons in videos. A video sequence is represented as a duplex spatiotemporal matrix. DSFN model containing a group of filters performs filtering at feature levels in both temporal and spatial dimensions, by which the model focuses on feature-level semantic information rather than image-level information as in the traditional filters. We propose sparse-orthogonal constraints to enforce the model to extract more discriminative features. DSFN seeks to characterize not only the appearance features but also dynamic information such as gaits embedded in video sequences and obtains performance improvement as a result. Experimental results show that the proposed method outperforms other comparison approaches. Chong Zheng, Ping Wei 0001, Nanning Zheng 0001 |
ICPR | 2 |
| 2020 | Multiscale Adaptation Fusion Networks for Depth CompletionabstractDepth completion is becoming a particularly important yet challenging problem with the growingly rapid progress of depth sensing technologies. Depth completion aims to complete sparse and noisy depth images to generate dense depth images. In this paper, we propose a multiscale adaptation fusion network (MAFN) for depth completion. The depth features are fused with RGB features at multiple scales with adaptation modules, where a neighbour attention mechanism is designed to adapt the local structures of the RGB image and the depth image. The fusion and completion process are unified under the encoder-decoder framework which is learned in an end-to-end way. By exploiting the detailed structural relationships of RGB images and depth images, our MAFN model can accurately complete and restore the invalid depth values on the sparse depth images. We test the proposed method on the challenging KITTI depth completion benchmark. The experimental results prove the effectiveness and strength of the proposed method. Yongchi Zhang, Ping Wei 0001, Nanning Zheng 0001 |
IJCNN | 2 |
| 2020 | A Slow-I-Fast-P Architecture for Compressed Video Action RecognitionabstractCompressed video action recognition has drawn growing attention for the storage and processing advantages of compressed videos over original raw videos. While the past few years have witnessed remarkable progress in this problem, most existing approaches rely on RGB frames from raw videos and require multi-step training. In this paper, we propose a novel Slow-I-Fast-P (SIFP) neural network model for compressed video action recognition. It consists of the slow I pathway receiving a sparse sampling I-frame clip and the fast P pathway receiving a dense sampling pseudo optical flow clip. An unsupervised estimation method and a new loss function are designed to generate pseudo optical flows in compressed videos. Our model eliminates the dependence on the traditional optical flows calculated from raw videos. The model is trained in an end-to-end way. The proposed method is evaluated on the challenging HMDB51 and UCF101 datasets. The extensive comparison results and ablation studies demonstrate the effectiveness and strength of the proposed method. Jiapeng Li 0003, Ping Wei 0001, Yongchi Zhang, Nanning Zheng 0001 |
ACM Multimedia | 2 |
| 2020 | Spatiotemporal neural networks for action recognition based on joint loss
Chao Jing, Ping Wei 0001, Hongbin Sun 0001, Nanning Zheng 0001 |
Neural Comput. Appl. | 2 |
| 2020 | Learning to infer human attention in daily activities
Zhixiong Nan, Tianmin Shu, Shu Wang 0002, Ping Wei 0001, Song-Chun Zhu, Nanning Zheng 0001 |
Pattern Recognit. | 5 |
| 2019 | Inverse Discriminative Networks for Handwritten Signature VerificationabstractHandwritten signature verification is an important technique for many financial, commercial, and forensic applications. In this paper, we propose an inverse discriminative network (IDN) for writer-independent handwritten signature verification, which aims to determine whether a test signature is genuine or forged compared to the reference signature. The IDN model contains four weight-shared neural network streams, of which two receiving the original signature images are the discriminative streams and the other two addressing the gray-inverted images form the inverse streams. Multiple paths of attention modules connect the discriminative streams and the inverse streams to propagate messages. With the inverse streams and the multi-path attention modules, the IDN model intensifies the effective information of signature verification. Since there was no proper Chinese signature dataset in the community, we collected a large-scale Chinese signature dataset with approximately 29,000 images of 749 individuals’ signatures. We test our method on the Chinese signature dataset and other three signature datasets of different languages: CEDAR, BHSig-B, and BHSig-H. Experiments prove the strength and potential of our method. Ping Wei 0001 |
CVPR | 1 |
| 2019 | Face Age Transformation with Progressive Residual Adversarial AutoencoderabstractFace age transformation is an important issue in many applications. While unidirectional and short-span face ageing has achieved remarkable progress, it remains a challenging problem to generate both younger-look and older-look face images over long age span. In this paper, we present a progressive residual adversarial autoencoder (PRAA) model for bidirectional and long-span face age transformation. Given an input face image, our model aims to synthesize face images of its younger looks (face rejuvenation) and older looks (face ageing). The PRAA contains adversarial generators and discriminators where age information is encoded as latent features. It adopts a residual-encoding strategy by which the original face images and the residual face images are jointly encoded. We adopt a progressive multi-scale method to train the network, by which our model can capture both the global structure and the local detail changes in face age transformation. We test our model on challenging data and the experimental results prove the strength of our method. Xuexiang Zhang, Ping Wei 0001, Nanning Zheng 0001 |
IJCNN | 2 |
| 2019 | Scene-Guided Region Proposal Re-ranking Method for On-road Vehicle Candidate GenerationabstractVehicle candidate generation is important for vehicle detection. Existing vehicle detection studies usually employ general-purpose region proposal methods to generate vehicle candidates, which do not consider the specificity of on-road vehicles in traffic scenes. In this paper, we propose a model to re-rank the candidates that are generated by general-purpose region proposal methods. Our model considers the specificity of on-road vehicle candidate generation in traffic scenes by encoding global-local semantic context and location-size geometric compatibility. In the experiments, we test our model on three art-of-the-state region proposal methods using two public datasets. The results show the significant performance improvement is gained after applying our model. Zhixiong Nan, Jiawei He 0002, Ping Wei 0001, Linhai Xu, Hongbin Sun 0001, Nanning Zheng 0001 |
IV | 4 |
| 2019 | Learning Composite Latent Structures for 3D Human Action Representation and Recognitionabstract3D human action representation and recognition are important issues in many multimedia applications. While latent state approaches have been widely used for action modeling, previous works assume the latent states of actions are single attribute. This assumption is inaccurate for representing structures of complex actions. In this paper, we propose that latent states have composite attributes and introduce a novel composite latent structure (CLS) model to represent and recognize 3D human actions with skeleton sequences. A human action is modeled with a hierarchical graph, which represents the action sequence as sequential atomic actions. An atomic action is represented as a composite latent state, which is composed of a latent semantic attribute and a latent geometric attribute. A discriminative EM-like algorithm is proposed to learn the model parameters and the composite latent structures of human actions. Given a 3D skeleton sequence, a composite attribute iterative programming algorithm is proposed to recognize the action and infer the action's latent temporal structure. We evaluate the proposed method on three challenging 3D action datasets-MSR 3D Action Dataset, Multiview 3D Event Dataset, and UTKinect-Action 3D Dataset. Extensive experimental results on these datasets demonstrate the effectiveness and advantage of the proposed method. Ping Wei 0001, Hongbin Sun 0001, Nanning Zheng 0001 |
IEEE Trans. Multim. | 1 |
| 2018 | Inferring Shared Attention in Social Scene VideosabstractThis paper addresses a new problem of inferring shared attention in third-person social scene videos. Shared attention is a phenomenon that two or more individuals simultaneously look at a common target in social scenes. Perceiving and identifying shared attention in videos plays crucial roles in social activities and social scene understanding. We propose a spatial-temporal neural network to detect shared attention intervals in videos and predict shared attention locations in frames. In each video frame, human gaze directions and potential target boxes are two key features for spatially detecting shared attention in the social scene. In temporal domain, a convolutional Long Short-Term Memory network utilizes the temporal continuity and transition constraints to optimize the predicted shared attention heatmap. We collect a new dataset VideoCoAtt1 from public TV show videos, containing 380 complex video sequences with more than 492,000 frames that include diverse social scenes for shared attention study. Experiments on this dataset show that our model can effectively infer shared attention in videos. We also empirically verify the effectiveness of different components in our model. Lifeng Fan, Yixin Chen 0003, Ping Wei 0001, Wenguan Wang, Song-Chun Zhu |
CVPR | 3 |
| 2018 | Where and Why Are They Looking? Jointly Inferring Human Attention and Intentions in Complex TasksabstractThis paper addresses a new problem - jointly inferring human attention, intentions, and tasks from videos. Given an RGB-D video where a human performs a task, we answer three questions simultaneously: 1) where the human is looking - attention prediction; 2) why the human is looking there - intention prediction; and 3) what task the human is performing - task recognition. We propose a hierarchical model of human-attention-object (HAO) which represents tasks, intentions, and attention under a unified framework. A task is represented as sequential intentions which transition to each other. An intention is composed of the human pose, attention, and objects. A beam search algorithm is adopted for inference on the HAO graph to output the attention, intention, and task results. We built a new video dataset of tasks, intentions, and attention. It contains 14 task classes, 70 intention categories, 28 object classes, 809 videos, and approximately 330,000 frames. Experiments show that our approach outperforms existing approaches. Ping Wei 0001, Yang Liu 0266, Tianmin Shu, Nanning Zheng 0001, Song-Chun Zhu |
CVPR | 1 |
| 2018 | Exploring the Potential of Using Semantic Context and Common Sense in On-Road Vehicle DetectionabstractVehicle detection is an important research topic for autonomous driving community. Since the great success of deep learning on object detection, almost all vehicle detection methods go along with this line. However, deep learning methods heavily rely on the training data, and the whole mechanism is like a “black box” Therefore, in this paper, we explore a vehicle detection method using traffic semantic context and human common sense instead of relying on the training data. To verify our idea, we compare our method with two classic machine learning methods as well as three state- of-the-art deep learning methods on a dataset collected in real traffics. The results show that our method outperforms others on this dataset. The deep learning methods may exceed ours after enlarging the training data or testing on more complicated datasets. However, the main contribution of this paper is providing inspiration for learning methods, and we believe their performance can be greatly improved after considering the idea of this paper. Zhixiong Nan, Menghan Pan, Xiao Wang 0002, Ping Wei 0001, Linhai Xu, Hongbin Sun 0001, Jingmin Xin, Nanning Zheng 0001 |
Intelligent Vehicles Symposium | 4 |
| 2017 | Jointly Recognizing Object Fluents and Tasks in Egocentric VideosabstractThis paper addresses the problem of jointly recognizing object fluents and tasks in egocentric videos. Fluents are the changeable attributes of objects. Tasks are goal-oriented human activities which interact with objects and aim to change some attributes of the objects. The process of executing a task is a process to change the object fluents over time. We propose a hierarchical model to represent tasks as concurrent and sequential object fluents. In a task, different fluents closely interact with each other both in spatial and temporal domains. Given an egocentric video, a beam search algorithm is applied to jointly recognizing the object fluents in each frame, and the task of the entire video. We collected a large scale egocentric video dataset of tasks and fluents. This dataset contains 14 categories of tasks, 25 object classes, 21 categories of object fluents, 809 video sequences, and approximately 333,000 video frames. The experimental results on this dataset prove the strength of our method. Yang Liu 0266, Ping Wei 0001, Song-Chun Zhu |
ICCV | 2 |
| 2017 | Monocular 3D Human Pose Estimation by Predicting Depth on JointsabstractThis paper aims at estimating full-body 3D human poses from monocular images of which the biggest challenge is the inherent ambiguity introduced by lifting the 2D pose into 3D space. We propose a novel framework focusing on reducing this ambiguity by predicting the depth of human joints based on 2D human joint locations and body part images. Our approach is built on a two-level hierarchy of Long Short-Term Memory (LSTM) Networks which can be trained end-to-end. The first level consists of two components: 1) a skeleton-LSTM which learns the depth information from global human skeleton features; 2) a patch-LSTM which utilizes the local image evidence around joint locations. The both networks have tree structure defined on the kinematic relation of human skeleton, thus the information at different joints is broadcast through the whole skeleton in a top-down fashion. The two networks are first pre-trained separately on different data sources and then aggregated in the second layer for final depth prediction. The empirical e-valuation on Human3.6M and HHOI dataset demonstrates the advantage of combining global 2D skeleton and local image patches for depth prediction, and our superior quantitative and qualitative performance relative to state-of-the-art methods. Xiaohan Nie, Ping Wei 0001, Song-Chun Zhu |
ICCV | 2 |
| 2017 | Predicting Human Activities Using Stochastic GrammarabstractThis paper presents a novel method to predict future human activities from partially observed RGB-D videos. Human activity prediction is generally difficult due to its non-Markovian property and the rich context between human and environments. We use a stochastic grammar model to capture the compositional structure of events, integrating human actions, objects, and their affordances. We represent the event by a spatial-temporal And-Or graph (ST-AOG). The ST-AOG is composed of a temporal stochastic grammar defined on sub-activities, and spatial graphs representing sub-activities that consist of human actions, objects, and their affordances. Future sub-activities are predicted using the temporal grammar and Earley parsing algorithm. The corresponding action, object, and affordance labels are then inferred accordingly. Extensive experiments are conducted to show the effectiveness of our model on both semantic event parsing and future activity prediction. Siyuan Qi, Siyuan Huang 0001, Ping Wei 0001, Song-Chun Zhu |
ICCV | 3 |
| 2017 | Inferring Human Attention by Learning Latent IntentionsabstractThis paper addresses the problem of inferring 3D human attention in RGB-D videos at scene scale. 3D human attention describes where a human is looking in 3D scenes. We propose a probabilistic method to jointly model attention, intentions, and their interactions. Latent intentions guide human attention which conversely reveals the intention features. This mutual interaction makes attention inference a joint optimization with latent intentions. An EM-based approach is adopted to learn the latent intentions and model parameters. Given an RGB-D video with 3D human skeletons, a joint-state dynamic programming algorithm is utilized to jointly infer the latent intentions, the 3D attention directions, and the attention voxels in scene point clouds. Experiments on a new 3D human attention dataset prove the strength of our method. Ping Wei 0001, Dan Xie 0005, Nanning Zheng 0001, Song-Chun Zhu |
IJCAI | 1 |
| 2017 | Modeling 4D Human-Object Interactions for Joint Event Segmentation, Recognition, and Object LocalizationabstractIn this paper, we present a 4D human-object interaction (4DHOI) model for solving three vision tasks jointly: i) event segmentation from a video sequence, ii) event recognition and parsing, and iii) contextual object localization. The 4DHOI model represents the geometric, temporal, and semantic relations in daily events involving human-object interactions. In 3D space, the interactions of human poses and contextual objects are modeled by semantic co-occurrence and geometric compatibility. On the time axis, the interactions are represented as a sequence of atomic event transitions with coherent objects. The 4DHOI model is a hierarchical spatial-temporal graph representation which can be used for inferring scene functionality and object affordance. The graph structures and parameters are learned using an ordered expectation maximization algorithm which mines the spatial-temporal structures of events from RGB-D video samples. Given an input RGB-D video, the inference is performed by a dynamic programming beam search algorithm which simultaneously carries out event segmentation, recognition, and object localization. We collected a large multiview RGB-D event dataset which contains 3,815 video sequences and 383,036 RGB-D frames captured by three RGB-D cameras. The experimental results on three challenging datasets demonstrate the strength of the proposed method. Ping Wei 0001, Yibiao Zhao, Nanning Zheng 0001, Song-Chun Zhu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2014 | Semantic propagation network with robust spatial context descriptors for multi-class object labeling
Ping Wei 0001, Yuehu Liu, Nanning Zheng 0001, Shaozhuo Zhai |
Neural Comput. Appl. | 1 |
| 2013 | Concurrent Action Detection with Structural PredictionabstractAction recognition has often been posed as a classification problem, which assumes that a video sequence only have one action class label and different actions are independent. However, a single human body can perform multiple concurrent actions at the same time, and different actions interact with each other. This paper proposes a concurrent action detection model where the action detection is formulated as a structural prediction problem. In this model, an interval in a video sequence can be described by multiple action labels. An detected action interval is determined both by the unary local detector and the relations with other actions. We use a wavelet feature to represent the action sequence, and design a composite temporal logic descriptor to describe the action relations. The model parameters are trained by structural SVM learning. Given a long video sequence, a sequential decision window search algorithm is designed to detect the actions. Experiments on our new collected concurrent action dataset demonstrate the strength of our method. Ping Wei 0001, Nanning Zheng 0001, Yibiao Zhao, Song-Chun Zhu |
ICCV | 1 |
| 2013 | Modeling 4D Human-Object Interactions for Event and Object RecognitionabstractRecognizing the events and objects in the video sequence are two challenging tasks due to the complex temporal structures and the large appearance variations. In this paper, we propose a 4D human-object interaction model, where the two tasks jointly boost each other. Our human-object interaction is defined in 4D space: i) the co occurrence and geometric constraints of human pose and object in 3D space, ii) the sub-events transition and objects coherence in 1D temporal dimension. We represent the structure of events, sub-events and objects in a hierarchical graph. For an input RGB-depth video, we design a dynamic programming beam search algorithm to: i) segment the video, ii) recognize the events, and iii) detect the objects simultaneously. For evaluation, we built a large-scale multiview 3D event dataset which contains 3815 video sequences and 383,036 RGBD frames captured by the Kinect cameras. The experiment results on this dataset show the effectiveness of our method. Ping Wei 0001, Yibiao Zhao, Nanning Zheng 0001, Song-Chun Zhu |
ICCV | 1 |
| 2011 | Fractal image coding using SSIMabstractSince Jacquin proposed original fractal image compression technique in 1990, fractal coding method has been developed into various schemes. Traditionally, fractal coding uses mean square error (MSE) to evaluate similarity of image blocks, but the similarity evaluated by MSE usually differs from human visual system (HVS). Compared with MSE, structural similarity (SSIM) is an image measure index which is more appropriate for the HVS. This paper proposes a new fractal coding scheme which uses structural similarity to measure the similarity between image blocks and compute these blocks' coefficients. The experiment results show that the proposed method generates higher quality images for the HVS than MSE scheme. Jianji Wang 0001, Yuehu Liu, Ping Wei 0001, Yaochen Li, Nanning Zheng 0001 |
ICIP | 3 |
| 2009 | A Statistical-Structural Constraint Model for Cartoon Face Wrinkle Representation and Generation
Ping Wei 0001, Yuehu Liu, Nanning Zheng 0001, Yang Yang 0066 |
ACCV (3) | 1 |
| 2008 | A Method for Deforming-Driven Exaggerated Facial Animation GenerationabstractThis paper proposes a method for automatically generating facial animation with exaggerated features from a facial image. According to facial diversity and identity, an exaggerated face can be determined by the neutral face and the exaggeration effect difference that is represented with the deforming parameters. The proposed method utilizes the central point and the feature vector to represent a facial component and then transforms these feature vectors under the control of deforming parameters to generate the exaggerated facial features. The exaggerated facial animation can be driven to generate by the sequence of the deforming parameters. Experimental results prove that the proposed method has the advantages of simplicity, flexibility and directness, and the generated facial animations are expressive. Ping Wei 0001, Yuehu Liu, Yuanqi Su |
CW | 1 |