EDBT 2026 Demo / reviewers in the wild / expert
Pengfei Ren 0001
dblp:31/10610-1
· DBLP profile ↗
49ranked-venue papers
10as first author
46since 2021 · last 2026
0000-0002-1691-6457ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 8 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 32 · 8 first-author · 29 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VALU: A Benchmark for Video Anomaly Temporal Localization and Understanding at Multiple Semantic LevelsabstractYixiao He, Menghao Zhang, Haifeng Sun, Jing Wang, Kangheng Lin, Jinghan Wang, Chenye Xu, Pengfei Ren, Qi Qi, Jingyu Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yixiao He, Menghao Zhang 0004, Haifeng Sun 0001, Jing Wang 0039, Kangheng Lin, Chenye Xu, Pengfei Ren 0001, Qi Qi 0001, Jingyu Wang 0001 |
ACL (1) | 8 |
| 2026 | InkFlow: Connected Handwriting Recognition for Natural Mid-Air Input in Mixed Reality
Xufeng Jian, Qi Qi 0001, Linpei Zhang, Haifeng Sun 0001, Pengfei Ren 0001, Guangtian Liu, Shule Cao, Jingyu Wang 0001 |
CHI | 5 |
| 2026 | Ubi Grip: Ubiquitous Grip-Based Tangible Object Utilization in Augmented RealityabstractTangible Augmented Reality (AR) enhances user immersion in virtual world by providing haptic feedback through physical proxy objects. However, existing approaches primarily focus on selecting proxy objects based on their global physical properties, neglecting the utilization of local features. Besides, the prevailing strategy of mapping one virtual object to a single dedicated physical proxy creates an inherent switching cost, limiting flexibility and efficiency. Additionally, due to challenges such as real-time performance, generalization and occlusion, the vision-based hand-object tracking remains a difficult task. In this paper, we propose Ubi Grip, an universal hand-object interaction framework for creating grip-based tangible AR applications based on the local graspable feature and a comprehensive hand-object interaction attributes methodology. We employ a lightweight object tracking method to perform tracking, utilizing a hand mask filter and transformation strategy to optimize object pose based on the hand-held properties. Moreover, we design a user-defined workflow for grasping tangible objects, allowing users to switch grips and map interactions. We evaluated our system through comprehensive algorithmic benchmarks and a user study. The benchmarks demonstrate our SOTA performance in object pose estimation and generalization, while the user study validates system usability, providing deeper insights. Xufeng Jian, Guangtian Liu, Xiayang Zhou, Haifeng Sun 0001, Qi Qi 0001, Pengfei Ren 0001, Shan Jiang 0008, Jing Wang 0039, Jianxin Liao, Jingyu Wang 0001 |
VR | 8 |
| 2026 | Enhancing MLLMs for Online Understanding in Video Services via Preference OptimizationabstractOnline video understanding is pivotal for emerging video streaming services. However, existing Multimodal Large Language Models (MLLMs) encounter significant challenges in this domain, specifically in maintaining holistic visual perception under strict token budgets and attaining grounded semantic reasoning amidst dynamic context changes. To mitigate these issues, we propose OV-DPO, a systematic multimodal preference optimization approach. To ensure holistic visual perception, we introduce a visual preference objective that compels the model to strictly ground its reasoning in clear visual evidence. This is supported by the QGDFR strategy, which dynamically reallocates resolution budgets to preserve critical visual details for chosen samples while intentionally degrading them for rejected ones to construct valid visual contrasts. Concurrently, to promote grounded semantic reasoning, we employ a textual preference objective. We design the SROVA framework to construct high-quality preference pairs, utilizing self-refinement to generate factual chosen responses and visual degradation to induce language-prior-driven hallucinations as rejected samples. By jointly optimizing these objectives alongside anchored and supervised fine-tuning terms, OV-DPO ensures training stability and robust alignment. Experimental results demonstrate that our approach significantly outperforms baselines such as SFT and DPO. Notably, utilizing only 29K high-quality samples, our 3B-scale model achieves results competitive with significantly larger-scale models on both online and offline benchmarks, offering an efficient solution for online video understanding. Qi Qi 0001, Yixiao He, Menghao Zhang 0004, Haifeng Sun 0001, Pengfei Ren 0001, Huazheng Wang, Jianxin Liao, Jingyu Wang 0001 |
IEEE Trans. Serv. Comput. | 5 |
| 2025 | Pose-Guided Temporal Enhancement for Robust Low-Resolution Hand Reconstructionabstract3D hand reconstruction is essential in non-contact human-computer interaction applications, but existing methods struggle with low-resolution images, which occur in slightly distant interactive scenes. Leveraging temporal information can mitigate the limitations of individual low-resolution images that lack detailed appearance information, enhancing the robustness and accuracy of hand reconstruction. Existing temporal methods typically use joint features to represent temporal information, avoiding interference from redundant background information. However, joint features excessively disregard the spatial context of visual features, limiting hand reconstruction accuracy. We propose to integrate temporal joint features with visual features to construct a robust low-resolution visual representation. We adopt Triplane Features, a dense representation with 3D spatial awareness, to bridge the gap between the joint features and visual features that are misaligned in terms of representation form and semantics. Triplane Features are obtained by orthogonally projecting joint features, embedding hand structure information into the 3D spatial context. Furthermore, we compress the spatial information of the three planes into a 2D dense feature thourgh Spatial-Aware Fusion to enhance the visual features. By using enhanced visual features enriched with temporal information for hand reconstruction, our method achieves competitive performance at much lower resolutions compared to state-of-the-art methods operating at high resolution on DexYCB, HanCo and H2O. Code is available at https://github.com/NewbieFan/Temp-LowRes-hand. Kaixin Fan, Pengfei Ren 0001, Jingyu Wang 0001, Haifeng Sun 0001, Qi Qi 0001, Zirui Zhuang, Jianxin Liao |
CVPR | 2 |
| 2025 | Unveiling Internal Reasoning Modes in LLMs: A Deep Dive into Latent Reasoning vs. Factual Shortcuts with Attribute Rate RatioabstractYiran Yang, Haifeng Sun, Jingyu Wang, Qi Qi, Zirui Zhuang, Huazheng Wang, Pengfei Ren, Jing Wang, Jianxin Liao. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Haifeng Sun 0001, Jingyu Wang 0001, Qi Qi 0001, Zirui Zhuang, Huazheng Wang, Pengfei Ren 0001, Jing Wang 0039, Jianxin Liao |
EMNLP | 7 |
| 2025 | Hierarchical-Aware Orthogonal Disentanglement Framework for Fine-Grained Skeleton-Based Action Recognition
Haochen Chang, Pengfei Ren 0001, Liang Xie 0012, Erwei Yin |
ICCV | 2 |
| 2025 | Prior-Aware Dynamic Temporal Modeling Framework for Sequential 3D Hand Pose Estimation
Pengfei Ren 0001, Jingyu Wang 0001, Haifeng Sun 0001, Qi Qi 0001, Menghao Zhang 0004, Lei Zhang 0094, Jing Wang 0039, Jianxin Liao |
ICCV | 1 |
| 2025 | Mitigating Object Hallucination in Large Vision-Language Models via Visual Attention Direct Preference OptimizationabstractLarge Vision-Language Models (LVLMs) suffer from severe object hallucinations, leading them to frequently generate outputs that do not correspond to the image content, significantly reducing the credibility and reliability of their responses. Recent research has attempted to enhance LVLMs by employing Direct Preference Optimization (DPO) to reduce hallucinations and improve response quality. However, these approaches typically utilize text-only preference response pairs, neglecting the influence of visual input in optimizing LVLMs. In this paper, we propose VA-DPO, a multimodal optimization objective. VA-DPO leverages the LVLMs' attention to select and corrupt critical parts of images, thereby constructing visual preference image pairs. This approach integrates both text and visual preference optimization objectives to achieve effective alignment optimization of LVLMs. We conduct extensive experiments on LVLMs of different sizes, and the results demonstrate that VA-DPO effectively reduces hallucinations across various tasks. Compared to other hallucination mitigation approaches, VA-DPO achieves more competitive results. Yixiao He, Haifeng Sun 0001, Qi Qi 0001, Zirui Zhuang, Pengfei Ren 0001, Huazheng Wang, Yafeng Nan, Jingyu Wang 0001 |
ICME | 5 |
| 2025 | Masked Self-Supervised Learning and Semantic Noise Separation for Video Anomaly DetectionabstractRecent progress in video anomaly detection assumes that anomalies cannot be effectively reconstructed because they remain unseen during training. However, we observe that most existing methods excessively rely on appearance features, resulting in the accurate reconstruction of anomalies with subtle short-term appearance variations, which we refer to as appearance confusion. Meanwhile, many approaches fail to exploit sufficient semantic distinction, resulting in motion confusion for anomalies with motion patterns similar to normal ones. In this paper, we propose a masked self-supervised learning-based framework, which effectively addresses the two confusions by exploring context-aware motion patterns and discriminative semantic normality representations. First, we introduce reconstructing multi-pattern masked spatiotemporal information to motivate the model to capture motion patterns that focus on long-term context. Then, we design a semantic noise separation network to address motion confusion, facilitating the construction of semantic normality boundaries through semantic-aware separation. Extensive experiments on the Avenue and ShanghaiTech datasets validate the effectiveness of our proposed method. Menghao Zhang 0004, Lei Zhang 0094, Qi Qi 0001, Haifeng Sun 0001, Pengfei Ren 0001, Bo He 0003, Jing Wang 0039, Jingyu Wang 0001 |
ICME | 6 |
| 2025 | A³-Net: Calibration-Free Multi-View 3D Hand Reconstruction for Enhanced Musical Instrument LearningabstractPrecise 3D hand posture is essential for learning musical instruments. Reconstructing highly precise 3D hand gestures enables learners to correct and master proper techniques through 3D simulation and Extended Reality. However, exsiting methods typically rely on precisely calibrated multi-camera systems, which are not easily deployable in everyday environments. In this paper, we focus on calibration-free multi-view 3D hand reconstruction in unconstrained scenarios. Establishing correspondences between multi-view images is particularly challenging without camera extrinsics. To address this, we propose A^3-Net, a multi-level alignment framework that utilizes 3D structural representations with hierarchical geometric and explicit semantic information as alignment proxies, facilitating multi-view feature interaction in both 3D geometric space and 2D visual space. Specifically, we first perfrom global geometric alignment to map multi-view features into a canonical space. Subsequently, we aggregate information into predefined sparse and dense proxies to further integrate cross-view semantics through mutual interaction. Finnaly, we perfrom 2D alignment to align projected 2D visual features with 2D observations. Our method achieves state-of-the-art results in the multi-view 3D hand reconstruction task, demonstrating the effectiveness of our proposed framework. Geng Chen 0006, Xufeng Jian, Pengfei Ren 0001, Jingyu Wang 0001, Haifeng Sun 0001, Qi Qi 0001, Jing Wang 0039, Jianxin Liao |
IJCAI | 4 |
| 2025 | Rule Meets Learning: Confidence-Aware Multi-View Fusion for Self-Supervised 3D Hand Pose EstimationabstractSelf-supervised 3D hand pose estimation methods can leverage labeled synthetic data along with unlabeled real-world data for model training, thereby alleviating the reliance on large-scale annotated datasets. Multi-view information fusion is a key factor in the success of these methods. Rule-based fixed fusion methods are simple, efficient, and generalizable, but they neglect the rich visual information in each view. Neural network-based learnable fusion methods can effectively model both intra- and inter-view semantic context, but they tend to overfit to the domain-specific feature of synthetic data and susceptible to interference of domain gaps. In this paper, we decompose multi-view fusion into two components: a learnable confidence estimation stage and a fixed confidence fusion stage. This design not only enables effective use of multi-view semantic cues but also ensures strong cross-domain generalization. To achieve accurate and robust confidence estimation, our method jointly exploits both multi-view pose consistency and pose-to-data consistency. Experiments on three public datasets demonstrate that our approach significantly outperforms existing state-of-the-art self-supervised 3D hand pose estimation methods. Pengfei Ren 0001, Jingyu Wang 0001, Haifeng Sun 0001, Qi Qi 0001, Jing Wang 0039, Jianxin Liao |
ACM Multimedia | 1 |
| 2025 | A Dual-Branch 3D Spatial-Aware Latent Diffusion for Realistic Depth Image SynthesisabstractSynthetic images serve as a promising alternative to real images in 3D hand pose estimation, providing accurate annotations at a lower cost. However, the domain gap between real and synthetic images constrains the generalization ability of hand pose estimation trained on synthetic data. Previous methods rely on Generative Adversarial Networks (GANs) for domain translation; however, they fail to achieve realistic depth synthesis due to instability and limited image quality. Diffusion models provide high-quality synthesis due to their stability and controllability. However, existing methods often ignore the 3D structure awareness in hand image generation. In this paper, we propose a Dual-Branch 3D Spatial-Aware Latent Diffusion (DSW-LD) for realistic depth image generation. The Global Structure Module (GSM) and the Local Geometry Module (LGM) complement each other, with GSM capturing global spatial structure through coarse-grained 3D joint features and LGM focusing on local geometric details using fine-grained 3D mesh representations. To maintain the global structure consistency, we adopt a layer-aware injection mechanism that enables the model to adaptively learn the optimal representation from fused 2D latent representations and 3D joint features. To explicitly align 3D and 2D features of local regions and enhance the flexibility of feature matching, we design a dynamic depth-aware interpolation to project 3D mesh features into 2D image space. Both quantitative and qualitative experimental results demonstrate the superiority of our method over the state-of-the-arts for realistic depth synthesis. Compared to training only on real depth images, our method enables the hand pose estimator to achieve significantly better performance with our synthetic data and less real data (10%). Shuang Hao 0017, Pengfei Ren 0001, Lei Zhang 0094, Haifeng Sun 0001, Pan Ting, Menghao Zhang 0004, Cong Liu 0046, Qi Qi 0001, Jianxin Liao, Jingyu Wang 0001 |
ACM Multimedia | 2 |
| 2025 | VIHand: Enhancing 3D Hand Pose Estimation with Visual-Inertial Benchmark
Pengfei Ren 0001, Liang Xie 0012, Yue Gao 0005, Erwei Yin |
ACM Multimedia | 2 |
| 2025 | Evaluating and Mitigating Object Hallucination in Large Vision-Language Models: Can They Still See Removed Objects?abstractYixiao He, Haifeng Sun, Pengfei Ren, Jingyu Wang, Huazheng Wang, Qi Qi, Zirui Zhuang, Jing Wang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yixiao He, Haifeng Sun 0001, Pengfei Ren 0001, Jingyu Wang 0001, Huazheng Wang, Qi Qi 0001, Zirui Zhuang, Jing Wang 0039 |
NAACL (Long Papers) | 3 |
| 2025 | Unified 2D-3D Discrete Priors for Noise-Robust and Calibration-Free Multiview 3D Human Pose EstimationabstractMulti-view 3D human pose estimation (HPE) leverages complementary information across views to improve accuracy and robustness. Traditional methods rely on camera calibration to establish geometric correspondences, which is sensitive to calibration accuracy and lacks flexibility in dynamic settings. Calibration-free approaches address these limitations by learning adaptive view interactions, typically leveraging expressive and flexible continuous representations. However, as the multiview interaction relationship is learned entirely from data without constraint, they are vulnerable to noisy input, which can propagate, amplify and accumulate errors across all views, severely corrupting the final estimated pose.
To mitigate this, we propose a novel framework that integrates a noise-resilient discrete prior into the continuous representation-based model. Specifically, we introduce the \textit{UniCodebook}, a unified, compact, robust, and discrete representation complementary to continuous features, allowing the model to benefit from robustness to noise while preserving regression capability.
Furthermore, we further propose an attribute-preserving and complementarity-enhancing Discrete-Continuous Spatial Attention (DCSA) mechanism to facilitate interaction between discrete priors and continuous pose features.
Extensive experiments on three representative datasets demonstrate that our approach outperforms both calibration-required and calibration-free methods, achieving state-of-the-art performance. Geng Chen 0006, Pengfei Ren 0001, Xufeng Jian, Haifeng Sun 0001, Menghao Zhang 0004, Qi Qi 0001, Zirui Zhuang, Jing Wang 0039, Jianxin Liao, Jingyu Wang 0001 |
NeurIPS | 2 |
| 2025 | Generalizable Hand-Object Modeling from Monocular RGB Images via 3D GaussiansabstractRecent advances in hand-object interaction modeling have employed implicit representations, such as Signed Distance Functions (SDF) and Neural Radiance Fields (NeRF) to reconstruct hands and objects with arbitrary topology and photo-realistic detail. However, these methods often rely on dense 3D surface annotations, or are tailored to short clips constrained in motion trajectories and scene contexts, limiting their generalization to diverse environments and movement patterns. In this work, we present HOGS, an adaptively perceptive 3D Gaussian Splatting (3DGS) framework for generalizable hand-object modeling from unconstrained monocular RGB images. By integrating photometric cues from the visual modality with the physically grounded structure of 3D Gaussians, HOGS disentangles inherent geometry from transient lighting and motion-induced appearance changes. This endows hand-object assets with the ability to generalize to unseen environments and dynamic motion patterns. Experiments on two challenging datasets demonstrate that HOGS outperforms state-of-the-art methods in monocular hand-object reconstruction and photo-realistic rendering. Pengfei Ren 0001, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Jing Wang 0039, Jianxin Liao, Jingyu Wang 0001 |
NeurIPS | 2 |
| 2025 | Do LVLMs Truly Understand Video Anomalies? Revealing Hallucination via Co-Occurrence PatternsabstractLarge Vision-Language Models (LVLMs) pretrained on large-scale multimodal data have shown promising capabilities in Video Anomaly Detection (VAD). However, their ability to reason about abnormal events based on scene semantics remains underexplored. In this paper, we investigate LVLMs’ behavior in VAD from a visual-textual co-occurrence perspective, focusing on whether their decisions are driven by statistical shortcuts between visual instances and textual phrases. By analyzing visual-textual co-occurrence in pretraining data and conducting experiments under different data settings, we reveal a hallucination phenomenon: LVLMs tend to rely on co-occurrence patterns between visual instances and textual phrases associated with either normality or abnormality, leading to incorrect predictions when these high-frequency objects appear in semantically mismatched contexts. To address this issue, we propose VAD-DPO, a direct preference optimization method supervised with counter-example pairs. By constructing visually similar but semantically contrasting video clips, VAD-DPO encourages the model to align its predictions with the semantics of scene rather than relying on co-occurrence patterns. Extensive experiments on six benchmark datasets demonstrate the effectiveness of VAD-DPO in enhancing both anomaly detection and reasoning performance, particularly in scene-dependent scenarios. Menghao Zhang 0004, Huazheng Wang, Pengfei Ren 0001, Kangheng Lin, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Lei Zhang 0094, Jianxin Liao, Jingyu Wang 0001 |
NeurIPS | 3 |
| 2025 | Towards Bare-Hand Interaction for Whiteboard Collaboration in Virtual RealityabstractWhiteboard collaboration in virtual reality (VR) is an important task in collaborative virtual environments. The current research mainly relies on the use of controllers or dedicated pens but additional devices will cause inconvenience to users. Bare-hand writing offers rich collaborative semantics through natural gestures but remains underexplored. This paper addresses challenges and solutions for bare-hand whiteboard collaboration. We analyze the input process and identify key challenges in determining pen-drop, writing, and pen-lift intentions while maintaining user control over their avatar. Our approach addresses two VR scenarios: one without and one with physical planes. The method for the first case is called Air-writing, which dynamically adjusts the distance between the avatar's torso and the virtual whiteboard during the processes of pen-drop and pen-lift to ensure a consistent writing experience in VR. The method for the second case is called Physical-writing, which allows users to write smoothly with passive haptic feedback and physical constraints provided by the real surface by remapping the whiteboard in VR with a plane in reality. A comprehensive user study is conducted to evaluate communication efficiency, input accuracy, collaboration efficiency, and user experience of the two methods. The experimental results indicate that bare-hand interaction improves communication efficiency by 8% over controllers and performs similarly to real-world whiteboard collaboration. The Physical-writing method also demonstrates higher accuracy and user satisfaction compared to the Air-writing method. Guangtian Liu, Haonan Su, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Pengfei Ren 0001, Jianxin Liao |
Proc. ACM Hum. Comput. Interact. | 7 |
| 2024 | Keypoint Fusion for RGB-D Based 3D Hand Pose EstimationabstractPrevious 3D hand pose estimation methods primarily rely on a single modality, either RGB or depth, and the comprehensive utilization of the dual modalities has not been extensively explored. RGB and depth data provide complementary information and thus can be fused to enhance the robustness of 3D hand pose estimation. However, there exist two problems for applying existing fusion methods in 3D hand pose estimation: redundancy of dense feature fusion and ambiguity of visual features. First, pixel-wise feature interactions introduce high computational costs and ineffective calculations of invalid pixels. Second, visual features suffer from ambiguity due to color and texture similarities, as well as depth holes and noise caused by frequent hand movements, which interferes with modeling cross-modal correlations. In this paper, we propose Keypoint-Fusion for RGB-D based 3D hand pose estimation, which leverages the unique advantages of dual modalities to mutually eliminate the feature ambiguity, and performs cross-modal feature fusion in a more efficient way. Specifically, we focus cross-modal fusion on sparse yet informative spatial regions (i.e. keypoints). Meanwhile, by explicitly extracting relatively more reliable information as disambiguation evidence, depth modality provides 3D geometric information for RGB feature pixels, and RGB modality complements the precise edge information lost due to the depth noise. Keypoint-Fusion achieves state-of-the-art performance on two challenging hand datasets, significantly decreasing the error compared with previous single-modal methods. Pengfei Ren 0001, Jingyu Wang 0001, Haifeng Sun 0001, Qi Qi 0001, Zirui Zhuang, Jianxin Liao |
AAAI | 2 |
| 2024 | Dynamic Support Information Mining for Category-Agnostic Pose EstimationabstractCategory-agnostic pose estimation (CAPE) aims to predict the pose of a query image based on few support images with pose annotations. Existing methods achieve the localization of arbitrary keypoints through similarity matching between support keypoint features and query image features. However, these methods primarily focus on mining information from the query images, neglecting the fact that support samples with keypoint annotations contain rich category-specific fine-grained semantic information and prior structural information. In this paper, we propose a Support-based Dynamic Perception Network (SDP-Net) for the robust and accurate CAPE. On the one hand, SDPNet models complex dependencies between support keypoints, constructing category-specific prior structure to guide the interaction of query keypoints. On the other hand, SDPNet extracts fine-grained semantic information from support samples, dynamically modulating the refinement process of query. Our method outperforms existing methods on MP-100 dataset by a large margin. Pengfei Ren 0001, Haifeng Sun 0001, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao |
CVPR | 1 |
| 2024 | Multi-Scale Video Anomaly Detection by Multi-Grained Spatio-Temporal Representation LearningabstractRecent progress in video anomaly detection suggests that the features of appearance and motion play crucial roles in distinguishing abnormal patterns from normal ones. However, we note that the effect of spatial scales of anomalies is ignored. The fact that many abnormal events occur in limited localized regions and severe background noise in-terferes with the learning of anomalous changes. Mean-while, most existing methods are limited by coarse-grained modeling approaches, which are inadequate for learning highly discriminative features to discriminate subtle differences between small-scale anomalies and normal patterns. To this end, this paper address multi-scale video anomaly detection by multi-grained spatiotemporal representation learning. We utilize video continuity to design three proxy tasks to perform feature learning at both coarse-grained and fine-grained levels, i.e., continuity judgment, discontinuity localization, and missing frame estimation. In particular, we formulate missing frame estimation as a contrastive learning task in feature space instead of a reconstruction task in RGB space to learn highly discriminative features. Experiments show that our proposed method outperforms state-of-the-art methods on four datasets, especially in scenes with small-scale anomalies. Menghao Zhang 0004, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Pengfei Ren 0001, Ruilong Ma, Jianxin Liao |
CVPR | 6 |
| 2024 | Coarse-to-Fine Implicit Representation Learning for 3D Hand-Object Reconstruction from a Single RGB-D Image
Pengfei Ren 0001, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Jianxin Liao |
ECCV (51) | 2 |
| 2024 | SO-Net: Model-Agnostic Sequential Hand Pose Optimization FrameworkabstractHand Pose Estimation (HPE) is a crucial technique for human-computer interaction perception. Recent works have shown that leveraging temporal information yields significant importance in the stability of the HPE system. However, existing sequential optimization methods are mainly designed for specific frameworks, limiting their applicability and generalization potential. To solve these problems, we propose a model-agnostic Sequential hand pose Optimization Network (SO-Net), which can be applicable to various HPE methods. Specifically, SO-Net first utilizes a feature-heterogeneous pre-embedding module that unifies multiple types of features in a coherent manner and facilitates their effective interactions. Then it adopts a transformer-based spatial-temporal network to capture the long-range information across frames. Finally, to enhance the robustness of the model, a dense refinement process is employed for further optimization. Our framework outperforms state-of-the-art methods on NYU and DexYCB datasets and also provides optimization advantages across various frameworks. Pengfei Ren 0001, Mingen Shu, Rui Chu, Jubiao Li, Jing Jin 0007 |
ICASSP | 2 |
| 2024 | GloveTyping: A Hand Gesture Recognition System for Text Input Using a Hierarchical Framework with Attention Mechanism
Tao Zhen, Pengfei Ren 0001, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
ICONIP (5) | 5 |
| 2024 | Safeguarding Sustainable Cities: Unsupervised Video Anomaly Detection through Diffusion-based Latent Pattern Learning
Menghao Zhang 0004, Jingyu Wang 0001, Qi Qi 0001, Pengfei Ren 0001, Haifeng Sun 0001, Zirui Zhuang, Lei Zhang 0094, Jianxin Liao |
IJCAI | 4 |
| 2024 | Enhanced Anomaly Detection in Dashcam Videos: Dual GAN Approach with Swin-Unet for Optical Flow and Region of Interest AnalysisabstractVideo anomaly detection plays a crucial role in the field of autonomous driving to ensure driving safety. Most existing video anomaly detection methods exhibit mediocre performance when analyzing video frames captured by dynamic cameras. To enhance their performance on dynamic cameras, we propose a video anomaly detection method based on Swin-Unet to separately predict optical flow and images cropped with ROI in the form of dual GAN. By integrating the predicted results of optical flow and ROI-cropped images, the model’s ability to learn dynamic information is significantly improved. The experimental results indicate that our proposed method outperforms state-of-the-art methods in terms of AUC on the Car Crash dataset and RetroTrucks dataset. Haodong Ru, Menghao Zhang 0004, Pengfei Ren 0001, Haifeng Sun 0001, Qi Qi 0001, Lejian Zhang, Jingyu Wang 0001 |
IJCNN | 4 |
| 2024 | Video Anomaly Detection via Progressive Learning of Multiple Proxy TasksabstractLearning multiple proxy tasks is a popular training strategy in semi-supervised video anomaly detection. However, the traditional method of learning multiple proxy tasks simultaneously is prone to suboptimal solutions, and simply executing multiple proxy tasks sequentially cannot ensure continuous performance improvement. In this paper, we thoroughly investigate the impact of task composition and training order on performance enhancement. We find that ensuring continuous performance improvement in multi-task learning requires different but continuous optimization objectives in different training phases. To this end, a training strategy based on progressive learning is proposed to enhance the multi-task learning in VAD. The learning objectives of the model in previous phases contribute to the training in subsequent phases. Specifically, we decompose video anomaly detection into three phases: perception, comprehension, and inference, continuously refining the learning objectives to enhance model performance. In the three phases, we perform the visual task, the semantic task and the open-set task in turn to train the model. The model learns different levels of features and focuses on different types of anomalies in different phases. Extensive experiments demonstrate the effectiveness of our method, highlighting that the benefits derived from the progressive learning transcend specific proxy tasks. Menghao Zhang 0004, Jingyu Wang 0001, Qi Qi 0001, Pengfei Ren 0001, Haifeng Sun 0001, Zirui Zhuang, Huazheng Wang, Lei Zhang 0094, Jianxin Liao |
ACM Multimedia | 4 |
| 2024 | Progressively global-local fusion with explicit guidance for accurate and robust 3d hand pose reconstruction
Kun Gao 0002, Pengfei Ren 0001, Tao Zhen, Liang Xie 0012, Zhongkui Li, Ye Yan 0001, Erwei Yin |
Knowl. Based Syst. | 3 |
| 2024 | SMR: Spatial-Guided Model-Based Regression for 3D Hand Pose and Mesh Reconstructionabstract3D hand reconstruction is an important technique for human-computer interaction. Interactive experience depends on the accuracy, efficiency, and robustness of the algorithm. Therefore, in this paper, we first propose a balanced framework called spatial-aware regression (SAR) to achieve precise and fast reconstruction. SAR can bridge convolutional networks and graph-structure networks more effectively than existing frameworks to fully exploit extracted spatial information using a novel spatial-aware initial graph building module. In addition, SAR uses adaptive-GCN to make keypoints interact efficiently and effectively; and regresses 2.5D belief maps to characterize uncertainty. SAR is highly flexible because it can predict an arbitrary number of keypoints and apply pose-guided refinement for coarse to fine regression. To produce more rational results for challenging cases and mitigate 3D label reliance, we also propose a more robust model-based framework called spatial-guided model-based regression (SMR) that is based on SAR. There are two critical designs of SMR: 1) it uses SAR to enhance the features with pose information to help the regression of hand model parameters; and 2) it regresses parameters in a spatially aware manner that is similar to SAR. Experiments demonstrate that the proposed frameworks surpass existing fully-supervised approaches on the FreiHAND, HO-3D, RHD, and STB datasets. Also, the performances of the proposed frameworks under weakly/self-supervised settings outperform other competitors. Meanwhile, the proposed frameworks are accurate and efficient. Haifeng Sun 0001, Xiaozheng Zheng, Pengfei Ren 0001, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Two Heads Are Better than One: Image-Point Cloud Network for Depth-Based 3D Hand Pose EstimationabstractDepth images and point clouds are the two most commonly used data representations for depth-based 3D hand pose estimation. Benefiting from the structuring of image data and the inherent inductive biases of the 2D Convolutional Neural Network (CNN), image-based methods are highly efficient and effective. However, treating the depth data as a 2D image inevitably ignores the 3D nature of depth data. Point cloud-based methods can better mine the 3D geometric structure of depth data. However, these methods suffer from the disorder and non-structure of point cloud data, which is computationally inefficient. In this paper, we propose an Image-Point cloud Network (IPNet) for accurate and robust 3D hand pose estimation. IPNet utilizes 2D CNN to extract visual representations in 2D image space and performs iterative correction in 3D point cloud space to exploit the 3D geometry information of depth data. In particular, we propose a sparse anchor-based "aggregation-interaction-propagation'' paradigm to enhance point cloud features and refine the hand pose, which reduces irregular data access. Furthermore, we introduce a 3D hand model to the iterative correction process, which significantly improves the robustness of IPNet to occlusion and depth holes. Experiments show that IPNet outperforms state-of-the-art methods on three challenging hand datasets. Pengfei Ren 0001, Jiachang Hao, Haifeng Sun 0001, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao |
AAAI | 1 |
| 2023 | Region-Aware Dynamic Filtering Network for 3D Hand Reconstructionabstract3D hand reconstruction from RGB image has attracted a lot of attention due to its crucial role in human-computer interaction. Nevertheless, it is still challenging to perform 3D hand reconstruction under conditions of hand-object interaction due to severe mutual occlusion. Previous methods usually adopt fixed convolution kernel to extract features. We argue that simply sharing the static filter for all regions is impertinent, given that the occlusion degree varies across different regions, resulting in inconsistent visual representations. To address this issue, we proposed Region-aware Dynamic Filtering Network (RDFNet), which dynamically generates convolution kernels based on the features of different regions, thereby adaptively extracting region-related information. Furthermore, we introduce a dynamic receptive field selection mechanism to determine the most appropriate scale for the convolution kernel. For the severely occluded regions, larger receptive field is needed to capture semantic-related features, while the visible regions are mainly concerned with their own local pattern to accumulate spatial-related features and avoid the interference of irrelevant information. Our proposed RDFNet outperforms state-of-the-art methods by a large margin on several challenging hand-object interaction datasets. Pengfei Ren 0001, Jingyu Wang 0001, Haifeng Sun 0001, Qi Qi 0001, Jing Wang 0039, Jianxin Liao |
ECAI | 2 |
| 2023 | Sample-Adapt Fusion Network for RGB-D Hand Detection in the WildabstractRGB and depth modalities provide complementary information, which can be effectively utilized to improve the performance of hand detection in the wild. Most existing fusion-based methods model the channel-wise or spatial-wise cross-modal correlation to exploit the complementary RGB-D information, in which the modeling operations are shared across all input samples. However, the input images show various modes due to the high diversity of scenes in the wild. This inter-sample variance cannot be effectively perceived by static modeling operations shared across all samples. To address this problem, we propose a Sample-Adapt Fusion Network (SAFNet) with Channel Dynamic Refinement Module (CDRM) and Spatial Dynamic Aggregation Module (SDAM) to adaptively model the channel-wise and spatial-wise cross-modal correlation. Specifically, we propose a Multi-kernel Attention Module (MAM) to individually generate attention maps for each input sample by applying learnable weighting operations to multiple convolutional kernels. Our method outperforms state-of-the-art methods on CUG Hand dataset. Pengfei Ren 0001, Cong Liu 0046, Jing Wang 0039, Haifeng Sun 0001, Qi Qi 0001, Jingyu Wang 0001 |
ICASSP | 2 |
| 2023 | Decoupled Iterative Refinement Framework for Interacting Hands Reconstruction from a Single RGB ImageabstractReconstructing interacting hands from a single RGB image is a very challenging task. On the one hand, severe mutual occlusion and similar local appearance between two hands confuse the extraction of visual features, resulting in the misalignment of estimated hand meshes and the image. On the other hand, there are complex spatial relationship between interacting hands, which significantly increases the solution space of hand poses and increases the difficulty of network learning. In this paper, we propose a decoupled iterative refinement framework to achieve pixel-alignment hand reconstruction while efficiently modeling the spatial relationship between hands. Specifically, we define two feature spaces with different characteristics, namely 2D visual feature space and 3D joint feature space. First, we obtain joint-wise features from the visual feature map and utilize a graph convolution network and a transformer to perform intra- and inter-hand information interaction in the 3D joint feature space, respectively. Then, we project the joint features with global information back into the 2D visual feature space in an obfuscation-free manner and utilize the 2D convolution for pixel-wise enhancement. By performing multiple alternate enhancements in the two feature spaces, our method can achieve an accurate and robust reconstruction of interacting hands. Our method outperforms all existing two-hand reconstruction methods by a large margin on the InterHand2.6M dataset. Pengfei Ren 0001, Xiaozheng Zheng, Zhou Xue, Haifeng Sun 0001, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao |
ICCV | 1 |
| 2023 | HaMuCo: Hand Pose Estimation via Multiview Collaborative Self-Supervised LearningabstractRecent advancements in 3D hand pose estimation have shown promising results, but its effectiveness has primarily relied on the availability of large-scale annotated datasets, the creation of which is a laborious and costly process. To alleviate the label-hungry limitation, we propose a self-supervised learning framework, HaMuCo, that learns a single-view hand pose estimator from multi-view pseudo 2D labels. However, one of the main challenges of self-supervised learning is the presence of noisy labels and the "groupthink" effect from multiple views. To overcome these issues, we introduce a cross-view interaction network that distills the single-view estimator by utilizing the cross-view correlated features and enforcing multi-view consistency to achieve collaborative learning. Both the single-view estimator and the cross-view interaction network are trained jointly in an end-to-end manner. Extensive experiments show that our method can achieve state-of-the-art performance on multi-view self-supervised hand pose estimation. Furthermore, the proposed cross-view interaction network can also be applied to hand pose estimation from multi-view input and outperforms previous methods under the same settings. Xiaozheng Zheng, Zhou Xue, Pengfei Ren 0001, Jingyu Wang 0001 |
ICCV | 4 |
| 2023 | SA-Fusion: Multimodal Fusion Approach for Web-based Human-Computer Interaction in the WildabstractWeb-based AR technology has broadened human-computer interaction scenes from traditional mechanical devices and flat screens to the real world, resulting in unconstrained environmental challenges such as complex backgrounds, extreme illumination, depth range differences, and hand-object interaction. The previous hand detection and 3D hand pose estimation methods are usually based on single modality such as RGB or depth data, which are not available in some scenarios in unconstrained environments due to the differences between the two modalities. To address this problem, we propose a multimodal fusion approach, named Scene-Adapt Fusion (SA-Fusion), which can fully utilize the complementarity of RGB and depth modalities in web-based HCI tasks. SA-Fusion can be applied in existing hand detection and 3D hand pose estimation frameworks to boost their performance, and can be further integrated into the prototyping AR system to construct a web-based interactive AR application for unconstrained environments. To evaluate the proposed multimodal fusion method, we conduct two user studies on CUG Hand and DexYCB dataset, to demonstrate its effectiveness in terms of accurately detecting hand and estimating 3D hand pose in unconstrained environments and hand-object interaction. Pengfei Ren 0001, Cong Liu 0046, Jing Wang 0039, Haifeng Sun 0001, Qi Qi 0001, Jingyu Wang 0001 |
WWW | 2 |
| 2023 | Pose-Guided Hierarchical Graph Reasoning for 3-D Hand Pose Estimation From a Single Depth ImageabstractEstimating 3-D hand pose estimation from a single depth image is important for human-computer interaction. Although depth-based 3-D hand pose estimation has made great progress in recent years, it is still difficult to deal with some complex scenes, especially the issues of serious self-occlusion and high self-similarity of fingers. Inspired by the fact that multipart context is critical to alleviate ambiguity, and constraint relations contained in the hand structure are important for the robust estimation, we attempt to explicitly model the correlations between different hand parts. In this article, we propose a pose-guided hierarchical graph convolution (PHG) module, which is embedded into the pixelwise regression framework to enhance the convolutional feature maps by exploring the complex dependencies between different hand parts. Specifically, the PHG module first extracts hierarchical fine-grained node features under the guidance of hand pose and then uses graph convolution to perform hierarchical message passing between nodes according to the hand structure. Finally, the enhanced node features are used to generate dynamic convolution kernels to generate hierarchical structure-aware feature maps. Our method achieves state-of-the-art performance or comparable performance with the state-of-the-art methods on five 3-D hand pose datasets: 1) HANDS 2019; 2) HANDS 2017; 3) NYU; 4) ICVL; and 5) MSRA. Pengfei Ren 0001, Haifeng Sun 0001, Jiachang Hao, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao |
IEEE Trans. Cybern. | 1 |
| 2023 | Fine-Grained Text-to-Video Temporal Grounding from Coarse BoundaryabstractText-to-video temporal grounding aims to locate a target video moment that semantically corresponds to the given sentence query in an untrimmed video. In this task, fully supervised works require text descriptions for each event along with its temporal segment coordinate for training, which is labor-consuming. Existing weakly supervised works require only video-sentence pairs but cannot achieve satisfactory performance. However, many available annotations in the form of coarse temporal boundaries for sentences are ignored and unexploited. These coarse boundaries are common in streaming media platform and can be collected in a mechanical manner. We propose a novel approach to perform fine-grained text-to-video temporal grounding from these coarse boundaries. We take dense video captioning as base task and leverage the trained captioning model to identify the relevance of each video frame to the sentence query according to the frame participation in event captioning. To quantify the frame participation in event captioning, we proposeevent activation sequence, a simple method that highlights the temporal regions which have high correlations to the text modality in videos. Experiments on modified ActivityNet Captions and a use case demonstrate the promising fine-grained performance of our approach. Jiachang Hao, Haifeng Sun 0001, Pengfei Ren 0001, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | Mining Multi-View Information: A Strong Self-Supervised Framework for Depth-based 3D Hand Pose and Mesh EstimationabstractIn this work, we study the cross-view information fusion problem in the task of self-supervised 3D hand pose estimation from the depth image. Previous methods usually adopt a hand-crafted rule to generate pseudo labels from multi-view estimations in order to supervise the network training in each view. However, these methods ignore the rich semantic information in each view and ignore the complex dependencies between different regions of different views. To solve these problems, we propose a cross-view fusion network to fully exploit and adaptively aggregate multi-view information. We encode diverse semantic information in each view into multiple compact nodes. Then, we introduce the graph convolution to model the complex dependencies between nodes and perform cross-view information interaction. Based on the cross-view fusion network, we propose a strong self-supervised framework for 3D hand pose and hand mesh estimation. Furthermore, we propose a pseudo multi-view training strategy to extend our framework to a more general scenario in which only single-view training data is used. Results on NYU dataset demonstrate that our method outperforms the previous self-supervised methods by 17.5% and 30.3% in multi-view and single-view scenarios. Meanwhile, our framework achieves comparable re-sults to several strongly supervised methods. Pengfei Ren 0001, Haifeng Sun 0001, Jiachang Hao, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao |
CVPR | 1 |
| 2022 | Can Shuffling Video Benefit Temporal Bias Problem: A Novel Training Framework for Temporal Grounding
Jiachang Hao, Haifeng Sun 0001, Pengfei Ren 0001, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao |
ECCV (36) | 3 |
| 2022 | Query-aware video encoder for video moment retrieval
Jiachang Hao, Haifeng Sun 0001, Pengfei Ren 0001, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao |
Neurocomputing | 3 |
| 2022 | A Dual-Branch Self-Boosting Framework for Self-Supervised 3D Hand Pose EstimationabstractAlthough 3D hand pose estimation has made significant progress in recent years with the development of the deep neural network, most learning-based methods require a large amount of labeled data that is time-consuming to collect. In this paper, we propose a dual-branch self-boosting framework for self-supervised 3D hand pose estimation from depth images. First, we adopt a simple yet effective image-to-image translation technology to generate realistic depth images from synthetic data for network pre-training. Second, we propose a dual-branch network to perform 3D hand model estimation and pixel-wise pose estimation in a decoupled way. Through a part-aware model-fitting loss, the network can be updated according to the fine-grained differences between the hand model and the unlabeled real image. Through an inter-branch loss, the two complementary branches can boost each other continuously during self-supervised learning. Furthermore, we adopt a refinement stage to better utilize the prior structure information in the estimated hand model for a more accurate and robust estimation. Our method outperforms previous self-supervised methods by a large margin without using paired multi-view images and achieves comparable results to strongly supervised methods. Besides, by adopting our regenerated pose annotations, the performance of the skeleton-based gesture recognition is significantly improved. Pengfei Ren 0001, Haifeng Sun 0001, Jiachang Hao, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao |
IEEE Trans. Image Process. | 1 |
| 2021 | Spatial Temporal Enhanced Contrastive and Pretext Learning for Skeleton-based Action RepresentationabstractIn this paper, we focus on unsupervised representation learning for skeleton-based action recognition. The critical issue of this task is extracting discriminative spatial-temporal information from skeleton sequences to form action representation. To better solve this, we propose a novel unsupervised framework named contrastive-pretext spatial-temporal network (CP-STN), aiming to achieve accurate action recognition by better exploiting discriminative spatial-temporal enhanced features from massive unlabeled data. We combine contrastive and pretext tasks learning paradigms in one framework by using asymmetric spatial and temporal augmentations to enable network extracting discriminative representations with spatial-temporal information fully. Furthermore, graph-based convolution is used as the backbone to explore natural spatial-temporal graph information in skeleton data. Extensive experimental results show that our CP-STN significantly boosts the performance of existing skeleton-based action representations learning networks and achieves state-of-the-art accuracy on two challenging benchmarks in both unsupervised and semi-supervised settings. Yiwen Zhan 0001, Pengfei Ren 0001, Haifeng Sun 0001, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao |
ACML | 3 |
| 2021 | Joint-Aware Regression: Rethinking Regression-Based Method for 3D Hand Pose Estimation
Xiaozheng Zheng, Pengfei Ren 0001, Haifeng Sun 0001, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao |
BMVC | 2 |
| 2021 | SAR: Spatial-Aware Regression for 3D Hand Pose and Mesh Reconstruction from a Monocular RGB Imageabstract3D hand reconstruction is a popular research topic in recent years, which has great potential for VR/AR applications. However, due to the limited computational resource of VR/AR equipment, the reconstruction algorithm must balance accuracy and efficiency to make the users have a good experience. Nevertheless, current methods are not doing well in balancing accuracy and efficiency. Therefore, this paper proposes a novel framework that can achieve a fast and accurate 3D hand reconstruction. Our framework relies on three essential modules, including spatial-aware initial graph building (SAIGB), graph convolutional network (GCN) based belief maps regression (GBBMR), and pose-guided refinement (PGR). At first, given image feature maps extracted by convolutional neural networks, SAIGB builds a spatial-aware and compact initial feature graph. Each node in this graph represents a vertex of the mesh and has vertex-specific spatial information that is helpful for accurate and efficient regression. After that, GBBMR first utilizes adaptive-GCN to introduce interactions between vertices to capture short-range and long-range dependencies between vertices efficiently and flexibly. Then, it maps vertices’ features to belief maps that can model the uncertainty of predictions for more accurate predictions. Finally, we apply PGR to compress the redundant vertices’ belief maps to compact-joints’ belief maps with the pose guidance and use these joints’ belief maps to refine previous predictions better to obtain more accurate and robust reconstruction results. Our method achieves state-of-the-art performance on four public benchmarks, FreiHAND, HO-3D, RHD, and STB. Moreover, our method can run at a speed of two to three times that of previous state-of-the-art methods. Our code is available at https://github.com/zxz267/SAR. Xiaozheng Zheng, Pengfei Ren 0001, Haifeng Sun 0001, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao |
ISMAR | 2 |
| 2021 | Spatial-aware stacked regression network for real-time 3D hand pose estimation
Pengfei Ren 0001, Haifeng Sun 0001, Weiting Huang, Jiachang Hao, Daixuan Cheng, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao |
Neurocomputing | 1 |
| 2020 | AWR: Adaptive Weighting Regression for 3D Hand Pose EstimationabstractIn this paper, we propose an adaptive weighting regression (AWR) method to leverage the advantages of both detection-based and regression-based method. Hand joint coordinates are estimated as discrete integration of all pixels in dense representation, guided by adaptive weight maps. This learnable aggregation process introduces both dense and joint supervision that allows end-to-end training and brings adaptability to weight maps, making network more accurate and robust. Comprehensive exploration experiments are conducted to validate the effectiveness and generality of AWR under various experimental settings, especially its usefulness for different types of dense representation and input modality. Our method outperforms other state-of-the-art methods on four publicly available datasets, including NYU, ICVL, MSRA and HANDS 2017 dataset. Weiting Huang, Pengfei Ren 0001, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001 |
AAAI | 2 |
| 2020 | Measuring Generalisation to Unseen Viewpoints, Articulations, Shapes and Objects for 3D Hand Pose Estimation Under Hand-Object Interaction
Anil Armagan, Guillermo Garcia-Hernando, Seungryul Baek, Shreyas Hampali, Mahdi Rad, Shipeng Xie, Mingxiu Chen, Boshen Zhang, Fu Xiong, Yang Xiao 0007, Zhiguo Cao 0001, Junsong Yuan 0001, Pengfei Ren 0001, Weiting Huang, Haifeng Sun 0001, Marek Hrúz, Jakub Kanis, Zdenek Krnoul, Qingfu Wan, Shile Li, Linlin Yang 0001, Dongheui Lee, Angela Yao, Weiguo Zhou, Sijia Mei, Adrian Spurr, Umar Iqbal 0001, Pavlo Molchanov 0001, Philippe Weinzaepfel, Romain Brégier, Grégory Rogez, Vincent Lepetit, Tae-Kyun Kim 0001 |
ECCV (23) | 14 |
| 2019 | SRN: Stacked Regression Network for Real-time 3D Hand Pose Estimation
Pengfei Ren 0001, Haifeng Sun 0001, Qi Qi 0001, Jingyu Wang 0001, Weiting Huang |
BMVC | 1 |