VLDB 2026 Research / reviewers in the wild / expert
Po-Chen Wu
dblp:12/10574
· DBLP profile ↗
20ranked-venue papers
7as first author
11since 2021 · last 2026
0009-0003-7204-5005ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Computer networks · 4 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-LEO Satellite Transmission with Terrestrial Coexistence: A Coverage-Enhanced and Interference-Mitigated Framework
Wang-Chuen Mao, Yu-Ting Li, Po-Chen Wu, Kai-Ten Feng, Feng Ouyang, Li-Hsiang Shen, Zhiguo Ding 0001, Jen-Ming Wu |
WCNC | 3 |
| 2026 | AERO: Adaptive Semi-One-Class Ratio-Fusion Learning for Device-Free Robust Wireless Sensing
Po-Chen Wu, Tzu-Hsun Huang, Zhong-Ting Tsai, Kai-Ten Feng, Zhi Ding 0001, Jen-Ming Wu |
WCNC | 1 |
| 2025 | PHD: Personalized 3D Human Body Fitting with Point DiffusionabstractWe introduce PHD, a novel approach for personalized 3D human mesh recovery (HMR) and body fitting that leverages user-specific shape information to improve pose estimation accuracy from videos. Traditional HMR methods are designed to be user-agnostic and optimized for generalization. While these methods often refine poses using constraints derived from the 2D image to improve alignment, this process compromises 3D accuracy by failing to jointly account for person-specific body shapes and the plausibility of 3D poses. In contrast, our pipeline decouples this process by first calibrating the user's body shape and then employing a personalized pose fitting process conditioned on that shape. To achieve this, we develop a body shape-conditioned 3D pose prior, implemented as a Point Diffusion Transformer, which iteratively guides the pose fitting via a Point Distillation Sampling loss. This learned 3D pose prior effectively mitigates errors arising from an over-reliance on 2D constraints. Consequently, our approach improves not only pelvis-aligned pose accuracy but also absolute pose accuracy -- an important metric often overlooked by prior work. Furthermore, our method is highly data-efficient, requiring only synthetic data for training, and serves as a versatile plug-and-play module that can be seamlessly integrated with existing 3D pose estimators to enhance their performance. Project page: https://phd-pose.github.io/ Hsuan-I Ho, Po-Chen Wu, Ivan Shugurov, Chengcheng Tang, Abhay Mittal, Sizhe An, Manuel Kaufmann, Linguang Zhang |
ICCV | 3 |
| 2025 | Multi-Head Reinforcement Learning Based Resource Allocation for Integrated Sensing and CommunicationsabstractIntegrated sensing and communication (ISAC) enables the simultaneous operation of communication and sensing functionalities, leading to improved efficiency and performance. Furthermore, multi-AP coordination (MAP-Co), which provides higher transmission coverage and capacity, has attracted significant attention. Therefore, in this paper, we propose an indoor ISAC scenario assisted by the MAP-Co scheme. By jointly optimizing transmit power, bandwidth, and subchannel selection, we aim to maximize transmission throughput while guaranteeing a minimum required human presence detection accuracy. To effectively solve the complicated task, we propose a novel multi-head self-attention-assisted reinforcement learning (MH-SARL) algorithm. In our algorithm, the communication head is designed to search for the optimal resource allocation policy to maximize transmission throughput, whilst the sensing head focuses on suggesting adequate spectrum utilization to satisfy the predefined sensing quality. Simulation results demonstrate the effectiveness of our proposed MH-SARL algorithm under different scenarios and parameter settings. Compared to other benchmarks in the open literature, MH-SARL can achieve the highest transmission rate of at least 5% increase with guaranteed sensing accuracy. Po-Chen Wu, Ting-Hui Wang, Li-Hsiang Shen, Kai-Ten Feng |
PIMRC | 1 |
| 2025 | Resource Allocation of Terrestrial-Satellite Service in Coexistence with Earth Exploration SatellitesabstractEarth exploration satellite service (EESS) plays a crucial role in environmental monitoring and weather forecasting by utilizing passive sensing technologies. However, the rapid expansion of terrestrial and satellite communication networks has introduced significant interference challenges, particularly in frequency bands that overlap with or are adjacent to EESS sensors. In this work, we develop a system model that explicitly characterizes EESS interference by considering reflected signal effects and spatial interference accumulation. Based on this model, we propose a EESS-aware resource allocation (EARA) framework that jointly optimizes power allocation and user association, while ensuring that interference to EESS sensors remains within acceptable limits. A non-convex joint optimization problem is formulated and efficiently solved leveraging the Lagrangian dual transform and Dinkelbach’s method. Simulation results demonstrate that the proposed EARA scheme achieves up to 26.3% higher sum rate compared to genetic algorithm and binary whale optimization algorithm, while strictly satisfying the ITU-defined interference threshold. This work establishes a foundation for future research on the coexistence of communication networks and passive Earth observation systems, offering practical strategies for interference mitigation and spectrum sharing in next-generation networks. Kai-Tse Wu, Po-Chen Wu, Li-Hsiang Shen, Kai-Ten Feng, Zhi Ding 0001, Jen-Ming Wu |
PIMRC | 2 |
| 2024 | Federated Reinforcement Learning for Multi-Dual-STAR-RIS Assisted DFRC-Enabled Multi-BS in ISAC SystemsabstractIntegrated sensing and communication (ISAC) has become a key technology in the sixth-generation (6G) wireless networks, catering to the growing need for ubiquitous sensing and communication tasks. Simultaneously transmitting and reflecting reconfigurable intelligent surface (STAR-RIS) can harness both reflective and refractive signals delivered. Due to orientation limitation of STAR-RISs, the multi-dual STAR-RISs (MD-STAR) is conceived to facilitate full-plane services in ISAC systems. In this paper, we intend to solve active beamforming of dual-function radar-communication (DFRC)-enabled BSs and passive beamforming of MD-STAR in ISAC systems, aiming for maximizing the achievable sum rate constrained by the maximum position error bound (PEB) as well as hardware limitation of MD-STAR. In order to solve this complex problem, we propose a two-layered multi-agent federated Q-learning (TMFQ) scheme. The inner layer Q-learning focuses on obtaining the solution of BSs and MD-STAR, whilst the outer layer Q-learning aims for optimizing the hyperparameters, including learning rate and discount rate of the inner-layer one. Additionally, we employ federated learning to facilitate information exchange between agents in the inner Q-learning. We evaluate our proposed TMFQ in terms of different numbers of MD-STAR elements, transmit antennas, and sensing targets. Benefiting from hyperparameter optimization of the inner layer Q-learning and information exchange of federated learning, the proposed TMFQ can achieve the highest rate compared to the other benchmarks, including Q-learning without hyperparameter optimization and without federated learning, heuristic algorithm, and conventional beamforming. Po-Chen Wu, Li-Hsiang Shen, Kai-Ten Feng, Ching-Yao Chan |
ICC | 1 |
| 2024 | Multi-Agent Deep Reinforcement Learning for Energy Efficient Multi-Hop STAR-RIS-Assisted TransmissionsabstractSimultaneously transmitting and reflecting reconfigurable intelligent surface (STAR-RIS) provides a promising way to expand coverage in wireless communications. However, limitation of single STAR-RIS inspire us to integrate the concept of multi-hop transmissions, as focused on RIS in existing research. Therefore, we propose the novel architecture of multi-hop STAR-RISs to achieve a wider range of full-plane service coverage. In this paper, we intend to solve active beamforming of the base station and passive beamforming of STAR-RISs, aiming for maximizing the energy efficiency constrained by hardware limitation of STAR-RISs. Furthermore, we investigate the impact of the on-off state of STAR-RIS elements on energy efficiency. To tackle the complex problem, a Multi-Agent Global and locAl deep Reinforcement learning (MAGAR) algorithm is designed. The global agent elevates the collaboration among local agents, which focus on individual learning. In numerical results, we observe the significant improvement of MAGAR compared to the other benchmarks, including Q-learning, multi-agent deep Q network (DQN) with golbal reward, and multi-agent DQN with local rewards. Moreover, the proposed architecture of multi-hop STAR-RISs achieves the highest energy efficiency compared to mode switching based STAR-RISs, conventional RISs and deployment without RISs or STAR-RISs. Pei-Hsiang Liao, Li-Hsiang Shen, Po-Chen Wu, Kai-Ten Feng |
VTC Fall | 3 |
| 2024 | D-STAR: Dual Simultaneously Transmitting and Reflecting Reconfigurable Intelligent Surfaces for Joint Uplink/Downlink TransmissionabstractThe joint uplink/downlink (JUD) design of simultaneously transmitting and reflecting reconfigurable intelligent surfaces (STAR-RIS) is conceived in support of both uplink (UL) and downlink (DL) users. Furthermore, the dual STAR-RISs (D-STAR) concept is conceived as a promising architecture for 360-degree full-plane service coverage, including UL/DL users located between the base station (BS) and the D-STAR as well as beyond. The corresponding regions are termed as primary (P) and secondary (S) regions. Both BS/users exist in the P-region, but only users are located in the S-region. The primary STAR-RIS (STAR-P) plays an important role in terms of tackling the P-region inter-user interference, the self-interference (SI) from the BS and from the reflective as well as refractive UL users imposed on the DL receiver. By contrast, the secondary STAR-RIS (STAR-S) aims for mitigating the S-region interferences. The non-linear and non-convex rate-maximization problem formulated is solved by alternating optimization amongst the decomposed convex sub-problems of the BS beamformer, and the D-STAR amplitude as well as phase shift configurations. We also propose a D-STAR based active beamforming and passive STAR-RIS amplitude/phase (DBAP) optimization scheme to solve the respective sub-problems by Lagrange dual with Dinkelbach’s transformation, alternating direction method of multipliers (ADMM) with successive convex approximation (SCA), and penalty convex-concave procedure (PCCP). Our simulation results reveal that the proposed D-STAR architecture outperforms the conventional single RIS, single STAR-RIS, and half-duplex networks. The proposed DBAP of D-STAR outperforms the state-of-the-art solutions found in the open literature for different numbers of quantization levels, geographic deployment, transmit power and for diverse numbers of transmit antennas, patch partitions as well as D-STAR elements. Li-Hsiang Shen, Po-Chen Wu, Chia-Jou Ku, Yu-Ting Li, Kai-Ten Feng, Yuanwei Liu, Lajos Hanzo |
IEEE Trans. Commun. | 2 |
| 2023 | Social Diffusion: Long-term Multiple Human Motion AnticipationabstractWe propose Social Diffusion, a novel method for short-term and long-term forecasting of the motion of multiple persons as well as their social interactions. Jointly forecasting motions for multiple persons involved in social activities is inherently a challenging problem due to the interdependencies between individuals. In this work, we leverage a diffusion model conditioned on motion histories and causal temporal convolutional networks to forecast individually and contextually plausible motions for all participants. The contextual plausibility is achieved via an order-invariant aggregation function. As a second contribution, we design a new evaluation protocol that measures the plausibility of social interactions which we evaluate on the Haggling dataset, which features a challenging social activity where people are actively taking turns to talk and switching their attention. We evaluate our approach on four datasets for multi-person forecasting where our approach outperforms the state-of-the-art in terms of motion realism and contextual plausibility. Julian Tanke, Linguang Zhang, Amy Zhao, Chengcheng Tang, Yujun Cai, Lezi Wang, Po-Chen Wu, Juergen Gall, Cem Keskin |
ICCV | 7 |
| 2022 | Neural Correspondence Field for Object Pose Estimation
Lin Huang 0004, Tomas Hodan, Lingni Ma, Linguang Zhang, Luan Tran, Christopher D. Twigg, Po-Chen Wu, Junsong Yuan 0001, Cem Keskin, Robert Wang 0002 |
ECCV (10) | 7 |
| 2022 | UmeTrack: Unified multi-view end-to-end hand tracking for VRabstractReal-time tracking of 3D hand pose in world space is a challenging problem and plays an important role in VR interaction. Existing work in this space are limited to either producing root-relative (versus world space) 3D pose or rely on multiple stages such as generating heatmaps and kinematic optimization to obtain 3D pose. Moreover, the typical VR scenario, which involves multi-view tracking from wide field of view (FOV) cameras is seldom addressed by these methods. In this paper, we present a unified end-to-end differentiable framework for multi-view, multi-frame hand tracking that directly predicts 3D hand pose in world space. We demonstrate the benefits of end-to-end differentiabilty by extending our framework with downstream tasks such as jitter reduction and pinch prediction. To demonstrate the efficacy of our model, we further present a new large-scale egocentric hand pose dataset that consists of both real and synthetic data. Experiments show that our system trained on this dataset handles various challenging interactive motions, and has been successfully applied to real-time VR applications. Shangchen Han, Po-Chen Wu, Linguang Zhang, Weiguang Si, Peizhao Zhang, Yujun Cai, Tomas Hodan, Randi Cabezas, Luan Tran, Muzaffer Akbay, Tsz-Ho Yu, Cem Keskin, Robert Wang 0002 |
SIGGRAPH Asia | 2 |
| 2018 | Direct pose estimation for planar objects
Po-Chen Wu, Hung-Yu Tseng, Ming-Hsuan Yang 0001, Shao-Yi Chien |
Comput. Vis. Image Underst. | 1 |
| 2017 | D-PET: A direct 6 DoF pose estimation and tracking system on graphics processing unitsabstractReal-time recovering an accurate 6 DoF pose of a known planar target is essential for augmented reality and robotics applications. Despite several pose estimation tracking systems have been proposed over recent years, there is still the need for a more efficient and more accurate solution for general planar objects. In this work, we develop an innovative GPU implementation of a real-time pose estimation and tracking system. It consists of a pose estimation unit and a pose tracker unit. While the former computes an initial pose of a target using direct method, the latter realizes accurate pose tracking with a hierarchical search scheme. Experiments on both synthetic and real datasets demonstrate that the proposed algorithm performs favorably with various planar targets. By implementing our method on an embedded GPU, the system achieves to work at 11 FPS and is suitable for real-time applications. Hung-Yu Tseng, Po-Chen Wu, Shao-Yi Chien |
ISCAS | 2 |
| 2017 | DodecaPen: Accurate 6DoF Tracking of a Passive StylusabstractWe propose a system for real-time six degrees of freedom (6DoF) tracking of a passive stylus that achieves sub-millimeter accuracy, which is suitable for writing or drawing in mixed reality applications. Our system is particularly easy to implement, requiring only a monocular camera, a 3D printed dodecahedron, and hand-glued binary square markers. The accuracy and performance we achieve are due to model-based tracking using a calibrated model and a combination of sparse pose estimation and dense alignment. We demonstrate the system performance in terms of speed and accuracy on a number of synthetic and real datasets, showing that it can be competitive with state-of-the-art multi-camera motion capture systems. We also demonstrate several applications of the technology ranging from 2D and 3D drawing in VR to general object manipulation and board games. Po-Chen Wu, Robert Wang 0002, Kenrick Kin, Christopher D. Twigg, Shangchen Han, Ming-Hsuan Yang 0001, Shao-Yi Chien |
UIST | 1 |
| 2016 | Direct 3D pose estimation of a planar targetabstractEstimating 3D pose of a known object from a given 2D image is an important problem with numerous studies for robotics and augmented reality applications. While the state-of-the-art Perspective-n-Point algorithms perform well in pose estimation, the success hinges on whether feature points can be extracted and matched correctly on targets with rich texture. In this work, we propose a robust direct method for 3D pose estimation with high accuracy that performs well on both textured and textureless planar targets. First, the pose of a planar target with respect to a calibrated camera is approximately estimated by posing it as a template matching problem. Next, the object pose is further refined and disambiguated with a gradient descent search scheme. Extensive experiments on both synthetic and real datasets demonstrate the proposed direct pose estimation algorithm performs favorably against state-of-the-art feature-based approaches in terms of robustness and accuracy under several varying conditions. Hung-Yu Tseng, Po-Chen Wu, Ming-Hsuan Yang 0001, Shao-Yi Chien |
WACV | 2 |
| 2015 | Real-time eye localization, blink detection, and gaze estimation system without infrared illuminationabstractGaze tracking systems have high potential to be used as natural user interface devices; however, the mainstream systems are designed with infrared illumination, which may be harmful for human eyes. In this paper, a real-time eye localization, blink detection, and gaze estimation system is proposed without infrared illumination. To deal with various lighting conditions and reflections on the iris, the proposed system is based on a continuously updated color model for robust iris detection. Moreover, the proposed algorithm employs both the simplified and the original eye images to achieve the balance between robustness and accuracy. Experimental results show that the proposed system can achieve the accuracy of 96.8% for blink detection and the accuracy of 1.973 degree for gaze estimation with the processing speed of 10-11fps. The performance is comparable to previous works with infrared illumination. Bo-Chun Chen, Po-Chen Wu, Shao-Yi Chien |
ICIP | 2 |
| 2014 | Stable pose tracking from a planar target with an analytical motion model in real-time applicationsabstractObject pose tracking from a camera is a well-developed method in computer vision. In theory, the pose can be determined uniquely from a calibrated camera. However, in practice, most real-time pose estimation algorithms experience pose ambiguity. We consider that pose ambiguity, i.e., the detection of two distinct local minima according to an error function, is caused by a geometric illusion. In this case, both ambiguous poses are plausible, but we cannot select the pose with the minimum error as the final pose. Thus, we developed a real-time algorithm for correct pose estimation for a planar target object using an analytical motion model. Our experimental results showed that the proposed algorithm effectively reduced the effects of pose jumping and pose jittering. To the best of our knowledge, this is the first approach to address the pose ambiguity problem using an analytical motion model in real-time applications. Po-Chen Wu, Yao-Hung Tsai, Shao-Yi Chien |
MMSP | 1 |
| 2012 | Stable Pose Estimation with a Motion Model in Real-Time ApplicationabstractEstimation of a object pose from camera is a well-developing topic in computer vision. In theory, the pose from a calibrated camera can be uniquely determined. But in practice, most of the real-time pose estimation algorithms suffer from pose ambiguity due to low accuracy of the target object. We think that pose ambiguity¡Xtwo distinct local minima of the according error function¡Xexist because of the phenomenon of geometric illusions. Both of the ambiguous poses are plausible. After obtaining the solution of two minima (pose candidates), we develop a real-time algorithm for stable pose estimation of a target objects with a motion model. In the experimental results, the proposed algorithm diminish the significance of pose jumping and pose jittering effectively. To the best of our knowledge, this is the first work to solve the pose ambiguity problem with motion model in real-time application. Po-Chen Wu, Jui-Hsin Lai, Ja-Ling Wu, Shao-Yi Chien |
ICME | 1 |
| 2012 | Tennis Real PlayabstractTennis Real Play (TRP) is an interactive tennis game system constructed with models extracted from videos of real matches. The key techniques proposed for TRP include player modeling and video-based player/court rendering. For player model creation, we propose the process for database normalization and the behavioral transition model of tennis players, which might be a good alternative for motion capture in the conventional video games. For player/court rendering, we propose the framework for rendering vivid game characters and providing the real-time ability. We can say that image-based rendering leads to a more interactive and realistic rendering. Experiments show that video games with vivid viewing effects and characteristic players can be generated from match videos without much user intervention. Because the player model can adequately record the ability and condition of a player in the real world, it can then be used to roughly predict the results of real tennis matches in the next days. The results of a user study reveal that subjects like the increased interaction, immersive experience, and enjoyment from playing TRP. Jui-Hsin Lai, Chieh-Li Chen, Po-Chen Wu, Chieh-Chi Kao, Min-Chun Hu 0001, Shao-Yi Chien |
IEEE Trans. Multim. | 3 |
| 2011 | Tennis real play: an interactive tennis game with models from real videosabstractTennis Real Play (TRP) is an interactive tennis game system constructed with models extracted from videos of real matches. The key techniques proposed for TRP include player modeling and video-based player/court rendering. For player model creation, we propose a database normalization process and a behavioral transition model of tennis players, which might be a good alternative for motion capture in the conventional video games. For player/court rendering, we propose a framework for rendering vivid game characters and providing the real-time ability. We can say that image-based rendering leads to a more interactive and realistic rendering. Experiments show that video games with vivid viewing effects and characteristic players can be generated from match videos without much user intervention. Because the player model can adequately record the ability and condition of a player in the real world, it can then be used to roughly predict the results of real tennis matches in the next days. The results of a user study reveal that subjects like the increased interaction, immersive experience, and enjoyment from playing TRP. Jui-Hsin Lai, Chieh-Li Chen, Po-Chen Wu, Chieh-Chi Kao, Shao-Yi Chien |
ACM Multimedia | 3 |