EDBT 2026 Demo / reviewers in the wild / expert
Jia-Wei Liu
dblp:85/3336
· DBLP profile ↗
21ranked-venue papers
3as first author
17since 2021 · last 2025
0000-0001-5438-2524ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 3 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 10 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Balanced Image Stylization with Style Matching Score
Liming Jiang 0001, Shuai Yang 0001, Jia-Wei Liu, Ivor W. Tsang, Zheng Shou 0001 |
ICCV | 4 |
| 2025 | Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person PerspectivesabstractWe present Ego-Exo4D, a diverse, large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g., sports, music, dance, bike repair). 740 participants from 13 cities worldwide performed these activities in 123 different natural scene contexts, yielding long-form captures from 1 to 42 minutes each and 1,286 hours of video combined. The multimodal nature of the dataset is unprecedented: the video is accompanied by multichannel audio, eye gaze, 3D point clouds, camera poses, IMU, and multiple paired language descriptions—including a novel “expert commentary” done by coaches and teachers and tailored to the skilled-activity domain. To push the frontier of first-person video understanding of skilled human activity, we also present a suite of benchmark tasks and their annotations, including fine-grained activity understanding, proficiency estimation, cross-view translation, and 3D hand/body pose. All resources are open sourced to fuel new research in the community. https://ego-exo4d-data.org/ Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Makoto Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, Eugene Byrne, Zachary Chavis, Joya Chen, Fu-Jen Chu, Sean Crane, Avijit Dasgupta, Jing Dong 0002, María Escobar, Cristhian Forigua, Abrham Gebreselasie, Sanjay Haresh, Jing Huang 0020, Md Mohaiminul Islam, Suyog Dutt Jain, Rawal Khirodkar, Devansh Kukreja, Kevin J. Liang, Jia-Wei Liu, Sagnik Majumder, Yongsen Mao, Effrosyni Mavroudi, Tushar Nagarajan, Francesco Ragusa, Santhosh K. Ramakrishnan, Luigi Seminara, Arjun Somayazulu, Yale Song, Shan Su, Zihui Xue, Jinxu Zhang, Angela Castillo, Changan Chen, Xinzhu Fu, Ryosuke Furuta, Cristina González, Prince Gupta, Jiabo Hu, Yifei Huang 0002, Yiming Huang 0011, Weslie Khoo, Anush Kumar, Robert Kuo, Sach Lakhavani, Miao Liu 0007, Mi Luo, Zhengyi Luo 0002, Brighid Meredith, Austin Miller, Oluwatumininu Oguntola, Xiaqing Pan, Penny Peng, Shraman Pramanick, Merey Ramazanova, Fiona Ryan, Kiran K. Somasundaram, Chenan Song, Audrey Southerland, Masatoshi Tateno, Takuma Yagi, Mingfei Yan, Xitong Yang, Zecheng Yu, Shengxin Cindy Zha, Chen Zhao 0002, Ziwei Zhao 0003, Zhifan Zhu 0001, Jeff Zhuo, Pablo Andrés Arbeláez, Gedas Bertasius, David Crandall, Dima Damen, Jakob J. Engel, Giovanni Maria Farinella, Antonino Furnari, Bernard Ghanem, Judy Hoffman, C. V. Jawahar, Richard A. Newcombe, Hyun Soo Park, James M. Rehg, Yoichi Sato 0001, Manolis Savva, Jianbo Shi, Mike Zheng Shout, Michael Wray |
Int. J. Comput. Vis. | 29 |
| 2025 | Correction: Instant3D: Instant Text-to-3D Generation
Ming Li 0073, Pan Zhou 0002, Jia-Wei Liu, Jussi Keppo, Shuicheng Yan, Xiangyu Xu 0002 |
Int. J. Comput. Vis. | 3 |
| 2025 | Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation
Junhao Zhang 0001, Jay Zhangjie Wu, Jia-Wei Liu, Rui Zhao 0001, Lingmin Ran, Yuchao Gu, Difei Gao, Zheng Shou 0001 |
Int. J. Comput. Vis. | 3 |
| 2025 | ColonNeRF: High-fidelity neural reconstruction of long colonoscopy
Yufei Shi 0003, Beijia Lu, Jia-Wei Liu, Ming Li 0073, Si Yong Yeo, Zheng Shou 0001 |
Neurocomputing | 3 |
| 2024 | VideoLLM-online: Online Video Large Language Model for Streaming VideoabstractRecent Large Language Models (LLMs) have been en-hanced with vision capabilities, enabling them to compre-hend images, videos, and interleaved vision-language con-tent. However, the learning methods of these large multi-modal models (LMMs) typically treat videos as predeter-mined clips, rendering them less effective and efficient at handling streaming video inputs. In this paper, we pro-pose a novel Learning-In- Video-Stream (LIVE) framework, which enables temporally aligned, long-context, and real-time dialogue within a continuous video stream. Our LIVE framework comprises comprehensive approaches to achieve video streaming dialogue, encompassing: (1) a training ob-jective designed to perform language modeling for contin-uous streaming inputs, (2) a data generation scheme that converts offline temporal annotations into a streaming di-alogue format, and (3) an optimized inference pipeline to speed up interactive chat in real-world video streams. With our LIVE framework, we develop a simplified model called VideoLLM-online and demonstrate its significant advan-tages in processing streaming videos. For instance, our VideoLLM-online-7B model can operate at over 10 FPS on an A100 GPU for a 5-minute video clip from Ego4D narration. Moreover, VideoLLM-online also showcases state-of-the-art performance on public offline video bench-marks, such as recognition, captioning, and forecasting. The code, model, data, and demo have been made available at showlab.github. iolvideollm-online. Joya Chen, Zhaoyang Lv, Qinghong Lin, Chenan Song, Difei Gao, Jia-Wei Liu, Ziteng Gao, Dongxing Mao, Zheng Shou 0001 |
CVPR | 7 |
| 2024 | Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person PerspectivesabstractWe present Ego-Exo4D, a diverse, large-scale multi-modal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured ego-centric and exocentric video of skilled human activities (e.g., sports, music, dance, bike repair). 740 participants from 13 cities worldwide performed these activities in 123 different natural scene contexts, yielding long-form captures from 1 to 42 minutes each and 1,286 hours of video combined. The multimodal nature of the dataset is un-precedented: the video is accompanied by multichannel audio, eye gaze, 3D point clouds, camera poses, IMU, and multiple paired language descriptions-including a novel “expert commentary” done by coaches and teachers and tailored to the skilled-activity domain. To push the frontier of first-person video understanding of skilled human activity, we also present a suite of benchmark tasks and their annotations, including fine-grained activity understanding, proficiency estimation, cross-view translation, and 3D hand/body pose. All resources are open sourced to fuel new research in the community. Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Makoto Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, Eugene Byrne, Zachary Chavis, Joya Chen, Fu-Jen Chu, Sean Crane, Avijit Dasgupta, Jing Dong 0002, María Escobar, Cristhian Forigua, Abrham Gebreselasie, Sanjay Haresh, Jing Huang 0020, Md Mohaiminul Islam, Suyog Dutt Jain, Rawal Khirodkar, Devansh Kukreja, Kevin J. Liang, Jia-Wei Liu, Sagnik Majumder, Yongsen Mao, Effrosyni Mavroudi, Tushar Nagarajan, Francesco Ragusa, Santhosh K. Ramakrishnan, Luigi Seminara, Arjun Somayazulu, Yale Song, Shan Su, Zihui Xue, Jinxu Zhang, Angela Castillo, Changan Chen, Xinzhu Fu, Ryosuke Furuta, Cristina González, Prince Gupta, Jiabo Hu, Yifei Huang 0002, Yiming Huang 0011, Weslie Khoo, Anush Kumar, Robert Kuo, Sach Lakhavani, Miao Liu 0007, Mi Luo, Zhengyi Luo 0002, Brighid Meredith, Austin Miller, Oluwatumininu Oguntola, Xiaqing Pan, Penny Peng, Shraman Pramanick, Merey Ramazanova, Fiona Ryan, Kiran K. Somasundaram, Chenan Song, Audrey Southerland, Masatoshi Tateno, Takuma Yagi, Mingfei Yan, Xitong Yang, Zecheng Yu, Shengxin Cindy Zha, Chen Zhao 0002, Ziwei Zhao 0003, Zhifan Zhu 0001, Jeff Zhuo, Pablo Andrés Arbeláez, Gedas Bertasius, Dima Damen, Jakob J. Engel, Giovanni Maria Farinella, Antonino Furnari, Bernard Ghanem, Judy Hoffman, C. V. Jawahar, Richard A. Newcombe, Hyun Soo Park, James M. Rehg, Yoichi Sato 0001, Manolis Savva, Jianbo Shi, Mike Zheng Shout, Michael Wray |
CVPR | 29 |
| 2024 | VideoSwap: Customized Video Subject Swapping with Interactive Semantic Point CorrespondenceabstractCurrent diffusion-based video editing primarily focuses on structure-preserved editing by utilizing various dense correspondences to ensure temporal consistency and motion alignment. However, these approaches are often in-effective when the target edit involves a shape change. To embark on video editing with shape change, we explore customized video subject swapping in this work, where we aim to replace the main subject in a source video with a target subject having a distinct identity and potentially different shape. In contrast to previous methods that rely on dense correspondences, we introduce the Video Swap framework that exploits semantic point correspondences, inspired by our observation that only a small number of semantic points are necessary to align the subject's motion trajectory and modify its shape. We also introduce various user-point interactions (e.g., removing points and dragging points) to address various semantic point correspondence. Extensive experiments demonstrate state-of-the-art video subject swapping results across a variety of real-world videos. Yuchao Gu, Yipin Zhou, Bichen Wu, Licheng Yu, Jia-Wei Liu, Rui Zhao 0001, Jay Zhangjie Wu, Junhao Zhang 0001, Zheng Shou 0001, Kevin Tang |
CVPR | 5 |
| 2024 | DynVideo-E: Harnessing Dynamic NeRF for Large-Scale Motion- and View-Change Human-Centric Video EditingabstractDespite recent progress in diffusion-based video editing, existing methods are limited to short-length videos due to the contradiction between long-range consistency and frame-wise editing. Prior attempts to address this challenge by introducing video-2D representations encounter significant difficulties with large motion- and view-change videos, especially in human-centric scenarios. To overcome this, we propose to introduce the dynamic Neural Radiance Fields (NeRF) as the innovative video representation, where the editing can be performed in the 3D spaces and propagated to the entire video via the deformation field. To provide consistent and controllable editing, we propose the image-based video-NeRF editing pipeline with a set of innovative designs, including multi-view multi-pose Score Distillation Sampling (SDS) from both the 2D personalized diffusion prior and 3D diffusion prior, reconstruction losses, text-guided local parts super-resolution, and style transfer. Extensive experiments demonstrate that our method dubbed as DynVideo-E, significantly outperforms SOTA approaches on two challenging datasets by a large margin of 50% ~ 95% for human preference. Code will be released at https://showlab.github.io/DynVideo-E/. Jia-Wei Liu, Yan-Pei Cao 0001, Jay Zhangjie Wu, Weijia Mao, Yuchao Gu, Rui Zhao 0001, Jussi Keppo, Ying Shan, Zheng Shou 0001 |
CVPR | 1 |
| 2024 | X- Adapter: Universal Compatibility of Plugins for Upgraded Diffusion ModelabstractWe introduce X-Adapter, a universal upgrader to enable the pretrained plug-and-play modules (e.g., ControlNet, LoRA) to work directly with the upgraded text-to-image diffusion model (e.g., SDXL) without further retraining. We achieve this goal by training an additional network to control the frozen upgraded model with the new text-image data pairs. In detail, X-Adapter keeps a frozen copy of the old model to preserve the connectors of different plugins. Additionally, X-Adapter adds trainable mapping layers that bridge the decoders from models of different versions for feature remapping. The remapped features will be used as guidance for the upgraded model. To enhance the guidance ability of X-Adapter, we employ a null-text training strategy for the upgraded model. After training, we also introduce a two-stage denoising strategy to align the initial latents of X Adapter and the upgraded model. Thanks to our strategies, X-Adapter demonstrates universal compatibility with various plugins and also enables plugins of different versions to work together, thereby expanding the functionalities of diffusion community. To verify the effectiveness of the proposed method, we conduct extensive experiments and the results show that X-Adapter may facilitate wider application in the upgraded foundational diffusion model. Project page at: https://showlab.github.io/X-Adapter/. Lingmin Ran, Xiaodong Cun, Jia-Wei Liu, Rui Zhao 0001, Song Zijie, Xintao Wang 0002, Jussi Keppo, Zheng Shou 0001 |
CVPR | 3 |
| 2024 | MagicAnimate: Temporally Consistent Human Image Animation using Diffusion ModelabstractThis paper studies the human image animation task, which aims to generate a video of a certain reference iden-tity following a particular motion sequence. Existing an-imation works typically employ the frame-warping technique to animate the reference image towards the target motion. Despite achieving reasonable results, these approaches face challenges in maintaining temporal consistency throughout the animation due to the lack of temporal modeling and poor preservation of reference identity. In this work, we introduce Magic/snimate, a diffusion-based framework that aims at enhancing temporal consistency, preserving reference image faithfully, and improving animation fidelity. To achieve this, we first develop a video diffusion model to encode temporal information. Second, to maintain the appearance coherence across frames, we introduce a novel appearance encoder to retain the intricate details of the reference image. Leveraging these two inno-vations, we further employ a simple video fusion technique to encourage smooth transitions for long video animation. Empirical results demonstrate the superiority of our method over baseline approaches on two benchmarks. Notably, our approach outperforms the strongest baseline by over 38% in terms of video fidelity on the challenging TikTok dancing dataset. Code and model will be made available at https://showlab.github.io/magicanimate. Zhongcong Xu, Jun Hao Liew, Hanshu Yan, Jia-Wei Liu, Jiashi Feng, Zheng Shou 0001 |
CVPR | 5 |
| 2024 | MotionDirector: Motion Customization of Text-to-Video Diffusion Models
Rui Zhao 0001, Yuchao Gu, Jay Zhangjie Wu, Junhao Zhang 0001, Jia-Wei Liu, Weijia Wu 0001, Jussi Keppo, Zheng Shou 0001 |
ECCV (56) | 5 |
| 2024 | Exocentric-to-Egocentric Video GenerationabstractWe introduce Exo2Ego-V, a novel exocentric-to-egocentric diffusion-based video generation method for daily-life skilled human activities where sparse 4-view exocentric viewpoints are configured 360° around the scene. This task is particularly challenging due to the significant variations between exocentric and egocentric viewpoints and high complexity of dynamic motions and real-world daily-life environments. To address these challenges, we first propose a new diffusion-based multi-view exocentric encoder to extract the dense multi-scale features from multi-view exocentric videos as the appearance conditions for egocentric video generation. Then, we design an exocentric-to-egocentric view translation prior to provide spatially aligned egocentric features as a concatenation guidance for the input of egocentric video diffusion model. Finally, we introduce the temporal attention layers into our egocentric video diffusion pipeline to improve the temporal consistency cross egocentric frames. Extensive experiments demonstrate that Exo2Ego-V significantly outperforms SOTA approaches on 5 categories from the Ego-Exo4D dataset with an average of 35% in terms of LPIPS. Our code and model will be made available on https://github.com/showlab/Exo2Ego-V. Jia-Wei Liu, Weijia Mao, Zhongcong Xu, Jussi Keppo, Zheng Shou 0001 |
NeurIPS | 1 |
| 2024 | Skinned Motion Retargeting with Dense Geometric Interaction PerceptionabstractCapturing and maintaining geometric interactions among different body parts is crucial for successful motion retargeting in skinned characters. Existing approaches often overlook body geometries or add a geometry correction stage after skeletal motion retargeting. This results in conflicts between skeleton interaction and geometry correction, leading to issues such as jittery, interpenetration, and contact mismatches. To address these challenges, we introduce a new retargeting framework, MeshRet, which directly models the dense geometric interactions in motion retargeting. Initially, we establish dense mesh correspondences between characters using semantically consistent sensors (SCS), effective across diverse mesh topologies. Subsequently, we develop a novel spatio-temporal representation called the dense mesh interaction (DMI) field. This field, a collection of interacting SCS feature vectors, skillfully captures both contact and non-contact interactions between body geometries. By aligning the DMI field during retargeting, MeshRet not only preserves motion semantics but also prevents self-interpenetration and ensures contact preservation. Extensive experiments on the public Mixamo dataset and our newly-collected ScanRet dataset demonstrate that MeshRet achieves state-of-the-art performance. Code available at https://github.com/abcyzj/MeshRet. Zijie Ye, Jia-Wei Liu, Shikun Sun, Zheng Shou 0001 |
NeurIPS | 2 |
| 2024 | Instant3D: Instant Text-to-3D Generation
Ming Li 0073, Pan Zhou 0002, Jia-Wei Liu, Jussi Keppo, Shuicheng Yan, Xiangyu Xu 0002 |
Int. J. Comput. Vis. | 3 |
| 2023 | STPrivacy: Spatio-Temporal Privacy-Preserving Action RecognitionabstractExisting methods of privacy-preserving action recognition (PPAR) mainly focus on frame-level (spatial) privacy removal through 2D CNNs. Unfortunately, they have two major drawbacks. First, they may compromise temporal dynamics in input videos, which are critical for accurate action recognition. Second, they are vulnerable to practical attacking scenarios where attackers probe for privacy from an entire video rather than individual frames. To address these issues, we propose a novel framework STPrivacy to perform video-level PPAR. For the first time, we introduce vision Transformers into PPAR by treating a video as a tubelet sequence, and accordingly design two complementary mechanisms, i.e., sparsification and anonymization, to remove privacy from a spatio-temporal perspective. In specific, our privacy sparsification mechanism applies adaptive token selection to abandon action-irrelevant tubelets. Then, our anonymization mechanism implicitly manipulates the remaining action-tubelets to erase privacy in the embedding space through adversarial learning. These mechanisms provide significant advantages in terms of privacy preservation for human eyes and action-privacy trade-off adjustment during deployment. We additionally contribute the first two large-scale PPAR benchmarks, VP-HMDB51 and VP-UCF101, to the community. Extensive evaluations on them, as well as two other tasks, validate the effectiveness and generalization capability of our framework. Ming Li 0073, Xiangyu Xu 0002, Hehe Fan, Pan Zhou 0002, Jun Liu 0036, Jia-Wei Liu, Jiahe Li 0009, Jussi Keppo, Zheng Shou 0001, Shuicheng Yan |
ICCV | 6 |
| 2023 | HOSNeRF: Dynamic Human-Object-Scene Neural Radiance Fields from a Single VideoabstractWe introduce HOSNeRF, a novel 360° free-viewpoint rendering method that reconstructs neural radiance fields for dynamic human-object-scene from a single monocular in-the-wild video. Our method enables pausing the video at any frame and rendering all scene details (dynamic humans, objects, and backgrounds) from arbitrary viewpoints. The first challenge in this task is the complex object motions in human-object interactions, which we tackle by introducing the new object bones into the conventional human skeleton hierarchy to effectively estimate large object deformations in our dynamic human-object model. The second challenge is that humans interact with different objects at different times, for which we introduce two new learnable object state embeddings that can be used as conditions for learning our human-object representation and scene representation, respectively. Extensive experiments show that HOSNeRF significantly outperforms SOTA approaches on two challenging datasets by a large margin of 40%~50% in terms of LPIPS. The code, data, and compelling examples of 360° free-viewpoint renderings from single videos: https://showlab.github.io/HOSNeRF. Jia-Wei Liu, Yan-Pei Cao 0001, Tianyuan Yang, Zhongcong Xu, Jussi Keppo, Ying Shan, Xiaohu Qie, Zheng Shou 0001 |
ICCV | 1 |
| 2017 | Designing an exergaming system for exercise bikes using kinect sensors and Google Earth
Shih-Yu Huang, Jen-Perng Yu, Jia-Wei Liu |
Multim. Tools Appl. | 4 |
| 2012 | A Low Complexity Interference Suppression Scheme for High Mobility STBC-OFDM SystemsabstractA novel low complexity scheme based on received data symbol pair reordering and optimal selection, is proposed for space time block coded (STBC) OFDM systems to suppress co-channel interference (CCI) and inter-carrier interference (ICI) caused by time-varying channels. The proposed scheme reorders the conventional received data symbol pairs to produce extra received data symbol pairs. Then, the STBC decoding is performed on all the received data symbol pairs, and the STBC decoding output signals are used to obtain multiple combined signals. Finally, the optimal combined signal with minimum interference is selected to suppress CCI and ICI, and the system performance will be improved as long as the channels are mutually uncorrelated. Our computer simulation results verify that the proposed low complexity scheme achieves good performance for high mobility 2×2 STBC-OFDM systems in the ITU Vehicular A (VA) channels. Chorng-Ren Sheu, Jia-Wei Liu, Chuan-Yuan Huang, Chia-Chi Huang |
VTC Spring | 2 |
| 2010 | A Low Complexity ICI Cancellation Scheme with Multi-Step Windowing and Modified SIC for High-Mobility OFDM SystemsabstractA novel low complexity scheme consists of multi-step windowing and modified SIC, is proposed for OFDM systems to reduce ICI caused by time-varying channel. Assuming that the maximum channel delay is smaller than the guard interval, a multi-step windowing module implemented by suitably shifting and summing the multiple weighted ISI-free received signal segments, is adopted to automatically alleviate ICI effect. Afterwards, a modified SIC equalizer with overlapped multi-block and parallel processing is adopted to further improve the BER performance and simultaneously reduce the iteration number for the SIC subblocks. Moreover, this work shows that the proposed windowing concentrates the ICI effect and determines the number of dominant ICI terms, and hence it helps to reduce the computation complexity in ICI reconstruction & cancellation. Computer simulation results verify that the proposed low complexity method achieves good performance with moderate iteration number for high mobility OFDM systems in the ITU Vehicular A channel. Chorng-Ren Sheu, Jia-Wei Liu, Chia-Chi Huang |
VTC Spring | 2 |
| 2008 | Frame synchronization, channel estimation scheme and signal compensation using regression method in OFDM systems
Jyh-Horng Wen, Gwo-Ruey Lee, Jia-Wei Liu, Te-Lung Kung |
Comput. Commun. | 3 |