Zhensong Zhang

dblp:130/7882 · DBLP profile ↗
← Back
33ranked-venue papers
3as first author
20since 2021 · last 2026
0009-0001-7911-7564ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 2 first-author · 14 since 2021Artificial intelligence and machine learning · 17 · 15 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SCENIC: Scene-Aware Semantic Navigation with Instruction-Guided Control
abstract
Synthesizing natural human motion that adapts to complex environments while allowing creative control remains a fundamental challenge in motion synthesis. Existing models often fall short, either by assuming flat terrain or lacking the ability to control motion semantics through text. To address these limitations, we introduce SCENIC, a diffusion model designed to generate human motion that adapts to dynamic terrains within virtual scenes while enabling semantic control through natural language. The key technical challenge lies in simultaneously reasoning about complex scene geometry while maintaining text control. This requires understanding both high-level navigation goals and fine-grained environmental constraints. The model must ensure physical plausibility and precise navigation across varied terrain, while also preserving user-specified text control, such as “carefully stepping over obstacles” or “walking upstairs like a zombie.” Our solution introduces a hierarchical scene reasoning approach. At the core of our method is a novel hierarchical scene reasoning framework. It combines two key components: a motion-scene cross-attention block that aligns the human body's motion features with local scene geometry, enabling precise low-level interactions; and a target point canonicalization module that provides global goal conditioning by normalizing target scene coordinates for high-level guidance. To ensure plausibility and naturalness, we leverage a pre-trained motion diffusion prior and apply scene-constrained diffusion noise optimization during sampling, enabling long-horizon motion generation that respects both scene structure and semantic text input. Experiments demonstrate that our novel diffusion model generates arbitrarily long human motions that both adapt to complex scenes with varying terrain surfaces and respond to textual prompts. Additionally, we show SCENIC can generalize to four real-scene datasets.
Sebastian Starke, Vladimir Guzov, Zhensong Zhang, Eduardo Pérez-Pellitero, Gerard Pons-Moll
3DV4
2026 Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent Disambiguation
abstract
The performance of egocentric AI agents is fundamentally limited by multimodal intent ambiguity. This challenge arises from a combination of underspecified language, imperfect visual data, and deictic gestures, which frequently leads to task failure. Existing monolithic Vision-Language Models (VLMs) struggle to resolve these multimodal ambiguous inputs, often failing silently or hallucinating responses. To address these ambiguities, we introduce the Plug-and-Play Clarifier, a zero-shot and modular framework that decomposes the problem into discrete, solvable sub-tasks. Specifically, our framework consists of three synergistic modules: (1) a text clarifier that uses dialogue-driven reasoning to interactively disambiguate linguistic intent, (2) a vision clarifier that delivers real-time guidance feedback, instructing users to adjust their positioning for improved capture quality, and (3) a cross-modal clarifier with grounding mechanism that robustly interprets 3D pointing gestures and identifies the specific objects users are pointing to. Extensive experiments demonstrate that our framework improves the intent clarification performance of small language models (4-8B) by approximately 30%, making them competitive with significantly larger counterparts. We also observe consistent gains when applying our framework to these larger models. Furthermore, our vision clarifier increases corrective guidance accuracy by over 20%, and our cross-modal clarifier improves semantic answer accuracy for referential grounding by 5%. Overall, our method provides a plug-and-play framework that effectively resolves multimodal ambiguity and significantly enhances user experience in egocentric interaction.
Weitong Cai, Shitong Sun, You He 0003, Jiankang Deng, Hang Zhang 0010, Jifei Song, Zhensong Zhang
AAAI9
2026 Egocentric Co-Pilot: Web-Native Smart-Glasses Agents for Assistive Egocentric AI
abstract
What if accessing the web did not require a screen, a stable desk, or even free hands? For people navigating crowded cities, living with low vision, or experiencing cognitive overload, smart glasses coupled with AI agents could turn the web into an always-on assistive layer over daily life. We present Egocentric Co-Pilot, a web-native neuro-symbolic framework that runs on smart glasses and uses a Large Language Model (LLM) to orchestrate a toolbox of perception, reasoning, and web tools. An egocentric reasoning core combines Temporal Chain-of-Thought with Hierarchical Context Compression to support long-horizon question answering and decision support over continuous first-person video, far beyond a single model's context window. Additionally, a lightweight multimodal intent layer maps noisy speech and gaze into structured commands. We further implement and evaluate a cloud-native WebRTC pipeline integrating streaming speech, video, and control messages into a unified channel for smart glasses and browsers. In parallel, we deploy an on-premise WebSocket baseline, exposing concrete trade-offs between local inference and cloud offloading in terms of latency, mobility, and resource use. Experiments on Egolife and HD-EPIC demonstrate competitive or state-of-the-art egocentric QA performance, and a human-in-the-loop study on smart glasses shows higher task completion and user satisfaction than leading commercial baselines. Taken together, these results indicate that web-connected egocentric co-pilots can be a practical path toward more accessible, context-aware assistance in everyday life. By grounding operation in web-native communication primitives and modular, auditable tool use, Egocentric Co-Pilot offers a concrete blueprint for assistive, always-on web agents that support education, accessibility, and social inclusion for people who may benefit most from contextual, egocentric AI.
Weitong Cai, Shitong Sun, Fengyi Fang, You He 0003, Yiqiao Xie, Jiankang Deng, Hang Zhang 0010, Jifei Song, Zhensong Zhang
WWW11
2026 An Interactive Conversational 3D Virtual Human
Richard Shaw, Youngkyoon Jang, Athanasios Papaioannou, Arthur Moreau, Helisa Dhamo, Zhensong Zhang, Eduardo Pérez-Pellitero
Int. J. Comput. Vis.6
2025 CaricatureBooth: Data-Free Interactive Caricature Generation in a Photo Booth
abstract
We present CaricatureBooth, a system that transforms caricature creation into a simple interactive experience – as easy as using a photo booth! A key challenge in caricature generation is two-fold: the scarcity of high-quality caricature data and the difficulty in enabling precise creative control over the exaggeration process while maintaining identity. Prior approaches either require large-scale caricature and photo data or lack intuitive mechanisms for users to guide the deformation without losing identity. We address the data scarcity by synthesising training data through Thin Plate Spline (TPS) deformation of standard face images. For creative control, we design a Bézier curve interface where users can easily manipulate facial features, with these edits then driving TPS transformations at inference time. When combined with a pre-trained ID-preserving diffusion model, our system maintains both identity preservation and creative flexibility. Through extensive experiments, we demonstrate that CaricatureBooth achieves state-of-the-art quality while making the joy of caricature creation as accessible as taking a photo – just walk in and walk out with your personalised caricature! Code is available at https://github.com/WinKawaks/CaricatureBooth.
Zhiyu Qu, Yunqi Miao, Zhensong Zhang, Jifei Song, Jiankang Deng, Yi-Zhe Song
CVPR3
2025 Identity-Preserving Audio-Driven Holistic Human Motion Video Generation
abstract
Generating realistic human motion videos is a pivotal challenge in advancing human-computer interaction. While existing approaches often focus on generating either head or gesture movements from audio, they lack unified control over full-body motion, frequently producing low-resolution and blurred outputs. Additionally, these methods struggle to maintain character identity throughout the generated content. In this paper, we introduce a novel framework that generates photorealistic, personalized human motion videos from audio by decoupling identity features. We integrate both visual features and voice timbre to enhance the preservation of character identity. Our approach follows a four-stage paradigm: (1) frame generation, (2) identity feature customization, (3) audio-motion modeling, and (4) motion-video rendering. Through the collaborative modeling of audio-motion and motion-video stages, our approach effectively maintains the consistency of character identity and background throughout the video, enhancing the realism and coherence of the generated video. Experimental results demonstrate that our framework delivers high-resolution videos with superior fidelity, establishing a new and effective baseline for holistic human motion video generation.
Haiwei Xue, Zhensong Zhang, Minglei Li 0001, Zonghong Dai, Zhiyong Wu 0001
ICASSP2
2025 Frequency-Guided Diffusion for Training-Free Text-Driven Image Translation
Zheng Gao 0003, Jifei Song, Zhensong Zhang, Jiankang Deng, Ioannis Patras
ICCV3
2025 VideoHumanMIB: Unlocking Appearance Decoupling for Video Human Motion In-betweening
abstract
We propose VideoHumanMIB, a novel framework for Video Human Motion In-betweening that enables seamless transitions between different motion video clips, facilitating the generation of longer and more natural digital human videos. While existing video frame interpolation methods work well for similar motions in adjacent frames, they often struggle with complex human movements, resulting in artifacts and unrealistic transitions. To address these challenges, we introduce a two-stage approach: First, we design an Appearance Reconstruction AutoEncoder to decouple appearance and motion information, extracting robust appearance-invariant features. Second, we develop an enhanced diffusion pretrained network that leverages both motion optical flow and human pose as guidance conditions, enabling the model to learn comprehensive latent distributions of possible motions. Rather than operating directly in pixel space, our model works in a learned latent space, allowing it to better capture the underlying motion dynamics. The framework is optimized with a dual-frame constraint loss and a motion flow loss to ensure temporal consistency and natural movement transitions. Extensive experiments demonstrate that our approach generates highly realistic transition sequences that significantly outperform existing methods, particularly in challenging scenarios with large motion variations. The proposed VideoHumanMIB establishes a new baseline for human motion synthesis and enables more natural and controllable digital human animation.
Haiwei Xue, Zhensong Zhang, Minglei Li 0001, Zonghong Dai, F. Richard Yu, Fei Ma 0006, Zhiyong Wu 0001
IJCAI2
2025 ViDAR: Video Diffusion-Aware 4D Reconstruction From Monocular Inputs
abstract
Dynamic Novel View Synthesis aims to generate photorealistic views of moving subjects from arbitrary viewpoints. This task is particularly challenging when relying on monocular video, where disentangling structure from motion is ill-posed and supervision is scarce. We introduce Video Diffusion-Aware Reconstruction (ViDAR), a novel 4D reconstruction framework that leverages personalised diffusion models to synthesise a pseudo multi-view supervision signal for training a Gaussian splatting representation. By conditioning on scene-specific features, ViDAR recovers fine-grained appearance details while mitigating artefacts introduced by monocular ambiguity. To address the spatio-temporal inconsistency of diffusion-based supervision, we propose a diffusion-aware loss function and a camera pose optimisation strategy that aligns synthetic views with the underlying scene geometry. Experiments on DyCheck, a challenging benchmark with extreme viewpoint variation, show that ViDAR outperforms all state-of-the-art baselines in visual quality and geometric consistency. We further highlight ViDAR’s strong improvement over baselines on dynamic regions and provide a new benchmark to compare performance in reconstructing motion-rich parts of the scene.
Michal Nazarczuk, Sibi Catley-Chandar, Thomas Tanay, Zhensong Zhang, Gregory Slabaugh, Eduardo Pérez-Pellitero
NeurIPS4
2025 Human Motion Video Generation: A Survey
abstract
Human motion video generation has garnered significant research interest due to its broad applications, enabling innovations such as photorealistic singing heads or dynamic avatars that seamlessly dance to music. However, existing surveys in this field focus on individual methods, lacking a comprehensive overview of the entire generative process. This paper addresses this gap by providing an in-depth survey of human motion video generation, encompassing over ten sub-tasks, and detailing the five key phases of the generation process: input, motion planning, motion video generation, refinement, and output. Notably, this is the first survey that discusses the potential of large language models in enhancing human motion video generation. Our survey reviews the latest developments and technological trends in human motion video generation across three primary modalities: vision, text, and audio. By covering over two hundred papers, we offer a thorough overview of the field and highlight milestone works that have driven significant technological breakthroughs. Our goal for this survey is to unveil the prospects of human motion video generation and serve as a valuable resource for advancing the comprehensive applications of digital humans.
Haiwei Xue, Xiangyang Luo 0002, Zhanghao Hu, Xin Zhang 0169, Xunzhi Xiang, Yuqin Dai, Jianzhuang Liu, Zhensong Zhang, Minglei Li 0001, Jian Yang 0003, Fei Ma 0006, Zhiyong Wu 0001, Changpeng Yang, Zonghong Dai, F. Richard Yu
IEEE Trans. Pattern Anal. Mach. Intell.8
2024 Low-Res Leads the Way: Improving Generalization for Super-Resolution by Self-Supervised Learning
abstract
For image super-resolution (SR), bridging the gap between the performance on synthetic datasets and real-world degradation scenarios remains a challenge. This work introduces a novel “Low-Res Leads the Way” (LWay) training framework, merging Supervised Pre-training with Self-supervised Learning to enhance the adaptability of SR models to real-world images. Our approach utilizes a low-resolution (LR) reconstruction network to extract degradation embeddings from LR images, merging them with super-resolved outputs for LR reconstruction. Leveraging unseen LR images for self-supervised learning guides the model to adapt its modeling space to the target domain, facili-tating fine-tuning of SR models without requiring paired high-resolution (HR) images. The integration of Discrete Wavelet Transform (DWT)further refines the focus on high-frequency details. Extensive evaluations show that our method significantly improves the generalization and de-tail restoration capabilities of SR models on unseen real-world datasets, outperforming existing methods. Our training regime is universally compatible, requiring no network architecture modifications, making it a practical solution for real-world SR applications.
Haoyu Chen 0003, Wenbo Li 0002, Jinjin Gu, Haoze Sun, Xueyi Zou, Zhensong Zhang, Youliang Yan, Lei Zhu 0003
CVPR7
2024 Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion Model
abstract
Co-speech gestures, if presented in the lively form of videos, can achieve superior visual effects in human-machine interaction. While previous works mostly gener-ate structural human skeletons, resulting in the omission of appearance information, we focus on the direct gener-ation of audio-driven co-speech gesture videos in this work. There are two main challenges: 1) A suitable motion feature is needed to describe complex human movements with crucial appearance information. 2) Gestures and speech exhibit inherent dependencies and should be temporally aligned even of arbitrary length. To solve these problems, we present a novel motion-decoupled framework to gener-ate co-speech gesture videos. Specifically, we first intro-duce a well-designed nonlinear TPS transformation to ob-tain latent motion features preserving essential appearance information. Then a transformer-based diffusion model is proposed to learn the temporal correlation between gestures and speech, and performs generation in the latent motion space, followed by an optimal motion selection mod-ule to produce long-term coherent and consistent gesture videos. For better visual perception, we further design a refinement network focusing on missing details of cer-tain areas. Extensive experimental results show that our proposed framework significantly outperforms existing approaches in both motion and video-related evaluations. Our code, demos, and more resources are available at https://github.com/thuhcsi/S2G-MDDiffusion.
Qiaochu Huang, Zhensong Zhang, Zhiyong Wu 0001, Minglei Li 0001, Songcen Xu
CVPR3
2024 Semantics-Aware Motion Retargeting with Vision-Language Models
abstract
Capturing and preserving motion semantics is essential to motion retargeting between animation characters. However, most of the previous works neglect the semantic information or rely on human-designed joint-level representations. Here, we present a novel Semantics-aware Motion reTargeting (SMT) method with the advantage of vision-language models to extract and maintain meaningful motion semantics. We utilize a differentiable module to ren-der 3D motions. Then the high-level motion semantics are incorporated into the motion retargeting process by feeding the vision-language model with the rendered images and aligning the extracted semantic embeddings. To en-sure the preservation of fine-grained motion details and high-level semantics, we adopt a two-stage pipeline consisting of skeleton-aware pretraining and fine-tuning with semantics and geometry constraints. Experimental results show the effectiveness of the proposed method in producing high-quality motion retargeting results while accurately preserving motion semantics. Project page can be found at https://sites.google.com/view/smtnet.
Zhike Chen, Haocheng Xu, Songcen Xu, Zhensong Zhang, Yue Wang 0020, Rong Xiong
CVPR7
2024 Conversational Co-Speech Gesture Generation via Modeling Dialog Intention, Emotion, and Context with Diffusion Models
abstract
Audio-driven co-speech human gesture generation has made remarkable advancements recently. However, most previous works only focus on single person audio-driven gesture generation. We aim at solving the problem of conversational co-speech gesture generation that considers multiple participants in a conversation, which is a novel and challenging task due to the difficulty of simultaneously incorporating semantic information and other relevant features from both the primary speaker and the interlocutor. To this end, we propose CoDiffuseGesture, a diffusion model-based approach for speech-driven interaction gesture generation via modeling bilateral conversational intention, emotion, and semantic context. Our method synthesizes appropriate interactive, speech-matched, high-quality gestures for conversational motions through the intention perception module and emotion reasoning module at the sentence level by a pretrained language model. Experimental results demonstrate the promising performance of the proposed method.
Haiwei Xue, Zhensong Zhang, Zhiyong Wu 0001, Minglei Li 0001, Zonghong Dai, Helen M. Meng
ICASSP3
2024 SCRREAM : SCan, Register, REnder And Map: A Framework for Annotating Accurate and Dense 3D Indoor Scenes with a Benchmark
abstract
Traditionally, 3d indoor datasets have generally prioritized scale over ground-truth accuracy in order to obtain improved generalization. However, using these datasets to evaluate dense geometry tasks, such as depth rendering, can be problematic as the meshes of the dataset are often incomplete and may produce wrong ground truth to evaluate the details. In this paper, we propose SCRREAM, a dataset annotation framework that allows annotation of fully dense meshes of objects in the scene and registers camera poses on the real image sequence, which can produce accurate ground truth for both sparse 3D as well as dense 3D tasks. We show the details of the dataset annotation pipeline and showcase four possible variants of datasets that can be obtained from our framework with example scenes, such as indoor reconstruction and SLAM, scene editing & object removal, human reconstruction and 6d pose estimation. Recent pipelines for indoor reconstruction and SLAM serve as new benchmarks. In contrast to previous indoor dataset, our design allows to evaluate dense geometry tasks on eleven sample scenes against accurately rendered ground truth depth maps.
Weihang Li, William Bittner, Nikolas Brasch, Jifei Song, Eduardo Pérez-Pellitero, Zhensong Zhang, Arthur Moreau, Nassir Navab, Benjamin Busam
NeurIPS8
2024 Towards a Unified Network for Robust Monocular Depth Estimation: Network Architecture, Training Strategy and Dataset
Mochu Xiang, Yuchao Dai, Zhensong Zhang
Int. J. Comput. Vis.6
2023 QPGesture: Quantization-Based and Phase-Guided Motion Matching for Natural Speech-Driven Gesture Generation
abstract
Speech-driven gesture generation is highly challenging due to the random jitters of human motion. In addition, there is an inherent asynchronous relationship between human speech and gestures. To tackle these challenges, we introduce a novel quantization-based and phase-guided motion matching framework. Specifically, we first present a gesture VQ-VAE module to learn a codebook to summarize meaningful gesture units. With each code representing a unique gesture, random jittering problems are alleviated effectively. We then use Levenshtein distance to align diverse gestures with different speech. Levenshtein distance based on audio quantization as a similarity metric of corresponding speech of gestures helps match more appropriate gestures with speech, and solves the alignment problem of speech and gestures well. Moreover, we introduce phase to guide the optimal gesture matching based on the semantics of context or rhythm of audio. Phase guides when text-based or speech-based gestures should be performed to make the generated gestures more natural. Extensive experiments show that our method outperforms recent approaches on speech-driven gesture generation. Our code, database, pre-trained models and demos are available at https://github.com/YoungSeng/QPGesture.
Zhiyong Wu 0001, Minglei Li 0001, Zhensong Zhang, Weihong Bao, Haolin Zhuang
CVPR4
2023 DiffuseStyleGesture: Stylized Audio-Driven Co-Speech Gesture Generation with Diffusion Models
abstract
The art of communication beyond speech there are gestures. The automatic co-speech gesture generation draws much attention in computer animation. It is a challenging task due to the diversity of gestures and the difficulty of matching the rhythm and semantics of the gesture to the corresponding speech. To address these problems, we present DiffuseStyleGesture, a diffusion model based speech-driven gesture generation approach. It generates high-quality, speech-matched, stylized, and diverse co-speech gestures based on given speeches of arbitrary length. Specifically, we introduce cross-local attention and self-attention to the gesture diffusion pipeline to generate better speech matched and realistic gestures. We then train our model with classifier-free guidance to control the gesture style by interpolation or extrapolation. Additionally, we improve the diversity of generated gestures with different initial gestures and noise. Extensive experiments show that our method outperforms recent approaches on speech-driven gesture generation. Our code, pre-trained models, and demos are available at https://github.com/YoungSeng/DiffuseStyleGesture.
Zhiyong Wu 0001, Minglei Li 0001, Zhensong Zhang, Weihong Bao, Long Xiao
IJCAI4
2023 UnifiedGesture: A Unified Gesture Synthesis Model for Multiple Skeletons
abstract
The automatic co-speech gesture generation draws much attention in computer animation. Previous works designed network structures on individual datasets, which resulted in a lack of data volume and generalizability across different motion capture standards. In addition, it is a challenging task due to the weak correlation between speech and gestures. To address these problems, we present UnifiedGesture, a novel diffusion model-based speech-driven gesture synthesis approach, trained on multiple gesture datasets with different skeletons. Specifically, we first present a retargeting network to learn latent homeomorphic graphs for different motion capture standards, unifying the representations of various gestures while extending the dataset. We then capture the correlation between speech and gestures based on a diffusion model architecture using cross-local attention and self-attention to generate better speech-matched and realistic gestures. To further align speech and gesture and increase diversity, we incorporate reinforcement learning on the discrete gesture units with a learned reward function. Extensive experiments show that UnifiedGesture outperforms recent approaches on speech-driven gesture generation in terms of CCA, FGD, and human-likeness.
Zilin Wang 0002, Zhiyong Wu 0001, Minglei Li 0001, Zhensong Zhang, Qiaochu Huang, Songcen Xu, Changpeng Yang, Zonghong Dai
ACM Multimedia5
2022 CLIFF: Carrying Location Information in Full Frames into Human Pose and Shape Estimation
Zhihao Li 0002, Jianzhuang Liu, Zhensong Zhang, Songcen Xu, Youliang Yan
ECCV (5)3
2020 Video super-resolution via pre-frame constrained and deep-feature enhanced sparse reconstruction
Qiuxia Lai, Yongwei Nie, Hanqiu Sun, Qiang Xu 0001, Zhensong Zhang, Mingyu Xiao 0001
Pattern Recognit.5
2020 Interactive Contour Extraction via Sketch-Alike Dense-Validation Optimization
abstract
We propose an interactive contour extraction method inspired by a skill often adopted in sketching: an artist usually sketches an object by first drawing lots of short, directional, and redundant strokes, then following these small strokes to draw the final outline of the object. Our method simulates this process. To extract a contour, our method relies on user interaction, which provides us with a narrow band containing the target contour. Then, we densely sample sub-bands from the whole band, with each sub-band containing a local segment of the target contour. We design a curve-centered coordinate system in which a dynamic programming algorithm is proposed to extract the local segment in each sub-band. The local segment is guaranteed to be as evident and smooth as possible, to mimic the strokes sketched by the artist. Finally, we integrate all local segments of all sub-bands together to obtain the whole target contour based on the weighted principal component analysis. Our method can extract high-quality object contours due to the dense validations among local segments. That is, even if one segment deviates from the right location, several other segments in its local neighborhood can correct it in the integration stage. Both quantitative experiments and a user study demonstrate the effectiveness of the proposed method.
Yongwei Nie, Ping Li 0016, Qing Zhang 0006, Zhensong Zhang, Guiqing Li, Hanqiu Sun
IEEE Trans. Circuits Syst. Video Technol.5
2020 Collision-Free Video Synopsis Incorporating Object Speed and Size Changes
abstract
This paper presents a new surveillance video synopsis method which performs much better than previous approaches in terms of both compression ratio and artifact. Previously, a surveillance video was usually compressed by shifting the moving objects of that video forward along the time axis, which inevitably yielded serious collision and chronological disorder artifacts between the shifted objects. The main observation of this paper is that these artifacts can be alleviated by changing the speed or size of the objects, since with varied speed and size the objects can move more flexibly to avoid collision points or to keep chronological relationships. Based on this observation, we propose a video synopsis method that performs object shifting, speed changing, and size scaling simultaneously. We show how to integrate the three heterogeneous operations into a single optimization framework and achieve high-quality synopsis results. Unlike previous approaches that usually use alternative optimization strategies to solve synopsis optimizations, we develop a Metropolis sampling algorithm to find the solution for our three-variable optimization problem. A variety of experiments demonstrate the effectiveness of our method.
Yongwei Nie, Zhenkai Li, Zhensong Zhang, Qing Zhang 0006, Tiezheng Ma, Hanqiu Sun
IEEE Trans. Image Process.3
2020 Multi-View Video Synopsis via Simultaneous Object-Shifting and View-Switching Optimization
abstract
We present a method for synopsizing multiple videos captured by a set of surveillance cameras with some overlapped field-of-views. Currently, object-based approaches that directly shift objects along the time axis are already able to compute compact synopsis results for multiple surveillance videos. The challenge is how to present the multiple synopsis results in a more compact and understandable way. Previous approaches show them side by side on the screen, which however is difficult for user to comprehend. In this paper, we solve the problem by joint object-shifting and camera view-switching. Firstly, we synchronize the input videos, and group the same object in different videos together. Then we shift the groups of objects along the time axis to obtain multiple synopsis videos. Instead of showing them simultaneously, we just show one of them at each time, and allow to switch among the views of different synopsis videos. In this view switching way, we obtain just a single synopsis results consisting of content from all the input videos, which is much easier for user to follow and understand. To obtain the best synopsis result, we construct a simultaneous object-shifting and view-switching optimization framework instead of solving them separately. We also present an alternative optimization strategy composed of graph cuts and dynamic programming to solve the unified optimization. Experiments demonstrate that our single synopsis video generated from multiple input videos is compact, complete, and easy to understand.
Zhensong Zhang, Yongwei Nie, Hanqiu Sun, Qing Zhang 0006, Qiuxia Lai, Guiqing Li, Mingyu Xiao 0001
IEEE Trans. Image Process.1
2020 Effective Video Stabilization via Joint Trajectory Smoothing and Frame Warping
abstract
Video stabilization is usually composed of three stages: feature trajectory extraction, trajectory smoothing, and frame warping. Most previous approaches view them as three separate stages. This paper proposes a method combining the last two stages, namely the trajectory smoothing and frame warping stages, into a single optimization framework. The novelty exists in the way of how we combine them: the trajectory smoothing part plays a major role while the frame warping part plays an auxiliary role. With this kind of design, we can conveniently increase the strength of the trajectory smoothing part by a robust first-order derivative term, which makes it possible to produce very aggressive stabilization effects. On the other hand, we adopt adaptive weighting mechanisms in the frame warping part, to follow the smoothed trajectories as much as possible while regularizing other places as similar as possible. Our method is robust to utilize both foreground and background features, and very short trajectories. The utilization of all these information in turn increases the accuracy of the proposed method. We also provide a simplified implementation of our method, which is less accurate but more efficient. Experiments on various kinds of videos demonstrate the effectiveness of our method.
Tiezheng Ma, Yongwei Nie, Qing Zhang 0006, Zhensong Zhang, Hanqiu Sun, Guiqing Li
IEEE Trans. Vis. Comput. Graph.4
2018 Temporal Coherent Video Super-resolution via Pre-frame-constrained Sparse Reconstruction
abstract
In this paper, we extend the sparse representation based image super-resolution method to process videos, mainly aiming at obtaining temporally consistent consecutive high-resolution (HR) video frames. In our formulation, the previous estimated HR frame is used to guide the sparse reconstruction of current low-resolution (LR) frame, which is able to obtain more consistent representations. We show that such guidance is robust and effective by incorporating with a non-rigid dense correspondence based motion compensation schema. We also propose a dictionary updating strategy which regularly updates the dictionaries that are critical for the sparse representation procedure using the newly reconstructed HR frames. To further preserve sharp edges and remove reconstruction errors, once a HR image is recovered, we refine it with a L0-norm based optimization that constrains the final HR output with relatively sparse gradients. Experimental results on natural videos demonstrated the effectiveness of our proposed method.
Qiuxia Lai, Yongwei Nie, Zhensong Zhang, Hanqiu Sun
CGI3
2018 Dynamic Video Stitching via Shakiness Removing
abstract
Stitching videos captured by hand-held mobile cameras can essentially enhance entertainment experience of ordinary users. However, such videos usually contain heavy shakiness and large parallax, which are challenging to stitch. In this paper, we propose a novel approach of video stitching and stabilization for videos captured by mobile devices. The main component of our method is a unified video stitching and stabilization optimization that computes stitching and stabilization simultaneously rather than does each one individually. In this way, we can obtain the best stitching and stabilization results relative to each other without any bias to one of them. To make the optimization robust, we propose a method to identify background of input videos, and also common background of them. This allows us to apply our optimization on background regions only, which is the key to handle large parallax problem. Since stitching relies on feature matches between input videos, and there inevitably exist false matches, we thus propose a method to distinguish between right and false matches, and encapsulate the false match elimination scheme and our optimization into a loop, to prevent the optimization from being affected by bad feature matches. We test the proposed approach on videos that are causally captured by smartphones when walking along busy streets, and use stitching and stability scores to evaluate the produced panoramic videos quantitatively. Experiments on a diverse of examples show that our results are much better than (challenging cases) or at least on par with (simple cases) the results of previous approaches.Stitching videos captured by hand-held mobile cameras can essentially enhance entertainment experience of ordinary users. However, such videos usually contain heavy shakiness and large parallax, which are challenging to stitch. In this paper, we propose a novel approach of video stitching and stabilization for videos captured by mobile devices. The main component of our method is a unified video stitching and stabilization optimization that computes stitching and stabilization simultaneously rather than does each one individually. In this way, we can obtain the best stitching and stabilization results relative to each other without any bias to one of them. To make the optimization robust, we propose a method to identify background of input videos, and also common background of them. This allows us to apply our optimization on background regions only, which is the key to handle large parallax problem. Since stitching relies on feature matches between input videos, and there inevitably exist false matches, we thus propose a method to distinguish between right and false matches, and encapsulate the false match elimination scheme and our optimization into a loop, to prevent the optimization from being affected by bad feature matches. We test the proposed approach on videos that are causally captured by smartphones when walking along busy streets, and use stitching and stability scores to evaluate the produced panoramic videos quantitatively. Experiments on a diverse of examples show that our results are much better than (challenging cases) or at least on par with (simple cases) the results of previous approaches.
Yongwei Nie, Tan Su, Zhensong Zhang, Hanqiu Sun, Guiqing Li
IEEE Trans. Image Process.3
2018 Corrections to "Dynamic Video Stitching via Shakiness Removing"
abstract
In[1], the biographies of Hanqiu Sun and Guiqing Li included incorrect information. The correct biographies are as follows.
Yongwei Nie, Tan Su, Zhensong Zhang, Hanqiu Sun, Guiqing Li
IEEE Trans. Image Process.3
2017 Homography Propagation and Optimization for Wide-Baseline Street Image Interpolation
abstract
Wide-baseline street image interpolation is useful but very challenging. Existing approaches either rely on heavyweight 3D reconstruction or computationally intensive deep networks. We present a lightweight and efficient method which uses simple homography computing and refining operators to estimate piecewise smooth homographies between input views. To achieve the goal, we show how to combine homography fitting and homography propagation together based on reliable and unreliable superpixel discrimination. Such a combination, other than using homography fitting only, dramatically increases the accuracy and robustness of the estimated homographies. Then, we integrate the concepts of homography and mesh warping, and propose a novel homography-constrained warping formulation which enforces smoothness between neighboring homographies by utilizing the first-order continuity of the warped mesh. This further eliminates small artifacts of overlapping, stretching, etc. The proposed method is lightweight and flexible, allows wide-baseline interpolation. It improves the state of the art and demonstrates that homography computation suffices for interpolation. Experiments on city and rural datasets validate the efficiency and effectiveness of our method.
Yongwei Nie, Zhensong Zhang, Hanqiu Sun, Tan Su, Guiqing Li
IEEE Trans. Vis. Comput. Graph.2
2016 A novel fast and memory efficient parallel MLCS algorithm for long and large-scale sequences alignments
abstract
Information usually can be abstracted as a character sequence over a finite alphabet. With the advent of the era of big data, the increasing length and size of the sequences from various application fields (e.g., biological sequences) result in the classical NP-hard problem, searching for the Multiple Longest Common Subsequences of multiple sequences (i.e., MLCS problem with many applications in the areas of bioinformatics, computational genomics, pattern recognition, etc.), becoming a research hotspot and facing severe challenges. In this paper, we firstly reveal that the leading dominant-point-based MLCS algorithms are very hard to apply to long and large-scale sequences alignments. To overcome their defects, based on the proposed problem-solving model and parallel topological sorting strategies, we present a novel efficient parallel MLCS algorithm. The comprehensive experiments on the benchmark datasets of both random and biological sequences demonstrate that both the time and space complexities of the proposed algorithm are only linearly related to the dominants from aligned sequences, and that the proposed algorithm greatly outperforms the existing state-of-the-art dominant-point-based MLCS algorithms, and hence it is very suitable for long and large-scale sequences alignments.
Yanni Li, Yuping Wang 0003, Zhensong Zhang
ICDE3
2015 Color correspondence of image warping using plane constraints
abstract
Image-based rendering (IBR) can render novel views from several existing photographs of the same scene. Current IBR methods can be roughly classified into reconstruction-based methods and warping-based methods. Warping-based methods [Liu et al. 2009; Chaurasia et al. 2011] directly warped the triangle mesh overlaid on the source image to make the source image match with the target one. Then the novel views between the two given images can be computed by interpolating and retexturing the intermediate mesh.
Zhensong Zhang, Yongwei Nie, Hanqiu Sun
VRST1
2014 Left and right hand distinction for multi-touch tabletop interactions
abstract
In multi-touch interactive systems, it is of great significance to distinguish which hand of the user is touching the surface in real time. Left-right hand distinction is essential for recognizing the multi-finger gestures and further fully exploring the potential of bimanual interaction. However, left-right hand distinction is beyond the capability of most existing multi-touch systems. In this paper, we present a new method for left and right hand distinction based on the human anatomy, work area, finger orientation and finger position. Considering the ergonomics principles of gesture designing, the body-forearm triangle model was proposed. Furthermore, a heuristic algorithm was introduced to group multi-touch contact points and then made left-right hand distinction. A dataset of 2880 images has been set up to evaluate the proposed left-right hand distinction method. The experimental results demonstrate that our method can guarantee the high recognition accuracy and real time performance in freely bimanual multi-touch interactions.
Zhensong Zhang, Fengjun Zhang, Hui Chen 0020, Jiasheng Liu, Hongan Wang, Guozhong Dai
IUI1
2013 Multi-objective optimization integration of query interfaces for the Deep Web based on attribute constraints
Yanni Li, Yuping Wang 0003, Zhensong Zhang
Data Knowl. Eng.4