EDBT 2026 Demo / reviewers in the wild / expert
Zhihang Zhong
dblp:259/7061
· DBLP profile ↗
27ranked-venue papers
10as first author
25since 2021 · last 2026
0000-0002-1801-8095ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 6 first-author · 22 since 2021Artificial intelligence and machine learning · 22 · 8 first-author · 21 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RacketVision: A Multiple Racket Sports Benchmark for Unified Ball and Racket AnalysisabstractWe introduce RacketVision, a novel dataset and benchmark for advancing computer vision in sports analytics, covering table tennis, tennis, and badminton. The dataset is the first to provide large-scale, fine-grained annotations for racket pose alongside traditional ball positions, enabling research into complex human-object interactions. It is designed to tackle three interconnected tasks: fine-grained ball tracking, articulated racket pose estimation, and predictive ball trajectory forecasting. Our evaluation of established baselines reveals a critical insight for multi-modal fusion: while naively concatenating racket pose features degrades performance, a Cross-Attention mechanism is essential to unlock their value, leading to trajectory prediction results that surpass strong unimodal baselines. RacketVision provides a versatile resource and a strong starting point for future research in dynamic object tracking, conditional motion forecasting, and multi-modal analysis in sports. Linfeng Dong, Yuchen Yang 0003, Wei Wang 0333, Yuenan Hou, Zhihang Zhong, Xiao Sun 0001 |
AAAI | 6 |
| 2026 | Velocity Disambiguation for Video Frame InterpolationabstractExisting video frame interpolation (VFI) methods blindly predict where each object is at a specific timestep $t$t ("time indexing"), which struggles to predict precise object movements. Given two images of a baseball, there are infinitely many possible trajectories: accelerating or decelerating, straight or curved. This often results in blurry frames as the method averages out these possibilities. Instead of forcing the network to learn this complicated time-to-location mapping implicitly together with predicting the frames, we provide the network with an explicit hint on how far the object has traveled between start and end frames, a novel approach termed "distance indexing". This method offers a clearer learning goal for models, reducing the uncertainty tied to object speeds. We further observed that, even with this extra guidance, objects can still be blurry especially when they are equally far from both input frames (i.e., halfway in-between), due to the directional ambiguity in long-range motion. To solve this, we propose an iterative reference-based estimation strategy that breaks down a long-range prediction into several short-range steps. When integrating our plug-and-play strategies into state-of-the-art learning-based models, they exhibit markedly sharper outputs and superior perceptual quality in arbitrary time interpolations, using a uniform distance indexing map in the same format as time indexing without requiring extra computation. Furthermore, we demonstrate that if additional latency is acceptable, a continuous map estimator can be employed to compute a pixel-wise dense distance indexing using multiple nearby frames. Combined with efficient multi-frame refinement, this extension can further disambiguate complex motion, thus enhancing performance both qualitatively and quantitatively. Additionally, the ability to manually specify distance indexing allows for independent temporal manipulation of each object, providing a novel tool for video editing tasks such as re-timing. Zhihang Zhong, Wei Wang 0333, Xiao Sun 0001, Yu Qiao 0001, Gurunandan Krishnan, Sizhuo Ma, Jian Wang 0100 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | MaskGaussian: Adaptive 3D Gaussian Representation from Probabilistic MasksabstractWhile 3D Gaussian Splatting (3DGS) has demonstrated remarkable performance in novel view synthesis and real-time rendering, the high memory consumption due to the use of millions of Gaussians limits its practicality. To mitigate this issue, improvements have been made by pruning unnecessary Gaussians, either through a hand-crafted criterion or by using learned masks. However, these methods deterministically remove Gaussians based on a snapshot of the pruning moment, leading to sub-optimized reconstruction performance from a long-term perspective. To address this issue, we introduce MaskGaussian, which models Gaussians as probabilistic entities rather than permanently removing them, and utilize them according to their probability of existence. To achieve this, we propose a masked-rasterization technique that enables unused yet probabilistically existing Gaussians to receive gradients, allowing for dynamic assessment of their contribution to the evolving scene and adjustment of their probability of existence. Hence, the importance of Gaussians iteratively changes and the pruned Gaussians are selected diversely. Extensive experiments demonstrate the superiority of the proposed method in achieving better rendering quality with fewer Gaussians than previous pruning methods, pruning over 60% of Gaussians on average with only a 0.02 PSNR decline. Our code can be found at: https://github.com/kaikai23/MaskGaussian Zhihang Zhong, Yifan Zhan, Xiao Sun 0001 |
CVPR | 2 |
| 2025 | DiffBody: Human Body Image Restoration with Generative Diffusion PriorabstractHuman body image restoration is crucial for various applications but remains challenging due to the limitations of generative models: General image restoration methods built on generative models may generate unnatural textures, noticeable structural misalignments, and significant loss of fine details. To address these shortcomings, we present DiffBody, a novel human body-aware diffusion model that incorporates domain-specific knowledge to significantly enhance restoration quality. Our approach adopts a two-stage framework: (1) a multi-branch joint diffusion model generates preliminary priors, including normal and depth maps supported by a robust reconstruction pre-processing step; (2) a restoration stage refines the output using a body-prior ControlNet and a color adapter, ensuring structural accuracy and color consistency. Extensive quantitative evaluations, qualitative evaluations, and user studies validate the superior performance of DiffBody in producing perceptually high-quality human body restoration results. Code is available at https://github.com/yimingz1218/DiffBody. Lionel Z. Wang, Sizhuo Ma, Xinjie Li 0002, Zhihang Zhong, Jian Wang 0100 |
ICCP | 6 |
| 2025 | CityGS-$\mathcal{X}$: A Scalable Architecture for Efficient and Geometrically Accurate Large-Scale Scene Reconstruction
Hao Li 0069, Zhengyu Zou, Zhihang Zhong, Dingwen Zhang, Junwei Han 0001 |
ICCV | 5 |
| 2025 | Sequential Gaussian Avatars with Hierarchical Motion Context
Wangze Xu, Yifan Zhan, Zhihang Zhong |
ICCV | 3 |
| 2025 | Towards Explicit Exoskeleton for the Reconstruction of Complicated 3D Human Avatars
Yifan Zhan, Qingtian Zhu, Muyao Niu, Mingze Ma, Jiancheng Zhao, Zhihang Zhong, Xiao Sun 0001, Yu Qiao 0001, Yinqiang Zheng |
ICCV | 6 |
| 2025 | Intrinsic Feature Rectification: Mitigating RAG Dependency by Addressing Information Loss in Image CaptioningabstractLarge Language Models (LLMs) have significantly advanced image captioning, yet their performance often degrades in limited-data settings due to overfitting in visual encoders, resulting in information loss during image encoding and poor generalization. While RAG is commonly used to boost performance in such cases, we argue that it primarily compensates for these visual encoding shortcomings rather than introducing truly novel knowledge. To address this issue at its root, we propose Intrinsic Feature Rectification (IFR), which enhances visual representations and reduces reliance on retrieval. IFR first aligns target image features with robust pre-trained CLIP representations and complementary auxiliary features from a vision-centric encoder. It then optionally fuses these features to construct a more holistic and generalizable visual representation, less prone to overfitting. Experiments show that IFR matches or exceeds the performance of RAG-based methods while eliminating retrieval overhead. Moreover, RAG offers little to no additional benefit when IFR is applied, supporting our hypothesis that RAG often serves as a fallback for weak intrinsic features. By improving visual understanding from the outset, IFR provides a more efficient solution for generalizable image captioning. Code will be available at https://github.com/haowuxc/IFR. Zhihang Zhong, Xiao Sun 0001 |
MMAsia | 2 |
| 2024 | IQ-VFI: Implicit Quadratic Motion Estimation for Video Frame InterpolationabstractAdvanced video frame interpolation (VFI) algorithms approximate intermediate motions between two input frames to synthesize intermediate frame. However, they struggle to handle complex scenarios with curvilinear motions since they overlook the latent acceleration information between the input frames. Moreover, the supervision of predicted motions is tricky because ground-truth motions are not available. To this end, we propose a novel frame-work for implicit quadratic video frame interpolation (IQ-VFI), which explores latent acceleration information and accurate intermediate motions via knowledge distillation. Specifically, the proposed IQ-VFI consists of an implicit acceleration estimation network (IANet) and a VFI back-bone, the former fully leverages spatio-temporal information to explore latent acceleration priors between two input frames, which is then used to progressively modulate linear motions from the latter into quadratic motions in coarse-to-fine manner. Furthermore, to encourage both components to distill more acceleration and motion cues oriented towards VFI, we propose a knowledge distillation strategy in which implicit acceleration distillation loss and implicit motion distillation loss are employed to adaptively guide latent acceleration priors and intermediate motions learning, respectively. Extensive experiments show that our proposed IQ-VFI can achieve state-of-the-art performances on various benchmark datasets. Mengshun Hu, Kui Jiang, Zhihang Zhong, Zheng Wang 0007, Yinqiang Zheng |
CVPR | 3 |
| 2024 | Fooling Polarization-Based Vision Using Locally Controllable Polarizing ProjectionabstractPolarization is a fundamental property of light that encodes abundant information regarding surface shape, material, illumination and viewing geometry. The computer vision community has witnessed a blossom of polarization-based vision applications, such as reflection removal, shape-from-polarization (SfP), transparent object segmentation and color constancy, partially due to the emergence of single-chip mono/color polarization sensors that make polarization data acquisition easier than ever. However, is polarization-based vision vulnerable to adversarial attacks? If so, is that possible to realize these adversarial attacks in the physical world, without being perceived by human eyes? In this paper, we warn the community of the vulnerability of polarization-based vision, which can be more serious than RGB-based vision. By adapting a commercial LCD projector, we achieve locally controllable polarizing projection, which is successfully utilized to fool state-of-the-art polarization-based vision algorithms for glass segmentation and SfP. Compared with existing physical attacks on RGB-based vision, which always suffer from the trade-off between attack efficacy and eye conceivability, the adversarial attackers based on polarizing projection are contact-free and visually imperceptible, since naked human eyes can rarely perceive the difference of viciously manipulated polarizing light and ordinary illumination. This poses unprecedented risks on polarization-based vision, for which due attentions should be paid and counter measures be considered. Zhuoxiao Li, Zhihang Zhong, Shohei Nobuhara, Ko Nishino, Yinqiang Zheng |
CVPR | 2 |
| 2024 | Within the Dynamic Context: Inertia-Aware 3D Human Modeling with Pose Sequence
Yifan Zhan, Zhihang Zhong, Wei Wang 0333, Xiao Sun 0001, Yu Qiao 0001, Yinqiang Zheng |
ECCV (49) | 3 |
| 2024 | KFD-NeRF: Rethinking Dynamic NeRF with Kalman Filter
Yifan Zhan, Zhuoxiao Li, Muyao Niu, Zhihang Zhong, Shohei Nobuhara, Ko Nishino, Yinqiang Zheng |
ECCV (45) | 4 |
| 2024 | Clearer Frames, Anytime: Resolving Velocity Ambiguity in Video Frame Interpolation
Zhihang Zhong, Gurunandan Krishnan, Xiao Sun 0001, Yu Qiao 0001, Sizhuo Ma, Jian Wang 0100 |
ECCV (33) | 1 |
| 2023 | Visibility Constrained Wide-Band Illumination Spectrum Design for Seeing-in-the-DarkabstractSeeing-in-the-dark is one of the most important and challenging computer vision tasks due to its wide applications and extreme complexities of in-the- wild scenarios. Existing arts can be mainly divided into two threads: 1) RGB-dependent methods restore information using degraded RGB inputs only (e.g., low-light enhancement), 2) RGB-independent methods translate images captured under auxiliary near-infrared (NIR) illuminants into RGB domain (e.g., NIR2RGB translation). The latter is very attractive since it works in complete darkness and the illuminants are visually friendly to naked eyes, but tends to be unstable due to its intrinsic ambiguities. In this paper, we try to robustify NIR2RGB translation by designing the optimal spectrum of auxiliary illumination in the wide-band VIS-NIR range, while keeping visual friendliness. Our core idea is to quantify the visibility constraint implied by the human vision system and incorporate it into the design pipeline. By modeling the formation process of images in the VIS-NIR range, the optimal multiplexing of a wide range of LEDs is automatically designed in a fully differentiable manner, within the feasible region defined by the visibility constraint. We also collect a substantially expanded VIS-NIR hyperspectral image dataset for experiments by using a customized 50-band filter wheel. Experimental results show that the task can be significantly improved by using the optimized wide-band illumination than using NIR only. Codes Available: https://github.com/MyNiuuu/VCSD. Muyao Niu, Zhuoxiao Li, Zhihang Zhong, Yinqiang Zheng |
CVPR | 3 |
| 2023 | Blur Interpolation Transformer for Real-World Motion from BlurabstractThis paper studies the challenging problem of recovering motion from blur, also known as joint deblurring and interpolation or blur temporal super-resolution. The challenges are twofold: 1) the current methods still leave considerable room for improvement in terms of visual quality even on the synthetic dataset, and 2) poor generalization to real-world data. To this end, we propose a blur interpolation transformer (BiT) to effectively unravel the underlying temporal correlation encoded in blur. Based on multi-scale residual Swin transformer blocks, we introduce dual-end temporal supervision and temporally symmetric ensembling strategies to generate effective features for time-varying motion rendering. In addition, we design a hybrid camera system to collect the first real-world dataset of one-to-many blur-sharp video pairs. Experimental results show that BiT has a significant gain over the state-of-the-art methods on the public dataset Adobe240. Besides, the proposed real-world dataset effectively helps the model generalize well to real blurry scenarios. Code and data are available at https://github.com/zzh-tech/Bi'T. Zhihang Zhong, Mingdeng Cao, Xiang Ji 0005, Yinqiang Zheng, Imari Sato |
CVPR | 1 |
| 2023 | Rethinking Video Frame Interpolation from Shutter Mode Induced DegradationabstractImage restoration from various motion-related degradations, like blurry effects recorded by a global shutter (GS) and jello effects caused by a rolling shutter (RS), has been extensively studied. It has been recently recognized that such degradations encode temporal information, which can be exploited for video frame interpolation (VFI), a more challenging task than pure restoration. However, these VFI researches are mainly grounded on experiments with synthetic data, rather than real data. More fundamentally, under the same imaging condition, it remains unknown which degradation will be more effective toward VFI. In this paper, we present the first real-world dataset for learning and benchmarking degraded video frame interpolation, named RD-VFI, and further explore the performance differences of three types of degradations, including GS blur, RS distortion, and an in-between effect caused by the rolling shutter with global reset (RSGR), thanks to our novel quad-axis imaging system. Moreover, we propose a unified Progressive Mutual Boosting Network (PMBNet) model to interpolate middle frames at arbitrary time for all shutter modes. Its disentanglement strategy and dual-stream correction enable us to adaptively deal with different degradations for VFI. Experimental results demonstrate that our PMBNet is superior to the respective state-of-the-art methods on all shutter modes. Xiang Ji 0005, Zhixiang Wang 0001, Zhihang Zhong, Yinqiang Zheng |
ICCV | 3 |
| 2023 | NIR-assisted Video Enhancement via Unpaired 24-hour DataabstractLow-light video enhancement in the visible (VIS) range is important yet technically challenging, and it is likely to become more tractable by introducing near-infrared (NIR) information for assistance, which in turn arouses a new challenge on how to obtain appropriate multispectral data for model training. In this paper, we defend the feasibility and superiority of NIR-assisted low-light video enhancement results by using unpaired 24-hour data for the first time, which significantly eases data collection and improves generalization performance on in-the-wild data. By accounting for different physical characteristics between unpaired daytime and nighttime videos, we first propose to turn daytime NIR & VIS into "nighttime mode". Specifically, we design a heuristic yet physics-inspired relighting algorithm to produce realistic pseudo nighttime NIR, and use a resampling strategy followed by a noiseGAN for nighttime VIS conversion. We further devise a temporal-aware network for video enhancement that extracts and fuses bi-directional temporal streams and is trained using real daytime videos and pseudo nighttime videos. We capture multi-spectral data using a co-axial camera and contribute Fulltime Multi-Spectral Video Dataset (FMSVD), the first dataset including aligned 24-hour NIR & VIS videos. Compared to alternative methods, we achieve significantly improved video quality as well as generalization ability on in-the-wild data in terms of both evaluation metrics and visual judgment. Codes and Data Available: https://github.com/MyNiuuu/NVEU. Muyao Niu, Zhihang Zhong, Yinqiang Zheng |
ICCV | 2 |
| 2023 | Event-guided Frame Interpolation and Dynamic Range Expansion of Single Rolling Shutter ImageabstractIn the presence of abrupt motion, the pushbroom scanning mechanism of a rolling shutter (RS) camera tends to bring undesirable distortion, which is recently shown to be beneficial for high-speed frame interpolation. Although promising results have been reported by using multiple consecutive RS frames, to interpolate intermediate distortion-free frames from a single RS image is still an open question, due to the existence of multiple motions that can account for the recorded distortion. Another limitation of RS cameras in complex dynamic scenarios lies in the dynamic range, since traditional ways of multiple exposure for high dynamic range (HDR) imaging will fail due to alignment issues. To deal with these two challenges simultaneously, we propose to use an event camera for assistance, which has much faster temporal response and wider dynamic range. Since there does not exist learning data for this brand new imaging setup, we first build a quad-axis imaging system to capture a realistic dataset called REG-HDR, with pairs of fully aligned RS image and its associated events, as well as their corresponding high-speed HDR GS images. We also propose a flow-based network for frame interpolation, compounded with an attention-based fusion network for dynamic range expansion. Experimental results have verified the effectiveness of our proposed algorithm and the superiority of using realistic data for this challenging dural-purpose enhancement task. Guixu Lin, Jin Han 0001, Mingdeng Cao, Zhihang Zhong, Yinqiang Zheng |
ACM Multimedia | 4 |
| 2023 | Real-World Video Deblurring: A Benchmark Dataset and an Efficient Recurrent Neural Network
Zhihang Zhong, Yinqiang Zheng, Imari Sato |
Int. J. Comput. Vis. | 1 |
| 2022 | Learning Adaptive Warping for RealWorld Rolling Shutter CorrectionabstractThis paper proposes the first real-world rolling shutter (RS) correction dataset, BS-RSC, and a corresponding model to correct the RS frames in a distorted video. Mobile devices in the consumer market with CMOS-based sensors for video capture often result in rolling shutter effects when relative movements occur during the video acquisition process, calling for RS effect removal techniques. However, current state-of-the-art RS correction methods often fail to remove RS effects in real scenarios since the motions are various and hard to model. To address this issue, we propose a real-world RS correction dataset BS-RSC. Real distorted videos with corresponding ground truth are recorded simultaneously via a well-designed beam-splitter-based acquisition system. BS-RSC contains various motions of both camera and objects in dynamic scenes. Further, an RS correction model with adaptive warping is proposed. Our model can warp the learned RS features into global shutter counterparts adaptively with predicted multiple displacement fields. These warped features are aggregated and then reconstructed into high-quality global shutter frames in a coarse-to-fine strategy. Experimental results demonstrate the effectiveness of the proposed method, and our dataset can improve the model's ability to remove the RS effects in the real world. The project is available at https://github.com/ljzycmd/BSRSC. Mingdeng Cao, Zhihang Zhong, Jiahao Wang 0005, Yinqiang Zheng, Yujiu Yang 0001 |
CVPR | 2 |
| 2022 | Efficient Video Deblurring Guided by Motion Magnitude
Yusheng Wang 0001, Yunfan Lu, Lin Wang 0025, Zhihang Zhong, Yinqiang Zheng, Atsushi Yamashita |
ECCV (19) | 5 |
| 2022 | Bringing Rolling Shutter Images Alive with Dual Reversed Distortion
Zhihang Zhong, Mingdeng Cao, Xiao Sun 0001, Zhirong Wu, Zhongyi Zhou, Yinqiang Zheng, Stephen Lin 0001, Imari Sato |
ECCV (7) | 1 |
| 2022 | Animation from Blur: Multi-modal Blur Decomposition with Motion Guidance
Zhihang Zhong, Xiao Sun 0001, Zhirong Wu, Yinqiang Zheng, Stephen Lin 0001, Imari Sato |
ECCV (19) | 1 |
| 2021 | Towards Rolling Shutter Correction and Deblurring in Dynamic ScenesabstractJoint rolling shutter correction and deblurring (RSCD) techniques are critical for the prevalent CMOS cameras. However, current approaches are still based on conventional energy optimization and are developed for static scenes. To enable learning-based approaches to address real-world RSCD problem, we contribute the first dataset, BS-RSCD, which includes both ego-motion and object-motion in dynamic scenes. Real distorted and blurry videos with corresponding ground truth are recorded simultaneously via a beam-splitter-based acquisition system.Since direct application of existing individual rolling shutter correction (RSC) or global shutter deblurring (GSD) methods on RSCD leads to undesirable results due to inherent flaws in the network architecture, we further present the first learning-based model (JCD) for RSCD. The key idea is that we adopt bi-directional warping streams for displacement compensation, while also preserving the non-warped deblurring stream for details restoration. The experimental results demonstrate that JCD achieves state-of-the-art performance on the realistic RSCD dataset (BS-RSCD) and the synthetic RSC dataset (Fastec-RS). The dataset and code are available at https://github.com/zzh-tech/RSCD. Zhihang Zhong, Yinqiang Zheng, Imari Sato |
CVPR | 1 |
| 2021 | Multistream Temporal Convolutional Network for Correct/Incorrect Patient Transfer Action Detection Using Body Sensor NetworkabstractThe development of body sensor networks (BSNs) with rich multimodal signals has enabled highly accurate fine-grained action detection, which is the cornerstone of many humancomputer interaction applications. However, in the case of consecutive fine-grained actions, most existing wearable sensor-based detection methods are constrained by sliding windows because of their limited temporal receptive fields, and existing sequence-to-sequence detection methods cannot effectively leverage the potential of multimodal information of wearable sensors. Herein, to give multimodal signals full play in fine-grained action detection, we propose a novel temporal convolutional network by designing a channel attention-based multistream structure. We apply it to a promising application for correct and incorrect patient transfer nursing action detection. A dataset is collected from a BSN on a patient when nurses perform patient transfer. Extensive experiments on our dataset and public dataset (C-MHAD) demonstrate that the proposed method is superior to the state-of-the-art methods, because it can strengthen the utilization of prediction features from the more convincing modal stream at each time frame. Zhihang Zhong, Chingszu Lin, Masako Kanai-Pak, Jukai Maeda, Yasuko Kitajima, Mitsuhiro Nakamura, Noriaki Kuwahara, Taiki Ogata, Jun Ota 0001 |
IEEE Internet Things J. | 1 |
| 2020 | Efficient Spatio-Temporal Recurrent Neural Network for Video Deblurring
Zhihang Zhong, Yinqiang Zheng |
ECCV (6) | 1 |
| 2020 | Multi-attention deep recurrent neural network for nursing action evaluation using wearable sensorabstractA nursing action evaluation system that can assess the performance of students practicing patient handling related nursing skills becomes an urgent need for solving the nursing educator shortage problem. Such an evaluation system should be designed with less hand-crafted procedures for its scalability. Additionally, realizing high accuracy of nursing action recognition, especially fine-grained action recognition remains a problem. This reflects in the recognition of the correct and incorrect methods when students perform a nursing action, and low accuracy of that would mislead the nursing students. We propose a multi-attention deep recurrent neural network (MA-DRNN) model for nursing action recognition by directly processing the raw acceleration and rotational speed signals from wearable sensors. Data samples of target nursing actions in a nursing skill called patient transfer were collected to train and compare the models. The experiment results show that the proposed model can reach approximately 96% recognition accuracy for four target fine-grained nursing action classes helped by the attention mechanism on time and layer domains, which outperforms the state-of-the-art models of wearable sensor-based HAR. Zhihang Zhong, Chingszu Lin, Taiki Ogata, Jun Ota 0001 |
IUI | 1 |