EDBT 2026 Demo / reviewers in the wild / expert
Ping Shi 0001
dblp:49/4271-1
· DBLP profile ↗
12ranked-venue papers
0as first author
10since 2021 · last 2026
0009-0003-7744-6246ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 9 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fine-grained aesthetic multi-attribute captioning with aligned vision-language representations
Yehui Liu, Minzheng Jia, Yongqiang Kong, Xin Jin 0015, Ping Shi 0001 |
J. Vis. Commun. Image Represent. | 7 |
| 2025 | LawDNet: Enhanced Audio-Driven Lip Synthesis via Local Affine Warping DeformationabstractIn the domain of photorealistic talking head generation, the fidelity of audio-driven lip motion synthesis is essential for realistic virtual interactions. Existing methods face two key challenges: a lack of vivacity due to limited diversity in generated lip poses and noticeable anamorphose motions caused by poor temporal coherence. To address these issues, we propose LawD-Net, a novel deep-learning architecture enhancing lip synthesis through a Local Affine Warping Deformation mechanism. This mechanism models the intricate lip movements in response to the audio input by controllable non-linear warping fields. These fields consist of local affine transformations focused on abstract keypoints within deep feature maps, offering a novel universal paradigm for feature warping in networks. Additionally, LawDNet incorporates a dual-stream discriminator for improved frame-to-frame continuity and employs face normalization techniques to handle pose and scene variations. Extensive evaluations demonstrate LawDNet’s superior robustness and lip movement dynamism performance compared to previous methods. Junli Deng, Yihao Luo, Xueting Yang, Siyou Li, Jinyang Guo 0002, Ping Shi 0001 |
ICASSP | 7 |
| 2025 | LkSFocalNets: Video Action Recognition With Large Kernel Selective Focal NetworksabstractVideo action recognition tasks face the trade-off challenge between computational cost and performance. Existing models often compromise their ability to capture extensive contextual information to reduce computational complexity. To address this issue, this paper proposes Large Kernel Selective Focal Networks (LkSFocalNet), an efficient network architecture. First, we propose the Kernel Selective Block (KSB) and the Video Focal Block (VFB). The KSB dynamically adjusts the receptive field to capture essential contextual information, while the VFB models spatio-temporal dependencies through video focus modulation. Additionally, we present the Depthwise Separable Convolution Gated Unit (DsCG), which reduces computational load while enhancing the network’s robustness. Finally, these three blocks are combined into a Large Kernel Selective Focal Module (LkSFM), which is stacked to form LkSFocalNet. Experimental results on four benchmark datasets demonstrate that LkSFocalNet outperforms baseline models and achieves state-of-the-art performance on the Diving-48 dataset, all while maintaining lower computational complexity. Ping Shi 0001, Qipei Li |
ICASSP | 2 |
| 2025 | DynaSplat: Dynamic-Static Gaussian Splatting with Hierarchical Motion Decomposition for Scene ReconstructionabstractReconstructing intricate, ever-changing environments remains a central ambition in computer vision—yet existing solutions often crumble before the complexity of real-world dynamics. We present DynaSplat, an approach that extends Gaussian Splatting to dynamic scenes by integrating dynamic-static separation and hierarchical motion modeling. First, we classify scene elements as static or dynamic through a novel fusion of deformation offset statistics and 2D motion flow consistency, refining our spatial representation to focus precisely where motion matters. We then introduce a hierarchical motion modeling strategy that captures both coarse global transformations and fine-grained local movements, enabling accurate handling of intricate, non-rigid motions. Finally, we integrate physically-based opacity estimation to ensure visually coherent reconstructions, even under challenging occlusions and perspective shifts. Extensive experiments on challenging datasets reveal that DynaSplat not only surpasses state-of-the-art alternatives in accuracy and realism but also provides a more intuitive, compact, and efficient route to dynamic scene reconstruction. Junli Deng, Ping Shi 0001, Qipei Li, Jinyang Guo 0002 |
ICME | 2 |
| 2025 | Research on Audio-Visual Quality Assessment Dataset and Method for User-Generated Omnidirectional VideoabstractIn response to the rising prominence of the Meta-Verse, omnidirectional videos (ODVs) have garnered notable interest, gradually shifting from professional-generated content (PGC) to user-generated content (UGC). However, the study of audio-visual quality assessment (AVQA) within ODVs remains limited. To address this, we construct a dataset of UGC omnidirectional audio and video (A/V) content. The videos are captured by five individuals using two different types of omnidirectional cameras, shooting 300 videos covering 10 different scene types. A subjective AVQA experiment is conducted on the dataset to obtain the Mean Opinion Scores (MOSs) of the A/V sequences. After that, to facilitate the development of UGC-ODV AVQA fields, we construct an effective AVQA baseline model on the proposed dataset, of which the baseline model consists of video feature extraction module, audio feature extraction and audiovisual fusion module. The experimental results demonstrate that our model achieves optimal performance on the proposed dataset. Da Pan 0001, Zelu Qi, Ping Shi 0001 |
ICME | 4 |
| 2025 | Gaussians on their Way: Wasserstein-Constrained 4D Gaussian Splatting with State-Space ModelingabstractAbstract Dynamic scene rendering has taken a leap forward with the rise of 4D Gaussian Splatting, but there is still one elusive challenge: how to make 3D Gaussians move through time as naturally as they would in the real world, all while keeping the motion smooth and consistent. In this paper, we present an approach that blends state‐space modeling with Wasserstein geometry, enabling a more fluid and coherent representation of dynamic scenes. We introduce a State Consistency Filter that merges prior predictions with the current observations, enabling Gaussians to maintain coherent trajectories over time. We also employ Wasserstein Consistency Constraint to ensure smooth, consistent updates of Gaussian parameters, reducing motion artifacts. Lastly, we leverage Wasserstein geometry to capture both translational motion and shape deformations, creating a more geometrically consistent model for dynamic scenes. Our approach models the evolution of Gaussians along geodesics on the manifold of Gaussian distributions, achieving smoother, more realistic motion and stronger temporal coherence. Experimental results show consistent improvements in rendering quality and efficiency. (see https://www.acm.org/publications/class-2012 ) Junli Deng, Ping Shi 0001 |
Comput. Graph. Forum | 2 |
| 2025 | Physics of Motion, Geometry of Cohesion: A Silky Gaussian Head Avatar FrameworkabstractWe presentPhysics of Motion, Geometry of Cohesion, a framework for creating high-fidelity, dynamic 3D Gaussian head avatars free from common motion and geometry artifacts. To capture the “Physics of Motion,” we introduce a physics-guided propagation module using second-order kinematics (means) and Lie group transformation (covariances) to generate plausible deformation priors. These priors inform a data-driven refinement network. For “Geometry of Cohesion,” we employ a hierarchical Optimal Transport (OT) regularization strategy. Grouping Gaussians by facial landmarks and using adaptive, hyperbolically weighted OT costs ensures spatiotemporal consistency while preserving local expressiveness. Experimental results demonstrate this synergistic approach effectively mitigates common artifacts like jitter and tearing, significantly reducing irregular deformations. This yields high-fidelity, dynamic avatars characterized by natural facial motion and a temporally coherent, “silky” visual quality. Junli Deng, Ping Shi 0001, Qipei Li, Jinyang Guo 0002 |
IEEE Signal Process. Lett. | 2 |
| 2024 | ESTGN: Enhanced Self-Mined Text Guided Super-Resolution Network for Superior Image Super ResolutionabstractIn this paper, we propose a novel Enhanced Self-mined Text Guided Super-resolution Network (ESTGN) for single image super-resolution (SISR). Unlike preceding methods, ESTGN autonomously mines task-related text from images and uses it to guide SR for high-frequency detail restoration. The proposed methods include the Self-mined Text Information Extraction Module, Multi-resolution Text-aware Gradient Balance Module, and Masked Text-conditioned Attention Module. Our method can fully leverage self-mined textual semantic information and enhance gradient propagation in text. We validate our method with extensive experiments on the benchmark dataset, where ESTGN significantly outperforms the baseline model and sets a new state-of-the-art. This work opens up a promising avenue for the integration of text information in image SR tasks. Qipei Li, Zefeng Ying, Da Pan 0001, Zhaoxin Fan, Ping Shi 0001 |
ICASSP | 5 |
| 2024 | DefocusSR: An Efficient Framework for Defocus Image Super-Resolution Guided by Depth InformationabstractExisting image super-resolution (SR) methods often cause oversharpening, especially in defocus images. However, we found that defocus regions and focus regions have different levels of difficulty in recovery. This provides an opportunity for efficient enhancements. In this paper, we propose DefocusSR, an efficient framework for defocus images in SR. DefocusSR comprises two modules: the Depth-guided Segmentation (DGS) and the Defocus-Aware Classify Enhance (DCE). In the DGS, we prompt MobileSAM with the depth of field information to accurately segment the input image and generate defocus maps which contain information about the locations of defocus areas. In the DCE, we crop the defocus map and classify them into defocus and focus patches based on a threshold. In practice, the defocus patches are input into the Efficient Blur Match SR Network (EBM-SR) while preserving the blur kernel to relieve the computation burden. The focus patches are processed using expensive operations. Therefore, DefocusSR combines defocus classification and SR in a unified framework. Experiments demonstrate that DefocusSR can accelerate most SR methods. Reduce the FLOPs of SR models by approximately 70% while preserving state-of-the-art SR performance. Qirong Liang, Da Pan 0001, Zefeng Ying, Ping Shi 0001 |
ICASSP | 4 |
| 2022 | A Full-reference Video Quality Assessment Method for 4K UHD Video based on Multi-Feature FusionabstractVideo quality assessment plays an important role in the quality control of video transmission and the development of video processing equipment and algorithms. With the popularity of UHD TV, the demand for UHD video quality assessment is becoming more and more urgent. In this paper, we propose a method for 4K UHD video quality assessment based on multi-feature fusion (MFF- VQA). First, we select eight frame-level features which could better reflect the perceived video quality through a series of ablation experiments. Then, we present a scheme which can fuse the eight features into a quality score. Experimental results show that, compared with other similar methods, the proposed method can achieve better performance even with lower algorithm complexity and fewer video frames. Yi Geng, Ping Shi 0001, Da Pan 0001 |
SNPD | 2 |
| 2019 | A Comprehensive Survey on Image Aesthetic Quality AssessmentabstractImage aesthetic quality assessment has demonstrated tremendous success in variety of application domains in recent years. This field has been growing so rapidly that various approaches have been proposed trying to solve this challenging problem. This report presents a comprehensive survey on image aesthetic quality assessment, mainly focus on the contributions and novelties of the existing approaches recently. In this work, we firstly illustrate datasets related to image aesthetics and investigate feature extractions. Then five different aesthetic tasks are reviewed, including aesthetic classification, aesthetic regression, aesthetic distribution, aesthetic factors and aesthetic description. In addition, we reviewed recent applications concerning image aesthetics. Finally, different evaluation criterions in different literatures are summarized. We hope the survey could serve as a comprehensive reference and be useful for those who are interested in exploiting image aesthetic for their research. Ping Shi 0001, Saike He, Da Pan 0001, Zefeng Ying |
ICIS | 2 |
| 2018 | Blind Predicting Similar Quality Map for Image Quality AssessmentabstractA key problem in blind image quality assessment (BIQA) is how to effectively model the properties of human visual system in a data-driven manner. In this paper, we propose a simple and efficient BIQA model based on a novel framework which consists of a fully convolutional neural network (FCNN) and a pooling network to solve this problem. In principle, FCNN is capable of predicting a pixel-by-pixel similar quality map only from a distorted image by using the intermediate similarity maps derived from conventional full-reference image quality assessment methods. The predicted pixel-by-pixel quality maps have good consistency with the distortion correlations between the reference and distorted images. Finally, a deep pooling network regresses the quality map into a score. Experiments have demonstrated that our predictions outperform many state-of-the-art BIQA methods. Da Pan 0001, Ping Shi 0001, Zefeng Ying, Sizhe Fu |
CVPR | 2 |