VLDB 2026 Research / reviewers in the wild / expert
Pengfei Xiong
dblp:48/8617
· DBLP profile ↗
20ranked-venue papers
2as first author
13since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 2 first-author · 13 since 2021Artificial intelligence and machine learning · 16 · 1 first-author · 11 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning Implicit Features with Flow-Infused Transformations for Realistic Virtual Try-On
Delong Zhang, Qiwei Huang, Yuanliu Liu, Wei-Shi Zheng 0001, Pengfei Xiong, Wei Zhang 0009 |
ICCV | 6 |
| 2024 | CPN: Complementary Proposal Network for Unconstrained Text DetectionabstractExisting methods for scene text detection can be divided into two paradigms: segmentation-based and anchor-based. While Segmentation-based methods are well-suited for irregular shapes, they struggle with compact or overlapping layouts. Conversely, anchor-based approaches excel for complex layouts but suffer from irregular shapes. To strengthen their merits and overcome their respective demerits, we propose a Complementary Proposal Network (CPN) that seamlessly and parallelly integrates semantic and geometric information for superior performance. The CPN comprises two efficient networks for proposal generation: the Deformable Morphology Semantic Network, which generates semantic proposals employing an innovative deformable morphological operator, and the Balanced Region Proposal Network, which produces geometric proposals with pre-defined anchors. To further enhance the complementarity, we introduce an Interleaved Feature Attention module that enables semantic and geometric features to interact deeply before proposal generation. By leveraging both complementary proposals and features, CPN outperforms state-of-the-art approaches with significant margins under comparable computation cost. Specifically, our approach achieves improvements of 3.6%, 1.3% and 1.0% on challenging benchmarks ICDAR19-ArT, IC15, and MSRA-TD500, respectively. Code for our method will be released. Longhuang Wu, Shangxuan Tian, Youxin Wang, Pengfei Xiong |
AAAI | 4 |
| 2023 | Token Mixing: Parameter-Efficient Transfer Learning from Image-Language to Video-LanguageabstractApplying large scale pre-trained image-language model to video-language tasks has recently become a trend, which brings two challenges. One is how to effectively transfer knowledge from static images to dynamic videos, and the other is how to deal with the prohibitive cost of fully fine-tuning due to growing model size. Existing works that attempt to realize parameter-efficient image-language to video-language transfer learning can be categorized into two types: 1) appending a sequence of temporal transformer blocks after the 2D Vision Transformer (ViT), and 2) inserting a temporal block into the ViT architecture. While these two types of methods only require fine-tuning the newly added components, there are still many parameters to update, and they are only validated on a single video-language task. In this work, based on our analysis of the core ideas of different temporal modeling components in existing approaches, we propose a token mixing strategy to enable cross-frame interactions, which enables transferring from the pre-trained image-language model to video-language tasks through selecting and mixing a key set and a value set from the input video samples. As token mixing does not require the addition of any components or modules, we can directly partially fine-tune the pre-trained image-language model to achieve parameter-efficiency. We carry out extensive experiments to compare our proposed token mixing method with other parameter-efficient transfer learning methods. Our token mixing method outperforms other methods on both understanding tasks and generation tasks. Besides, our method achieves new records on multiple video-language tasks. The code is available at https://github.com/yuqi657/video_language_model. Yuqi Liu 0003, Luhui Xu, Pengfei Xiong, Qin Jin |
AAAI | 3 |
| 2023 | Video-Text as Game Players: Hierarchical Banzhaf Interaction for Cross-Modal Representation LearningabstractContrastive learning-based video-language representation learning approaches, e.g., CLIP, have achieved outstanding performance, which pursue semantic interaction upon pre-defined video-text pairs. To clarify this coarse-grained global interaction and move a step further, we have to encounter challenging shell-breaking interactions for fine-grained cross-modal learning. In this paper, we creatively model video-text as game players with multivariate cooperative game theory to wisely handle the uncertainty during fine-grained semantic interaction with diverse granularity, flexible combination, and vague intensity. Concretely, we propose Hierarchical Banzhaf Interaction (HBI) to value possible correspondence between video frames and text words for sensitive and explainable cross-modal contrast. To efficiently realize the cooperative game of multiple video frames and multiple text words, the proposed method clusters the original video frames (text words) and computes the Banzhaf Interaction between the merged tokens. By stacking token merge modules, we achieve cooperative games at different semantic levels. Extensive experiments on commonly used text-video retrieval and video-question answering bench-marks with superior performances justify the efficacy of our HBI. More encouragingly, it can also serve as a visualization tool to promote the understanding of cross-modal interaction, which have a far-reaching impact on the community. Project page is available at https://jpthu17.github.io/HBI/. Peng Jin 0001, Jinfa Huang, Pengfei Xiong, Shangxuan Tian, Chang Liu 0030, Xiangyang Ji, Li Yuan 0007, Jie Chen 0001 |
CVPR | 3 |
| 2023 | Uncertainty Guided Adaptive Warping for Robust and Efficient Stereo MatchingabstractCorrelation based stereo matching has achieved outstanding performance, which pursues cost volume between two feature maps. Unfortunately, current methods with a fixed model do not work uniformly well across various datasets, greatly limiting their real-world applicability. To tackle this issue, this paper proposes a new perspective to dynamically calculate correlation for robust stereo matching. A novel Uncertainty Guided Adaptive Correlation (UGAC) module is introduced to robustly adapt the same model for different scenarios. Specifically, a variance-based uncertainty estimation is employed to adaptively adjust the sampling area during warping operation. Additionally, we improve the traditional non-parametric warping with learnable parameters, such that the position-specific weights can be learned. We show that by empowering the recurrent network with the UGAC module, stereo matching can be exploited more robustly and effectively. Extensive experiments demonstrate that our method achieves state-of-the-art performance over the ETH3D, KITTI, and Middlebury datasets when employing the same fixed model over these datasets without any retraining procedure. To target real-time applications, we further design a lightweight model based on UGAC, which also outperforms other methods over KITTI benchmarks with only 0.6 M parameters. Junpeng Jing, Jiankun Li, Pengfei Xiong, Jiangyu Liu, Shuaicheng Liu, Xin Deng 0002, Mai Xu, Lai Jiang 0004, Leonid Sigal |
ICCV | 3 |
| 2023 | Transferring Image-CLIP to Video-Text Retrieval via Temporal RelationsabstractWe present a novel network to transfer the image-language pre-trained model to video-text retrieval in an end-to-end manner. Leading approaches in the domain of video-and-language learning try to distill the spatio-temporal video features and multi-modal interaction between videos and language from a large-scale video-text dataset. Differently, we leverage the pre-trained image-language model, and simplify it as a two-stage framework including co-learning of image and text, and enhancing temporal relations between video frames and video-text respectively. Specifically, based on the spatial semantics captured by Contrastive Language-Image Pre-training (CLIP) model, our model involves a Temporal Difference Block (TDB) to capture motions at fine temporal video frames, and a Temporal Alignment Block (TAB) to re-align the tokens of video clips and phrases and enhance the cross-modal correlation. These two temporal blocks efficiently realize video-language learning and enable the proposed model to scale well on comparatively small datasets. We conduct extensive experimental studies including ablation studies and comparisons with existing SOTA methods, and our proposed approach outperforms them on the popularly-employed text-to-video and video-to-text retrieval benchmarks, including MSR-VTT, MSVD, LSMDC, and VATEX. Han Fang 0002, Pengfei Xiong, Luhui Xu, Wenhan Luo |
IEEE Trans. Multim. | 2 |
| 2022 | Shrinking Temporal Attention in Transformers for Video Action RecognitionabstractSpatiotemporal modeling in an unified architecture is key for video action recognition. This paper proposes a Shrinking Temporal Attention Transformer (STAT), which efficiently builts spatiotemporal attention maps considering the attenuation of spatial attention in short and long temporal sequences. Specifically, for short-term temporal tokens, query token interacts with them in a fine-grained manner in dealing with short-range motion. It then shrinks to a coarse attention in neighborhood for long-term tokens, to provide larger receptive field for long-range spatial aggregation. Both of them are composed in a short-long temporal integrated block to build visual appearances and temporal structure concurrently with lower costly in computation. We conduct thorough ablation studies, and achieve state-of-the-art results on multiple action recognition benchmarks including Kinetics400 and Something-Something v2, outperforming prior methods with 50% less FLOPs and without any pretrained model. Bonan Li, Pengfei Xiong, Congying Han, Tiande Guo |
AAAI | 2 |
| 2022 | Practical Stereo Matching via Cascaded Recurrent Network with Adaptive CorrelationabstractWith the advent of convolutional neural networks, stereo matching algorithms have recently gained tremendous progress. However, it remains a great challenge to accurately extract disparities from real-world image pairs taken by consumer-level devices like smartphones, due to practical complicating factors such as thin structures, non-ideal rectification, camera module inconsistencies and various hard-case scenes. In this paper, we propose a set of innovative designs to tackle the problem of practical stereo matching: 1) to better recover fine depth details, we design a hierarchical network with recurrent refinement to update disparities in a coarse-to-fine manner, as well as a stacked cascaded architecture for inference; 2) we propose an adaptive group correlation layer to mitigate the impact of erroneous rectification; 3) we introduce a new synthetic dataset with special attention to difficult cases for better generalizing to real-world scenes. Our results not only rank 1ston both Middlebury and ETH3D benchmarks, outperforming existing state-of-the-art methods by a notable margin, but also exhibit high-quality details for real-life photos, which clearly demonstrates the efficacy of our contributions. Jiankun Li, Peisen Wang, Pengfei Xiong, Ziwei Yan, Jiangyu Liu, Haoqiang Fan, Shuaicheng Liu |
CVPR | 3 |
| 2022 | Aesthetic Text Logo Synthesis via Content-aware Layout InferringabstractText logo design heavily relies on the creativity and expertise of professional designers, in which arranging element layouts is one of the most important procedures. However, few attention has been paid to this task which needs to take many factors (e.g., fonts, linguistics, topics, etc.) into consideration. In this paper, we propose a content-aware layout generation network which takes glyph images and their corresponding text as input and synthesizes aesthetic layouts for them automatically. Specifically, we develop a dual-discriminator module, including a sequence discriminator and an image discriminator, to evaluate both the character placing trajectories and rendered shapes of synthesized text logos, respectively. Furthermore, we fuse the information of linguistics from texts and visual semantics from glyphs to guide layout prediction, which both play important roles in professional layout design. To train and evaluate our approach, we construct a dataset named as TextLogo3K, consisting of about 3,500 text logo images and their pixel-level annotations. Experimental studies on this dataset demonstrate the effectiveness of our approach for synthesizing visually-pleasing text logos and verify its superiority against the state of the art. Guo Pu, Wenhan Luo, Yexin Wang, Pengfei Xiong, Hongwen Kang, Zhouhui Lian |
CVPR | 5 |
| 2022 | DIP: Deep Inverse Patchmatch for High-Resolution Optical FlowabstractRecently, the dense correlation volume method achieves state-of-the-art performance in optical flow. However, the correlation volume computation requires a lot of memory, which makes prediction difficult on high-resolution images. In this paper, we propose a novel Patchmatch-based framework to work on high-resolution optical flow estimation. Specifically, we introduce the first end-to-end Patchmatch based deep learning optical flow. It can get high-precision results with lower memory benefiting from propagation and local search of Patchmatch. Furthermore, a new inverse propagation is proposed to decouple the complex operations of propagation, which can significantly reduce calculations in multiple iterations. At the time of submission, our method ranks 1st on all the metrics on the popular KITTI2015 [28] benchmark, and ranks 2ndon EPE on the Sintel [7] clean benchmark among published optical flow methods. Experiment shows our method has a strong cross-dataset generalization ability that the F1-all achieves 13.73%, reducing 21% from the best published result 17.4% on KITTI2015. What's more, our method shows a good details preserving result on the high-resolution dataset DAVIS [1] and consumes 2× less memory than RAFT [36]. Code will be available at github.com/zihuarheng/DIP Zihua Zheng, Ni Nie, Pengfei Xiong, Jiangyu Liu, Jiankun Li |
CVPR | 4 |
| 2022 | TS2-Net: Token Shift and Selection Transformer for Text-Video Retrieval
Yuqi Liu 0003, Pengfei Xiong, Luhui Xu, Shengming Cao, Qin Jin |
ECCV (14) | 2 |
| 2021 | Practical Wide-Angle Portraits Correction With Deep Structured ModelsabstractWide-angle portraits often enjoy expanded views. However, they contain perspective distortions, especially noticeable when capturing group portrait photos, where the background is skewed and faces are stretched. This paper introduces the first deep learning based approach to remove such artifacts from freely-shot photos. Specifically, given a wide-angle portrait as input, we build a cascaded network consisting of a LineNet, a ShapeNet, and a transition module (TM), which corrects perspective distortions on the background, adapts to the stereographic projection on facial regions, and achieves smooth transitions between these two projections, accordingly. To train our network, we build the first perspective portrait dataset with a large diversity in identities, scenes and camera modules. For the quantitative evaluation, we introduce two novel metrics, line consistency and face congruence. Compared to the previous state-of-the-art approach, our method does not require camera distortion parameters. We demonstrate that our approach significantly outperforms the previous state-of-the-art approach both qualitatively and quantitatively. Shan Zhao 0010, Pengfei Xiong, Jiangyu Liu, Haoqiang Fan, Shuaicheng Liu |
CVPR | 3 |
| 2021 | Hierarchical Fusion for Practical Ghost-free High Dynamic Range ImagingabstractGhosting artifacts and missing content due to the over-/under-saturated regions caused by misalignments are generally considered as the two key challenges in high dynamic range (HDR) imaging for dynamic scenes. However, previous CNN-based methods directly reconstruct the HDR image from the input low dynamic range (LDR) images, with implicit ghost removal and multi-exposure image fusion in an end-to-end network structure. In this paper, we decompose HDR imaging into ghost-free image fusion and ghost-based image restoration, and propose a novel practical Hierarchical Fusion Network (HFNet), which contains three sub-networks: Mask Fusion Network, Mask Compensation Network, and Refine Network. Specifically, LDR images are linearly fused in Mask Fusion Network ignoring the misaligned regions. Then the ghost regions of fusion image are restored with mask compensation. Finally, all these results are refined in the third network. This strategy of divide and rule makes the proposed method significantly more tiny than previous methods. Experiments on different datasets show that superior performance of HFNet with 9x fewer FLOPs, 4x fewer parameters and 3x faster inference speed than the existing methods while providing comparable accuracy. And it achieves state-of-the-art quantitative and qualitative results while applied with similar FLOPs. Pengfei Xiong |
ACM Multimedia | 1 |
| 2020 | Local Context Attention for Salient Object Segmentation
Pengfei Xiong, Zhengyi Lv, Kuntao Xiao, Yuwen He |
ACCV (1) | 2 |
| 2020 | TP-LSD: Tri-Points Based Line Segment Detector
Siyu Huang, Fangbo Qin, Pengfei Xiong, Yijia He, Xiao Liu 0042 |
ECCV (27) | 3 |
| 2019 | DFANet: Deep Feature Aggregation for Real-Time Semantic SegmentationabstractThis paper introduces an extremely efficient CNN architecture named DFANet for semantic segmentation under resource constraints. Our proposed network starts from a single lightweight backbone and aggregates discriminative features through sub-network and sub-stage cascade respectively. Based on the multi-scale feature propagation, DFANet substantially reduces the number of parameters, but still obtains sufficient receptive field and enhances the model learning ability, which strikes a balance between the speed and segmentation performance. Experiments on Cityscapes and CamVid datasets demonstrate the superior performance of DFANet with 8$\times$ less FLOPs and 2$\times$ faster than the existing state-of-the-art real-time semantic segmentation methods while providing comparable accuracy. Specifically, it achieves 70.3\% Mean IOU on the Cityscapes test dataset with only 1.7 GFLOPs and a speed of 160 FPS on one NVIDIA Titan X card, and 71.3\% Mean IOU with 3.4 GFLOPs while inferring on a higher resolution image. Hanchao Li, Pengfei Xiong, Haoqiang Fan, Jian Sun 0001 |
CVPR | 2 |
| 2019 | Deep Fusion Network for Image CompletionabstractDeep image completion usually fails to harmonically blend the restored image into existing content, especially in the boundary area. This paper handles this problem from a new perspective of creating a smooth transition and proposes a concise Deep Fusion Network (DFNet). Firstly, a fusion block is introduced to generate a flexible alpha composition map for combining known and unknown regions. The fusion block not only provides a smooth fusion between restored and existing content but also provides an attention map to make network focus more on the unknown pixels. In this way, it builds a bridge for structural and texture information, so that information can be naturally propagated from the known region into completion. Furthermore, fusion blocks are embedded into several decoder layers of the network. Accompanied by the adjustable loss constraints on each layer, more accurate structure information is achieved. We qualitatively and quantitatively compare our method with other state-of-the-art methods on Places2 and CelebA datasets. The results show the superior performance of DFNet, especially in the aspects of harmonious texture transition, texture detail and semantic structural consistency. Pengfei Xiong, Renhe Ji, Haoqiang Fan |
ACM Multimedia | 2 |
| 2018 | Pyramid Attention Network for Semantic Segmentation
Hanchao Li, Pengfei Xiong, Jie An 0002, Lingxue Wang |
BMVC | 2 |
| 2017 | Tiny Transform Net for Mobile Image StylizationabstractArtistic stylization is an image transformation problem that renders an image in the style of another one. Existing methods either regard image style transfer as an optimization of perceptual loss function based on a pre-trained network, or train a feed forward network that achieves style transfer through one forward propagation. However, time-consuming optimization processes or relatively large feed forward networks are unacceptable for mobile application. In this work we propose a tiny transform net to accomplish image stylization on mobile devices. The advantages of our proposed architecture come from that: (i) The size of the carefully designed network is less than 40KB, which is more than 166 times smaller than the current popular network; (ii) Progressive training is put forward to keep the training stable, which is implemental to achieve semantics aware stylization; (iii) Deep convolutional network inference algorithm is reconstructed on mobile platform to reduce the overhead of storage and time. In addition, well-trained tiny transform nets and demo application will be made available. Shilun Lin, Pengfei Xiong |
ICMR | 2 |
| 2010 | Initialization and Pose Alignment in Active Shape ModelabstractIn this paper, we propose a new algorithm for shape initialization and 3D pose alignment in Active Shape Model (ASM). Instead of initializing with average shape in previous works, we build a scatter data interpolation model from key points to obtain the initial shape, which ensures shape initialized around face organs. These key points are chosen from organs of face shape and located with a strong classifier firstly. Then they are utilized to build a Radial Basis Function (RBF) model to deform the average shape as initial shape. Besides, to cope with variety face poses, we define a 3D general shape to align face shapes in 3D instead of 2D alignment in Classic ASM. With the accurate 3D rotation angles iteratively calculated by Levenberg-Marquardt (LM) algorithm, shapes can be aligned to standard shape more reliably. Experiments and comparisons on FERET show that both shape initialization and 3D pose alignment of our algorithm greatly improve the location accuracy. Pengfei Xiong, Lei Huang 0002, Changping Liu |
ICPR | 1 |