VLDB 2026 Research / reviewers in the wild / expert
Yangfu Li
dblp:330/4201
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0002-4087-2060ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MSA2: Multi-Task Framework With Structure-Aware and Style-Adaptive Character Representation for Open-Set Chinese Text Recognition
Yangfu Li, Hongjian Zhan, Yujie Xiong, Yue Lu 0001 |
ICCV | 1 |
| 2025 | Why 1 + 1 < 1 in Visual Token Pruning: Beyond Naive Integration via Multi-Objective Balanced CoveringabstractExisting visual token pruning methods target prompt alignment and visual preservation with static strategies, overlooking the varying relative importance of these objectives across tasks, which leads to inconsistent performance. To address this, we derive the first closed-form error bound for visual token pruning based on the Hausdorff distance, uniformly characterizing the contributions of both objectives. Moreover, leveraging $\epsilon$-covering theory, we reveal an intrinsic trade-off between these objectives and quantify their optimal attainment levels under a fixed budget. To practically handle this trade-off, we propose Multi-Objective Balanced Covering (MoB), which reformulates visual token pruning as a bi-objective covering problem. In this framework, the attainment trade-off reduces to budget allocation via greedy radius trading. MoB offers a provable performance bound and linear scalability with respect to the number of input visual tokens, enabling adaptation to challenging pruning scenarios. Extensive experiments show that MoB preserves 96.4\% of performance for LLaVA-1.5-7B using only 11.1\% of the original visual tokens and accelerates LLaVA-Next-7B by 1.3-1.5$\times$ with negligible performance loss. Additionally, evaluations on Qwen2-VL and Video-LLaVA confirm that MoB integrates seamlessly into advanced MLLMs and diverse vision-language tasks. The code will be made available soon. Yangfu Li, Hongjian Zhan, Yujie Xiong, Yue Lu 0001 |
NeurIPS | 1 |
| 2024 | FaRE: A Feature-Aware Radical Encoding Strategy for Zero-Shot Chinese Character Recognition
Hongjian Zhan, Yangfu Li, Yujie Xiong, Yue Lu 0001 |
ACCV (1) | 2 |
| 2024 | LK-Net: Efficient Large Kernel ConvNet for Document Enhancement
Qijun Shi, Hongjian Zhan, Yangfu Li, Weijun Zou, Huasheng Li, Umapada Pal 0001, Yue Lu 0001 |
ICPR (21) | 3 |
| 2024 | Free Lunch: Frame-level Contrastive Learning with Text Perceiver for Robust Scene Text Recognition in Lightweight Models
Hongjian Zhan, Yangfu Li, Yujie Xiong, Umapada Pal 0001, Yue Lu 0001 |
ACM Multimedia | 2 |
| 2024 | DS-TDNN: Dual-Stream Time-Delay Neural Network With Global-Aware Filter for Speaker VerificationabstractConventional time-delay neural networks (TDNNs) struggle to handle long-range context, their ability to represent speaker information is therefore limited for long utterances. Existing solutions either depend on increasing model complexity or try to strike a balance between local features and global context to address this issue. To effectively leverage the long-term dependencies of audio signals and constrain model complexity, we introduce a novel module called Global-aware Filter layer (GF layer) in this work, which employs a set of learnable transform-domain filters between a 1D discrete Fourier transform and its inverse transform to capture global context. Additionally, we develop a dynamic filtering strategy and a sparse regularization method to enhance the performance of the GF layer and prevent overfitting. Based on the GF layer, we present a dual-stream TDNN architecture called DS-TDNN for automatic speaker verification (ASV), which utilizes two unique branches to extract both local and global features in parallel and employs an efficient strategy to fuse different-scale information. Experiments on the Voxceleb and SITW databases demonstrate that the DS-TDNN achieves a relative improvement of 10% together with a relative decline of 20% in computational cost over the ECAPA-TDNN in the speaker verification task. This improvement becomes more evident as the utterance's duration grows. Furthermore, the DS-TDNN also beats popular deep residual models and attention-based systems on utterances of arbitrary length. Yangfu Li, Jiapan Gan, Xiaodan Lin, Yingqiang Qiu, Hongjian Zhan, Hui Tian 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | DeflickerCycleGAN: Learning to Detect and Remove Flickers in a Single ImageabstractEliminating the flickers in digital images captured by rolling shutter cameras is a fundamental and important task in computer vision applications. The flickering effect in a single image stems from the mechanism of asynchronous exposure of rolling shutters employed by cameras equipped with CMOS sensors. In an artificial lighting environment, the light intensity captured at different time intervals varies due to the fluctuation of the power grid, ultimately resulting in the flickering artifact in the image. Up to date, there are few studies related to single image deflickering. Further, it is even more challenging to remove flickers without a priori information, e.g., camera parameters or paired images. To address these challenges, we propose an unsupervised framework termed DeflickerCycleGAN, which is trained on unpaired images for end-to-end single image deflickering. Besides the cycle-consistency loss to maintain the similarity of image contents, we meticulously design another two novel loss functions, i.e., gradient loss and flicker loss, to reduce the risk of edge blurring and color distortion. Moreover, we provide a strategy to determine whether an image contains flickers or not without extra training, which leverages an ensemble methodology based on the output of two previously trained markovian discriminators. Extensive experiments on both synthetic and real datasets show that our proposed DeflickerCycleGAN not only achieves excellent performance on flicker removal in a single image but also shows high accuracy and competitive generalization ability on flicker detection, compared to that of a well-trained classifier based on ResNet50. Xiaodan Lin, Yangfu Li, Jianqing Zhu, Huanqiang Zeng |
IEEE Trans. Image Process. | 2 |
| 2022 | A Cony-Attention Network for Detecting the Presence of ENF Signal in Short-Duration AudioabstractDetecting the presence of the electric network frequency (ENF) signal in audio recordings is a prerequisite of applying the ENF criterion that plays an essential role in numerous forensic applications. However, existing detection methods are powerless to handle short-duration audio recordings that have attracted considerable attention due to the popularity of voice messaging apps. This paper proposes a novel deep learning-based approach for ENF detection in short audio recordings, reducing the minimum operating range of audio duration to 1/10 of the state-of-the-art methods. Meanwhile, a convolutional attention network termed Conv-AttNet is proposed to improve the detection performance of convolutional neural networks (CNN) through the attention mechanism. Experiments on both synthetic and real-world audio recordings reveal that Conv-AttNet is able to detect the ENF signal buried in only 2 seconds of audio recordings, surpassing both matched filtering and typical CNN like ResNet50. In addition, the detection accuracy can be further increased by utilizing audio recordings of longer duration. Yangfu Li, Xiaodan Lin, Yingqiang Qiu, Huanqiang Zeng |
MMSP | 1 |