VLDB 2026 Research / reviewers in the wild / expert
Yao Luo
dblp:124/9704
· DBLP profile ↗
18ranked-venue papers
5as first author
13since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Why Does the Effective Context Length of LLMs Fall Short?abstractAdvancements in distributed training and efficient attention mechanisms have significantly expanded the context window sizes of large language models (LLMs). However, recent work reveals that the effective context lengths of open-source LLMs often fall short, typically not exceeding half of their training lengths. In this work, we attribute this limitation to the left-skewed frequency distribution of relative positions formed in LLMs pretraining and post-training stages, which impedes their ability to effectively gather distant information.
To address this challenge, we introduce Shifted Rotray Position Embedding (STRING). STRING shifts well-trained positions to overwrite the original ineffective positions during inference, enhancing performance within their existing training lengths.
Experimental results show that without additional training, STRING dramatically improves the performance of the latest large-scale models, such as Llama3.1 70B and Qwen2 72B, by over 10 points on popular long-context benchmarks RULER and InfiniteBench, establishing new state-of-the-art results for open-source LLMs. Compared to commercial models, Llama 3.1 70B with STRING even achieves better performance than GPT-4-128K and clearly surpasses Claude 2 and Kimi-chat. Chenxin An, Jun Zhang 0003, Ming Zhong 0005, Lei Li 0039, Shansan Gong, Yao Luo, Jingjing Xu 0001, Lingpeng Kong |
ICLR | 6 |
| 2025 | FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence InferenceabstractLarge language models (LLMs) encounter computational challenges during long-sequence inference, especially in the attention pre-filling phase, where the complexity grows quadratically with the prompt length. Previous efforts to mitigate these challenges have relied on fixed sparse attention patterns or identifying sparse attention patterns based on limited cases. However, these methods lacked the flexibility to efficiently adapt to varying input demands. In this paper, we introduce FlexPrefill, a Flexible sparse Pre-filling mechanism that dynamically adjusts sparse attention patterns and computational budget in real-time to meet the specific requirements of each input and attention head. The flexibility of our method is demonstrated through two key innovations: 1) Query-Aware Sparse Pattern Determination: By measuring Jensen-Shannon divergence, this component adaptively switches between query-specific diverse attention patterns and predefined attention patterns. 2) Cumulative-Attention Based Index Selection: This component dynamically selects query-key indexes to be computed based on different attention patterns, ensuring the sum of attention scores meets a predefined threshold.
FlexPrefill adaptively optimizes the sparse pattern and sparse ratio of each attention head based on the prompt, enhancing efficiency in long-sequence inference tasks. Experimental results show significant improvements in both speed and accuracy over prior methods, providing a more flexible and efficient solution for LLM inference. Xunhao Lai, Jianqiao Lu, Yao Luo, Yiyuan Ma |
ICLR | 3 |
| 2025 | Model Merging in Pre-training of Large Language ModelsabstractModel merging has emerged as a promising technique for enhancing large language models, though its application in large-scale pre-training remains relatively unexplored. In this paper, we present a comprehensive investigation of model merging techniques during the pre-training process. Through extensive experiments with both dense and Mixture-of-Experts (MoE) architectures ranging from millions to over 100 billion parameters, we demonstrate that merging checkpoints trained with constant learning rates not only achieves significant performance improvements but also enables accurate prediction of annealing behavior. These improvements lead to both more efficient model development and significantly lower training costs. Our detailed ablation studies on merging strategies and hyperparameters provide new insights into the underlying mechanisms while uncovering novel applications. Through comprehensive experimental analysis, we offer the open-source community practical pre-training guidelines for effective model merging. Yunshui Li, Yiyuan Ma, Chaoyi Zhang, Jianqiao Lu, Ziwen Xu, Mengzhao Chen, Minrui Wang, Shiyi Zhan, Xunhao Lai, Yao Luo, Xingyan Bin, Hongbin Ren, Mingji Han, Wenhao Hao, Bairen Yi, LingJun Liu, Bole Ma, Xiaoying Jia 0005 |
NeurIPS | 13 |
| 2025 | Heavy-Haul Train Braking Simulation With Fluid Dynamics-Based Air Braking SystemabstractThe dynamics of heavy-haul trains during braking are highly intricate, directly impacting their operational safety. A heavy-haul train longitudinal-vertical coupled dynamics model (HTLVDM) considering the air braking system is established in this paper. Based on the fluid dynamics theory, a detailed air braking system is established, which considers the characteristics of brake pipes, reservoirs and brake waves. The HTLVDM in-corporates locomotive and wagon components, along with nonlinear hysteresis characteristics of the coupler draft gear. Thus, the air braking system, which delivers pressure signals from the brake pipe to the brake cylinder via a control valve, is coupled to the locomotive and wagon dynamics system through a brake shoe. This model facilitates studying the dynamic performance of the entire system under varied train formations and braking strategies. Using this model, we compare and analyze the impacts of the fluid dynamics model (FDM) and empirical model (EM) of the air braking system on train dynamics behavior. Results reveal significant differences in train dynamic behavior between FDM and EM simulations, highlighting the necessity of employing the FDM for evaluating train operation safety. Kaizhong Liu, Yao Luo |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Dual-view Pyramid Network for Video Frame InterpolationabstractVideo frame interpolation is a critical component of video streaming, a vibrant research area dealing with requests of both service providers and users. However, existing methods cannot handle changing video resolutions while improving user perceptual quality. We aim to unleash the multifaceted knowledge yielded by the hierarchical views at multiple scales in a pyramid network. Specifically, we build a dual-view pyramid network by introducing pyramidal dual-view correspondence matching. It compels each scale to actively seek knowledge in view of both the current scale and a coarser scale, conducting robust correspondence matching by considering neighboring scales. Meanwhile, an auxiliary multi-scale collaborative supervision is devised to enforce the exchange of knowledge among scales and thus reduce error propagation from coarse to fine scales. Based on the robust capture of video dynamics via pyramidal dual-view correspondence matching, we further construct a pyramidal refinement module that formulates frame refinement as progressive latent representation generations by developing flow-guided cross-scale attention for feature fusion among frames. The proposed method is able to improve the perceptual quality on several benchmarks of varying video resolutions, while keeping low distortion and a compact model size. Yao Luo, Ming Yang 0014, Jinhui Tang 0001 |
ACM Multimedia | 1 |
| 2024 | Genotypic-phenotypic landscape computation based on first principle and deep learningabstractThe relationship between genotype and fitness is fundamental to evolution, but quantitatively mapping genotypes to fitness has remained challenging. We propose the Phenotypic-Embedding theorem (P-E theorem) that bridges genotype-phenotype through an encoder-decoder deep learning framework. Inspired by this, we proposed a more general first principle for correlating genotype-phenotype, and the P-E theorem provides a computable basis for the application of first principle. As an application example of the P-E theorem, we developed the Co-attention based Transformer model to bridge Genotype and Fitness model, a Transformer-based pre-train foundation model with downstream supervised fine-tuning that can accurately simulate the neutral evolution of viruses and predict immune escape mutations. Accordingly, following the calculation path of the P-E theorem, we accurately obtained the basic reproduction number (${R}_0$) of SARS-CoV-2 from first principles, quantitatively linked immune escape to viral fitness and plotted the genotype-fitness landscape. The theoretical system we established provides a general and interpretable method to construct genotype-phenotype landscapes, providing a new paradigm for studying theoretical and computational biology. Yuexing Liu, Yao Luo, Ruikun He |
Briefings Bioinform. | 2 |
| 2023 | SVMV: Spatiotemporal Variance-Supervised Motion Volume for Video Frame InterpolationabstractHigh-performance video frame interpolation is challenging for complex scenes with diverse motion and occlusion characteristics. Existing methods, deploying off-the-shelf flow estimators to acquire initial characterizations refined by multiple subsequent models, often require heavy network architectures that are not practical for resource constrained systems. We investigate the unary potentials of the characterizations to improve efficiency. Specifically, we design a lightweight neural network to construct motion volumes via ensembles of offset approximations, and propose a spatiotemporal variance-aware loss to supervise the network learning. For network compactness, our spatiotemporal variance-supervised motion volume (SVMV) utilizes shared spatiotemporal representations via correlations among approximations, of which the diversifications are exploited to better leverage the network’s expressiveness through the spatiotemporal variances of motions and occlusions within the time interval to be interpolated. Experiments on publicly available datasets show that our method performs favorably against existing methods with a more compact network and less runtime. Yao Luo, Jinshan Pan, Jinhui Tang 0001 |
ICASSP | 1 |
| 2022 | Forensic Analysis of JPEG-Domain Enhanced Images via Coefficient Likelihood ModelingabstractJPEG-domain enhancement improves the visual quality of JPEG images by directly manipulating the decoded DCT (discrete cosine transform) coefficients, which inevitably leads to mixed compression and enhancement artifacts. Existing forensic methods that merely consider JPEG artifacts are unsuitable to address such mixed artifacts and hence suffer a considerable performance decline in compression parameter estimation and lack the ability to estimate the enhancement parameter. This work attempts to explore the characterization of the mixed artifacts, and to further estimate both the enhancement and compression parameters of JPEG-domain enhanced images. First, a statistical likelihood function is proposed to characterize the periodicity of DCT coefficients, which can measure how well an enhanced image is de-enhanced back to its JPEG compressed version given the compression and enhancement parameters. The proposed likelihood function reaches its maximum if the parameters match their true values. Then, a forensic method of enhancement detection and parameter estimation is developed based on the proposed likelihood function for two kinds of classical JPEG-domain enhancement. Specifically, JPEG-domain enhanced images are detected by thresholding a scalar feature computed upon the likelihoods, and the enhancement and compression parameters are estimated by locating the maximal likelihood. In addition, mathematical proof of the de-enhancement feasibility is provided. Experimental results demonstrate that the proposed method outperforms the compared methods in both enhancement detection and parameter estimation. Jianquan Yang, Guopu Zhu, Yao Luo, Sam Kwong, Xinpeng Zhang 0001, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Bi-Directional Pseudo-Three-Dimensional Network for Video Frame InterpolationabstractRecent video frame interpolation methods have employed the curvilinear motion model to accommodate nonlinear motion among frames. The effectiveness of such model often hinges on motion estimation and occlusion detection, and therefore is greatly challenged when these methods are used to handle dynamic scenes that contain complex motions and occlusions. We address the challenges by proposing a bi-directional pseudo-three-dimensional network to exploit the correlation between motion estimation and depth-related occlusion estimation that considers the third dimension: depth. Specifically, the network exploits the correlation by learning shared multi-scale spatiotemporal representations, and by coupling the estimations, in both the past and future directions, to synthesize intermediate frames through a bi-directional pseudo-three-dimensional warping layer, where adaptive convolution kernels are estimated progressively from the coalescence of motion and depth-related occlusion estimations across multiple scales to acquire nonlocal and adaptive neighborhoods. The proposed network utilizes a novel multi-task collaborative learning strategy, which facilitates the supervised learning of video frame interpolation using complementary self-supervisory signals from motion and depth-related occlusion estimations. Across various benchmark datasets, the proposed method outperforms state-of-the-art methods in terms of accuracy, model size and runtime performance. Yao Luo, Jinshan Pan, Jinhui Tang 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | A Fusion Framework to Enhance sEMG-Based Gesture Recognition Using TD and FD Features
Yao Luo, Tao Luo 0010, Qianchen Xia, Huijiong Yan, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
ICONIP (6) | 1 |
| 2021 | Bi-branch network for dynamic scene deblurring
Yao Luo, Zhong-Hui Duan, Jinhui Tang 0001 |
Comput. Vis. Image Underst. | 1 |
| 2021 | No-reference omnidirectional video quality assessment based on generative adversarial networks
Jiefeng Guo, Yao Luo |
Multim. Tools Appl. | 2 |
| 2021 | 3D Face Reconstruction From A Single Image Assisted by 2D Face Images in the Wildabstract3D face reconstruction from a single image is an important task in many multimedia applications. Recent works typically learn a CNN-based 3D face model that regresses coefficients of a 3D Morphable Model (3DMM) from 2D images to perform 3D face reconstruction. However, the shortage of training data with 3D annotations considerably limits performance of these methods. To alleviate this issue, we propose a novel 2D-Assisted Learning (2DAL) method that can effectively use “in the wild” 2D face images with noisy landmark information to substantially improve 3D face model learning. Specifically, taking the sparse 2D facial landmark heatmaps as additional information, 2DAL introduces four novel self-supervision schemes that view the 2D landmark and 3D landmark prediction as a self-mapping process, including the landmark self-prediction consistency for 2D and 3D faces respectively, cycle-consistency over the 2D landmark prediction and self-critic over the predicted 3DMM coefficients based on landmark prediction. Using these four self-supervision schemes, 2DAL significantly relieves the demands for the the conventional paired 2D-to-3D annotations and gives much higher-quality 3D face models without requiring any additional 3D annotations. Experiments on AFLW2000-3D, AFLW-LFPA and Florence benchmarks show that our method outperforms state-of-the-arts for both 3D face reconstruction and dense face alignment by a large margin. Xiaoguang Tu, Jian Zhao 0006, Mei Xie, Zihang Jiang, Akshaya Balamurugan, Yao Luo, Yang Zhao 0003, Lingxiao He, Zheng Ma 0005, Jiashi Feng |
IEEE Trans. Multim. | 6 |
| 2020 | Improved InSAR Layover and Shadow Detection using Multi-FeatureabstractLayover and shadow areas in SAR image cause InSAR interferometric phase unwrapping errors in adjacent area. Therefore, they should be detected and marked. Existing detection methods mostly use single feature, exhibiting poor universality in complicated scenes. In this paper, a joint detection method integrating multi-feature is proposed. It uses local frequency estimation and improved eigenvalue decomposition to make preliminary judgments, and then performs joint detection on the results of both. Its effectiveness is experimentally demonstrated from the detected results of simulated and real data: compared with existing methods, the joint detection greatly reduces false-alarm in problem areas detection while improves the accuracy. Huaping Xu, Yao Luo |
IGARSS | 4 |
| 2019 | Poster: A Calibration-free Gaze based Mobile Gesture Control System
Xipeng Ma, Chengkun Jiang, Yao Luo, Qilong Zhao, Meng Jin 0002, Yuan He 0004 |
EWSN | 3 |
| 2019 | TVV: Real-Time Visual Identity and Tracking with Edge Computing
Junchen Guo, Chunya Liu, Yao Luo, Meng Jin 0002, Ziqiang Zhou, Zhoubin Liu |
EWSN | 5 |
| 2018 | IoT for the Power Industry: Recent Advances and Future Directions with PavatarabstractThe development of Internet-of-Things (IoT) technologies in recent years brings us unprecedented opportunities for innovations in the power industry. This demo abstract introduces our research and practice with Pavatar - IoT for the power industry. Pavatar includes a series of system deployments in the core sections of Global Energy Internet (GEI), for the purposes of automatic surveillance and remote diagnosis of ultra-high-voltage converter stations (UHVCSs). Pavatar incorporates technologies like lower-power or battery-free sensing, cross-technology communication, edge computing, machine learning, and enhances the user experience with 3D virtual reality. The deployed system significantly reduces the manpower cost and enhances the operational efficiency of the UHVCS. Yuan He 0004, Junchen Guo, Haozhen Liu, Qilong Zhao, Xiaolong Zheng 0002, Meng Jin 0002, Chunya Liu, Yao Luo, Songzhen Yang, Chengkun Jiang, Xiuzhen Guo |
SenSys | 11 |
| 2015 | K-nearest neighbor based structural twin support vector machine
Xianli Pan, Yao Luo, Yitian Xu |
Knowl. Based Syst. | 2 |