Xinyue Li 0001

dblp:33/5667-1 · also XinYue Li 0001 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
12since 2021 · last 2026
0000-0001-7362-0532ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Lumina-mGPT: Flexible Photorealistic Autoregressive Text-to-Image Generation
Yi Xin 0003, Shitian Zhao, Le Zhuo, Weifeng Lin, Xinyue Li 0001, Guangtao Zhai, Xiaohong Liu 0001, Hongsheng Li 0001, Yu Qiao 0001, Peng Gao 0007
Int. J. Comput. Vis.6
2026 LMVQ: Label-Free Metric-Learning for General AI-Generated Video Quality Assessment
abstract
The recent rapid development of video generation technology has led to a significant demand for quality assessment of the latest AI-generated videos. However, current supervised approaches depend on expensive and quickly outdated human scores, and label-free methods overlook the general distortions of AI-generated videos. To address these limitations, we introduce LMVQ, a Label-free Metric-learning framework for general AI-generated Video Quality assessment of three dimensions, spatial, temporal, and alignment. The LMVQ is the first to introduce sample degradations specially designed for AIGC-specific distortions, and constructs a comprehensive training set through two complementary sample generation strategies. It then employs two synergistic modules, the Intra-Quality Token Transformer (IQ-Trans), which explicitly refines dimension-specific quality representations, and the Inter-Quality Mixture of Experts (IQ-MoE), which fuses interactions across multiple quality dimensions. Finally, a Multi-Proxy Metric-Learning (MPML) strategy aligns the learned representations with multi-dimensional quality scores and constrains the model to learn discriminative quality-aware representations. Extensive experiments on four public AIGC-VQA benchmarks show that MPML outperforms previous label-free methods by over 20%, and greatly narrows the gap with supervised methods. This provides a scalable, adaptive foundation for evaluating the ever-evolving quality of AI-generated videos.
Xinyue Li 0001, Wei Sun 0029, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai
IEEE Trans. Circuits Syst. Video Technol.2
2025 Lumina-Image 2.0: a Unified and Efficient Image Generative Framework
abstract
We introduce Lumina-Image 2.0, an advanced text-to-image generation framework that achieves significant progress compared to previous work, Lumina-Next. Lumina-Image 2.0 is built upon two key principles: (1) Unification - it adopts a unified architecture (Unified Next-DiT) that treats text and image tokens as a joint sequence, enabling natural cross-modal interactions and allowing seamless task expansion. Besides, since high-quality captioners can provide semantically well-aligned text-image training pairs, we introduce a unified captioning system, Unified Captioner (UniCap), specifically designed for T2I generation tasks. UniCap excels at generating comprehensive and accurate captions, accelerating convergence and enhancing prompt adherence. (2) Efficiency - to improve the efficiency of our proposed model, we develop multi-stage progressive training strategies and introduce inference acceleration techniques without compromising image quality. Extensive evaluations on academic benchmarks and public text-to-image arenas show that Lumina-Image 2.0 delivers strong performances even with only 2.6B parameters, highlighting its scalability and design efficiency. We have released our training details, code, and models at https://github.com/Alpha-VLLM/Lumina-Image-2.0.
Le Zhuo, Yi Xin 0003, Ruoyi Du, Zhen Li 0026, Yiting Lu, Xinyue Li 0001, Will Beddow, Erwann Millon, Victor Perez 0005, Wenhai Wang, Yu Qiao 0001, Bo Zhang 0069, Xiaohong Liu 0001, Hongsheng Li 0001, Chang Xu 0002, Peng Gao 0007
ICCV8
2025 LPerceptual Quality Assessment of AI Generated Content Videos: a Dataset and Benchmark
abstract
In recent years, artificial intelligence (AI) driven video generation has garnered significant attention due to advancements in large language model techniques. Thus, there is a great demand to explore the effectiveness of video quality assessment (VQA) models in evaluating the perceptual quality of AI-generated content (AIGC) videos and in optimizing video generation techniques. Therefore, in this paper, we try to systemically investigate the AIGC-VQA problem from both subjective and objective quality assessment perspectives. For the subjective perspective, we construct a Large-scale Generated Video Quality assessment (LGVQ) dataset, consisting of 2,808 AIGC videos generated by 6 video generation models using 468 carefully selected text prompts. We evaluate the perceptual quality of AIGC videos from three dimensions: spatial quality, temporal quality, and text-to-video alignment, which hold the utmost importance for current video generation techniques. For the objective perspective, we establish a benchmark for evaluating existing quality assessment metrics on the LGVQ dataset, which fully demonstrates the performance of current mainstream VQA methods in evaluating AIGV quality. We hope that this work can contribute to the advancement of AIGC video generation technology as well as the evaluation techniques for AIGC videos. The LGVQ dataset will release publicly.
Wei Sun 0029, Xinyue Li 0001, Jun Jia, Xiongkuo Min, Chunyi Li 0001, Zhongpeng Ji, Fengyu Sun, Shangling Jui, Guangtao Zhai
ISCAS3
2025 Human-Activity AGV Quality Assessment: A Benchmark Dataset and an Objective Evaluation Metric
abstract
AI-driven video generation techniques have made significant progress in recent years. However, AI-generated videos (AGVs) involving human activities often exhibit substantial visual and semantic distortions, hindering the practical application of video generation technologies in real-world scenarios. To address this challenge, we conduct a pioneering study on human activity AGV quality assessment, focusing on visual quality evaluation and the identification of semantic distortions. First, we construct the AI-Generated Human activity Video Quality Assessment (Human-AGVQA) dataset, consisting of 6,000 AGVs derived from 15 popular text-to-video (T2V) models using 400 text prompts that describe diverse human activities. We conduct a subjective study to evaluate the human appearance quality, action continuity quality, and overall video quality of AGVs, and identify semantic issues of human body parts. Based on Human-AGVQA, we benchmark the performance of T2V models and analyze their strengths and weaknesses in generating different categories of human activities. Second, we develop an objective evaluation metric, named AI-Generated Human activity Video Quality metric (GHVQ), to automatically analyze the quality of human activity AGVs. GHVQ systematically extracts human-focused quality features, AI-generated content-aware quality features, and temporal continuity features, making it a comprehensive and explainable quality metric for human activity AGVs. The extensive experimental results show that GHVQ outperforms existing quality metrics on the Human-AGVQA dataset by a large margin, demonstrating its efficacy in assessing the quality of human activity AGVs. The Human-AGVQA dataset and GHVQ metric will be released at https://github.com/zczhang-sjtu/GHVQ.git.
Wei Sun 0029, Xinyue Li 0001, Qihang Ge, Jun Jia, Zhongpeng Ji, Fengyu Sun, Shangling Jui, Xiongkuo Min, Guangtao Zhai
ACM Multimedia3
2025 Sekai: A Video Dataset towards World Exploration
abstract
Video generation techniques have made remarkable progress, promising to be the foundation of interactive world exploration.However, existing video generation datasets are not well-suited for world exploration training as they suffer from some limitations: limited locations, short duration, static scenes, and a lack of annotations about exploration and the world.In this paper, we introduce Sekai (meaning "world" in Japanese), a high-quality first-person view worldwide video dataset with rich annotations for world exploration. It consists of over 5,000 hours of walking or drone view (FPV and UVA) videos from over 100 countries and regions across 750 cities. We develop an efficient and effective toolbox to collect, pre-process and annotate videos with location, scene, weather, crowd density, captions, and camera trajectories.Comprehensive analyses and experiments demonstrate the dataset’s scale, diversity, annotation quality, and effectiveness for training video generation models.We believe Sekai will benefit the area of video generation and world exploration, and motivate valuable applications.
Zhen Li 0026, Chuanhao Li 0001, Xiaofeng Mao, Shaoheng Lin, Ming Li 0010, Shitian Zhao, Zhaopan Xu, Xinyue Li 0001, Yukang Feng, Zizhen Li, Fanrui Zhang, Jiaxin Ai, Yuwei Wu 0001, Tong He 0001, Yunde Jia, Kaipeng Zhang
NeurIPS8
2025 Situation-adaptive neural network for fast pre-computing image enhancement
Xinyue Li 0001, Huiyu Duan, Jia Wang 0004, Xiaohong Liu 0001, Guangtao Zhai
Sci. China Inf. Sci.1
2025 Quality-Guided Skin Tone Enhancement for Portrait Photography
abstract
In recent years, learning-based color and tone enhancement methods for photos have become increasingly popular. However, most learning-based image enhancement methods just learn a mapping from one distribution to another based on one dataset, lacking the ability to adjust images continuously and controllably. It is important to enable the learning-based enhancement models to adjust an image continuously, since in many cases we may want to get a slighter or stronger enhancement effect rather than one fixed adjusted result. In this paper, we propose a quality-guided image enhancement paradigm that enables image enhancement models to learn the distribution of images with various quality ratings. By learning this distribution, image enhancement models can associate image features with their corresponding perceptual qualities, which can be used to adjust images continuously according to different quality scores. To validate the effectiveness of our proposed method, a subjective quality assessment experiment is first conducted, focusing on skin tone adjustment in portrait photography. Guided by the subjective quality ratings obtained from this experiment, our method can adjust the skin tone corresponding to different quality requirements. Furthermore, an experiment conducted on 10 natural raw images corroborates the effectiveness of our model in situations with fewer subjects and fewer shots, and also demonstrates its general applicability to natural images.
Shiqi Gao, Huiyu Duan, Xinyue Li 0001, Yicong Peng, Qihang Xu, Yuanyuan Chang, Jia Wang 0004, Xiongkuo Min, Guangtao Zhai
IEEE Trans. Multim.3
2025 Benchmarking Multi-dimensional AIGC Video Quality Assessment: A Dataset and Unified Model
abstract
In recent years, AI-driven video generation has gained significant attention due to great advancements in visual and language generative techniques. Consequently, there is a growing need for accurate Video Quality Assessment (VQA) metrics to evaluate the perceptual quality of AI-generated content (AIGC) videos and optimize video generation models. However, assessing the quality of AIGC videos remains a significant challenge because these videos often exhibit highly complex distortions, such as unnatural actions and irrational objects. To address this challenge, we systematically investigate the AIGC-VQA problem in this article, considering both subjective and objective quality assessment perspectives. For the subjective perspective, we construct the L arge-scale G enerated V ideo Q uality Assessment (LGVQ) dataset, consisting of \(2,\!808\) AIGC videos generated by six video generation models using 468 carefully curated text prompts. Unlike previous subjective VQA experiments, we evaluate the perceptual quality of AIGC videos from three critical dimensions: spatial quality, temporal quality, and text-video alignment, which hold utmost importance for current video generation techniques. For the objective perspective, we establish a benchmark for evaluating existing quality assessment metrics on the LGVQ dataset. Our findings show that current metrics perform poorly on this dataset, highlighting a gap in effective evaluation tools. To bridge this gap, we propose the U nify G enerated V ideo Q uality Assessment (UGVQ) model, designed to accurately evaluate the multi-dimensional quality of AIGC videos. The UGVQ model integrates the visual and motion features of videos with the textual features of their corresponding prompts, forming a unified quality-aware feature representation tailored to AIGC videos. Experimental results demonstrate that UGVQ achieves state-of-the-art performance on the LGVQ dataset across all three quality dimensions, validating its effectiveness as an accurate quality metric for AIGC videos. We hope that our benchmark can promote the development of AIGC-VQA studies. Both the LGVQ dataset and the UGVQ model are publicly available on https://github.com/zczhang-sjtu/UGVQ.git .
Wei Sun 0029, Xinyue Li 0001, Jun Jia, Xiongkuo Min, Chunyi Li 0001, Zijian Chen 0001, Puyi Wang, Fengyu Sun, Shangling Jui, Guangtao Zhai
ACM Trans. Multim. Comput. Commun. Appl.3
2023 Ranked Similarity Weighting and Top-nk Sampling in Deep Metric Learning
abstract
Deep metric learning has been widely used in many visual tasks. Its key idea is to increase the similarity of positive samples and decrease the similarity of negative samples through network training. To achieve this purpose, many studies excessively extend the distance between the query sample and hard negative samples. This may compress the distance between similar samples of other classes, causing these samples to cluster together. We call this phenomenon Negative Sample Aggregation. To address this problem, first, we propose a weighting method based on the Ranking Similarity of sample pairs, short for RS. The proposed weighting method can not only enlarge the distance between the query sample and hard negative samples, but also maintain the embedding distribution of proximal negative samples. Second, we propose a Top-nk sampling method, which can dynamically adjust the sampling strategy according to the distribution of a dataset. It solves the problem that the descent direction of the network gradient is inconsistent with the optimization target. The effectiveness of our methods is evaluated by extensive experiments on four public datasets and compared with that of other state-of-the-art methods. The results show that the proposed method obtains excellent performance, reaching 67.8% on CUB-200-2011 and 85.2% on Cars-196 at Recall@1.
Jian Wang 0130, Xinyue Li 0001, Wei Song 0007, Weiqi Guo
IEEE Trans. Multim.2
2022 Multi-Hierarchy Proxy Structure for Deep Metric Learning
abstract
Mainstream methods for deep metric learning can be divided into pair-based and proxy-based methods. In recent years, proxy-based methods have attracted wide attention for their low training complexity and fast network convergence. Most proxy-based studies assign only one proxy per class to capture the features of the class, this leads to ignoring the hidden hierarchy and regular aggregation of features within the class. However, these details are meaningful for capturing features of the class. Therefore, we propose a multi-hierarchy proxy (MHP) structure to extract the hierarchical details and regular features hidden in the embedding space. At the same time, we design a layerwise merging similarity operator to reasonably measure the similarity between samples and classes. Our MHP method maintains the low time complexity of the proxy-based method and can be easily integrated into existing proxy-based losses. The effectiveness of our method is evaluated by extensive experiments on three public datasets and compared with state-of-the-art methods. The results show that the proposed MHP method can significantly improve the performance of proxy-based methods, reaching 69.8% on CUB-2002011 and 87.4% on Cars-196 dataset at Recall@1.
Jian Wang 0130, Xinyue Li 0001, Wei Song 0007, Weiqi Guo
ICASSP2
2021 A Ranked Similarity Loss Function with pair Weighting for Deep Metric Learning
abstract
Metric learning is a widely-used method for image retrieval. The object of metric learning is to limit the distance between similar samples and increase the distance between samples of different classes through learning. Many studies tend to pay more attention to keep the distance between positive and negative samples, but ignore the distance between different classes of negative samples. In fact, query samples should be separated from negative samples of different classes by different distances. To address these problems, we propose to build a ranked similarity loss function with pair weighting (dubbed RMS loss). The proposed RMS loss can keep a distance between samples of different classes by weighting the negative samples according to the sorting order. Meanwhile, it further widens the distance between positive and negative samples by different processing of similarity of positive pairs and negative pairs. The effectiveness of our method is evaluated by extensive experiments on four public datasets and compared with state-of-the-art methods. The results show the proposed method obtains new performance on four public datasets, e.g., reaching 67.4% on CUB200 at Recall@1.
Jian Wang 0130, Dongmei Huang 0001, Wei Song 0007, Quanmiao Wei, Xinyue Li 0001
ICASSP6