Shijie Xuyang

dblp:346/1097 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Deep learning architectures and training · 22% Efficient and distributed learning · 22% Language models and text generation · 22%
Computer graphics and multimedia
2 papers
Image and video processing · 70% Multimedia analysis and retrieval · 30%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › efficient training
compute-optimal training
0.912025
Predictable Scale (Part II) - Farseer: A Refined Scaling Law in LLMs · NeurIPS 2025
Natural language and speech › Language models and text generation
large language model training
0.912025
Predictable Scale (Part II) - Farseer: A Refined Scaling Law in LLMs · NeurIPS 2025
Machine learning › Deep learning architectures and training
scaling laws
0.912025
Predictable Scale (Part II) - Farseer: A Refined Scaling Law in LLMs · NeurIPS 2025
Computer vision › Image recognition and object detection
scene text recognition
0.712023
STIRER: A Unified Model for Low-Resolution Scene Text Image Recovery and Recognition · ACM Multimedia 2023
Image and video processing › super-resolution
image super-resolution
0.712023
STIRER: A Unified Model for Low-Resolution Scene Text Image Recovery and Recognition · ACM Multimedia 2023
Image and video processing › super-resolution › image super-resolution
scene text image super-resolution
0.712023
STIRER: A Unified Model for Low-Resolution Scene Text Image Recovery and Recognition · ACM Multimedia 2023
Computer vision › Vision and language
visual question answering
0.612022
Towards Video Text Visual Question Answering: Benchmark and Baseline · NeurIPS 2022
Multimedia analysis and retrieval › video understanding
multimodal video understanding
0.612022
Towards Video Text Visual Question Answering: Benchmark and Baseline · NeurIPS 2022
Computer vision › Image recognition and object detection › text recognition
optical character recognition
0.212022
Towards Video Text Visual Question Answering: Benchmark and Baseline · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

swin transformer · 1.3feature encoder-decoder · 1.3multimodal transformer fusion · 1.1answer generation · 1.1OCR token extraction · 1.1scaling law fitting · 0.9loss surface modeling · 0.9
YearPublicationVenuePosition
2026 Anime-2026: A Large-scale Anime Character Dataset for Anime-related AI Tasks
abstract
Anime, as a popular medium, has attracted hundreds of millions audience, especially the young followers. In recent years, various anime-related AI tasks like anime character classification (ACC), retrieval (ACR), tag-prediction (ACTP), question answering (ACQA), and generation (ACG) have been proposed to meet the requirements of various applications. However, there is still a lack of large-scale datasets for these tasks. Such a situation is definitely not beneficial to anime-related academic research and industrial applications. In this paper, to boost anime-related AI technical research and application development, we present Anime-2026, a new and large-scale anime character dataset, which can support various anime-related AI tasks, including ACC, ACR, ACTP, ACQA and ACG etc. Anime-2026 consists of 1.5M anime character images, 14k different characters, 16k unique semantic keyword tags, 10k question-answer pairs, and 4k manually designed text queries by crowdsourcing for the ACR and ACG tasks. Furthermore, to assess the dataset, we re-implement a number of generic and anime-specific AI baseline models, and conduct extensive experiments to evaluate these models on Anime-2026. In summary, as a general benchmark dataset, Anime-2026 provides the largest free anime character resource to support future anime-related AI research and development. We expect that Anime-2026 will promote the R&D of new and more advanced models and methods of various anime-related AI tasks. The dataset is available on https://huggingface.co/datasets/miaojiemiao/Anime-2026.
Shijie Xuyang, Bingzhe Yu, Minyi Zhao, Guangze Li, Jihong Guan, Shuigeng Zhou
ICMR1
2025 Predictable Scale (Part II) - Farseer: A Refined Scaling Law in LLMs
abstract
Training Large Language Models (LLMs) is prohibitively expensive, creating a critical scaling gap where insights from small-scale experiments often fail to transfer to resource-intensive production systems, thereby hindering efficient innovation. To bridge this, we introduce Farseer, a novel and refined scaling law offering enhanced predictive accuracy across scales. By systematically constructing a model loss surface $L(N,D)$, Farseer achieves a significantly better fit to empirical data than prior laws (e.g., \Chinchilla's law). Our methodology yields accurate, robust, and highly generalizable predictions, demonstrating excellent extrapolation capabilities, outperforming Chinchilla's law, whose extrapolation error is 433\% higher. This allows for the reliable evaluation of competing training strategies across all $(N,D)$ settings, enabling conclusions from small-scale ablation studies to be confidently extrapolated to predict large-scale performance. Furthermore, Farseer provides new insights into optimal compute allocation, better reflecting the nuanced demands of modern LLM training. To validate our approach, we trained an extensive suite of approximately 1,000 LLMs across diverse scales and configurations, consuming roughly 3 million NVIDIA H100 GPU hours. To foster further research, we are comprehensively open-sourcing all code, data, results (https://github.com/Farseer-Scaling-Law/Farseer), all training logs (https://wandb.ai/billzid/Farseer?nw=nwuserbillzid), all models used in scaling law fitting (https://huggingface.co/Farseer-Scaling-Law).
Houyi Li, Wenzhen Zheng, Zhenyu Ding, Haoying Wang, Shijie Xuyang, Ning Ding 0006, Shuigeng Zhou, Xiangyu Zhang 0005, Daxin Jiang
NeurIPS7
2023 STIRER: A Unified Model for Low-Resolution Scene Text Image Recovery and Recognition
abstract
Though scene text recognition (STR) from high-resolution (HR) images has achieved significant success in the past years, text recognition from low-resolution (LR) images is still a challenging task. This inspires the study on scene text image super-resolution (STISR) to generate super-resolution (SR) images based on the LR images, then STR is performed on the generated SR images, which eventually boosts the recognition performance. However, existing methods have two major drawbacks: 1) STISR models may generate imperfect SR images, which mislead the subsequent recognition. 2) As the STISR models are optimized for high recognition accuracy, the fidelity of SR images may be degraded. Consequently, neither the recognition performance of STR nor the fidelity of STISR is desirable. In this paper, a novel model called STIRER (the abbreviation of Scene Text Image REcovery and Recognition) is proposed to effectively and simultaneously recover and recognize LR scene text images under a unified framework. Concretely, STIRER consists of a feature encoder to obtain pixel features and two dedicated decoders to generate SR images and recognize texts respectively based on the encoded features and the raw LR images. We propose a progressive scene text swin transformer architecture as the encoder to enrich the representations of the pixel features for better recovery and recognition. Extensive experiments on two LR datasets show the superiority of our model to the existing methods on recognition performance, super-resolution fidelity and computational cost. The STIRER Code is available in https://github.com/zhaominyiz/STIRER.
Minyi Zhao, Shijie Xuyang, Jihong Guan, Shuigeng Zhou
ACM Multimedia2
2022 Towards Video Text Visual Question Answering: Benchmark and Baseline
abstract
There are already some text-based visual question answering (TextVQA) benchmarks for developing machine's ability to answer questions based on texts in images in recent years. However, models developed on these benchmarks cannot work effectively in many real-life scenarios (e.g. traffic monitoring, shopping ads and e-learning videos) where temporal reasoning ability is required. To this end, we propose a new task named Video Text Visual Question Answering (ViteVQA in short) that aims at answering questions by reasoning texts and visual information spatiotemporally in a given video. In particular, on the one hand, we build the first ViteVQA benchmark dataset named M4-ViteVQA --- the abbreviation of Multi-category Multi-frame Multi-resolution Multi-modal benchmark for ViteVQA, which contains 7,620 video clips of 9 categories (i.e., shopping, traveling, driving, vlog, sport, advertisement, movie, game and talking) and 3 kinds of resolutions (i.e., 720p, 1080p and 1176x664), and 25,123 question-answer pairs. On the other hand, we develop a baseline method named T5-ViteVQA for the ViteVQA task. T5-ViteVQA consists of five transformers. It first extracts optical character recognition (OCR) tokens, question features, and video representations via two OCR transformers, one language transformer and one video-language transformer, respectively. Then, a multimodal fusion transformer and an answer generation module are applied to fuse multimodal information and generate the final prediction. Extensive experiments on M4-ViteVQA demonstrate the superiority of T5-ViteVQA to the existing approaches of TextVQA and VQA tasks. The ViteVQA benchmark is available in https://github.com/bytedance/VTVQA.
Minyi Zhao, Bingjia Li, Wanqing Li 0007, Shijie Xuyang, Zhihang Yu, Xinkun Yu, Guangze Li, Aobotao Dai, Shuigeng Zhou
NeurIPS7