EDBT 2026 Demo / reviewers in the wild / expert
Shijie Xuyang
dblp:346/1097
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Deep learning architectures and training · 22% Efficient and distributed learning · 22% Language models and text generation · 22% | |
| Computer graphics and multimedia
2 papers |
Image and video processing · 70% Multimedia analysis and retrieval · 30% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning › efficient training
compute-optimal training |
0.9 | 1 | 2025 | Predictable Scale (Part II) - Farseer: A Refined Scaling Law in LLMs · NeurIPS 2025 |
Natural language and speech › Language models and text generation
large language model training |
0.9 | 1 | 2025 | Predictable Scale (Part II) - Farseer: A Refined Scaling Law in LLMs · NeurIPS 2025 |
Machine learning › Deep learning architectures and training
scaling laws |
0.9 | 1 | 2025 | Predictable Scale (Part II) - Farseer: A Refined Scaling Law in LLMs · NeurIPS 2025 |
Computer vision › Image recognition and object detection
scene text recognition |
0.7 | 1 | 2023 | STIRER: A Unified Model for Low-Resolution Scene Text Image Recovery and Recognition · ACM Multimedia 2023 |
Image and video processing › super-resolution
image super-resolution |
0.7 | 1 | 2023 | STIRER: A Unified Model for Low-Resolution Scene Text Image Recovery and Recognition · ACM Multimedia 2023 |
Image and video processing › super-resolution › image super-resolution
scene text image super-resolution |
0.7 | 1 | 2023 | STIRER: A Unified Model for Low-Resolution Scene Text Image Recovery and Recognition · ACM Multimedia 2023 |
Computer vision › Vision and language
visual question answering |
0.6 | 1 | 2022 | Towards Video Text Visual Question Answering: Benchmark and Baseline · NeurIPS 2022 |
Multimedia analysis and retrieval › video understanding
multimodal video understanding |
0.6 | 1 | 2022 | Towards Video Text Visual Question Answering: Benchmark and Baseline · NeurIPS 2022 |
Computer vision › Image recognition and object detection › text recognition
optical character recognition |
0.2 | 1 | 2022 | Towards Video Text Visual Question Answering: Benchmark and Baseline · NeurIPS 2022 |
Methods — techniques the papers use, named apart from their topics
swin transformer · 1.3feature encoder-decoder · 1.3multimodal transformer fusion · 1.1answer generation · 1.1OCR token extraction · 1.1scaling law fitting · 0.9loss surface modeling · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Anime-2026: A Large-scale Anime Character Dataset for Anime-related AI TasksabstractAnime, as a popular medium, has attracted hundreds of millions audience, especially the young followers. In recent years, various anime-related AI tasks like anime character classification (ACC), retrieval (ACR), tag-prediction (ACTP), question answering (ACQA), and generation (ACG) have been proposed to meet the requirements of various applications. However, there is still a lack of large-scale datasets for these tasks. Such a situation is definitely not beneficial to anime-related academic research and industrial applications. In this paper, to boost anime-related AI technical research and application development, we present Anime-2026, a new and large-scale anime character dataset, which can support various anime-related AI tasks, including ACC, ACR, ACTP, ACQA and ACG etc. Anime-2026 consists of 1.5M anime character images, 14k different characters, 16k unique semantic keyword tags, 10k question-answer pairs, and 4k manually designed text queries by crowdsourcing for the ACR and ACG tasks. Furthermore, to assess the dataset, we re-implement a number of generic and anime-specific AI baseline models, and conduct extensive experiments to evaluate these models on Anime-2026. In summary, as a general benchmark dataset, Anime-2026 provides the largest free anime character resource to support future anime-related AI research and development. We expect that Anime-2026 will promote the R&D of new and more advanced models and methods of various anime-related AI tasks. The dataset is available on https://huggingface.co/datasets/miaojiemiao/Anime-2026. Shijie Xuyang, Bingzhe Yu, Minyi Zhao, Guangze Li, Jihong Guan, Shuigeng Zhou |
ICMR | 1 |
| 2025 | Predictable Scale (Part II) - Farseer: A Refined Scaling Law in LLMsabstractTraining Large Language Models (LLMs) is prohibitively expensive, creating a critical scaling gap where insights from small-scale experiments often fail to transfer to resource-intensive production systems, thereby hindering efficient innovation.
To bridge this, we introduce Farseer, a novel and refined scaling law offering enhanced predictive accuracy across scales. By systematically constructing a model loss surface $L(N,D)$, Farseer achieves a significantly better fit to empirical data than prior laws (e.g., \Chinchilla's law).
Our methodology yields accurate, robust, and highly generalizable predictions, demonstrating excellent extrapolation capabilities, outperforming Chinchilla's law, whose extrapolation error is 433\% higher.
This allows for the reliable evaluation of competing training strategies across all $(N,D)$ settings, enabling conclusions from small-scale ablation studies to be confidently extrapolated to predict large-scale performance.
Furthermore, Farseer provides new insights into optimal compute allocation, better reflecting the nuanced demands of modern LLM training. To validate our approach, we trained an extensive suite of approximately 1,000 LLMs across diverse scales and configurations, consuming roughly 3 million NVIDIA H100 GPU hours. To foster further research, we are comprehensively open-sourcing all code, data, results (https://github.com/Farseer-Scaling-Law/Farseer), all training logs (https://wandb.ai/billzid/Farseer?nw=nwuserbillzid), all models used in scaling law fitting (https://huggingface.co/Farseer-Scaling-Law). Houyi Li, Wenzhen Zheng, Zhenyu Ding, Haoying Wang, Shijie Xuyang, Ning Ding 0006, Shuigeng Zhou, Xiangyu Zhang 0005, Daxin Jiang |
NeurIPS | 7 |
| 2023 | STIRER: A Unified Model for Low-Resolution Scene Text Image Recovery and RecognitionabstractThough scene text recognition (STR) from high-resolution (HR) images has achieved significant success in the past years, text recognition from low-resolution (LR) images is still a challenging task. This inspires the study on scene text image super-resolution (STISR) to generate super-resolution (SR) images based on the LR images, then STR is performed on the generated SR images, which eventually boosts the recognition performance. However, existing methods have two major drawbacks: 1) STISR models may generate imperfect SR images, which mislead the subsequent recognition. 2) As the STISR models are optimized for high recognition accuracy, the fidelity of SR images may be degraded. Consequently, neither the recognition performance of STR nor the fidelity of STISR is desirable. In this paper, a novel model called STIRER (the abbreviation of Scene Text Image REcovery and Recognition) is proposed to effectively and simultaneously recover and recognize LR scene text images under a unified framework. Concretely, STIRER consists of a feature encoder to obtain pixel features and two dedicated decoders to generate SR images and recognize texts respectively based on the encoded features and the raw LR images. We propose a progressive scene text swin transformer architecture as the encoder to enrich the representations of the pixel features for better recovery and recognition. Extensive experiments on two LR datasets show the superiority of our model to the existing methods on recognition performance, super-resolution fidelity and computational cost. The STIRER Code is available in https://github.com/zhaominyiz/STIRER. Minyi Zhao, Shijie Xuyang, Jihong Guan, Shuigeng Zhou |
ACM Multimedia | 2 |
| 2022 | Towards Video Text Visual Question Answering: Benchmark and BaselineabstractThere are already some text-based visual question answering (TextVQA) benchmarks for developing machine's ability to answer questions based on texts in images in recent years. However, models developed on these benchmarks cannot work effectively in many real-life scenarios (e.g. traffic monitoring, shopping ads and e-learning videos) where temporal reasoning ability is required. To this end, we propose a new task named Video Text Visual Question Answering (ViteVQA in short) that aims at answering questions by reasoning texts and visual information spatiotemporally in a given video. In particular, on the one hand, we build the first ViteVQA benchmark dataset named M4-ViteVQA --- the abbreviation of Multi-category Multi-frame Multi-resolution Multi-modal benchmark for ViteVQA, which contains 7,620 video clips of 9 categories (i.e., shopping, traveling, driving, vlog, sport, advertisement, movie, game and talking) and 3 kinds of resolutions (i.e., 720p, 1080p and 1176x664), and 25,123 question-answer pairs. On the other hand, we develop a baseline method named T5-ViteVQA for the ViteVQA task. T5-ViteVQA consists of five transformers. It first extracts optical character recognition (OCR) tokens, question features, and video representations via two OCR transformers, one language transformer and one video-language transformer, respectively. Then, a multimodal fusion transformer and an answer generation module are applied to fuse multimodal information and generate the final prediction. Extensive experiments on M4-ViteVQA demonstrate the superiority of T5-ViteVQA to the existing approaches of TextVQA and VQA tasks. The ViteVQA benchmark is available in https://github.com/bytedance/VTVQA. Minyi Zhao, Bingjia Li, Wanqing Li 0007, Shijie Xuyang, Zhihang Yu, Xinkun Yu, Guangze Li, Aobotao Dai, Shuigeng Zhou |
NeurIPS | 7 |