Zilin Xiao

dblp:330/7498 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0001-8686-718XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Language models and text generation · 36% Reinforcement learning · 18% Efficient and distributed learning · 18%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 77% Software maintenance and evolution · 23%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model reasoning
0.912025
RAST: Reasoning Activation in LLMs via Small-model Transfer · NeurIPS 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
RAST: Reasoning Activation in LLMs via Small-model Transfer · NeurIPS 2025
Machine learning › Reinforcement learning
reinforcement learning for reasoning
0.912025
RAST: Reasoning Activation in LLMs via Small-model Transfer · NeurIPS 2025
Information retrieval › image retrieval › web image search
image re-ranking
0.912025
LOCORE: Image Re-ranking with Long-Context Sequence Modeling · CVPR 2025
Information retrieval
image retrieval
0.912025
LOCORE: Image Re-ranking with Long-Context Sequence Modeling · CVPR 2025
Information retrieval › reranking
listwise reranking
0.912025
LOCORE: Image Re-ranking with Long-Context Sequence Modeling · CVPR 2025
Computer vision › Vision and language
visual entity recognition
0.812024
Grounding Language Models for Visual Entity Recognition · ECCV (11) 2024
Natural language and speech › Information extraction and text analysis
entity linking
0.712023
Instructed Language Models with Retrievers Are Powerful Entity Linkers · EMNLP 2023
Natural language and speech › Language models and text generation
instruction tuning
0.712023
Instructed Language Models with Retrievers Are Powerful Entity Linkers · EMNLP 2023
Natural language and speech › Language models and text generation
mathematical reasoning
0.312025
RAST: Reasoning Activation in LLMs via Small-model Transfer · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

sliding window · 0.9reinforcement learning · 0.9probability distribution transfer · 0.9long-context sequence model · 0.9large language model · 0.9language model · 0.8sequence-to-sequence training · 0.7potential mention retriever · 0.7in-context learning · 0.7
YearPublicationVenuePosition
2025 LOCORE: Image Re-ranking with Long-Context Sequence Modeling
abstract
We introduce LoCoRe, Long-Context Re-ranker, a model that takes as input local descriptors corresponding to an image query and a list of gallery images and outputs similarity scores between the query and each gallery image. This model is used for image retrieval, where typically a first ranking is performed with an efficient similarity measure, and then a shortlist of top-ranked images is re-ranked based on a more fine-grained similarity model. Compared to existing methods that perform pair-wise similarity estimation with local descriptors or list-wise re-ranking with global descriptors, LoCoRe is the first method to perform list-wise re-ranking with local descriptors. To achieve this, we leverage efficient long-context sequence models to effectively capture the dependencies between query and gallery images at the local-descriptor level. During testing, we process long shortlists with a sliding window strategy that is tailored to overcome the context size limitations of sequence models. Our approach achieves superior performance compared with other re-rankers on established image retrieval benchmarks of landmarks ($\mathcal{R}{\text{Oxf}}$ and $\mathcal{R}{\text{Par}}$), products (SOP), fashion items (In-Shop), and bird species (CUB-200) while having comparable latency to the pair-wise local descriptor re-rankers.
Zilin Xiao, Pavel Suma, Ayush Sachdeva, Hao-Jen Wang, Giorgos Kordopatis-Zilos, Giorgos Tolias, Vicente Ordonez
CVPR1
2025 RepoGraph: Enhancing AI Software Engineering with Repository-level Code Graph
abstract
Large Language Models (LLMs) excel in code generation yet struggle with modern AI software engineering tasks. Unlike traditional function-level or file-level coding tasks, AI software engineering requires not only basic coding proficiency but also advanced skills in managing and interacting with code repositories. However, existing methods often overlook the need for repository-level code understanding, which is crucial for accurately grasping the broader context and developing effective solutions. On this basis, we present RepoGraph, a plug-in module that manages a repository-level structure for modern AI software engineering solutions. RepoGraph offers the desired guidance and serves as a repository-wide navigation for AI software engineers. We evaluate RepoGraph on the SWE-bench by plugging it into four different methods of two lines of approaches, where RepoGraph substantially boosts the performance of all systems, leading to a new state-of-the-art among open-source frameworks. Our analyses also demonstrate the extensibility and flexibility of RepoGraph by testing on another repo-level coding benchmark, CrossCodeEval. Our code is available at https://github.com/ozyyshr/RepoGraph.
Siru Ouyang, Wenhao Yu 0002, Kaixin Ma, Zilin Xiao, Zhihan Zhang 0001, Mengzhao Jia, Jiawei Han 0001, Hongming Zhang 0009, Dong Yu 0001
ICLR4
2025 RAST: Reasoning Activation in LLMs via Small-model Transfer
abstract
Reinforcement learning (RL) has become a powerful approach for improving the reasoning capabilities of large language models (LLMs), as evidenced by recent successes such as OpenAI's o1 and Deepseek-R1. However, applying RL at scale remains intimidatingly resource-intensive, requiring multiple model copies and extensive GPU workloads. On the other hand, while being powerful, recent studies suggest that RL does not fundamentally endow models with new knowledge; rather, it primarily reshapes the model's output distribution to activate reasoning capabilities latent in the base model. Building on this insight, we hypothesize that the changes in output probabilities induced by RL are largely model-size invariant, opening the door to a more efficient paradigm: training a small model with RL and transferring its induced probability shifts to larger base models. To verify our hypothesis, we conduct a token-level analysis of decoding trajectories and find high alignment in RL-induced output distributions across model scales, validating our hypothesis. Motivated by this, we propose RAST, a simple yet effective method that transfers reasoning behaviors by injecting RL-induced probability adjustments from a small RL-trained model into larger models. Experiments across multiple mathematical reasoning benchmarks show that RAST substantially and consistently enhances the reasoning capabilities of base models while requiring significantly lower GPU memory than direct RL training, sometimes even yielding better performance than the RL-trained counterparts. Our findings offer new insights into the nature of RL-driven reasoning and practical strategies for scaling its benefits without incurring its full computational cost. The project page of RAST is available at https://ozyyshr.github.io/RAST/.
Siru Ouyang, Zilin Xiao, Minhao Jiang, Yu Meng 0001, Jiawei Han 0001
NeurIPS3
2024 Grounding Language Models for Visual Entity Recognition
Zilin Xiao, Ming Gong 0001, Paola Cascante-Bonilla, Xingyao Zhang 0003, Jie Wu 0018, Vicente Ordonez
ECCV (11)1
2023 Instructed Language Models with Retrievers Are Powerful Entity Linkers
abstract
Generative approaches powered by large language models (LLMs) have demonstrated emergent abilities in tasks that require complex reasoning abilities.Yet the generative nature still makes the generated content suffer from hallucinations, thus unsuitable for entity-centric tasks like entity linking (EL) requiring precise entity predictions over a large knowledge base.We present Instructed Generative Entity Linker (INSGENEL), the first approach that enables casual language models to perform entity linking over knowledge bases.Several methods to equip language models with EL capability were proposed in this work, including (i) a sequence-to-sequence training EL objective with instruction-tuning, (ii) a novel generative EL framework based on a light-weight potential mention retriever that frees the model from heavy and non-parallelizable decoding, achieving 4× speedup without compromise on linking metrics.INSGENEL outperforms previous generative alternatives with +6.8 F1 points gain on average, also with a huge advantage in training data efficiency and training compute consumption.In addition, our skillfully engineered incontext learning (ICL) framework for EL still lags behind INSGENEL significantly, reaffirming that the EL task remains a persistent hurdle for general LLMs.
Zilin Xiao, Ming Gong 0001, Jie Wu 0018, Xingyao Zhang 0003, Linjun Shou, Daxin Jiang
EMNLP1
2022 Exploring the Effectiveness of Appearance Descriptor in DeepSORT
abstract
Tracking-by-detection approaches have demonstrated their strength in addressing Multiple Object Tracking (MOT) problems. DeepSORT, one of the classical tracking-by-detection MOT methods, relies on a deep appearance descriptor to extract global appearance features of identities. Although the appearance descriptor acts as a key component of such tracking-by-detection methods, which is responsible for modeling appearance information, the relationship between it and tracking performance remains unclear, especially whether further improvements to it will be reflected in the tracking performance. To explore that, extensive experiments are conducted on the appearance descriptor by applying various traditional optimization methods. Furthermore, we propose an Evolutionary Neural Architecture Search (ENAS) strategy for the appearance descriptor named Genetic-SORT to assist exploration. The experimental results demonstrate that tracking performance fails to follow the improvements applied on the appearance descriptor and even shows a negative correlation, which is contrary to our intuition.
Zilin Xiao, Yanan Sun 0001
IJCNN1