Ziyang Xiao

dblp:126/1033 · DBLP profile ↗
← Back
14ranked-venue papers
8as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 5 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DeepOR: A Deep Reasoning Foundation Model for Optimization Modeling
abstract
Optimization modeling plays a critical role in supporting optimal decision-making across various domains. Previous works have demonstrated that large language models (LLMs) tailored for optimization modeling have significantly automated and simplified this process. However, these models typically employ a straightforward input-output paradigm and struggle with challenging instances. In contrast, recent advances in general-purpose reasoning LLMs (RLLMs), such as DeepSeek-R1, have shown impressive capabilities in complex domains like mathematics and coding. In this paper, we introduce DeepOR, the first RLLM specifically designed for optimization modeling. Instead of directly outputting solutions, DeepOR explicitly performs multiple intermediate reasoning steps. To adapt a base LLM into an RLLM, we begin by synthesizing long chain-of-thought (CoT) data guided by a flowchart, which is automatically generated using a self-exploration algorithm. Once the training data are prepared, we employ supervised fine-tuning on the base LLM to endow it with reasoning capabilities tailored for optimization modeling. To fully leverage the model's reasoning potential, we further apply reinforcement learning with reward-shaping derived from solver feedback. Experimental results on benchmarks confirm that DeepOR consistently and significantly outperforms existing state-of-the-art approaches.
Ziyang Xiao, Yuan Jessica Wang, Xiongwei Han, Shisi Guan, Jingyan Zhu, Jingrong Xie 0001, Lilin Xu, Han Wu 0004, Wing Yin Yu, Zehua Liu, Xiaojin Fu, Gang Chen 0001, Dongxiang Zhang
AAAI1
2025 A Survey of Optimization Modeling Meets LLMs: Progress and Future Directions
abstract
By virtue of its great utility in solving real-world problems, optimization modeling has been widely employed for optimal decision-making across various sectors, but it requires substantial expertise from operations research professionals. With the advent of large language models (LLMs), new opportunities have emerged to automate the procedure of mathematical modeling. This survey presents a comprehensive and timely review of recent advancements that cover the entire technical stack, including data synthesis and fine-tuning for the base model, inference frameworks, benchmark datasets, and performance evaluation. In addition, we conducted an in-depth analysis on the quality of benchmark datasets, which was found to have a surprisingly high error rate. We cleaned the datasets and constructed a new leaderboard with fair performance evaluation in terms of base LLM model and datasets. We also build an online portal that integrates resources of cleaned datasets, code and paper repository to benefit the community. Finally, we identify limitations in current methodologies and outline future research opportunities.
Ziyang Xiao, Jingrong Xie 0001, Lilin Xu, Shisi Guan, Jingyan Zhu, Xiongwei Han, Xiaojin Fu, WingYin Yu, Han Wu 0004, Qingcan Kang, Jiahui Duan, Tao Zhong 0004, Mingxuan Yuan, Yuan Wang 0003, Gang Chen 0001, Dongxiang Zhang
IJCAI1
2024 Chain-of-Experts: When LLMs Meet Complex Operations Research Problems
abstract
Large language models (LLMs) have emerged as powerful techniques for various NLP tasks, such as mathematical reasoning and plan generation. In this paper, we study automatic modeling and programming for complex operation research (OR) problems, so as to alleviate the heavy dependence on domain experts and benefit a spectrum of industry sectors. We present the first LLM-based solution, namely Chain-of-Experts (CoE), a novel multi-agent cooperative framework to enhance reasoning capabilities. Specifically, each agent is assigned a specific role and endowed with domain knowledge related to OR. We also introduce a conductor to orchestrate these agents via forward thought construction and backward reflection mechanism. Furthermore, we release a benchmark dataset (ComplexOR) of complex OR problems to facilitate OR research and community development. Experimental results show that CoE significantly outperforms the state-of-the-art LLM-based approaches both on LPWP and ComplexOR.
Ziyang Xiao, Dongxiang Zhang, Yangjun Wu, Lilin Xu, Yuan Jessica Wang, Xiongwei Han, Xiaojin Fu, Tao Zhong 0004, Mingli Song, Gang Chen 0001
ICLR1
2024 PREACT: Predictive Resource Allocation for Bursty Workloads in a Co-located Data Center
abstract
Co-locating online latency-critical (LC) services with best-effort (BE) batch jobs in the same server has been widely adopted by modern data centers to improve resource utilization. Various approaches have been proposed to maximize the resources allocated to the BE jobs without SLO (service level objective) violation. However, when facing bursty workloads, existing solutions suffer from poor performance because they cannot react promptly to the sudden and sharp increase of LC service requests. Consequently, these methods result in either a high violation rate of the SLO constraint or low resource utilization caused by conservative allocation strategies.
Dingyu Yang, Ziyang Xiao, Dongxiang Zhang, Shuhao Zhang 0001, Jian Cao 0001, Gang Chen 0001
ICPP2
2024 Enhancing LLM Reasoning via Vision-Augmented Prompting
abstract
Verbal and visual-spatial information processing are two critical subsystems that activate different brain regions and often collaborate together for cognitive reasoning. Despite the rapid advancement of LLM-based reasoning, the mainstream frameworks, such as Chain-of-Thought (CoT) and its variants, primarily focus on the verbal dimension, resulting in limitations in tackling reasoning problems with visual and spatial clues. To bridge the gap, we propose a novel dual-modality reasoning framework called Vision-Augmented Prompting (VAP). Upon receiving a textual problem description, VAP automatically synthesizes an image from the visual and spatial clues by utilizing external drawing tools. Subsequently, VAP formulates a chain of thought in both modalities and iteratively refines the synthesized image. Finally, a conclusive reasoning scheme based on self-alignment is proposed for final result generation. Extensive experiments are conducted across four versatile tasks, including solving geometry problems, Sudoku, time series prediction, and travelling salesman problem. The results validated the superiority of VAP over existing LLMs-based reasoning frameworks.
Ziyang Xiao, Dongxiang Zhang, Xiongwei Han, Xiaojin Fu, Wing Yin Yu, Tao Zhong 0004, Sai Wu, Yuan Wang 0003, Jianwei Yin, Gang Chen 0001
NeurIPS1
2024 Reliable monitoring and prediction method for transmission lines based on FBG and LSTM
Shanyong Cai, Aobo Fan, Ziyang Xiao, Luming Li
Adv. Eng. Informatics7
2023 Neural TSP Solver with Progressive Distillation
abstract
Travelling salesman problem (TSP) is NP-Hard with exponential search space. Recently, the adoption of encoder-decoder models as neural TSP solvers has emerged as an attractive topic because they can instantly obtain near-optimal results for small-scale instances. Nevertheless, their training efficiency and solution quality degrade dramatically when dealing with large-scale problems. To address the issue, we propose a novel progressive distillation framework, by adopting curriculum learning to train TSP samples in increasing order of their problem size and progressively distilling high-level knowledge from small models to large models via a distillation loss. In other words, the trained small models are used as the teacher network to guide action selection when training large models. To accelerate training speed, we also propose a Delaunary-graph based action mask and a new attention-based decoder to reduce decoding cost. Experimental results show that our approach establishes clear advantages over existing encoder-decoder models in terms of training effectiveness and solution quality. In addition, we validate its usefulness as an initial solution generator for the state-of-the-art TSP solvers, whose probability of obtaining the optimal solution can be further improved in such a hybrid manner.
Dongxiang Zhang, Ziyang Xiao, Yuan Wang 0003, Mingli Song, Gang Chen 0001
AAAI2
2023 A deep reinforcement learning agent for geometry online tutoring
Ziyang Xiao, Dongxiang Zhang
Knowl. Inf. Syst.1
2023 DoveDB: A Declarative and Low-Latency Video Database
abstract
Concerning the usability and efficiency to manage video data generated from large-scale cameras, we demonstrate DoveDB, a declarative and low-latency video database. We devise a more comprehensive video query language called VMQL to improve the expressiveness of previous SQL-like languages, which are augmented with functionalities for model-oriented management and deployment. We also propose a light-weight ingestion scheme to extract tracklets of all the moving objects and build semantic indexes to facilitate efficient query processing. For user interaction, we construct a simulation environment with 120 cameras deployed in a road network and demonstrate three interesting scenarios. Using VMQL, users are allowed to 1) train a visual model using SQL-like statement and deploy it on dozens of target cameras simultaneously for online inference; 2) submit multi-object tracking (MOT) requests on target cameras, store the ingested results and build semantic indexes; and 3) issue an aggregation or top- k query on the ingested cameras and obtain the response within milliseconds. A preliminary video introduction of DoveDB is available at https://www.youtube.com/watch?v=N139dEyvAJk
Ziyang Xiao, Dongxiang Zhang, Zepeng Li 0002, Sai Wu, Kian-Lee Tan, Gang Chen 0001
Proc. VLDB Endow.1
2015 Inverse Sparse Tracker With a Locally Weighted Distance Metric
abstract
Sparse representation has been recently extensively studied for visual tracking and generally facilitates more accurate tracking results than classic methods. In this paper, we propose a sparsity-based tracking algorithm that is featured with two components: 1) an inverse sparse representation formulation and 2) a locally weighted distance metric. In the inverse sparse representation formulation, the target template is reconstructed with particles, which enables the tracker to compute the weights of all particles by solving only one l1 optimization problem and thereby provides a quite efficient model. This is in direct contrast to most previous sparse trackers that entail solving one optimization problem for each particle. However, we notice that this formulation with normal Euclidean distance metric is sensitive to partial noise like occlusion and illumination changes. To this end, we design a locally weighted distance metric to replace the Euclidean one. Similar ideas of using local features appear in other works, but only being supported by popular assumptions like local models could handle partial noise better than holistic models, without any solid theoretical analysis. In this paper, we attempt to explicitly explain it from a mathematical view. On that basis, we further propose a method to assign local weights by exploiting the temporal and spatial continuity. In the proposed method, appearance changes caused by partial occlusion and shape deformation are carefully considered, thereby facilitating accurate similarity measurement and model update. The experimental validation is conducted from two aspects: 1) self validation on key components and 2) comparison with other state-of-the-art algorithms. Results over 15 challenging sequences show that the proposed tracking algorithm performs favorably against the existing sparsity-based trackers and the other state-of-the-art methods.
Dong Wang 0004, Huchuan Lu, Ziyang Xiao, Ming-Hsuan Yang 0001
IEEE Trans. Image Process.3
2014 L2-RLS-Based Object Tracking
abstract
In this paper, we present a robust and fast tracking algorithm in which object tracking is achieved by solving ℓ2-regularized least square (ℓ2-RLS) problems in a Bayesian inference framework. First, the changing appearance of the tracked target is modeled with PCA basis vectors and square templates, which makes the tracker not only exploit the strength of subspace representation but also explicitly take partial occlusion into consideration. They can together represent both the intact and corrupted objects well. Second, we adopt the ℓ2-regularized least square method to solve the proposed representation model. Compared with the complex ℓ1-based algorithm, it provides a very fast performance without the loss of accuracy in handling the tracking problem. In addition, a novel likelihood function and a refined update scheme further help to improve the robustness of our tracker. Both qualitative and quantitative evaluations on several challenging image sequences demonstrate that the proposed method performs favorably against several state-of-the-art tracking algorithms.
Ziyang Xiao, Huchuan Lu, Dong Wang 0004
IEEE Trans. Circuits Syst. Video Technol.1
2014 Visual Tracking via Discriminative Sparse Similarity Map
abstract
In this paper, we cast the tracking problem as finding the candidate that scores highest in the evaluation model based upon a matrix called discriminative sparse similarity map (DSS map). This map demonstrates the relationship between all the candidates and the templates, and it is constructed based on the solution to an innovative optimization formulation named multitask reverse sparse representation formulation, which searches multiple subsets from the whole candidate set to simultaneously reconstruct multiple templates with minimum error. A customized APG method is derived for getting the optimum solution (in matrix form) within several iterations. This formulation allows the candidates to be evaluated accurately in parallel rather than one-by-one like most sparsity-based trackers do and meanwhile considers the relationship between candidates, therefore it is more superior in terms of cost-performance ratio. The discriminative information containing in this map comes from a large template set with multiple positive target templates and hundreds of negative templates. A Laplacian term is introduced to keep the coefficients similarity level in accordance with the candidates similarities, thereby making our tracker more robust. A pooling approach is proposed to extract the discriminative information in the DSS map for easily yet effectively selecting good candidates from bad ones and finally get the optimum tracking results. Plenty experimental evaluations on challenging image sequences demonstrate that the proposed tracking algorithm performs favorably against the state-of-the-art methods.
Bohan Zhuang, Huchuan Lu, Ziyang Xiao, Dong Wang 0004
IEEE Trans. Image Process.3
2013 Fast and effective color-based object tracking by boosted color distribution
Dong Wang 0004, Huchuan Lu, Ziyang Xiao, Yen-Wei Chen 0001
Pattern Anal. Appl.3
2012 Object tracking with L2-RLS
Ziyang Xiao, Huchuan Lu, Dong Wang 0004
ICPR1