Siyang Gao

dblp:136/9876 · DBLP profile ↗
← Back
16ranked-venue papers
1as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Distributed semantic trajectory similarity join
abstract
Similarity join is a fundamental operation for managing semantic trajectory data. In massive data scenarios, distributed paradigms can be utilized to process similarity joins for huge amounts of semantic trajectory data, but they face the challenge of locally aware partitioning of the data. To address this problem, we propose a distributed similarity join framework based on trajectory segments. The semantic trajectory data are partitioned by segments, and semantically similar trajectories are stored in the same partition to improve the local similarity of the partitions. We design global and local indexes to efficiently manage partitioned data. We develop a filtering validation framework to enhance similarity query performance by pruning irrelevant trajectories based on the temporal, spatial, and semantic distances of trajectory segments. Extensive experiments on three real-world datasets demonstrate that our method achieves superior scalability and query efficiency, compared to other methods.
Ruijie Tian, Siyang Gao, Fayadh Alenezi, Kemal Polat
Inf. Process. Manag.3
2025 Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging
abstract
Vision-Language Models (VLMs) combine visual perception with the general capabilities, such as reasoning, of Large Language Models (LLMs). However, the mechanisms by which these two abilities can be combined and contribute remain poorly understood. In this work, we explore to compose perception and reasoning through model merging that connects parameters of different models. Unlike previous works that often focus on merging models of the same kind, we propose merging models across modalities, enabling the incorporation of the reasoning capabilities of LLMs into VLMs. Through extensive experiments, we demonstrate that model merging offers a successful pathway to transfer reasoning abilities from LLMs to VLMs in a training-free manner. Moreover, we utilize the merged models to understand the internal mechanism of perception and reasoning and how merging affects it. We find that perception capabilities are predominantly encoded in the early layers of the model, whereas reasoning is largely facilitated by the middle-to-late layers. After merging, we observe that all layers begin to contribute to reasoning, whereas the distribution of perception abilities across layers remains largely unchanged. These observations shed light on the potential of model merging as a tool for multimodal integration and interpretation.
Shiqi Chen 0002, Jinghan Zhang 0006, Tongyao Zhu, Wei Liu 0131, Siyang Gao, Miao Xiong, Manling Li, Junxian He
ICML5
2025 Why Is Spatial Reasoning Hard for VLMs? An Attention Mechanism Perspective on Focus Areas
abstract
Large Vision Language Models (VLMs) have long struggled with spatial reasoning tasks. Surprisingly, even simple spatial reasoning tasks, such as recognizing “under” or “behind” relationships between only two objects, pose significant challenges for current VLMs. We believe it is crucial to use the lens of mechanism interpretability, opening up the model and diving into model’s internal states to examine the interactions between image and text tokens during spatial reasoning. Our analysis of attention behaviors reveals significant differences in how VLMs allocate attention to image versus text. By tracing the areas of images that receive the highest attention scores throughout intermediate layers, we observe a notable pattern: errors often coincide with attention being misdirected towards irrelevant objects within the image. Moreover, such attention patterns exhibit substantial differences between familiar (e.g., “on the left side of ”) and unfamiliar (e.g.,“in front of ”) spatial relationships. Motivated by these findings, we propose ADAPTVIS based on inference-time confidence scores to sharpen the attention on highly relevant regions when the model exhibits high confidence, while smoothing and broadening the attention window to consider a wider context when confidence is lower. This training-free decoding method shows significant improvement (e.g., up to a 50 absolute point improvement) on spatial reasoning benchmarks such as WhatsUp and VSR with negligible additional cost.
Shiqi Chen 0002, Tongyao Zhu, Ruochen Zhou, Jinghan Zhang 0006, Siyang Gao, Juan Carlos Niebles, Mor Geva, Junxian He, Jiajun Wu 0001, Manling Li
ICML5
2024 In-Context Sharpness as Alerts: An Inner Representation Perspective for Hallucination Mitigation
abstract
Large language models (LLMs) frequently hallucinate, e.g., making factual errors, yet our understanding of why they make these errors remains limited. In this study, we aim to understand the underlying mechanisms of LLM hallucinations from the perspective of *inner representations*. We discover a pattern associated with hallucinations: correct generations tend to have *sharper* context activations in the hidden states of the in-context tokens, compared to that of the incorrect generations. Leveraging this signal, we propose an entropy-based metric to quantify the *sharpness* among the in-context hidden states and incorporate it into the decoding process, i.e, use the entropy value to adjust the next token prediction distribution to improve the factuality and overall quality of the generated text. Experiments on knowledge-seeking datasets (Natural Questions, HotpotQA, TriviaQA) and hallucination benchmark (TruthfulQA) demonstrate our consistent effectiveness, e.g., up to 8.6 absolute points on TruthfulQA. We believe this study can improve our understanding of hallucinations and serve as a practical solution for hallucination mitigation.
Shiqi Chen 0002, Miao Xiong, Junteng Liu, Zhengxuan Wu, Teng Xiao, Siyang Gao, Junxian He
ICML6
2023 FELM: Benchmarking Factuality Evaluation of Large Language Models
abstract
Assessing factuality of text generated by large language models (LLMs) is an emerging yet crucial research area, aimed at alerting users to potential errors and guiding the development of more reliable LLMs. Nonetheless, the evaluators assessing factuality necessitate suitable evaluation themselves to gauge progress and foster advancements. This direction remains under-explored, resulting in substantial impediments to the progress of factuality evaluators. To mitigate this issue, we introduce a benchmark for Factuality Evaluation of large Language Models, referred to as FELM. In this benchmark, we collect responses generated from LLMs and annotate factuality labels in a fine-grained manner. Contrary to previous studies that primarily concentrate on the factuality of world knowledge (e.g. information from Wikipedia), FELM focuses on factuality across diverse domains, spanning from world knowledge to math and reasoning. Our annotation is based on text segments, which can help pinpoint specific factual errors. The factuality annotations are further supplemented by predefined error types and reference links that either support or contradict the statement. In our experiments, we investigate the performance of several LLM-based factuality evaluators on FELM, including both vanilla LLMs and those augmented with retrieval mechanisms and chain-of-thought processes. Our findings reveal that while retrieval aids factuality evaluation, current LLMs are far from satisfactory to faithfully detect factual errors.
Shiqi Chen 0002, Yiran Zhao 0006, Jinghan Zhang 0006, I-Chun Chern, Siyang Gao, Pengfei Liu 0003, Junxian He
NeurIPS5
2023 Improving the Knowledge Gradient Algorithm
abstract
The knowledge gradient (KG) algorithm is a popular policy for the best arm identification (BAI) problem. It is built on the simple idea of always choosing the measurement that yields the greatest expected one-step improvement in the estimate of the best mean of the arms. In this research, we show that this policy has limitations, causing the algorithm not asymptotically optimal. We next provide a remedy for it, by following the manner of one-step look ahead of KG, but instead choosing the measurement that yields the greatest one-step improvement in the probability of selecting the best arm. The new policy is called improved knowledge gradient (iKG). iKG can be shown to be asymptotically optimal. In addition, we show that compared to KG, it is easier to extend iKG to variant problems of BAI, with the $\epsilon$-good arm identification and feasible arm identification as two examples. The superior performances of iKG on these problems are further demonstrated using numerical examples.
Le Yang 0012, Siyang Gao, Chin Pang Ho
NeurIPS2
2023 Fast Bellman Updates for Wasserstein Distributionally Robust MDPs
abstract
Markov decision processes (MDPs) often suffer from the sensitivity issue under model ambiguity. In recent years, robust MDPs have emerged as an effective framework to overcome this challenge. Distributionally robust MDPs extend the robust MDP framework by incorporating distributional information of the uncertain model parameters to alleviate the conservative nature of robust MDPs. This paper proposes a computationally efficient solution framework for solving distributionally robust MDPs with Wasserstein ambiguity sets. By exploiting the specific problem structure, the proposed framework decomposes the optimization problems associated with distributionally robust Bellman updates into smaller subproblems, which can be solved efficiently. The overall complexity of the proposed algorithm is quasi-linear in both the numbers of states and actions when the distance metric of the Wasserstein distance is chosen to be $L_1$, $L_2$, or $L_{\infty}$ norm, and so the computational cost of distributional robustness is substantially reduced. Our numerical experiments demonstrate that the proposed algorithms outperform other state-of-the-art solution methods.
Zhuodong Yu, Shaohang Xu, Siyang Gao, Chin Pang Ho
NeurIPS4
2023 Convergence Analysis of Stochastic Kriging-Assisted Simulation with Random Covariates
abstract
We consider performing simulation experiments in the presence of covariates. Here, covariates refer to some input information other than system designs to the simulation model that can also affect the system performance. To make decisions, decision makers need to know the covariate values of the problem. Traditionally in simulation-based decision making, simulation samples are collected after the covariate values are known; in contrast, as a new framework, simulation with covariates starts the simulation before the covariate values are revealed and collects samples on covariate values that might appear later. Then, when the covariate values are revealed, the collected simulation samples are directly used to predict the desired results. This framework significantly reduces the decision time compared with the traditional way of simulation. In this paper, we follow this framework and suppose there are a finite number of system designs. We adopt the metamodel of stochastic kriging (SK) and use it to predict the system performance of each design and the best design. The goal is to study how fast the prediction errors diminish with the number of covariate points sampled. This is a fundamental problem in simulation with covariates and helps quantify the relationship between the offline simulation efforts and the online prediction accuracy. Particularly, we adopt measures of the maximal integrated mean squared error (IMSE) and integrated probability of false selection (IPFS) for assessing errors of the system performance and the best design predictions. Then, we establish convergence rates for the two measures under mild conditions. Last, these convergence behaviors are illustrated numerically using test examples. History: Accepted by Bruno Tuffin, area editor for simulation. Funding: This work was supported in part by Singapore Ministry of Education Academic Research Funds [Tier 1 Grants R-155-000-201-114 and A-0004822-00-00], the City University of Hong Kong [Grants 7005269 and 7005568], and the National Natural Science Foundation of China [Grant 72091211]. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2022.1263 ) as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2021.0329 ) at ( http://dx.doi.org/10.5281/zenodo.7344997 ).
Cheng Li 0063, Siyang Gao, Jianzhong Du
INFORMS J. Comput.2
2022 On the Finite-Time Performance of the Knowledge Gradient Algorithm
abstract
The knowledge gradient (KG) algorithm is a popular and effective algorithm for the best arm identification (BAI) problem. Due to the complex calculation of KG, theoretical analysis of this algorithm is difficult, and existing results are mostly about the asymptotic performance of it, e.g., consistency, asymptotic sample allocation, etc. In this research, we present new theoretical results about the finite-time performance of the KG algorithm. Under independent and normally distributed rewards, we derive lower bounds and upper bounds for the probability of error and simple regret of the algorithm. With these bounds, existing asymptotic results become simple corollaries. We also show the performance of the algorithm for the multi-armed bandit (MAB) problem. These developments not only extend the existing analysis of the KG algorithm, but can also be used to analyze other improvement-based algorithms. Last, we use numerical experiments to further demonstrate the finite-time behavior of the KG algorithm.
Yanwen Li, Siyang Gao
ICML2
2022 Texture image retrieval using DNST domain local neighborhood intensity pattern
Xiangyang Wang 0001, Hongying Yang, Siyang Gao, Panpan Niu
Multim. Tools Appl.3
2021 Wafer Defect Inspection Optimization With Partial Coverage - A Numerical Approach
abstract
Electron beam inspection (EBI) with high resolution is a promising technique to improve the defect inspection on the surface of patterned wafer. However, high resolution usually means long inspection time, which results in the low throughput and limitation of EBI applied in practice. This study aims to optimize the inspection time of EBI by reducing the total number of inspection regions without loss of the accuracy. We first refine this defect inspection optimization problem as a partial congruent square cover problem. Then, we propose two novel mixed-integer linear programming models for this problem. To deal with the large-scale problems, an approximation algorithm is developed to obtain the high-quality solutions. This approximation algorithm efficiently utilizes the linear programming (LP) rounding technique and greedy strategy based on the proposed model. Compared with the existing algorithms in the literature, numerical results show the superiority of the proposed model and algorithm.Note to Practitioners—Defect inspection is a key process in wafer fabrication for identifying and inspecting the patterning defects generated during the complicated fabrication processes. Electron beam inspection (EBI) takes place of optical inspection gradually as the design rules keep shrinking and the circuits are more susceptible to nanoscale killer defects. Low throughput is the main drawback of EBI and the improvements on throughput have far-reaching significance on the high volume manufacturing of semiconductor products. This study aims to reduce the number of inspection regions to improve the total inspection time in the EBI process. Considering that the inspection time for each inspection region is constant, less inspection regions means less total inspection time. However, the positions of inspection regions are arbitrary across the continuous planar space, putting pressure on modeling and solving the problem. A preprocessing algorithm is designed to discover a limited number of candidate positions of inspection regions without loss of optimality, which greatly simplifies the problem. For dealing with large-scale instances, an approximation algorithm combining linear programming (LP)-rounding technique and greedy strategy is designed to get near optimal solutions. Since the number of inspection regions is one key factor determining the total inspection time, the proposed optimization methods have a potential to be applicable to enhance the efficiency of advanced EBI platforms, such as ASML HMI eP series.
Ming Qin, Zhongshun Shi, Weiwei Chen 0003, Siyang Gao, Leyuan Shi
IEEE Trans Autom. Sci. Eng.4
2019 Advancing Constrained Ranking and Selection With Regression in Partitioned Domains
abstract
Ranking and selection (R&S) procedures are powerful tools to enhance the efficiency of simulation-based optimization. In this paper, we consider the R&S problem subject to stochastic constraints and seek to improve the selection efficiency by incorporating the information from across the domain into quadratic regression metamodels. To better fulfill the quadratic assumption of the regression metamodel used in this paper, we divide the solution space into adjacent partitions such that the underlying functions of both the objective and constraint measures in each partition are approximately quadratic with homogeneous noise. Using the large deviations theory, we characterize the asymptotically optimal allocation rule by maximizing the rate at which the probability of false selection tends to zero. Numerical experiments demonstrate that our approach dramatically improves the selection efficiency by 50%-90% on some typical selection examples compared with the existing approaches.
Fei Gao 0012, Siyang Gao, Hui Xiao 0001, Zhongshun Shi
IEEE Trans Autom. Sci. Eng.2
2018 Optimizing HIV Interventions for Multiplex Social Networks via Partition-Based Random Search
abstract
There are multiple modes for human immunodeficiency virus (HIV) transmissions, each of which is usually associated with a certain key population (e.g., needle sharing among people who inject drugs). Recent field studies revealed the merging trend of multiple key populations, making HIV intervention difficult because of the existence of multiple transmission modes in such complex multiplex social networks. In this paper, we aim to address this challenge by developing a multiplex social network framework to capture the multimode transmission across two key populations. Based on the multiplex social network framework, we propose a new random search method, named partition-based random search with network and memory prioritization (PRS-NMP), to identify the optimal subset of high-value individuals in the social network for interventions. Numerical experiments demonstrated that the proposed PRS-NMP-based interventions could effectively reduce the scale of HIV transmissions. The performance of PRS-NMP-based interventions is consistently better than the benchmark nested partitions method and network-based metrics.
Qingpeng Zhang, Lu Zhong, Siyang Gao, Xiaoming Li 0008
IEEE Trans. Cybern.3
2017 A Sequential Budget Allocation Framework for Simulation Optimization
abstract
Many problems in automation and manufacturing are most suitable to be modeled as simulation optimization problems. Solving these problems typically involves two efforts: one is to explore the solution space, and the other is to exploit the performance values of the sampled solutions. When the amount of computing budget is limited, we need to know how to balance these two efforts in order to obtain the best result. In this study, we derive two measures to quantify the marginal contribution of exploring the search space and exploiting the performance values. A sequential budget allocation framework is designed by keeping the two measures approximately the same at each iteration. Numerical experiments on both continuous and discrete simulation optimization problems demonstrate that our new approach can significantly enhance the computing efficiency.
Siyang Gao, Loo Hay Lee, Chun-Hung Chen, Leyuan Shi
IEEE Trans Autom. Sci. Eng.1
2017 Simulation Optimization for Medical Staff Configuration at Emergency Department in Hong Kong
abstract
Medical staff configuration is a critical problem in the management of an emergency department (ED) in Hong Kong (HK). Given the service requirements by HK government, it is imperative for the hospital managers to develop medical staff configuration in a cost-and-time-effective way. In this paper, the medical staff configuration problem in ED is modeled as minimizing the total labor cost while satisfying the service quality requirements. To solve this issue, we propose a highly efficient search method, called random boundary generation with feasibility detection (RBG-FD). The random boundary generation (RBG) is applied to efficiently identify good-quality solutions based on the objective value. The feasibility detection (FD) procedure is used to retain the probability of correct feasibility detection of each sampled solution at the desired level, which intrinsically allocates a reasonable number of simulation replications. To estimate the performance measures of the ED, a discrete-event simulation model is developed to reflect the patient flow. Using these techniques, the efficiency of identifying the optimal staff configuration can be significantly improved. A case study is performed in a public hospital in HK. The numerical results indicate significantly higher practicability and efficiency of the proposed method with different patient arrival rates and service constraints. Note to Practitioners This paper seeks to solve the problem of minimizing the medical staff cost constrained by certain service requirements [i.e., patients' waiting times for treatment] at an emergency department in Hong Kong. In our formulation, these service requirements are characterized by some stochastic constraints. Most of the existing random search methods concentrate on the computing efforts in the neighborhood of the best-so-far solutions in order to obtain good-quality solutions. Due to the special structure of this problem and ease of computing the objective values, we proposed an efficient random search approach that iteratively identifies a solution with a better objective value than that of the current best solution. Experimental studies demonstrate the significantly higher efficiency of this method. In order to obtain the same solution quality, it is able to reduce the computational time by 90% compared with some existing approaches in the literature.
Hainan Guo, Siyang Gao, Kwok-Leung Tsui, Tie Niu
IEEE Trans Autom. Sci. Eng.2
2014 An Optimal Sample Allocation Strategy for Partition-Based Random Search
abstract
Partition-based random search (PRS) provides a class of effective algorithms for global optimization. In each iteration of a PRS algorithm, the solution space is partitioned into subsets which are randomly sampled and evaluated. One subset is then determined to be the promising subset for further partitioning. In this paper, we propose the problem of allocating samples to each subset so that the samples are utilized most efficiently. Two types of sample allocation problems are discussed, with objectives of maximizing the probability of correctly selecting the promising subset$(P\{CSPS\})$given a sample budget and minimizing the required sample size to achieve a satisfied level of$P\{CSPS\}$, respectively. An extreme value-based prospectiveness criterion is introduced and an asymptotically optimal solution to the two types of sample allocation problems is developed. The resulting optimal sample allocation strategy (OSAS) is an effective procedure for the existing PRS algorithms by intelligently utilizing the limited computing resources. Numerical tests confirm that OSAS is capable of increasing the$P\{CSPS\}$in each iteration and subsequently improving the performance of PRS algorithms.
Weiwei Chen 0003, Siyang Gao, Chun-Hung Chen, Leyuan Shi
IEEE Trans Autom. Sci. Eng.2