Boshi Wang

dblp:216/7905 · DBLP profile ↗
← Back
18ranked-venue papers
8as first author
17since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 6 first-author · 12 since 2021Computer networks · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
abstract
The advancements of language language models (LLMs) have piqued growing interest in developing LLM-based language agents to automate scientific discovery end-to-end, which has sparked both excitement and skepticism about the true capabilities of such agents. In this work, we argue that for an agent to fully automate scientific discovery, it must be able to complete all essential tasks in the workflow. Thus, we call for rigorous assessment of agents on individual tasks in a scientific workflow before making bold claims on end-to-end automation. To this end, we present ScienceAgentBench, a new benchmark for evaluating language agents for data-driven scientific discovery. To ensure the scientific authenticity and real-world relevance of our benchmark, we extract 102 tasks from 44 peer-reviewed publications in four disciplines and engage nine subject matter experts to validate them. We unify the target output for every task to a self-contained Python program file and employ an array of evaluation metrics to examine the generated programs, execution results, and costs. Each task goes through multiple rounds of manual validation by annotators and subject matter experts to ensure its annotation quality and scientific plausibility. We also propose two effective strategies to mitigate data contamination concerns. Using our benchmark, we evaluate five open-weight and proprietary LLMs, each with three frameworks: direct prompting, OpenHands, and self-debug. Given three attempts for each task, the best-performing agent can only solve 32.4% of the tasks independently and 34.3% with expert-provided knowledge. These results underscore the limited capacities of current language agents in generating code for data-driven discovery, let alone end-to-end automation for scientific research.
Ziru Chen, Yuting Ning, Qianheng Zhang, Boshi Wang, Botao Yu, Yifei Li 0005, Zeyi Liao, Zitong Lu, Vishal Dey, Mingyi Xue 0001, Frazier N. Baker, Benjamin Burns, Daniel Adu-Ampratwum, Xuhui Huang, Xia Ning, Song Gao 0001, Yu Su 0001, Huan Sun 0001
ICLR5
2025 Resource Allocation in Wideband Cooperative ISAC Systems
abstract
This paper investigates the resource allocation problem for multi-user wideband cooperative integrated sensing and communication (ISAC) networks based on orthogonal frequency-division multiplexing (OFDM) waveforms. In order to balance sensing and communication performance with limited spectrum resources in this wideband cell-free system, we aim to maximize the sum rate, encompassing both communication and radar rates, while adhering to constraints related to access point (AP) power and spectrum resources. We utilize alternate optimization (AO) methods to optimize power and spectrum resources separately. For power optimization, we employ the fractional programming (FP) algorithm to convert the problem into a convex one, which can be quickly solved by the primal-dual subgradient (PDS) method. As for subcarrier allocation optimization, we derive its closed-form solution. Simulation results indicate that the communication and sensing performance of the cell-free ISAC system outperforms that of the conventional centralized ISAC system.
Chenhan Yuan, Boshi Wang, Zhiyuan Yu 0007, Cunhua Pan, Hong Ren
VTC2025-Spring2
2025 Beamforming Design for Double-Active-RIS-Aided Communication Systems With Inter-Excitation
abstract
In this paper, we investigate a double-active-reconfigurable intelligent surface (RIS)-aided downlink wireless communication system, where a multi-antenna base station (BS) serves multiple single-antenna users with both double reflection and single reflection links. Due to the signal amplification capability of active RISs, they can effectively mitigate the multiplicative fading effect. However, this also induces signal bouncing between the two active RISs that cannot be ignored. This phenomenon is termed as the “inter-excitation” effect and is characterized in the received signal by proposing a feedback-type model. Based on the signal model, we formulate a weighted sum rate (WSR) maximization problem by jointly optimizing the beamforming matrix at the BS and the reflecting coefficient matrices at the two active RISs, subject to power constraints at the BS and active RISs, as well as the maximum amplification gain constraints of the active RISs. To solve this non-convex problem, we first transform the problem into a more tractable form using the fractional programming (FP) method. Then, by introducing auxiliary variables, the problem can be converted into an equivalent form that can be solved by using a penalty dual decomposition (PDD) algorithm. Furthermore, the power scaling order of the signal-to-noise ratio (SNR) in double-active-RIS-aided system considering inter-excitation effect is derived. Finally, simulation results indicate that the proposed scheme outperforms benchmark schemes with single active RIS and double passive RISs in terms of achievable rate. Furthermore, the results demonstrate that the proposed scheme can enhance the WSR by 30% compared to scenarios that do not take this effect into account when the maximum amplification gain is 40 dB. Additionally, the proposed scheme is capable of achieving high WSR performance at most locations where double active RISs are deployed between the BS and the users, thereby providing greater flexibility in their deployment.
Boshi Wang, Cunhua Pan, Hong Ren, Zhiyuan Yu 0007, Yang Zhang 0114, Gui Zhou
IEEE Trans. Wirel. Commun.1
2024 LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error
abstract
Tools are essential for large language models (LLMs) to acquire up-to-date information and take consequential actions in external environments.Existing work on tool-augmented LLMs primarily focuses on the broad coverage of tools and the flexibility of adding new tools.However, a critical aspect that has surprisingly been understudied is simply how accurately an LLM uses tools for which it has been trained.We find that existing LLMs, including GPT-4 and open-source LLMs specifically fine-tuned for tool use, only reach a correctness rate in the range of 30% to 60%, far from reliable use in practice.We propose a biologically inspired method for tool-augmented LLMs, simulated trial and error (STE), that orchestrates three key mechanisms for successful tool use behaviors in the biological system: trial and error, imagination, and memory.Specifically, STE leverages an LLM's 'imagination' to simulate plausible scenarios for using a tool, after which the LLM interacts with the tool to learn from its execution feedback.Both short-term and long-term memory are employed to improve the depth and breadth of the exploration, respectively.Comprehensive experiments on Tool-Bench show that STE substantially improves tool learning for LLMs under both in-context learning and fine-tuning settings, bringing a boost of 46.7% to Mistral-Instruct-7B and enabling it to outperform GPT-4.We also show effective continual learning of tools via a simple experience replay strategy.1 * Work done as an intern at Microsoft Semantic Machines. 1 Code and data available at https://github.com/ microsoft/simulated-trial-and-error.
Boshi Wang, Hao Fang 0002, Jason Eisner, Benjamin Van Durme, Yu Su 0001
ACL (1)1
2024 How Trustworthy are Open-Source LLMs? An Assessment under Malicious Demonstrations Shows their Vulnerabilities
abstract
Lingbo Mo, Boshi Wang, Muhao Chen, Huan Sun. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Lingbo Mo, Boshi Wang, Muhao Chen 0001, Huan Sun 0001
NAACL-HLT2
2024 Grokking of Implicit Reasoning in Transformers: A Mechanistic Journey to the Edge of Generalization
abstract
We study whether transformers can learn to *implicitly* reason over parametric knowledge, a skill that even the most capable language models struggle with. Focusing on two representative reasoning types, composition and comparison, we consistently find that transformers *can* learn implicit reasoning, but only through *grokking*, i.e., extended training far beyond overfitting. The levels of generalization also vary across reasoning types: when faced with out-of-distribution examples, transformers fail to systematically generalize for composition but succeed for comparison. We delve into the model's internals throughout training, conducting analytical experiments that reveal: 1) the mechanism behind grokking, such as the formation of the generalizing circuit and its relation to the relative efficiency of generalizing and memorizing circuits, and 2) the connection between systematicity and the configuration of the generalizing circuit. Our findings guide data and training setup to better induce implicit reasoning and suggest potential improvements to the transformer architecture, such as encouraging cross-layer knowledge sharing. Furthermore, we demonstrate that for a challenging reasoning task with a large search space, GPT-4-Turbo and Gemini-1.5-Pro based on non-parametric memory fail badly regardless of prompting styles or retrieval augmentation, while a fully grokked transformer can achieve near-perfect accuracy, showcasing the power of parametric memory for complex reasoning.
Boshi Wang, Xiang Yue, Yu Su 0001, Huan Sun 0001
NeurIPS1
2024 Reconfigurable Intelligent Surface-Aided Dual-Function Radar and Communication System With MU-MIMO Communication
abstract
In this paper, we investigate an reconfigurable intelligent surface (RIS)-aided integrated sensing and communication (ISAC) system. Our objective is to maximize the achievable sum rate of the multi-antenna communication users through the joint active and passive beamforming. Weighted minimum mean-square error (WMMSE) method is used to reformulate the original problem into an equivalent one. Then, we utilize an alternating optimization (AO) approach to separate the optimization variables and decompose this challenging problem into two subproblems. Given reflecting coefficients, a penalty-based algorithm is utilized to deal with the transmit power and the non-convex radar signal-to-noise ratio (SNR) constraints. For the given beamforming matrix of the BS, we apply majorization-minimization (MM) to transform the problem into a quadratic constraint quadratic programming (QCQP) problem, which is ultimately solved using a semidefinite relaxation (SDR)-based algorithm. Simulation results illustrate the advantage of deploying RIS in the considered multi-user MIMO (MU-MIMO) ISAC systems.
Yasheng Jin, Zhiyuan Yu 0007, Ruisong Weng, Boshi Wang, Hong Ren, Cunhua Pan
WCNC4
2024 Transmission Design for Double Cooperative Active RIS-Aided Communication
abstract
Reconfigurable intelligent surfaces (RISs) have emerged as a disruptive technology that can reconfigure wireless communication environments cost-effectively. In order to fully unveil the potential of RIS-aided wireless communications, some existing contributions considered the double cooperative passive RISs to achieve a higher capacity scaling orders. However, due to the multiplicative fading effect, the double cooperative passive RISs performs poorly when deployed far from the BS and user respectively. To address this issue, we investigate double cooperative active RISs which are equipped with amplifiers and can overcome the severe path loss caused by the multiplicative fading. Specifically, we aim to maximize the downlink achievable rate subject to the transmit power constraints of the base station (BS) and the double active RISs. The formulated problem is solved by using an alternating optimization (AO) algorithm based on the majorization-minimization (MM) algorithm and the fractional programming (FP) method. Simulation results demonstrate that much better rate performance can be achieved by adopting active RIS compared to passive RIS in the double RIS-aided systems. Meanwhile, deploying them appropriately far away from the BS and user and more elements assigned to the active RIS near the user will achieve better performance.
Boshi Wang, Cunhua Pan, Hong Ren, Gui Zhou, Zhiyuan Yu 0007
WCNC1
2024 Active RIS-Aided ISAC Systems: Beamforming Design and Performance Analysis
abstract
This paper considers an active reconfigurable intelligent surface (RIS)-aided integrated sensing and communication (ISAC) system. We aim to maximize radar signal-to-interference-plus-noise-ratio (SINR) by jointly optimizing the beamforming matrix at a dual-function radar-communication (DFRC) base station (BS) and the reflecting coefficients at an active RIS subject to the quality of service (QoS) constraints of communication user equipments (UEs) and the transmit power constraints of active RIS and DFRC BS. To tackle the optimization problem, the majorization-minimization (MM) algorithm is applied to address the nonconvex radar SINR objective function, and the resulting quartic problem is solved by developing an semidefinite relaxation (SDR)-based approach. Moreover, we derive the scaling order of the radar SINR with a large number of reflecting elements. Next, the transmit power allocation problem and the deployment strategy of the active RIS are studied with a moderate number of reflecting elements. Finally, we validate the potential of the active RIS in ISAC systems compared to passive RIS. Additionally, we deliberate on several open problems that remain for future research.
Zhiyuan Yu 0007, Hong Ren, Cunhua Pan, Gui Zhou, Boshi Wang, Mianxiong Dong, Jiangzhou Wang
IEEE Trans. Commun.5
2023 Towards Understanding Chain-of-Thought Prompting: An Empirical Study of What Matters
abstract
Boshi Wang, Sewon Min, Xiang Deng, Jiaming Shen, You Wu, Luke Zettlemoyer, Huan Sun. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Boshi Wang, Sewon Min, Xiang Deng 0001, You Wu 0001, Luke Zettlemoyer, Huan Sun 0001
ACL (1)1
2023 A Retrieve-and-Read Framework for Knowledge Graph Link Prediction
abstract
Knowledge graph (KG) link prediction aims to infer new facts based on existing facts in the KG. Recent studies have shown that using the graph neighborhood of a node via graph neural networks (GNNs) provides more useful information compared to just using the query information. Conventional GNNs for KG link prediction follow the standard message-passing paradigm on the entire KG, which leads to superfluous computation, over-smoothing of node representations, and also limits their expressive power. On a large scale, it becomes computationally expensive to aggregate useful information from the entire KG for inference. To address the limitations of existing KG link prediction frameworks, we propose a novel retrieve-and-read framework, which first retrieves a relevant subgraph context for the query and then jointly reasons over the context and the query with a high-capacity reader. As part of our exemplar instantiation for the new framework, we propose a novel Transformer-based GNN as the reader, which incorporates graph-based attention structure and cross-attention between query and context for deep fusion. This simple yet effective design enables the model to focus on salient context information relevant to the query. Empirical results on two standard KG link prediction datasets demonstrate the competitive performance of the proposed method. Furthermore, our analysis yields valuable insights for designing improved retrievers within the framework.
Vardaan Pahuja, Boshi Wang, Hugo Latapie, Jayanth Srinivasa, Yu Su 0001
CIKM2
2023 Mind2Web: Towards a Generalist Agent for the Web
abstract
We introduce Mind2Web, the first dataset for developing and evaluating generalist agents for the web that can follow language instructions to complete complex tasks on any website. Existing datasets for web agents either use simulated websites or only cover a limited set of websites and tasks, thus not suitable for generalist web agents. With over 2,000 open-ended tasks collected from 137 websites spanning 31 domains and crowdsourced action sequences for the tasks, Mind2Web provides three necessary ingredients for building generalist web agents: 1) diverse domains, websites, and tasks, 2) use of real-world websites instead of simulated and simplified ones, and 3) a broad spectrum of user interaction patterns. Based on Mind2Web, we conduct an initial exploration of using large language models (LLMs) for building generalist web agents. While the raw HTML of real-world websites are often too large to be fed to LLMs, we show that first filtering it with a small LM significantly improves the effectiveness and efficiency of LLMs. Our solution demonstrates a decent level of performance, even on websites or entire domains the model has never seen before, but there is still a substantial room to improve towards truly generalizable agents. We open-source our dataset, model implementation, and trained models (https://osu-nlp-group.github.io/Mind2Web) to facilitate further research on building a generalist agent for the web.
Xiang Deng 0001, Yu Gu 0016, Boyuan Zheng 0001, Samual Stevens, Boshi Wang, Huan Sun 0001, Yu Su 0001
NeurIPS6
2022 Scalable Learning with Incremental Probabilistic PCA
abstract
Incremental class learning is the classification problem of learning a model where instances from new object classes are added sequentially, and it is desired that the model be retrained only on the new classes with minimal training on the old classes. One major problem facing class incremental learning is catastrophic forgetting, where the updated model forgets the old classes and focuses only on the new classes. This paper proposes a simple and novel incremental class learning method that uses a self-supervised pretrained feature extractor to obtain meaningful features and trains Probabilistic PCA models on the extracted features for each class separately. The Mahalanobis distance is used to obtain the classification result, and an equivalent equation is derived to make the approach computationally affordable. Experiments on standard and large datasets show that the proposed approach outperforms existing state of the art incremental learning methods by a large margin. The fact that the model is trained on each class separately makes it applicable to training on very large datasets such as the whole ImageNet with more than 10,000 classes.
Boshi Wang, Adrian Barbu
IEEE Big Data1
2022 Iteratively Prompt Pre-trained Language Models for Chain of Thought
abstract
While Pre-trained Language Models (PLMs) internalize a great amount of world knowledge, they have been shown incapable of recalling these knowledge to solve tasks requiring complex & multi-step reasoning.Similar to how humans develop a "chain of thought" for these tasks, how can we equip PLMs with such abilities?In this work, we explore an iterative prompting framework, a new prompting paradigm which progressively elicits relevant knowledge from PLMs for multi-step inference.We identify key limitations of existing prompting methods, namely they are either restricted to queries with a single identifiable relation/predicate, or being agnostic to input contexts, which makes it difficult to capture variabilities across different inference steps.We propose an iterative context-aware prompter, which addresses these limitations by learning to dynamically synthesize prompts conditioned on the current step's contexts.Experiments on three datasets involving multi-step reasoning show the effectiveness of the iterative scheme and the context-aware prompter design. 1
Boshi Wang, Xiang Deng 0001, Huan Sun 0001
EMNLP1
2022 Automatic Loss Function Search for Predict-Then-Optimize Problems with Strong Ranking Property
Boshi Wang, Jialin Yi, Hang Dong 0004, Bo Qiao 0001, Chuan Luo 0002, Qingwei Lin
ICLR1
2021 Homomorphic Sensing: Sparsity and Noise
abstract
\emph{Unlabeled sensing} is a recent problem encompassing many data science and engineering applications and typically formulated as solving linear equations whose right-hand side vector has undergone an unknown permutation. It was generalized to the \emph{homomorphic sensing} problem by replacing the unknown permutation with an unknown linear map from a given finite set of linear maps. In this paper we present tighter and simpler conditions for the homomorphic sensing problem to admit a unique solution. We show that this solution is locally stable under noise, while under a sparsity assumption it remains unique under less demanding conditions. Sparsity in the context of unlabeled sensing leads to the problem of \textit{unlabeled compressed sensing}, and a consequence of our general theory is the existence under mild conditions of a unique sparsest solution. On the algorithmic level, we solve unlabeled compressed sensing by an iterative algorithm validated by synthetic data experiments. Finally, under the unifying homomorphic sensing framework we connect unlabeled sensing to other important practical problems.
Liangzu Peng, Boshi Wang, Manolis C. Tsakiris
ICML2
2021 Predictive Job Scheduling under Uncertain Constraints in Cloud Computing
abstract
Capacity management has always been a great challenge for cloud platforms due to massive, heterogeneous on-demand instances running at different times. To better plan the capacity for the whole platform, a class of cloud computing instances have been released to collect computing demands beforehand. To use such instances, users are allowed to submit jobs to run for a pre-specified uninterrupted duration in a flexible range of time in the future with a discount compared to the normal on-demand instances. Proactively scheduling those pre-collected job requests considering the capacity status over the platform can greatly help balance the computing workloads along time. In this work, we formulate the scheduling problem for these pre-collected job requests under uncertain available capacity as a Prediction + Optimization problem with uncertainty in constraints, and propose an effective algorithm called Controlling under Uncertain Constraints (CUC), where the predicted capacity guides the optimization of job scheduling and job scheduling results are leveraged to improve the prediction of capacity through Bayesian optimization. The proposed formulation and solution are commonly applicable for proactively scheduling problems in cloud computing. Our extensive experiments on three public, industrial datasets shows that CUC has great potential for supporting high reliability in cloud platforms.
Hang Dong 0004, Boshi Wang, Bo Qiao 0001, Wenqian Xing, Chuan Luo 0002, Si Qin, Qingwei Lin, Dongmei Zhang 0001, Gurpreet Virdi, Thomas Moscibroda
IJCAI2
2020 Designing Context-Sensitive Norm Inverse Reinforcement Learning Framework for Norm-Compliant Autonomous Agents
abstract
Human behaviors are often prohibited, or permitted by social norms. Therefore, if autonomous agents interact with humans, they also need to reason about various legal rules, social and ethical social norms, so they would be trusted and accepted by humans. Inverse Reinforcement Learning (IRL) can be used for the autonomous agents to learn social norm-compliant behavior via expert demonstrations. However, norms are context-sensitive, i.e. different norms get activated in different contexts. For example, the privacy norm is activated for a domestic robot entering a bathroom where a person may be present, whereas it is not activated for the robot entering the kitchen. Representing various contexts in the state space of the robot, as well as getting expert demonstrations under all possible tasks and contexts is extremely challenging. Inspired by recent work on Modularized Normative MDP (MNMDP) and early work on context-sensitive RL, we propose a new IRL framework, Context-Sensitive Norm IRL (CNIRL). CNIRL treats states and contexts separately, and assumes that the expert determines the priority of every possible norm in the environment, where each norm is associated with a distinct reward function. The agent chooses the action to maximize its cumulative rewards. We present the CNIRL model and show that its computational complexity is scalable in the number of norms. We also show via two experimental scenarios that CNIRL can handle problems with changing context spaces.
Yue Guo 0003, Boshi Wang, Dana Hughes 0001, Michael Lewis 0001, Katia P. Sycara
RO-MAN2