Sinong Wang

dblp:140/0795 · DBLP profile ↗
← Back
23ranked-venue papers
6as first author
14since 2021 · last 2025
0009-0008-5329-9620ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 5 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Improving Model Factuality with Fine-grained Critique-based Evaluator
abstract
Yiqing Xie, Wenxuan Zhou, Pradyot Prakash, Di Jin, Yuning Mao, Quintin Fettes, Arya Talebzadeh, Sinong Wang, Han Fang, Carolyn Rose, Daniel Fried, Hejia Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yiqing Xie, Pradyot Prakash, Yuning Mao, Quintin Fettes, Arya Talebzadeh, Sinong Wang, Carolyn P. Rosé, Daniel Fried
ACL (1)8
2025 Think Smarter not Harder: Adaptive Reasoning with Inference Aware Optimization
abstract
Solving mathematics problems has been an intriguing capability of large language models, and many efforts have been made to improve reasoning by extending reasoning length, such as through self-correction and extensive long chain-of-thoughts. While promising in problem-solving, advanced long reasoning chain models exhibit an undesired single-modal behavior, where trivial questions require unnecessarily tedious long chains of thought. In this work, we propose a way to allow models to be aware of inference budgets by formulating it as utility maximization with respect to an inference budget constraint, hence naming our algorithm Inference Budget-Constrained Policy Optimization (IBPO). In a nutshell, models fine-tuned through IBPO learn to ``understand'' the difficulty of queries and allocate inference budgets to harder ones. With different inference budgets, our best models are able to have a $4.14$\% and $5.74$\% absolute improvement ($8.08$\% and $11.2$\% relative improvement) on MATH500 using $2.16$x and $4.32$x inference budgets respectively, relative to LLaMA3.1 8B Instruct. These improvements are approximately $2$x those of self-consistency under the same budgets.
Zishun Yu, Tengyu Xu, Di Jin 0005, Karthik Abinav Sankararaman, Zhouhao Zeng, Eryk Helenowski, Sinong Wang, Hao Ma 0001
ICML10
2024 Representation Deficiency in Masked Language Modeling
abstract
Masked Language Modeling (MLM) has been one of the most prominent approaches for pretraining bidirectional text encoders due to its simplicity and effectiveness. One notable concern about MLM is that the special $\texttt{[MASK]}$ symbol causes a discrepancy between pretraining data and downstream data as it is present only in pretraining but not in fine-tuning. In this work, we offer a new perspective on the consequence of such a discrepancy: We demonstrate empirically and theoretically that MLM pretraining allocates some model dimensions exclusively for representing $\texttt{[MASK]}$ tokens, resulting in a representation deficiency for real tokens and limiting the pretrained model's expressiveness when it is adapted to downstream data without $\texttt{[MASK]}$ tokens. Motivated by the identified issue, we propose MAE-LM, which pretrains the Masked Autoencoder architecture with MLM where $\texttt{[MASK]}$ tokens are excluded from the encoder. Empirically, we show that MAE-LM improves the utilization of model dimensions for real token representations, and MAE-LM consistently outperforms MLM-pretrained models on the GLUE and SQuAD benchmarks.
Yu Meng 0001, Jitin Krishnan, Sinong Wang, Qifan Wang 0001, Yuning Mao, Marjan Ghazvininejad, Jiawei Han 0001, Luke Zettlemoyer
ICLR3
2024 LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models
abstract
Chi Han, Qifan Wang, Hao Peng, Wenhan Xiong, Yu Chen, Heng Ji, Sinong Wang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Chi Han, Qifan Wang 0001, Hao Peng 0009, Wenhan Xiong, Yu Chen 0022, Heng Ji 0001, Sinong Wang
NAACL-HLT7
2024 Effective Long-Context Scaling of Foundation Models
abstract
Wenhan Xiong, Jingyu Liu, Igor Molybog, Hejia Zhang, Prajjwal Bhargava, Rui Hou, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, Madian Khabsa, Han Fang, Yashar Mehdad, Sharan Narang, Kshitiz Malik, Angela Fan, Shruti Bhosale, Sergey Edunov, Mike Lewis, Sinong Wang, Hao Ma. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Wenhan Xiong, Igor Molybog, Prajjwal Bhargava, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, Madian Khabsa, Yashar Mehdad, Sharan Narang, Kshitiz Malik, Angela Fan, Shruti Bhosale, Sergey Edunov, Mike Lewis, Sinong Wang, Hao Ma 0001
NAACL-HLT20
2023 MUSTIE: Multimodal Structural Transformer for Web Information Extraction
abstract
Qifan Wang, Jingang Wang, Xiaojun Quan, Fuli Feng, Zenglin Xu, Shaoliang Nie, Sinong Wang, Madian Khabsa, Hamed Firooz, Dongfang Liu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Qifan Wang 0001, Jingang Wang, Xiaojun Quan, Fuli Feng, Zenglin Xu, Shaoliang Nie, Sinong Wang, Madian Khabsa, Hamed Firooz, Dongfang Liu
ACL (1)7
2023 Defending Against Patch-based Backdoor Attacks on Self-Supervised Learning
abstract
Recently, self-supervised learning (SSL) was shown to be vulnerable to patch-based data poisoning backdoor attacks. It was shown that an adversary can poison a small part of the unlabeled data so that when a victim trains an SSL model on it, the final model will have a back-door that the adversary can exploit. This work aims to defend self-supervised learning against such attacks. We use a three-step defense pipeline, where we first train a model on the poisoned data. In the second step, our proposed defense algorithm (PatchSearch) uses the trained model to search the training data for poisoned samples and removes them from the training set. In the third step, a final model is trained on the cleaned-up training set. Our results show that PatchSearch is an effective defense. As an example, it improves a model's accuracy on images containing the trigger from 38.2% to 63.7% which is very close to the clean model's accuracy, 64.6%. More-over, we show that PatchSearch outperforms baselines and state-of-the-art defense approaches including those using additional clean, trusted data. Our code is available at https://github.com/UCDvision/PatchSearch
Ajinkya Tejankar, Maziar Sanjabi, Qifan Wang 0001, Sinong Wang, Hamed Firooz, Hamed Pirsiavash, Liang Tan 0005
CVPR4
2023 APrompt: Attention Prompt Tuning for Efficient Adaptation of Pre-trained Language Models
abstract
Qifan Wang, Yuning Mao, Jingang Wang, Hanchao Yu, Shaoliang Nie, Sinong Wang, Fuli Feng, Lifu Huang, Xiaojun Quan, Zenglin Xu, Dongfang Liu. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Qifan Wang 0001, Yuning Mao, Jingang Wang, Hanchao Yu, Shaoliang Nie, Sinong Wang, Fuli Feng, Lifu Huang, Xiaojun Quan, Zenglin Xu, Dongfang Liu
EMNLP6
2022 Learning to Generate Question by Asking Question: A Primal-Dual Approach with Uncommon Word Generation
abstract
Automatic question generation (AQG) is the task of generating a question from a given passage and an answer.Most existing AQG methods aim at encoding the passage and the answer to generate the question.However, limited work has focused on modeling the correlation between the target answer and the generated question.Moreover, unseen or rare word generation has not been studied in previous works.In this paper, we propose a novel approach which incorporates question generation with its dual problem, question answering, into a unified primal-dual framework.Specifically, the question generation component consists of an encoder that jointly encodes the answer with the passage, and a decoder that produces the question.The question answering component then re-asks the generated question on the passage to ensure that the target answer is obtained.We further introduce a knowledge distillation module to improve the model generalization ability.We conduct an extensive set of experiments on SQuAD and HotpotQA benchmarks.Experimental results demonstrate the superior performance of the proposed approach over several state-of-the-art methods.
Qifan Wang 0001, Xiaojun Quan, Fuli Feng, Dongfang Liu, Zenglin Xu, Sinong Wang, Hao Ma 0001
EMNLP7
2022 IDPG: An Instance-Dependent Prompt Generation Method
abstract
Zhuofeng Wu, Sinong Wang, Jiatao Gu, Rui Hou, Yuxiao Dong, V.G.Vinod Vydiswaran, Hao Ma. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Zhuofeng Wu 0001, Sinong Wang, Jiatao Gu, Yuxiao Dong, V. G. Vinod Vydiswaran, Hao Ma 0001
NAACL-HLT2
2022 Sparse Distillation: Speeding Up Text Classification by Using Bigger Student Models
abstract
Qinyuan Ye, Madian Khabsa, Mike Lewis, Sinong Wang, Xiang Ren, Aaron Jaech. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Qinyuan Ye, Madian Khabsa, Mike Lewis, Sinong Wang, Xiang Ren 0001, Aaron Jaech
NAACL-HLT4
2021 On the Influence of Masking Policies in Intermediate Pre-training
abstract
Current NLP models are predominantly trained through a two-stage "pre-train then fine-tune" pipeline.Prior work has shown that inserting an intermediate pre-training stage, using heuristic masking policies for masked language modeling (MLM), can significantly improve final performance.However, it is still unclear (1) in what cases such intermediate pre-training is helpful, (2) whether hand-crafted heuristic objectives are optimal for a given task, and (3) whether a masking policy designed for one task is generalizable beyond that task.In this paper, we perform a large-scale empirical study to investigate the effect of various masking policies in intermediate pre-training with nine selected tasks across three categories.Crucially, we introduce methods to automate the discovery of optimal masking policies via direct supervision or meta-learning.We conclude that the success of intermediate pre-training is dependent on appropriate pre-train corpus, selection of output format (i.e., masked spans or full sentence), and clear understanding of the role that MLM plays for the downstream task.In addition, we find our learned masking policies outperform the heuristic of masking named entities on TriviaQA, and policies learned from one task can positively transfer to other tasks in certain cases, inviting future research in this direction.
Qinyuan Ye, Belinda Z. Li, Sinong Wang, Benjamin Bolte, Hao Ma 0001, Scott Yih, Xiang Ren 0001, Madian Khabsa
EMNLP (1)3
2021 On Unifying Misinformation Detection
abstract
Nayeon Lee, Belinda Z. Li, Sinong Wang, Pascale Fung, Hao Ma, Wen-tau Yih, Madian Khabsa. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Nayeon Lee, Belinda Z. Li, Sinong Wang, Pascale Fung, Hao Ma 0001, Scott Yih, Madian Khabsa
NAACL-HLT3
2021 Luna: Linear Unified Nested Attention
abstract
The quadratic computational and memory complexities of the Transformer's attention mechanism have limited its scalability for modeling long sequences. In this paper, we propose Luna, a linear unified nested attention mechanism that approximates softmax attention with two nested linear attention functions, yielding only linear (as opposed to quadratic) time and space complexity. Specifically, with the first attention function, Luna packs the input sequence into a sequence of fixed length. Then, the packed sequence is unpacked using the second attention function. As compared to a more traditional attention mechanism, Luna introduces an additional sequence with a fixed length as input and an additional corresponding output, which allows Luna to perform attention operation linearly, while also storing adequate contextual information. We perform extensive evaluations on three benchmarks of sequence modeling tasks: long-context sequence modelling, neural machine translation and masked language modeling for large-scale pretraining. Competitive or even better experimental results demonstrate both the effectiveness and efficiency of Luna compared to a variety of strong baseline methods including the full-rank attention and other efficient sparse and dense attention methods.
Xuezhe Ma, Xiang Kong, Sinong Wang, Chunting Zhou, Jonathan May, Hao Ma 0001, Luke Zettlemoyer
NeurIPS3
2020 To Pretrain or Not to Pretrain: Examining the Benefits of Pretrainng on Resource Rich Tasks
abstract
Pretraining NLP models with variants of Masked Language Model (MLM) objectives has recently led to a significant improvements on many tasks.This paper examines the benefits of pretrained models as a function of the number of training samples used in the downstream task.On several text classification tasks, we show that as the number of training examples grow into the millions, the accuracy gap between finetuning BERT-based model and training vanilla LSTM from scratch narrows to within 1%.Our findings indicate that MLM-based models might reach a diminishing return point as the supervised data size increases significantly.
Sinong Wang, Madian Khabsa, Hao Ma 0001
ACL1
2019 Computation Efficient Coded Linear Transform
abstract
In large-scale distributed linear transform problems, coded computation plays an important role to reduce the delay caused by slow machines. However, existing coded schemes could end up destroying the significant sparsity that exists in large-scale machine learning problems, and in turn increase the computational delay. In this paper, we propose a coded computation strategy, referred to as diagonal code, that achieves the optimum recovery threshold and the optimum computation load. Furthermore, by leveraging the ideas from random proposal graph theory, we design a random code that achieves a constant computation load, which significantly outperforms the existing best known result. We apply our schemes to the distributed gradient descent problem and demonstrate the advantage of the approach over current fastest coded schemes.
Sinong Wang, Jiashang Liu, Ness Shroff, Pengyu Yang
AISTATS1
2018 A Near-Optimal Control Policy in Cloud Systems with Renewable Sources and Time-Dependent Energy Price
abstract
The cost of energy usage is of significant concern in cloud/data center systems that support a large number of servers. A simple way to reduce energy consumption and the electricity bill is to turn some of the servers off during periods of under utilization. However, turning a server back on typically consumes a lot of energy. Another way to reduce the energy cost is to equip cloud systems with renewable resources and batteries. Most works in the literature have focused on one or the other approach. In this work, we propose a joint server on-off and energy control policy, which determines the servers' on-off status, as well as the energy purchasing behavior, by taking electricity price, renewable resources, possible future tasks and turn-on cost into account. The server on-off control component is proved to be optimal in terms of energy consumption minimization. The joint policy is shown to be arbitrarily close to the optimal solution in terms of electricity bill minimization, in the case where the battery has infinite capacity with stable energy level. Simulation results show that even under reasonable battery size, a significant electricity cost reduction is achieved with the proposed policy.
Jiashang Liu, Ness Shroff, Prasun Sinha, Sinong Wang
IEEE CLOUD5
2018 Coded Sparse Matrix Multiplication
abstract
In a large-scale and distributed matrix multiplication problem $C=A^{\intercal}B$, where $C\in\mathbb{R}^{r\times t}$, the coded computation plays an important role to effectively deal with “stragglers” (distributed computations that may get delayed due to few slow or faulty processors). However, existing coded schemes could destroy the significant sparsity that exists in large-scale machine learning problems, and could result in much higher computation overhead, i.e., $O(rt)$ decoding time. In this paper, we develop a new coded computation strategy, we call sparse code, which achieves near optimal recovery threshold, low computation overhead, and linear decoding time $O(nnz(C))$. We implement our scheme and demonstrate the advantage of the approach over both uncoded and current fastest coded strategies.
Sinong Wang, Jiashang Liu, Ness Shroff
ICML1
2018 UCBoost: A Boosting Approach to Tame Complexity and Optimality for Stochastic Bandits
abstract
In this work, we address the open problem of finding low-complexity near-optimal multi-armed bandit algorithms for sequential decision making problems. Existing bandit algorithms are either sub-optimal and computationally simple (e.g., UCB1) or optimal and computationally complex (e.g., kl-UCB). We propose a boosting approach to Upper Confidence Bound based algorithms for stochastic bandits, that we call UCBoost. Specifically, we propose two types of UCBoost algorithms. We show that UCBoost(D) enjoys O(1) complexity for each arm per round as well as regret guarantee that is 1/e-close to that of the kl-UCB algorithm. We propose an approximation-based UCBoost algorithm, UCBoost(epsilon), that enjoys a regret guarantee epsilon-close to that of kl-UCB as well as O(log(1/epsilon)) complexity for each arm per round. Hence, our algorithms provide practitioners a practical way to trade optimality with computational complexity. Finally, we present numerical results which show that UCBoost(epsilon) can achieve the same regret performance as the standard kl-UCB while incurring only 1% of the computational cost of kl-UCB.
Fang Liu 0020, Sinong Wang, Swapna Buccapatnam, Ness Shroff
IJCAI2
2017 Non-Additive Security Games
abstract
Security agencies have found security games to be useful models to understand how to better protect their assets. The key practical elements in this work are: (i) the attacker can simultaneously attack multiple targets, and (ii) different targets exhibit different types of dependencies based on the assets being protected (e.g., protection of critical infrastructure, network security, etc.). However, little is known about the computational complexity of these problems, especially when there exist dependencies among the targets. Moreover, previous security game models do not in general scale well. In this paper, we investigate a general security game where the utility function is defined on a collection of subsets of all targets, and provide a novel theoretical framework to show how to compactly represent such a game, efficiently compute the optimal (minimax) strategies, and characterize the complexity of this problem. We apply our theoretical framework to the network security game. We characterize settings under which we find a polynomial time algorithm for computing optimal strategies. In other settings we prove the problem is NP-hard and provide an approximation algorithm.
Sinong Wang, Fang Liu 0020, Ness Shroff
AAAI1
2017 A New Alternating Direction Method for Linear Programming
abstract
It is well known that, for a linear program (LP) with constraint matrix $\mathbf{A}\in\mathbb{R}^{m\times n}$, the Alternating Direction Method of Multiplier converges globally and linearly at a rate $O((\|\mathbf{A}\|_F^2+mn)\log(1/\epsilon))$. However, such a rate is related to the problem dimension and the algorithm exhibits a slow and fluctuating ``tail convergence'' in practice. In this paper, we propose a new variable splitting method of LP and prove that our method has a convergence rate of $O(\|\mathbf{A}\|^2\log(1/\epsilon))$. The proof is based on simultaneously estimating the distance from a pair of primal dual iterates to the optimal primal and dual solution set by certain residuals. In practice, we result in a new first-order LP solver that can exploit both the sparsity and the specific structure of matrix $\mathbf{A}$ and a significant speedup for important problems such as basis pursuit, inverse covariance matrix estimation, L1 SVM and nonnegative matrix factorization problem compared with current fastest LP solvers.
Sinong Wang, Ness Shroff
NIPS1
2014 A predictive methodology for truthful double spectrum auctions in cognitive radio networks
abstract
Auction is often applied in cognitive radio networks due to its efficiency and fairness properties. An important issue in designing an auction mechanism is how to utilize the limited spectrum resource in an efficient manner. In order to achieve this goal, we propose a predictive double spectrum auction model in this paper. Our auction model first obtains the bidding range from statistical analysis, and then separates the interval into independent states and employees a Markovian prediction based algorithm to generate guidelines for the bidding range of primary and secondary users, respectively. Comparing with existing approaches, our proposed auction model is more efficient in spectrum utilization and satisfies the economic properties. Extensive simulation results show that our work achieves an utilization ratio up to 91%.
Zhe Liu 0024, Sinong Wang, Weijie Wu, Xiaohua Tian, Changle Li, Xinbing Wang
GLOBECOM2
2013 A New Approach to Multi-objective Virtual Machine Placement in Virtualized Data Center
abstract
In this paper, a virtual machine placement model to maximize resource utilization, balance multi-dimensional resources use and minimize communication traffic simultaneously within the data center is proposed. The multi-objective problem is simplified by employing average valued inequality and positional constraints. The improved genetic algorithm with local heuristic method and elitism strategy is developed to solve the problem. The simulation results show that performance gains in all aspects can be achieved by the proposed model and algorithm compared to the existing algorithms.
Sinong Wang, Huaxi Gu
NAS1