Silei Xu

dblp:143/9693 · DBLP profile ↗
← Back
13ranked-venue papers
7as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 3 since 2021Databases, data management, data science and information retrieval · 5 · 4 first-authorSystems, architecture and hardware · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Language models and text generation · 38% Question answering and dialogue systems · 27% Information extraction and text analysis · 15%
Network and information security
2 papers
Security and privacy of machine learning · 75% Privacy and data protection · 25%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Storage systems · 100%
Databases, data mining, and information retrieval
3 papers
Knowledge graphs · 38% Recommender systems · 37% Database theory · 25%
Theoretical computer science
1 paper
Mathematical optimization · 50% Approximation and online algorithms · 50%

Topics — the 21 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
semantic parsing
1.122023
Fine-tuned LLMs Know More, Hallucinate Less with Few-Shot Sequence-to-Sequence Semantic Parsing over Wikidata · EMNLP 2023
AutoQA: From Databases To QA Semantic Parsers With Only Synthetic Training Data · EMNLP (1) 2020
Natural language and speech › Language models and text generation › instruction following
instruction hierarchy
0.912025
Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy · ICLR 2025
Natural language and speech › Language models and text generation
large language model safety
0.912025
Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy · ICLR 2025
Security and privacy of machine learning › adversarial defense
prompt injection defense
0.912025
Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy · ICLR 2025
Natural language and speech › Language models and text generation
preference optimization
0.812024
WPO: Enhancing RLHF with Weighted Preference Optimization · EMNLP 2024
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.812024
WPO: Enhancing RLHF with Weighted Preference Optimization · EMNLP 2024
Natural language and speech › Question answering and dialogue systems
knowledge base question answering
0.712023
Fine-tuned LLMs Know More, Hallucinate Less with Few-Shot Sequence-to-Sequence Semantic Parsing over Wikidata · EMNLP 2023
Natural language and speech › Question answering and dialogue systems
open-domain question answering
0.412020
Localizing Open-Ontology QA Semantic Parsers in a Day Using Machine Translation · EMNLP (1) 2020
Natural language and speech › Question answering and dialogue systems
question answering over structured data
0.412020
AutoQA: From Databases To QA Semantic Parsers With Only Synthetic Training Data · EMNLP (1) 2020
Machine learning › Generative modeling
synthetic training data
0.412020
AutoQA: From Databases To QA Semantic Parsers With Only Synthetic Training Data · EMNLP (1) 2020
Natural language and speech › Question answering and dialogue systems
natural language interface
0.312017
Almond: The Architecture of an Open, Crowdsourced, Privacy-Preserving, Programmable Virtual Assistant · WWW 2017
Human-AI interaction › intelligent assistant
virtual assistants
0.312017
Almond: The Architecture of an Open, Crowdsourced, Privacy-Preserving, Programmable Virtual Assistant · WWW 2017
Natural language and speech › Language models and text generation
instruction following
0.312025
Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy · ICLR 2025
Knowledge graphs › semantic web
wikidata
0.212023
Fine-tuned LLMs Know More, Hallucinate Less with Few-Shot Sequence-to-Sequence Semantic Parsing over Wikidata · EMNLP 2023
Storage systems › storage reliability
erasure coding
0.212014
Single Disk Failure Recovery forX-Code-Based Parallel Storage Systems · IEEE Trans. Computers 2014
Storage systems › distributed storage
parallel storage system
0.212014
Single Disk Failure Recovery forX-Code-Based Parallel Storage Systems · IEEE Trans. Computers 2014
Storage systems
storage reliability
0.212014
Single Disk Failure Recovery forX-Code-Based Parallel Storage Systems · IEEE Trans. Computers 2014
Approximation and online algorithms › approximation algorithms › combinatorial approximation algorithms
greedy approximation
0.212014
Product selection problem: improve market share by learning consumer behavior · KDD 2014
Mathematical optimization › submodular optimization
submodular maximization
0.212014
Product selection problem: improve market share by learning consumer behavior · KDD 2014
Natural language and speech › Machine translation
neural machine translation
0.112020
Localizing Open-Ontology QA Semantic Parsers in a Day Using Machine Translation · EMNLP (1) 2020
Database theory
database schema
0.112020
AutoQA: From Databases To QA Semantic Parsers With Only Synthetic Training Data · EMNLP (1) 2020

Methods — techniques the papers use, named apart from their topics

instructional segment embedding · 1.7few-shot sequence-to-sequence semantic parsing · 1.3LLM fine-tuning · 1.3template-based parsing · 0.9paraphrasing · 0.9weighted preference optimization · 0.8reinforcement learning · 0.8domain-specific language · 0.6crowdsourcing · 0.6neural machine translation · 0.4few-shot learning · 0.4XLMR-LSTM · 0.4submodularity analysis · 0.4greedy algorithm · 0.4trace-driven simulation · 0.2integer linear programming · 0.2
YearPublicationVenuePosition
2025 Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy
abstract
Large Language Models (LLMs) are susceptible to security and safety threats, such as prompt injection, prompt extraction, and harmful requests. One major cause of these vulnerabilities is the lack of an instruction hierarchy. Modern LLM architectures treat all inputs equally, failing to distinguish between and prioritize various types of instructions, such as system messages, user prompts, and data. As a result, lower-priority user prompts may override more critical system instructions, including safety protocols. Existing approaches to achieving instruction hierarchy, such as delimiters and instruction-based training, do not address this issue at the architectural level. We introduce the $\textbf{I}$nstructional $\textbf{S}$egment $\textbf{E}$mbedding (ISE) technique, inspired by BERT, to modern large language models, which embeds instruction priority information directly into the model. This approach enables models to explicitly differentiate and prioritize various instruction types, significantly improving safety against malicious prompts that attempt to override priority rules. Our experiments on the Structured Query and Instruction Hierarchy benchmarks demonstrate an average robust accuracy increase of up to 15.75\% and 18.68\%, respectively. Furthermore, we observe an improvement in the instruction-following capability of up to 4.1\% on AlpacaEval. Overall, our approach offers a promising direction for enhancing the safety and effectiveness of LLM architectures.
Shujian Zhang, Kaiqiang Song, Silei Xu, Sanqiang Zhao, Ravi Agrawal, Sathish Reddy Indurthi, Chong Xiang 0001, Prateek Mittal, Wenxuan Zhou 0005
ICLR4
2024 WPO: Enhancing RLHF with Weighted Preference Optimization
abstract
Wenxuan Zhou, Ravi Agrawal, Shujian Zhang, Sathish Reddy Indurthi, Sanqiang Zhao, Kaiqiang Song, Silei Xu, Chenguang Zhu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Wenxuan Zhou 0005, Ravi Agrawal, Shujian Zhang, Sathish Reddy Indurthi, Sanqiang Zhao, Kaiqiang Song, Silei Xu
EMNLP7
2023 Fine-tuned LLMs Know More, Hallucinate Less with Few-Shot Sequence-to-Sequence Semantic Parsing over Wikidata
abstract
While large language models (LLMs) can answer many questions correctly, they can also hallucinate and give wrong answers.Wikidata, with its over 12 billion facts, can be used to ground LLMs to improve their factuality.This paper presents WikiWebQuestions, a highquality question answering benchmark for Wikidata.Ported over from WebQuestions for Freebase, it consists of real-world data with SPARQL annotation.This paper presents a few-shot sequence-tosequence semantic parser for Wikidata.We modify SPARQL to use the unique domain and property names instead of their IDs.We train the parser to use either the results from an entity linker or mentions in the query.We fine-tune LLaMA by adding the few-shot training data to that used to fine-tune Alpaca.Our experimental results demonstrate the effectiveness of this methodology, establishing a strong baseline of 76% and 65% answer accuracy in the dev and test sets of WikiWeb-Questions, respectively.By pairing our semantic parser with GPT-3, we combine verifiable results with qualified GPT-3 guesses to provide useful answers to 96% of the questions in dev.We also show that our method outperforms the state-of-the-art for the QALD-7 Wikidata dataset by 3.6% in F1 score. 1 * Equal contribution 1 Code, data, and model are available at https://github.com/stanford-oval/ wikidata-emnlp23
Silei Xu, Shicheng Liu, Theo Culhane, Elizaveta Pertseva, Meng-Hsi Wu, Sina J. Semnani, Monica S. Lam
EMNLP1
2020 Schema2QA: High-Quality and Low-Cost Q&A Agents for the Structured Web
abstract
Building a question-answering agent currently requires large annotated datasets, which are prohibitively expensive. This paper proposes Schema2QA, an open-source toolkit that can generate a Q&A system from a database schema augmented with a few annotations for each field. The key concept is to cover the space of possible compound queries on the database with a large number of in-domain questions synthesized with the help of a corpus of generic query templates. The synthesized data and a small paraphrase set are used to train a novel neural network based on the BERT pretrained model. We use Schema2QA to generate Q&A systems for five Schema.org domains, restaurants, people, movies, books and music, and obtain an overall accuracy between 64% and 75% on crowdsourced questions for these domains. Once annotations and paraphrases are obtained for a Schema.org schema, no additional manual effort is needed to create a Q&A agent for any website that uses the same schema. Furthermore, we demonstrate that learning can be transferred from the restaurant to the hotel domain, obtaining a 64% accuracy on crowdsourced questions with no manual effort. Schema2QA achieves an accuracy of 60% on popular restaurant questions that can be answered using Schema.org. Its performance is comparable to Google Assistant, 7% lower than Siri, and 15% higher than Alexa. It outperforms all these assistants by at least 18% on more complex, long-tail questions.
Silei Xu, Giovanni Campagna, Jian Li 0054, Monica S. Lam
CIKM1
2020 Localizing Open-Ontology QA Semantic Parsers in a Day Using Machine Translation
abstract
We propose Semantic Parser Localizer (SPL), a toolkit that leverages Neural Machine Translation (NMT) systems to localize a semantic parser for a new language.Our methodology is to (1) generate training data automatically in the target language by augmenting machine-translated datasets with local entities scraped from public websites, (2) add a fewshot boost of human-translated sentences and train a novel XLMR-LSTM semantic parser, and (3) test the model on natural utterances curated using human translators.We assess the effectiveness of our approach by extending the current capabilities of Schema2QA, a system for English Question Answering (QA) on the open web, to 10 new languages for the restaurants and hotels domains.Our models achieve an overall test accuracy ranging between 61% and 69% for the hotels domain and between 64% and 78% for restaurants domain, which compares favorably to 69% and 80% obtained for English parser trained on gold English data and a few examples from validation set.We show our approach outperforms the previous state-of-theart methodology by more than 30% for hotels and 40% for restaurants with localized ontologies for the subset of languages tested.Our methodology enables any software developer to add a new language capability to a QA system for a new domain, leveraging machine translation, in less than 24 hours.Our code is released open-source. 1 Language Country Examples Hotels English I want a hotel near times square that has at least 1000 reviews. ArabicGerman Ich möchte ein hotel in der nähe von marienplatz, das mindestens 1000 bewertungen hat.Spanish Busco un hotel cerca de puerto banús que tenga al menos 1000 comentarios.Farsi Finnish Haluan paikan helsingin tuomiokirkko läheltä hotellin, jolla on vähintään 1000 arvostelua.
Mehrad Moradshahi, Giovanni Campagna, Sina J. Semnani, Silei Xu, Monica S. Lam
EMNLP (1)4
2020 AutoQA: From Databases To QA Semantic Parsers With Only Synthetic Training Data
abstract
We propose AutoQA, a methodology and toolkit to generate semantic parsers that answer questions on databases, with no manual effort.Given a database schema and its data, AutoQA automatically generates a large set of high-quality questions for training that covers different database operations.It uses automatic paraphrasing combined with templatebased parsing to find alternative expressions of an attribute in different parts of speech.It also uses a novel filtered auto-paraphraser to generate correct paraphrases of entire sentences.We apply AutoQA to the Schema2QA dataset and obtain an average logical form accuracy of 62.9% when tested on natural questions, which is only 6.4% lower than a model trained with expert natural language annotations and paraphrase data collected from crowdworkers.To demonstrate the generality of AutoQA, we also apply it to the Overnight dataset.AutoQA achieves 69.8% answer accuracy, 16.4% higher than the state-of-the-art zero-shot models and only 5.2% lower than the same model trained with human data.
Silei Xu, Sina J. Semnani, Giovanni Campagna, Monica S. Lam
EMNLP (1)1
2019 Genie: a generator of natural language semantic parsers for virtual assistant commands
abstract
To understand diverse natural language commands, virtual assistants today are trained with numerous labor-intensive, manually annotated sentences. This paper presents a methodology and the Genie toolkit that can handle new compound commands with significantly less manual effort. We advocate formalizing the capability of virtual assistants with a Virtual Assistant Programming Language (VAPL) and using a neural semantic parser to translate natural language into VAPL code. Genie needs only a small realistic set of input sentences for validating the neural model. Developers write templates to synthesize data; Genie uses crowdsourced paraphrases and data augmentation, along with the synthesized data, to train a semantic parser. We also propose design principles that make VAPL languages amenable to natural language translation. We apply these principles to revise ThingTalk, the language used by the Almond virtual assistant. We use Genie to build the first semantic parser that can support compound virtual assistants commands with unquoted free-form parameters. Genie achieves a 62% accuracy on realistic user inputs. We demonstrate Genie’s generality by showing a 19% and 31% improvement over the previous state of the art on a music skill, aggregate functions, and access control.
Giovanni Campagna, Silei Xu, Mehrad Moradshahi, Richard Socher, Monica S. Lam
PLDI2
2018 Brassau: automatic generation of graphical user interfaces for virtual assistants
abstract
This paper presents Brassau, a graphical virtual assistant that converts natural language commands into GUIs. A virtual assistant with a GUI has the following benefits compared to text or speech based virtual assistants: users can monitor multiple queries simultaneously, it is easy to re-run complex commands, and user can adjust settings using multiple modes of interaction. Brassau introduces a novel template-based approach that leverages a large corpus of images to make GUIs visually diverse and interesting. Brassau matches a command from the user to an image to create a GUI. This approach decouples the commands from GUIs and allows for reuse of GUIs across multiple commands. In our evaluation, users prefer the widgets produced by Brassau over plain GUIs.
Michael Fischer 0010, Giovanni Campagna, Silei Xu, Monica S. Lam
MobileHCI3
2017 Almond: The Architecture of an Open, Crowdsourced, Privacy-Preserving, Programmable Virtual Assistant
abstract
This paper presents the architecture of Almond, an open, crowdsourced, privacy-preserving and programmable virtual assistant for online services and the Internet of Things (IoT). Included in Almond is Thingpedia, a crowdsourced public knowledge base of natural language interfaces and open APIs. Our proposal addresses four challenges in virtual assistant technology: generality, interoperability, privacy, and usability. Generality is addressed by crowdsourcing Thingpedia, while interoperability is provided by ThingTalk, a high-level domain-specific language that connects multiple devices or services via open APIs. For privacy, user credentials and user data are managed by our open-source ThingSystem, which can be run on personal phones or home servers. Finally, we address usability by providing a natural language interface, whose capability can be extended via training with the help of a menu-driven interface.
Giovanni Campagna, Rakesh Ramesh, Silei Xu, Michael Fischer 0010, Monica S. Lam
WWW3
2016 Product Selection Problem: Improve Market Share by Learning Consumer Behavior
abstract
It is often crucial for manufacturers to decide what products to produce so that they can increase their market share in an increasingly fierce market. To decide which products to produce, manufacturers need to analyze the consumers’ requirements and how consumers make their purchase decisions so that the new products will be competitive in the market. In this paper, we first present a general distance-based product adoption model to capture consumers’ purchase behavior. Using this model, various distance metrics can be used to describe different real life purchase behavior. We then provide a learning algorithm to decide which set of distance metrics one should use when we are given some accessible historical purchase data. Based on the product adoption model, we formalize the k most marketable products (or k- MMP ) selection problem and formally prove that the problem is NP-hard . To tackle this problem, we propose an efficient greedy-based approximation algorithm with a provable solution guarantee. Using submodularity analysis, we prove that our approximation algorithm can achieve at least 63% of the optimal solution. We apply our algorithm on both synthetic datasets and real-world datasets (TripAdvisor.com), and show that our algorithm can easily achieve five or more orders of speedup over the exhaustive search and achieve about 96% of the optimal solution on average. Our experiments also demonstrate the robustness of our distance metric learning method, and illustrate how one can adopt it to improve the accuracy of product selection.
Silei Xu, John C. S. Lui
ACM Trans. Knowl. Discov. Data1
2014 Product selection problem: improve market share by learning consumer behavior
abstract
It is often crucial for manufacturers to decide what products to produce so that they can increase their market share in an increasingly fierce market. To decide which products to produce, manufacturers need to analyze the consumers' requirements and how consumers make their purchase decisions so that the new products will be competitive in the market. In this paper, we first present a general distance-based product adoption model to capture consumers' purchase behavior. Using this model, various distance metrics can be used to describe different real life purchase behavior. We then provide a learning algorithm to decide which set of distance metrics one should use when we are given some historical purchase data. Based on the product adoption model, we formalize the k most marketable products (or k-MMP) selection problem and formally prove that the problem is NP-hard. To tackle this problem, we propose an efficient greedy-based approximation algorithm with a provable solution guarantee. Using submodularity analysis, we prove that our approximation algorithm can achieve at least 63% of the optimal solution. We apply our algorithm on both synthetic datasets and real-world datasets (TripAdvisor.com), and show that our algorithm can easily achieve five or more orders of speedup over the exhaustive search and achieve about 96% of the optimal solution on average. Our experiments also show the significant impact of different distance metrics on the results, and how proper distance metrics can improve the accuracy of product selection.
Silei Xu, John C. S. Lui
KDD1
2014 A provable algorithmic approach to product selection problems for market entry and sustainability
abstract
Given the globalized economy, how to process the heterogeneous web data so to extract customers' purchase behavior is crucial to manufacturers who want to enter or sustain in a competitive market. To maximize the sales, manufacturers not only need to decide what products to produce so to meet diverse customers' requirements, but at the same time, compete with competitors' products. In this paper, we present a general framework for the following product selection problems: (1) k-BSP problem, which is for a manufacturer to enter a competitive market, and (2) k-BBP problem, which is for a manufacturer to sustain in a competitive market. We propose several product adoption models to describe the complex purchase behavior of customers, and formally show that these problems are NP-hard in general. To tackle these problems, we propose computationally efficient greedy-based approximation algorithms. Based on the submodularity analysis, we prove that our algorithms can guarantee a (1--1/e)-approximation ratio as compared to the optimal solutions. We perform large scale data analysis to show the efficiency and accuracy of our framework. In our experiments, we observe 1,300 to 250,000 times speedup as compared to the exhaustive algorithms, and our solutions can achieve on average 96% of solution quality as compared to the optimal solutions. Finally, we apply our algorithms on web dataset to show the impact of customers' different purchase behavior on the results of product selection.
Silei Xu, Yishi Lin, Hong Xie 0004, John C. S. Lui
SSDBM1
2014 Single Disk Failure Recovery forX-Code-Based Parallel Storage Systems
abstract
In modern parallel storage systems (e.g., cloud storage and data centers), it is important to provide data availability guarantees against disk (or storage node) failures via redundancy coding schemes. One coding scheme is X-code, which is double-fault tolerant while achieving the optimal update complexity. When a disk/node fails, recovery must be carried out to reduce the possibility of data unavailability. We propose an X-code-based optimal recovery scheme called minimum-disk-read-recovery (MDRR), which minimizes the number of disk reads for single-disk failure recovery. We make several contributions. First, we show that MDRR provides optimal single-disk failure recovery and reduces about 25 percent of disk reads compared to the conventional recovery approach. Second, we prove that any optimal recovery scheme for X-code cannot balance disk reads among different disks within a single stripe in general cases. Third, we propose an efficient logical encoding scheme that issues balanced disk read in a group of stripes for any recovery algorithm (including the MDRR scheme). Finally, we implement our proposed recovery schemes and conduct extensive testbed experiments in a networked storage system prototype. Experiments indicate that MDRR reduces around 20 percent of recovery time of the conventional approach, showing that our theoretical findings are applicable in practice.
Silei Xu, Runhui Li, Patrick P. C. Lee, Yunfeng Zhu, Liping Xiang, Yinlong Xu 0001, John C. S. Lui
IEEE Trans. Computers1