VLDB 2026 Research / reviewers in the wild / expert
Xinyu Pi
dblp:243/8713
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | UOUO: Uncontextualized Uncommon Objects for Measuring Knowledge Horizons of Vision Language ModelsabstractXinyu Pi, Mingyuan Wu, Jize Jiang, Haozhen Zheng, Beitong Tian, ChengXiang Zhai, Klara Nahrstedt, Zhiting Hu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Xinyu Pi, Mingyuan Wu, Jize Jiang, Haozhen Zheng, Beitong Tian, ChengXiang Zhai, Klara Nahrstedt, Zhiting Hu |
EMNLP | 1 |
| 2023 | S/C: Speeding up Data Materialization with Bounded MemoryabstractWith data pipeline tools and the expressiveness of SQL, managing interdependent materialized views (MVs) are becoming increasingly easy. These MVs are updated repeatedly upon new data ingestion (e.g., daily), from which database admins can observe performance metrics (e.g., refresh time of each MV, size on disk) in a consistent way for different types of updates (full vs. incremental) and for different systems (single node, distributed, cloud-hosted). One missed opportunity is that existing data systems treat those MV updates as independent SQL statements without fully exploiting their dependency information and performance metrics. However, if we know that the result of a SQL statement will be consumed immediately after for subsequent operations, those subsequent operations do not have to wait until the early results are fully materialized on storage because the results are already readily available in memory. Of course, this may come at a cost because keeping those results in memory (even temporarily) will reduce the amount of available memory; thus, our decision should be careful.In this paper, we introduce a new system, called S/C, which tackles this problem through efficient creation and update of a set of MVs with acyclic dependencies among them. S/C judiciously uses bounded memory to reduce the end-to-end MV refresh time by short-circuiting expensive reads and writes; S/C’s objective function accurately estimates the time savings from keeping intermediate data in memory for particular periods. Our solution jointly optimizes an MV refresh order, what data to keep in memory, and when to release the data from memory. At a high level, S/C still materializes all data exactly as defined in MV definitions; thus, it does not impact any service-level agreements. In our experiments with TPC-DS datasets (up to 1TB), we show that S/C's optimization can speedup end-to-end runtime by 1.04×–5.08× with (only) 1.6GB memory. Zhaoheng Li, Xinyu Pi, Yongjoo Park |
ICDE | 2 |
| 2023 | Light field reconstruction via attention maps of hybrid networks
Anzhi Wang, Xinyu Pi |
Vis. Comput. | 5 |
| 2022 | Towards Robustness of Text-to-SQL Models Against Natural and Realistic Adversarial Table PerturbationabstractThe robustness of Text-to-SQL parsers against adversarial perturbations plays a crucial role in delivering highly reliable applications.Previous studies along this line primarily focused on perturbations in the natural language question side, neglecting the variability of tables.Motivated by this, we propose the Adversarial Table Perturbation (ATP) as a new attacking paradigm to measure the robustness of Textto-SQL models.Following this proposition, we curate ADVETA, the first robustness evaluation benchmark featuring natural and realistic ATPs.All tested state-of-the-art models experience dramatic performance drops on ADVETA, revealing models' vulnerability in real-world practices.To defend against ATP, we build a systematic adversarial training example generation framework tailored for better contextualization of tabular data.Experiments show that our approach not only brings the best robustness improvement against tableside perturbations but also substantially empowers models against NL-side perturbations.We release our benchmark and code at: https://github.com/microsoft/ContextualSP. Xinyu Pi, Yan Gao 0002, Zhoujun Li 0001, Jian-Guang Lou |
ACL (1) | 1 |
| 2022 | Reasoning Like Program ExecutorsabstractReasoning over natural language is a longstanding goal for the research community.However, studies have shown that existing language models are inadequate in reasoning.To address the issue, we present POET, a novel reasoning pre-training paradigm.Through pretraining language models with programs and their execution results, POET empowers language models to harvest the reasoning knowledge possessed by program executors via a data-driven approach.POET is conceptually simple and can be instantiated by different kinds of program executors.In this paper, we showcase two simple instances POET-Math and POET-Logic, in addition to a complex instance, POET-SQL.Experimental results on six benchmarks demonstrate that POET can significantly boost model performance in natural language reasoning, such as numerical reasoning, logical reasoning, and multi-hop reasoning.POET opens a new gate on reasoningenhancement pre-training, and we hope our analysis would shed light on the future research of reasoning like program executors. Xinyu Pi, Qian Liu 0033, Bei Chen 0008, Morteza Ziyadi, Zeqi Lin, Qiang Fu 0015, Yan Gao 0002, Jian-Guang Lou, Weizhu Chen |
EMNLP | 1 |
| 2022 | LogiGAN: Learning Logical Reasoning via Adversarial Pre-trainingabstractWe present LogiGAN, an unsupervised adversarial pre-training framework for improving logical reasoning abilities of language models. Upon automatic identification of logical reasoning phenomena in massive text corpus via detection heuristics, we train language models to predict the masked-out logical statements. Inspired by the facilitation effect of reflective thinking in human learning, we analogically simulate the learning-thinking process with an adversarial Generator-Verifier architecture to assist logic learning. LogiGAN implements a novel sequential GAN approach that (a) circumvents the non-differentiable challenge of the sequential GAN by leveraging the Generator as a sentence-level generative likelihood scorer with a learning objective of reaching scoring consensus with the Verifier; (b) is computationally feasible for large-scale pre-training with arbitrary target length. Both base and large size language models pre-trained with LogiGAN demonstrate obvious performance improvement on 12 datasets requiring general reasoning abilities, revealing the fundamental role of logic in broad reasoning, as well as the effectiveness of LogiGAN. Ablation studies on LogiGAN components reveal the relative orthogonality between linguistic and logic abilities and suggest that reflective thinking's facilitation effect might also generalize to machine learning. Xinyu Pi, Wanjun Zhong, Yan Gao 0002, Nan Duan 0001, Jian-Guang Lou |
NeurIPS | 1 |
| 2021 | REFORM: Fast and Adaptive Solution for Subteam ReplacementabstractSubteam Replacement: given a team of people embedded in a social network to complete a certain task, and a subset of members (i.e., subteam) in this team which have become unavailable, find another set of people who can perform the subteam’s role in the larger team. We conjecture that a good candidate subteam should have high skill and structural similarity with the replaced subteam while sharing a similar connection with the larger team as a whole. Based on this conjecture, we propose a novel graph kernel which evaluates the goodness of candidate subteams in this holistic way freely adjustable to the need of the situation. To tackle the significant computational difficulties, we equip our kernel with a fast approximation algorithm which (a) employs effective pruning strategies, (b) exploits the similarity between candidate team structures to reduce kernel computations, and (c) features a solid theoretical bound on the quality of the obtained solution. We extensively test our solution on both synthetic and real datasets to demonstrate its effectiveness and efficiency. Our proposed graph kernel outputs more human-agreeable recommendations compared to metrics used in previous work, and our algorithm consistently outperforms alternative choices by finding nearoptimal solutions while scaling linearly with the size of the replaced subteam. Zhaoheng Li, Xinyu Pi, Mingyuan Wu, Hanghang Tong |
IEEE BigData | 2 |
| 2021 | Subteam Replacement: Problem Definition and Fast SolutionabstractIn settings such as corporate management where team structure is highly volatile and large-scale personnel changes are commonplace, the ability to simultaneously replace multiple team members in a team is highly appreciated. We define the problem of Subteam Replacement to address this observation: given a team of people embedded in a social network to complete a certain task, and a subset of members - subteam - in this team which has become unavailable, find another set of people which can perform the subteam's role in the larger team. We propose a holistic evaluation metric and scalable solution for Subteam Replacement with strong theoretical guarantees and perform quantitative evaluations on both generated and real datasets. Zhaoheng Li, Xinyu Pi, Mingyuan Wu |
SIGMOD Conference | 2 |