VLDB 2026 Research / reviewers in the wild / expert
Xinyi Hou
dblp:73/7287
· DBLP profile ↗
13ranked-venue papers
5as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-User Boolean Keyword Searchable Encryption With Fine-Grained Access Control for Cloud StorageabstractABSTRACT Searchable Encryption (SE) enables users to perform searches on encrypted data while preserving data privacy. Since cloud servers are platforms that provide services for a large number of users, and data owners require access control over their data, SE schemes that support multi‐user settings and access control are therefore more suitable for cloud storage. However, in existing SE schemes that support multi‐user settings and access control, most only support single‐keyword or conjunctive keyword searches, and the search time grows linearly with the total amount of data. These limitations negatively impact both the accuracy and efficiency of search operations. This work proposes an SE scheme specifically designed for multi‐user settings. Data owners can enforce fine‐grained access control policies, while a specialized retrieval structure allows the cloud to assist users in performing Boolean keyword searches with improved efficiency. The search complexity of the proposed scheme is , where denotes the number of files relevant to the queried keyword. We demonstrate the scheme's effectiveness and practicality through performance analysis. Xinyi Hou, Ye Su 0001, Jing Qin 0002, Jixin Ma 0001 |
Concurr. Comput. Pract. Exp. | 1 |
| 2025 | Unsupervised Domain Adaptive Hand Mesh Reconstruction of 2D Images in the Wild
Xinyi Hou, Huayi Zhou 0001, Yue Ding 0001, Hongtao Lu 0001 |
ICANN (2) | 1 |
| 2025 | GPT Store Mining and AnalysisabstractAs an important extension of the ChatGPT ecosystem, GPT Store has developed into an active market hosting more than 3 million customized ChatGPTs (GPTs).Despite its large scale, the current academic community still has obvious limitations in its understanding of the ecosystem of this platform.Based on a complete dataset of more than 700,000 GPTs, this paper has achieved a multidimensional analysis of GPT Store.We first systematically examined the platform operation mechanism, covering core elements such as the classification system, interaction mode, and evaluation system.We also comprehensively analyzed the security risks, such as data leakage and jailbreak in GPT Store.Finally, through a user study, this work revealed the behavioral characteristics and experience pain points in real usage scenarios.Based on these findings, we provide operational platform optimization suggestions, including functional improvement, security enhancement, and interaction improvement.This study not only constructs an analytical framework for the GPT Store ecosystem but also provides empirical evidence and optimization directions for its future development. Dongxun Su, Yanjie Zhao 0001, Xinyi Hou, Shenao Wang 0001, Haoyu Wang 0001 |
Internetware | 3 |
| 2025 | CLIP-MT: Multi-Modal Knowledge-Driven Adaptive Scale Feature Allocation for Multi-Task Dense PredictionabstractRecent advancements in visual multi-task learning (MTL) have sparked significant interest. However, existing dense prediction MTL methods predominantly rely on single-modality image data, limiting their performance due to the absence of complementary knowledge from other modalities. Additionally, different dense tasks exhibit heterogeneous preferences during information decoding, posing a critical challenge in effectively allocating multi-scale encoded features. To address these limitations, we propose CLIP-MT, a Multi-Modal Knowledge-Driven Adaptive Scale Feature Allocation for Multi-Task Dense Prediction. Specifically, to enrich task-shared image features with multi-modal knowledge, we introduce a novel CLIP-Guided Global Feature Enhancer (CGGF), which leverages aligned text-image information to augment object-level representations through a dual-path feature fusion architecture. Furthermore, to tackle the task-specific scale preference problem, we design an Adaptive Scale Selection Gate (ASSG), a learnable gating mechanism that dynamically selects high- or low-scale features based on task-specific demands. Finally, we integrate multi-modal and multi-scale information through a Task-Aware Feature Fusion Module (TAFF). Extensive experiments on the NYUDv2 and PASCAL-Context datasets demonstrate that CLIP-MT achieves state-of-the-art performance, outperforming existing methods across multiple dense prediction tasks. Shalayiding Sirejiding, Yue Ding 0001, Xinyi Hou, Shaokai Wu, Qichen He, Hongtao Lu 0001 |
ACM Multimedia | 4 |
| 2025 | On the (In)Security of LLM App StoresabstractLLM app stores have seen rapid growth, leading to the proliferation of numerous custom LLM apps. However, this expansion raises security concerns. In this study, we propose a three-layer concern framework to identify the potential security risks of LLM apps, i.e., LLM apps with abusive potential, LLM apps with malicious intent, and LLM apps with backdoors. Over five months, we collected 786,036 LLM apps from six major app stores: GPT Store, FlowGPT, Poe, Coze, Cici, and Character.AI. Our research integrates static and dynamic analysis, and uses a complementary approach to detect harmful content, combining a self-refining LLM-based toxic content detector with rule-based pattern matching. Additionally, we constructed a large-scale toxic word dictionary (i.e., ToxicDict) comprising over 31,783 entries. We used these methods to uncover that 15,414 apps had misleading descriptions, 1,366 collected sensitive personal information against their privacy policies, and 15,996 generated harmful content such as hate speech, self-harm, extremism, etc. Additionally, we evaluated the potential for LLM apps to facilitate malicious activities, finding that 616 apps could be used for malware generation, phishing, etc. We reported these security risks to relevant platforms, including OpenAI and Quora, which acknowledged and appreciated our findings. The platforms are actively investigating the flagged apps; as of the submission of this paper, 1,643 apps have been removed from the GPT Store. Xinyi Hou, Yanjie Zhao 0001, Haoyu Wang 0001 |
SP | 1 |
| 2025 | Large Language Model Supply Chain: A Research AgendaabstractThe rapid advancement of large language models (LLMs) has revolutionized artificial intelligence, introducing unprecedented capabilities in natural language processing and multimodal content generation. However, the increasing complexity and scale of these models have given rise to a multifaceted supply chain that presents unique challenges across infrastructure, foundation models, and downstream applications. This article provides the first comprehensive research agenda of the LLM supply chain, offering a structured approach to identify critical challenges and opportunities through the dual lenses of software engineering (SE) and security and privacy (S&P). We begin by establishing a clear definition of the LLM supply chain, encompassing its components and dependencies. We then analyze each layer of the supply chain, presenting a vision for robust and secure LLM development, reviewing the current state of practices and technologies, and identifying key challenges and research opportunities. This work aims to bridge the existing research gap in systematically understanding the multifaceted issues within the LLM supply chain, offering valuable insights to guide future efforts in this rapidly evolving domain. Shenao Wang 0001, Yanjie Zhao 0001, Xinyi Hou, Haoyu Wang 0001 |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2025 | LLM App Store Analysis: A Vision and RoadmapabstractThe rapid growth and popularity of large language model (LLM) app stores have created new opportunities and challenges for researchers, developers, users, and app store managers. As the LLM app ecosystem continues to evolve, it is crucial to understand the current landscape and identify potential areas for future research and development. This article presents a forward-looking analysis of LLM app stores, focusing on key aspects such as data mining, security risk identification, development assistance, and market dynamics. Our comprehensive examination extends to the intricate relationships between various stakeholders and the technological advancements driving the ecosystem’s growth. We explore the ethical considerations and potential societal impacts of widespread LLM app adoption, highlighting the need for responsible innovation and governance frameworks. By examining these aspects, we aim to provide a vision for future research directions and highlight the importance of collaboration among stakeholders to address the challenges and opportunities within the LLM app ecosystem. The insights and recommendations provided in this article serve as a foundation for driving innovation, ensuring responsible development, and creating a thriving, user-centric LLM app landscape. Yanjie Zhao 0001, Xinyi Hou, Shenao Wang 0001, Haoyu Wang 0001 |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2024 | GPTZoo: A Large-scale Dataset of GPTs for the Research CommunityabstractThe rapid advancements in Large Language Models (LLMs) have revolutionized natural language processing, with GPTs, customized versions of ChatGPT available on the GPT Store, emerging as a prominent technology for specific domains and tasks. To support academic research on GPTs, we introduce GPTZoo, a large-scale dataset comprising 730,420 GPT instances. Each instance includes rich metadata with 21 attributes describing its characteristics, as well as instructions, knowledge files, and third-party services utilized during its development. GPTZoo aims to provide researchers with a comprehensive and readily available resource to study the real-world applications, performance, and potential of GPTs. To facilitate efficient retrieval and analysis of GPTs, we also developed an automated command-line interface (CLI) that supports keyword-based searching of the dataset. To promote open research and innovation, the GPTZoo dataset will undergo continuous updates, and we are granting researchers public access to GPTZoo and its associated tools. Xinyi Hou, Yanjie Zhao 0001, Shenao Wang 0001, Haoyu Wang 0001 |
ASE | 1 |
| 2024 | Models Are Codes: Towards Measuring Malicious Code Poisoning Attacks on Pre-trained Model HubsabstractThe proliferation of pre-trained models (PTMs) and datasets has led to the emergence of centralized model hubs like Hugging Face, which facilitate collaborative development and reuse. However, recent security reports have uncovered vulnerabilities and instances of malicious attacks within these platforms, highlighting growing security concerns. This paper presents the first systematic study of malicious code poisoning attacks on pre-trained model hubs, focusing on the Hugging Face platform. We conduct a comprehensive threat analysis, develop a taxonomy of model formats, and perform root cause analysis of vulnerable formats. While existing tools like Fickling and ModelScan offer some protection, they face limitations in semantic-level analysis and comprehensive threat detection. To address these challenges, we propose MalHug, an end-to-end pipeline tailored for Hugging Face that combines dataset loading script extraction, model deserialization, in-depth taint analysis, and heuristic pattern matching to detect and classify malicious code poisoning attacks in datasets and models. In collaboration with Ant Group, a leading financial technology company, we have implemented and deployed MalHug on a mirrored Hugging Face instance within their infrastructure, where it has been operational for over three months. During this period, MalHug has monitored more than 705K models and 176K datasets, uncovering 91 malicious models and 9 malicious dataset loading scripts. These findings reveal a range of security threats, including reverse shell, browser credential theft, and system reconnaissance. This work not only bridges a critical gap in understanding the security of the PTM supply chain but also provides a practical, industry-tested solution for enhancing the security of pre-trained model hubs. Shenao Wang 0001, Yanjie Zhao 0001, Xinyi Hou, Kailong Wang 0001, Peiming Gao, Haoyu Wang 0001 |
ASE | 4 |
| 2024 | ChatGPT Chats Decoded: Uncovering Prompt Patterns for Superior Solutions in Software Development LifecycleabstractThe advent of Large Language Models (LLMs) like ChatGPT has markedly transformed software development, aiding tasks from code generation to issue resolution with their human-like text generation. Nevertheless, the effectiveness of these models greatly depends on the nature of the prompts given by developers. Therefore, this study delves into the DevGPT dataset, a rich collection of developer-ChatGPT dialogues, to unearth the patterns in prompts that lead to effective problem resolutions. The underlying motivation for this research is to enhance the collaboration between human developers and AI tools, thereby improving productivity and problem-solving efficacy in software development. Utilizing a combination of textual analysis and data-driven approaches, this paper seeks to identify the attributes of prompts that are associated with successful interactions, providing crucial insights for the strategic employment of ChatGPT in software engineering environments. Liangxuan Wu, Yanjie Zhao 0001, Xinyi Hou, Tianming Liu 0002, Haoyu Wang 0001 |
MSR | 3 |
| 2024 | Large Language Models for Software Engineering: A Systematic Literature ReviewabstractLarge Language Models (LLMs) have significantly impacted numerous domains, including Software Engineering (SE). Many recent publications have explored LLMs applied to various SE tasks. Nevertheless, a comprehensive understanding of the application, effects, and possible limitations of LLMs on SE is still in its early stages. To bridge this gap, we conducted a Systematic Literature Review (SLR) on LLM4SE, with a particular focus on understanding how LLMs can be exploited to optimize processes and outcomes. We selected and analyzed 395 research articles from January 2017 to January 2024 to answer four key Research Questions (RQs). In RQ1, we categorize different LLMs that have been employed in SE tasks, characterizing their distinctive features and uses. In RQ2, we analyze the methods used in data collection, pre-processing, and application, highlighting the role of well-curated datasets for successful LLM for SE implementation. RQ3 investigates the strategies employed to optimize and evaluate the performance of LLMs in SE. Finally, RQ4 examines the specific SE tasks where LLMs have shown success to date, illustrating their practical contributions to the field. From the answers to these RQs, we discuss the current state-of-the-art and trends, identifying gaps in existing research, and highlighting promising areas for future study. Our artifacts are publicly available at https://github.com/security-pride/LLM4SE_SLR . Xinyi Hou, Yanjie Zhao 0001, Yue Liu 0011, Zhou Yang 0003, Kailong Wang 0001, Li Li 0029, Xiapu Luo, David Lo 0001, John C. Grundy, Haoyu Wang 0001 |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2023 | Key-aggregate searchable encryption supporting conjunctive queries for flexible data sharing in the cloud
Jinlu Liu, Bo Zhao 0027, Jing Qin 0002, Xinyi Hou, Jixin Ma 0001 |
Inf. Sci. | 4 |
| 2023 | Fully automatic identification of post-treatment infarct lesions after endovascular therapy based on non-contrast computed tomography
Ximing Nie, Xiran Liu, Weibin Gu, Xinyi Hou, Yufei Wei, Qixuan Lu, Haiwei Bai, Jiaping Chen, Tianhang Liu, Hongyi Yan, Miao Wen, Yuesong Pan, Chao Huang 0002, Long Wang 0015 |
Neural Comput. Appl. | 6 |