VLDB 2026 Research / reviewers in the wild / expert
Liyao Li
dblp:133/3737
· DBLP profile ↗
15ranked-venue papers
6as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 7 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Security and privacy · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Invariant Latent Space Perspective on Language Model InversionabstractLanguage model inversion (LMI), i.e., recovering hidden prompts from outputs, emerges as a concrete threat to user privacy and system security. We recast LMI as reusing the LLM's own latent space and propose the Invariant Latent Space Hypothesis (ILSH): (1) diverse outputs from the same source prompt should preserve consistent semantics (source invariance), and (2) inputoutput cyclic mappings should be self-consistent within a shared latent space (cyclic invariance). Accordingly, we present Inv2A, which treats the LLM as an invariant decoder and learns only a lightweight inverse encoder that maps outputs to a denoised pseudo-representation. When multiple outputs are available, they are sparsely concatenated at the representation layer to increase information density. Training proceeds in two stages: contrastive alignment (source invariance) and supervised reinforcement (cyclic invariance). An optional training-free neighborhood search can refine local performance. Across 9 datasets covering user and system prompt scenarios, Inv2A outperforms baselines by an average of 4.77% BLEU score while reducing dependence on large inverse corpora. Our analysis further shows that prevalent defenses provide limited protection, underscoring the need for stronger strategies. Wentao Ye, Haobo Wang 0001, Xinpeng Ti, Zhiqing Xiao, Hao Chen 0081, Liyao Li, Lei Feng 0006, Sai Wu, Junbo Zhao 0002 |
AAAI | 7 |
| 2026 | Reinforcement Learning with Verbalized Probabilities for LLM ClassificationabstractWhile Large Language Models (LLMs) excel at many reasoning tasks, their native inability to produce calibrated, multi-class probability distributions limits their use in high-stakes Web applications like content moderation and fraud detection. Existing methods to elicit probabilities from LLMs either sacrifice their crucial Chain-of-Thought (CoT) reasoning capabilities or suffer from poor calibration. To address this, we introduce a new paradigm, Verbalized Probability Distribution, and a novel training framework, RLVP (Reinforcement Learning with Verbalized Probabilities). RLVP fine-tunes an LLM to generate both an interpretable CoT and a complete, verbalized probability distribution. We overcome the ''insufficient reward granularity'' problem in standard Reinforcement Learning (RL) for classification by using soft probabilities from expert tabular models as a dense reward curriculum. Through large-scale joint training on 169 tabular tasks, we demonstrate that a single RLVP-trained model can surpass a strong, task-specific XGBoost baseline on up to 55% of tasks. More importantly, the trained model achieves state-of-the-art few-shot performance on unseen, heterogeneous Web benchmarks that mix structured data with free text, achieving performance comparable to or superior than expert models trained on the same limited data. This showcases a strong capability for generalization and knowledge transfer to complex Web data. Our work presents a viable path toward building general-purpose, probabilistically-sound, and interpretable foundation models for the Web. Liyao Li, Hao Chen 0081, Jiaming Tian, Wentao Ye, Lirong Gao, Chao Ye 0002, Ningtao Wang, Yu Cheng 0005, Haobo Wang 0001, Gang Chen 0001, Junbo Zhao 0002 |
WWW | 1 |
| 2026 | Toward real-world Table Agents: capabilities, workflows, and design principles for LLM-based table intelligence
Jiaming Tian, Liyao Li, Wentao Ye, Haobo Wang 0001, Lingxin Wang, Lihua Yu, Zujie Ren, Gang Chen 0001, Junbo Zhao 0002 |
World Wide Web (WWW) | 2 |
| 2025 | From Signal-based to Impedance-based Sensing: A paradigm Shift for Plug-and-Play, Mobile, and Sensitive Battery-free SensingabstractBattery-free sensing has revolutionized IoT applications, but current solutions relying on signal variations between transmitted and backscattered signals remain vulnerable to environmental dynamics and deployment variations. This paper promotes a paradigm shift: inferring targets through antenna impedance variations instead of signal fluctuations, thereby eliminating the impact of unpredictable wireless communication. We demonstrate the effectiveness of this paradigm by reimplementing three existing applications: RIO [1], Keystub [2], and RF-EATS [3]. Compared to original signal-based implementations, our approach shows significant improvements in accuracy and robustness across diverse environments. Furthermore, By integrating antenna engineering with advanced materials science, we also transform antennas into innovative sensors for pressure, temperature, and UV light sensing. This interdisciplinary methodology pushes the boundaries of battery-free sensing, opening new avenues for IoT applications. Liyao Li, Bozhao Shang, Jie Xiong 0001, Wenyao Xu, Xiaojiang Chen, Yaxiong Xie |
MobiCom | 1 |
| 2025 | Table as a Modality for Large Language ModelsabstractTo migrate the remarkable successes of Large Language Models (LLMs), the community has made numerous efforts to generalize them to the table reasoning tasks for the widely deployed tabular data. Despite that, in this work, by showing a probing experiment on our proposed StructQA benchmark, we postulate that even the most advanced LLMs (such as GPTs) may still fall short of coping with tabular data. More specifically, the current scheme often simply relies on serializing the tabular data, together with the meta information, then inputting them through the LLMs. We argue that the loss of structural information is the root of this shortcoming. In this work, we further propose TAMO, which bears an ideology to treat the tables as an independent modality integrated with the text tokens. The resulting model in TAMO is a multimodal framework consisting of a hypergraph neural network as the global table encoder seamlessly integrated with the mainstream LLM. Empirical results on various benchmarking datasets, including HiTab, WikiTQ, WikiSQL, FeTaQA, and StructQA, have demonstrated significant improvements on generalization with an average relative gain of **42.65%**. Liyao Li, Chao Ye 0002, Wentao Ye, Haobo Wang 0001, Jiaming Tian, Yiming Zhang 0023, Ningtao Wang, Gang Chen 0001, Junbo Zhao 0002 |
NeurIPS | 1 |
| 2024 | Cyclops: A Nanomaterial-based, Battery-Free Intraocular Pressure (IOP) Monitoring System inside Contact Lens
Liyao Li, Bozhao Shang, Jie Xiong 0001, Xiaojiang Chen, Yaxiong Xie |
NSDI | 1 |
| 2024 | Towards Cross-Table Masked Pretraining for Web Data MiningabstractTabular data pervades the landscape of the World Wide Web, playing a foundational role in the digital architecture that underpins online information. Given the recent influence of large-scale pretrained models like ChatGPT and SAM across various domains, exploring the application of pretraining techniques for mining tabular data on the web has emerged as a highly promising research direction. Indeed, there have been some recent works around this topic where most (if not all) of them are limited in the scope of a fixed-schema/single table. Due to the scale of the dataset and the parameter size of the prior models, we believe that we have not reached the ''BERT moment'' for the ubiquitous tabular data. The development on this line significantly lags behind the counterpart research domains such as natural language processing. In this work, we first identify the crucial challenges behind tabular data pretraining, particularly overcoming the cross-table hurdle. As a pioneering endeavor, this work mainly (i)-contributes a high-quality real-world tabular dataset, (ii)-proposes an innovative, generic, and efficient cross-table pretraining framework, dubbed as CM2, where the core to it comprises a semantic-aware tabular neural network that uniformly encodes heterogeneous tables without much restriction and (iii)-introduces a novel pretraining objective --- prompt Masked Table Modeling (pMTM) --- inspired by NLP but intricately tailored to scalable pretraining on tables. Our extensive experiments demonstrate CM2's state-of-the-art performance and validate that cross-table pretraining can enhance various downstream tasks. Chao Ye 0002, Guoshan Lu, Haobo Wang 0001, Liyao Li, Sai Wu, Gang Chen 0001, Junbo Zhao 0002 |
WWW | 4 |
| 2024 | DBUP: Dynamic blockchain UTXO processing for storage efficiency optimization
Liyao Li, Qifeng Nian, Lailong Luo, Deke Guo |
Comput. Networks | 3 |
| 2024 | A Reputation Awareness Randomization Consensus Mechanism in Blockchain SystemsabstractBlockchain, as an emerging technology, has gained widespread research in academia and industry due to its decentralization and traceability. As an important form of blockchain, consortium chains are often applied in the Internet of Things (IoT) to ensure the authenticity and reliability of data. Within consortium chains, the practical Byzantine fault tolerance (PBFT) method is a key technology for ensuring the data consistency. It plays a central role in enhancing the system performance, security, and scalability. However, with the increase in the number of user nodes and the diversification of application scenarios, PBFT faces significant challenges in maintaining performance and security, particularly due to the increased communication overhead, longer consensus latency (CL), and risks of malicious attacks on the leader node. To overcome these challenges, this article proposes a new blockchain consensus mechanism, namely the reputation awareness randomization consensus mechanism in the blockchain systems (RARCs). This mechanism first builds an evaluation model for the nodes, dividing them into ordinary nodes and candidate nodes through the reputation assessment. Second, it constructs a consensus node selection strategy to select the high-quality consensus nodes from the candidate nodes. Finally, RARC establishes a leader node randomization selection mechanism, increasing the unpredictability of the leader node and reducing the probability of the malicious attacks. Through the theoretical analysis and simulation experiments, we demonstrate that the RARC can significantly reduce the CL, enhance the throughput, and increase the unpredictability of the leader node, thereby improving the performance and security of the blockchain systems. Yongtao Sun, Deke Guo, Lailong Luo, Liyao Li, Qifeng Nian, Shi Zhu, Fangliao Yang |
IEEE Internet Things J. | 5 |
| 2023 | Learning a Data-Driven Policy Network for Pre-Training Automated Feature Engineering
Liyao Li, Haobo Wang 0001, Liangyu Zha, Qingyi Huang, Sai Wu, Gang Chen 0001, Junbo Zhao 0002 |
ICLR | 1 |
| 2022 | SmartLens: sensing eye activities using zero-power contact lensabstractAs the most important organs of sense, human eyes perceive 80% information from our surroundings. Eyeball movement is closely related to our brain health condition. Eyeball movement and eye blink are also widely used as an efficient human-computer interaction scheme for paralyzed individuals to communicate with others. Traditional methods mainly use intrusive EOG sensors or cameras to capture eye activity information. In this work, we propose a system named SmartLens to achieve eye activity sensing using zero-power contact lens. To make it happen, we develop dedicated antenna design which can be fitted in an extremely small space and still work efficiently to reach a working distance more than 1 m. To accurately track eye movements in the presence of strong self-interference, we employ another tag to track the user's head movement and cancel it out to support sensing a walking or moving user. Comprehensive experiments demonstrate the effectiveness of the proposed system. At a distance of 1.4 m, the proposed system can achieve an average accuracy of detecting the basic eye movement and blink at 89.63% and 82%, respectively. Liyao Li, Yaxiong Xie, Jie Xiong 0001, Ziyu Hou, Yingchun Zhang, Qing We, Dingyi Fang, Xiaojiang Chen |
MobiCom | 1 |
| 2022 | Multi-level reversible data hiding for crypto-imagery via a block-wise substitution-transposition cipher
Xu Wang 0027, Liyao Li, Ching-Chun Chang, Yongfeng Huang 0001 |
J. Inf. Secur. Appl. | 2 |
| 2020 | An efficient general data hiding scheme based on image interpolation
Yong-qing Chen, Wei-jiao Sun, Liyao Li, Chin-Chen Chang 0001, Xu Wang 0027 |
J. Inf. Secur. Appl. | 3 |
| 2019 | Tagtag: material sensing with commodity RFIDabstractMaterial sensing is an essential ingredient for many IoT applications. While hyperspectral camera, infrared, X-Ray, and Radar provide potential solutions for material identification, high cost is the major concern limiting their applications. In this paper, we explore the capability of employing RF signals for fine-grained material sensing with commodity RFID device. The key reason for our system to work is that the tag antenna's impedance is changed when it is close or attached to a target. The amount of impedance change is dependent on the target's material type, thus enabling us to utilize the impedance-related phase change available at commodity RFID devices for material sensing. Several key challenges are addressed before we turn the idea into a functional system: (i) the random tag-reader distance causes an additional unknown phase change on top of the phase change caused by the target material; (ii) the tag rotations cause phase shifts and (iii) for conductive liquid, there exists liquid reflection which interferes with the impedance-caused phase change. We address these challenges with novel solutions. Comprehensive experiments show high identification accuracies even for very similar materials such as Pepsi and Coke. Binbin Xie, Jie Xiong 0001, Xiaojiang Chen, Eugene Chai, Liyao Li, Zhanyong Tang, Dingyi Fang |
SenSys | 5 |
| 2013 | Architecture design for management as a service cloud
Yuding Wang, Liyao Li |
IM | 5 |