Liming Dong 0002

dblp:204/5556-2 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Scaling Collaborative Filtering with Multimodal Contrastive Fine-tuning
abstract
Scaling laws have enabled large language models(LLMs) to achieve remarkable performance and strong generalization across diverse language understanding tasks, including few-shot, in-context, and zero-shot learning. While prior studies in large-scale collaborative filtering(CF) have revealed clear relationships between model performance and scaling factors such as data size and model capacity, little attention has been given to how heterogeneous datasets can be synergistically combined for recommender systems(RS). In particular, it remains unclear whether systematically integrating diverse recommendation datasets can yield scaling behaviors analogous to those observed in LLMs, while simultaneously addressing challenges such as cold-start recommendation and cross-domain transfer. In this paper, we present RecCLIP, a multimodal framework that reformulates user--item interactions as visual representations compatible with vision--language models(VLMs). RecCLIP compresses interaction signals and employs prompt-based ranking to enable unified representation across heterogeneous data sources. Extensive experiments reveal consistent power-law scaling trends with respect to data size, and demonstrate that RecCLIP achieves superior performance in both cold-start and cross-domain transfer scenarios. Our findings underscore the importance of data-centric design in recommender systems and provide practical insights into scaling them effectively.The code for replication is available at https://github.com/jinliwei-1/RecCLIP.
Dan Luo 0004, Lixin Zou, Chenliang Li 0005, Xiangyang Luo 0001, Xixun Lin, Liming Dong 0002
WWW7
2026 Learning Discrete Identifiers and Dense Vectors for Generative Retrieval
abstract
Generative retrieval presents a promising approach to information retrieval, streamlining both indexing and retrieval processes through end-to-end optimization. This method typically involves assigning a unique identifier to each document, with the retrieval goal being the generation of the correct document identifier in response to a query. Although generative retrieval has demonstrated empirical success in various tasks, designing an effective document identifier remains a challenge. Previous studies have either depended excessively on one-to-one discrete identifiers, leading to increased retrieval latency and loss of semantics in documents or have used retrieval-agnostic dense document identifiers, which can hinder performance. To this end, we propose to integrate the benefits of generative retrieval and dense retrieval using an encoder-decoder-based pre-trained language model. Particularly, the decoder, i.e., the discrete identifier, functions as a coarse retriever, effectively reducing the retrieval space in an end-to-end manner. As a complement, the encoder, i.e., the dense vector, serves as a fine-grained retriever, efficiently and precisely ranking documents in a condensed space. Accordingly, we introduce a three-stage end-to-end learning framework that optimizes identifiers and vectors. Extensive experiments reveal that the proposed method exceeds the current models in terms of effectiveness and time efficiency, across both small and larger corpus sets.
Yunfan Xie, Lixin Zou, Xiangyang Luo 0001, Hengyi Cai, Chaoran Zhang 0001, Liming Dong 0002, Xixun Lin, Chenliang Li 0005
ACM Trans. Inf. Syst.6
2026 Erratum: Learning Discrete Identifiers and Dense Vectors for Generative Retrieval
abstract
This is an erratum for the article “Learning Discrete Identifiers and Dense Vectors for Generative Retrieval” published in ACM Trans. Inf. Syst. 44, 2, Article 42 (December 2025), 24 pages.
Yunfan Xie, Lixin Zou, Xiangyang Luo 0001, Hengyi Cai, Chaoran Zhang 0001, Liming Dong 0002, Xixun Lin, Chenliang Li 0005
ACM Trans. Inf. Syst.6
2025 Mitigating Language Confusion through Inference-time Intervention
abstract
Although large language models (LLMs) trained on extensive multilingual corpora exhibit impressive language transfer, they often fail to respond in the user’s desired language due to corpus imbalances, an embarrassingly simple problem known as the language confusion. However, existing solutions like in-context learning and supervised fine-tuning (SFT) have drawbacks: in-context learning consumes context window space, diminishing attention as text lengthens, while SFT requires extensive, labor-intensive data collection. To overcome these limitations, we propose the language-sensitive intervention (LSI), a novel, lightweight, and label-free approach. Specifically, we analyze language confusion from a causal perspective, revealing that the training corpus’s language distribution acts as a confounder, disadvantaging languages that are underrepresented in the dataset. Then, we identify a language-sensitive dimension in the LLM’s residual stream, i.e., the language vector, which allows us to estimate the average causal effect of prompts on this dimension. During inference, we directly intervene on the language vector to generate responses in the desired language.To further advance research on this issue, we introduce a new benchmark that detects language confusion and assesses content quality. Experimental results demonstrate that our method effectively mitigates language confusion without additional complex mechanisms. Our code is available at https://github.com/SoseloX/LSI.
Yunfan Xie, Lixin Zou, Dan Luo 0004, Chenliang Li 0005, Liming Dong 0002, Xiangyang Luo 0001
COLING6
2025 Flow Matching Based Sequential Recommender Model
abstract
Generative models, particularly diffusion model, have emerged as powerful tools for sequential recommendation. However, accurately modeling user preferences remains challenging due to the noise perturbations inherent in the forward and reverse processes of diffusion-based methods. Towards this end, this study introduces FMRec, a Flow Matching based model that employs a straight flow trajectory and a modified loss tailored for the recommendation task. Additionally, from the diffusion-model perspective, we integrate a reconstruction loss to improve robustness against noise perturbations, thereby retaining user preferences during the forward process. In the reverse process, we employ a deterministic reverse sampler, specifically an ODE-based updating function, to eliminate unnecessary randomness, thereby ensuring that the generated recommendations closely align with user needs. Extensive evaluations on four benchmark datasets reveal that FMRec achieves an average improvement of 6.53% over state-of-the-art methods. The replication code is available at https://github.com/FengLiu-1/FMRec.
Lixin Zou, Xiangyu Zhao 0001, Liming Dong 0002, Dan Luo 0004, Xiangyang Luo 0001, Chenliang Li 0005
IJCAI5
2025 Blockchain and Deep Reinforcement Learning Empowered Data Storage in Internet of Things
abstract
With the advancement of Internet of Things (IoT) technology, various applications have greatly enhanced convenience in daily life, but there are still many security issues. Due to centralized cloud storage and limited computational resources, sensitive IoT data face security risks such as malicious tampering, privacy leakage, and data loss. The integration of blockchain and IoT provides an effective solution for ensuring data security. However, existing blockchain systems suffer from high storage demands, long consensus time, and vulnerability to malicious attacks. To address these issues, we propose a secure storage scheme for IoT data based on blockchain and Deep Reinforcement Learning (DRL). Specifically, our solution can be divided into two parts. The first part focuses on data generation, hierarchical encryption transmission, storage and verification. A capability-differentiated identity authentication method is introduced to ensure the trustworthiness of entities in the system. Besides, the integration of blockchain and the InterPlanetary File System (IPFS) ensures reliable and decentralized data storage. The second part introduces a consensus nodes selection mechanism based on DRL. This mechanism optimally selects consensus nodes according to different conditions, achieving a higher consensus success rate and lower consensus latency compared to traditional methods. Simulation results demonstrate that our scheme maintains a low computational overhead of 14.97ms, and obtains an average response delay for a single storage request of 0.42s, which proves its effectiveness in IoT environments.
Qinghua Gong, Jinnan Zhang, Xueguang Yuan, Liming Dong 0002
IEEE Internet Things J.6
2020 Marviq: Quality-Aware Geospatial Visualization of Range-Selection Queries Using Materialization
abstract
We study the problem of efficient spatial visualization on a large data set stored in a database using SQL queries with ad-hoc range conditions on numerical attributes, for example, a spatial scatterplot of taxi pickup events in New York between 1/1/2015 and 3/10/2015. We present a novel middleware-based technique called Marviq. It divides the selection-attribute domain into intervals, and precomputes and stores a visualization for each interval. These results are called MVS and stored as tables in the database. We can compute an exact visualization for a request by accessing MVS and retrieving additional records from the base table. To further reduce the latter time, we present algorithms for using MVS to compute an approximate visualization that satisfies a user-specified similarity threshold. We show a family of functions with certain properties that can use this technique. We present an improvement by dividing the MVS intervals into smaller intervals and materializing low-resolution visualization for these intervals. We report the results of an extensive evaluation of Marviq, including a user study, and show its high performance in both space and time.
Liming Dong 0002, Qiushi Bai, Taewoo Kim 0001, Taiji Chen, Weidong Liu 0001, Chen Li 0001
SIGMOD Conference1
2017 Replica-Aware Partitioning Design in Parallel Database Systems
Liming Dong 0002, Weidong Liu 0001, Renchuan Li, Weiguo Zhao 0003
Euro-Par1