VLDB 2026 Research / reviewers in the wild / expert
Xiaopeng Li 0002
dblp:45/1827-2
· DBLP profile ↗
9ranked-venue papers
3as first author
4since 2021 · last 2024
0000-0003-4916-1131ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 3 since 2021Computer networks · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Language models and text generation · 59% Efficient and distributed learning · 35% Representation and self-supervised learning · 5% | |
| Software engineering, system software, and programming languages
3 papers |
Program synthesis and code generation · 78% Debugging and program repair · 22% | |
| Databases, data mining, and information retrieval
1 paper |
Recommender systems · 100% |
Topics — the 15 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Program synthesis and code generation
code generation with language models |
1.4 | 2 | 2024 | LeDex: Training LLMs to Better Self-Debug and Explain Code · NeurIPS 2024 Towards Greener Yet Powerful Code Generation via Quantization: An Empirical Study · ESEC/SIGSOFT FSE 2023 |
Debugging and program repair
automated program repair |
0.8 | 1 | 2024 | LeDex: Training LLMs to Better Self-Debug and Explain Code · NeurIPS 2024 |
Natural language and speech › Language models and text generation › pre-trained language model
causal language model |
0.7 | 1 | 2023 | ContraCLM: Contrastive Learning For Causal Language Model · ACL (1) 2023 |
Natural language and speech › Language models and text generation
code generation |
0.7 | 1 | 2023 | Multi-lingual Evaluation of Code Generation Models · ICLR 2023 |
Machine learning › Efficient and distributed learning
model compression |
0.7 | 1 | 2023 | Towards Greener Yet Powerful Code Generation via Quantization: An Empirical Study · ESEC/SIGSOFT FSE 2023 |
Natural language and speech › Language models and text generation › evaluation of language models
multilingual evaluation |
0.7 | 1 | 2023 | Multi-lingual Evaluation of Code Generation Models · ICLR 2023 |
Machine learning › Efficient and distributed learning › model compression
quantization |
0.7 | 1 | 2023 | Towards Greener Yet Powerful Code Generation via Quantization: An Empirical Study · ESEC/SIGSOFT FSE 2023 |
Program synthesis and code generation
code generation evaluation |
0.7 | 1 | 2023 | Multi-lingual Evaluation of Code Generation Models · ICLR 2023 |
Program synthesis and code generation › code generation with language models
multilingual code generation |
0.7 | 1 | 2023 | Multi-lingual Evaluation of Code Generation Models · ICLR 2023 |
Recommender systems
collaborative filtering |
0.3 | 1 | 2017 | Collaborative Variational Autoencoder for Recommender Systems · KDD 2017 |
Recommender systems
content-based recommendation |
0.3 | 1 | 2017 | Collaborative Variational Autoencoder for Recommender Systems · KDD 2017 |
Recommender systems › collaborative filtering
hybrid recommendation |
0.3 | 1 | 2017 | Collaborative Variational Autoencoder for Recommender Systems · KDD 2017 |
Recommender systems › generative recommendation
variational autoencoder-based recommendation |
0.3 | 1 | 2017 | Collaborative Variational Autoencoder for Recommender Systems · KDD 2017 |
Natural language and speech › Language models and text generation
large language model |
0.2 | 1 | 2024 | LeDex: Training LLMs to Better Self-Debug and Explain Code · NeurIPS 2024 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.2 | 1 | 2023 | ContraCLM: Contrastive Learning For Causal Language Model · ACL (1) 2023 |
Methods — techniques the papers use, named apart from their topics
supervised fine-tuning · 1.5reinforcement learning · 1.5execution verification · 1.5quantization · 1.3large language model · 1.3empirical study · 1.3contrastive learning · 0.7variational autoencoder · 0.3collaborative filtering · 0.3bayesian generative model · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | LeDex: Training LLMs to Better Self-Debug and Explain CodeabstractIn the domain of code generation, self-debugging is crucial. It allows LLMs to refine their generated code based on execution feedback. This is particularly important because generating correct solutions in one attempt proves challenging for complex tasks. Prior works on self-debugging mostly focus on prompting methods by providing LLMs with few-shot examples, which work poorly on small open-sourced LLMs. In this work, we propose LeDex, a training framework that significantly improves the self-debugging capability of LLMs. Intuitively, we observe that a chain of explanations on the wrong code followed by code refinement helps LLMs better analyze the wrong code and do refinement. We thus propose an automated pipeline to collect a high-quality dataset for code explanation and refinement by generating a number of explanations and refinement trajectories from the LLM itself or a larger teacher model and filtering via execution verification. We perform supervised fine-tuning (SFT) and further reinforcement learning (RL) on both success and failure trajectories with a novel reward design considering code explanation and refinement quality. SFT improves the pass@1 by up to 15.92\% and pass@10 by 9.30\% over four benchmarks. RL training brings additional up to 3.54\% improvement on pass@1 and 2.55\% improvement on pass@10. The trained LLMs show iterative refinement ability and can keep refining code continuously. Lastly, our human evaluation shows that the LLMs trained with our framework generate more useful code explanations and help developers better understand bugs in source code. Nan Jiang 0012, Xiaopeng Li 0002, Shiqi Wang 0002, Qiang Zhou 0009, Soneya Binta Hossain, Baishakhi Ray, Xiaofei Ma 0001, Anoop Deoras |
NeurIPS | 2 |
| 2023 | ContraCLM: Contrastive Learning For Causal Language ModelabstractNihal Jain, Dejiao Zhang, Wasi Uddin Ahmad, Zijian Wang, Feng Nan, Xiaopeng Li, Ming Tan, Ramesh Nallapati, Baishakhi Ray, Parminder Bhatia, Xiaofei Ma, Bing Xiang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Nihal Jain, Dejiao Zhang, Wasi Uddin Ahmad, Zijian Wang 0002, Feng Nan, Xiaopeng Li 0002, Ramesh Nallapati, Baishakhi Ray, Parminder Bhatia, Xiaofei Ma 0001, Bing Xiang |
ACL (1) | 6 |
| 2023 | Multi-lingual Evaluation of Code Generation Models
Ben Athiwaratkun, Sanjay Krishna Gouda, Zijian Wang 0002, Xiaopeng Li 0002, Wasi Uddin Ahmad, Shiqi Wang 0002, Qing Sun 0013, Mingyue Shang, Sujan K. Gonugondla, Hantian Ding, Nathan Fulton, Arash Farahani, Siddhartha Jain 0001, Robert Giaquinto, Haifeng Qian, Murali Krishna Ramanathan, Ramesh Nallapati |
ICLR | 4 |
| 2023 | Towards Greener Yet Powerful Code Generation via Quantization: An Empirical StudyabstractML-powered code generation aims to assist developers to write code in a more productive manner by intelligently generating code blocks based on natural language prompts. Recently, large pretrained deep learning models have pushed the boundary of code generation and achieved impressive performance. However, the huge number of model parameters poses a significant challenge to their adoption in a typical software development environment, where a developer might use a standard laptop or mid-size server to develop code. Such large models cost significant resources in terms of memory, latency, dollars, as well as carbon footprint. Xiaokai Wei, Sujan K. Gonugondla, Shiqi Wang 0002, Wasi Uddin Ahmad, Baishakhi Ray, Haifeng Qian, Xiaopeng Li 0002, Zijian Wang 0002, Qing Sun 0013, Ben Athiwaratkun, Mingyue Shang, Murali Krishna Ramanathan, Parminder Bhatia, Bing Xiang |
ESEC/SIGSOFT FSE | 7 |
| 2018 | Visual Background Recommendation for Dance Performances Using Deep Matrix FactorizationabstractThe stage background is one of the most important features for a dance performance, as it helps to create the scene and atmosphere. In conventional dance performances, the background images are usually selected or designed by professional stage designers according to the theme and the style of the dance. In new media dance performances, the stage effects are usually generated by media editing software. Selecting or producing a dance background is quite challenging and is generally carried out by skilled technicians. The goal of the research reported in this article is to ease this process. Instead of searching for background images from the sea of available resources, dancers are recommended images that they are more likely to use. This work proposes the idea of a novel system to recommend images based on content-based social computing. The core part of the system is a probabilistic prediction model to predict a dancer’s interests in candidate images through social platforms. Different from traditional collaborative filtering or content-based models, the model proposed here effectively combines a dancer’s social behaviors (rating action, click action, etc.) with the visual content of images shared by the dancer using deep matrix factorization (DMF). With the help of such a system, dancers can select from the recommended images and set them as the backgrounds of their dance performances through a media editor. According to the experiment results, the proposed DMF model outperforms the previous methods, and when the dataset is very sparse, the proposed DMF model shows more significant results. Jiqing Wen, James She, Xiaopeng Li 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2017 | Collaborative Variational Autoencoder for Recommender SystemsabstractModern recommender systems usually employ collaborative filtering with rating information to recommend items to users due to its successful performance. However, because of the drawbacks of collaborative-based methods such as sparsity, cold start, etc., more attention has been drawn to hybrid methods that consider both the rating and content information. Most of the previous works in this area cannot learn a good representation from content for recommendation task or consider only text modality of the content, thus their methods are very limited in current multimedia scenario. This paper proposes a Bayesian generative model called collaborative variational autoencoder (CVAE) that considers both rating and content for recommendation in multimedia scenario. The model learns deep latent representations from content data in an unsupervised manner and also learns implicit relationships between items and users from both content and rating. Unlike previous works with denoising criteria, the proposed CVAE learns a latent distribution for content in latent space instead of observation space through an inference network and can be easily extended to other multimedia modalities other than text. Experiments show that CVAE is able to significantly outperform the state-of-the-art recommendation methods with more robust performance. Xiaopeng Li 0002, James She |
KDD | 1 |
| 2017 | An Efficient Computation Framework for Connection Discovery using Shared ImagesabstractWith the advent and popularity of the social network, social graphs become essential to improve services and information relevance to users for many social media applications to predict follower/followee relationship, community membership, and so on. However, the social graphs could be hidden by users due to privacy concerns or kept by social media. Recently, connections discovered from user-shared images using machine-generated labels are proved to be more accessible alternatives to social graphs. But real-time discovery is difficult due to high complexity, and many applications are not possible. This article proposes an efficient computation framework for connection discovery using user-shared images, which is suitable for any image processing and computer vision techniques for connection discovery on the fly. The framework includes the architecture of online computation to facilitate real-time processing, offline computation for a complete processing, and online/offline communication. The proposed framework is implemented to demonstrate its effectiveness by speeding up connection discovery through user-shared images. By studying 300K+ user-shared images from two popular social networks, it is proven that the proposed computation framework reduces 90% of runtime with a comparable accurate with existing frameworks. Ming Cheung 0001, Xiaopeng Li 0002, James She |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2017 | A Distributed Streaming Framework for Connection Discovery Using Shared VideosabstractWith the advances in mobile devices and the popularity of social networks, users can share multimedia content anytime, anywhere. One of the most important types of emerging content is video, which is commonly shared on platforms such as Instagram and Facebook. User connections, which indicate whether two users are follower/followee or have the same interests, are essential to improve services and information relevant to users for many social media applications. But they are normally hidden due to users’ privacy concerns or are kept confidential by social media sites. Using user-shared content is an alternative way to discover user connections. This article proposes to use user-shared videos for connection discovery with the Bag of Feature Tagging method and proposes a distributed streaming computation framework to facilitate the analytics. Exploiting the uniqueness of shared videos, the proposed framework is divided into Streaming processing and Online and Offline Computation. With experiments using a dataset from Twitter, it has been proved that the proposed method using user-shared videos for connection discovery is feasible. And the proposed computation framework significantly accelerates the analytics, reducing the processing time to only 32% for follower/followee recommendation. It has also been proved that comparable performance can be achieved with only partial data for each video and leads to more efficient computation. Xiaopeng Li 0002, Ming Cheung 0001, James She |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2016 | Connection discovery using shared images by Gaussian relational topic modelabstractSocial graphs, representing online friendships among users, are one of the fundamental types of data for many applications, such as recommendation, virality prediction and marketing in social media. However, this data may be unavailable due to the privacy concerns of users, or kept private by social network operators, which makes such applications difficult. Inferring users' interests and discovering users' connections through their shared multimedia content has attracted more and more attention in recent years. This paper proposes a Gaussian relational topic model for connection discovery using user shared images in social media. The proposed model not only models users' interests as latent variables through their shared images, but also considers the connections between users as a result of their shared images. It explicitly relates user shared images to user connections in a hierarchical, systematic and supervisory way and provides an end-to-end solution for the problem. This paper also derives efficient variational inference and learning algorithms for the posterior of the latent variables and model parameters. It is demonstrated through experiments with over 200k images from Flickr that the proposed method significantly outperforms the methods in previous works. Xiaopeng Li 0002, Ming Cheung 0001, James She |
IEEE BigData | 1 |