VLDB 2026 Research / reviewers in the wild / expert
Congrui Li
dblp:203/0229
· DBLP profile ↗
5ranked-venue papers
2as first author
4since 2021 · last 2025
0009-0001-5138-0336ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Graph learning · 50% Generative modeling · 44% Representation and self-supervised learning · 6% | |
| Theoretical computer science
2 papers |
Mathematical optimization · 100% | |
| Network and information security
1 paper |
Digital forensics and information hiding · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 100% | |
| Computer graphics and multimedia
1 paper |
Image and video coding · 100% |
Topics — the 14 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Graph learning
graph foundation model |
0.9 | 1 | 2025 | OPTFM: A Scalable Multi-View Graph Transformer for Hierarchical Pre-Training in Combinatorial Optimization · NeurIPS 2025 |
Machine learning › Graph learning › graph neural network
graph transformer |
0.9 | 1 | 2025 | OPTFM: A Scalable Multi-View Graph Transformer for Hierarchical Pre-Training in Combinatorial Optimization · NeurIPS 2025 |
Mathematical optimization
combinatorial optimization |
0.9 | 1 | 2025 | OPTFM: A Scalable Multi-View Graph Transformer for Hierarchical Pre-Training in Combinatorial Optimization · NeurIPS 2025 |
Machine learning › Generative modeling › diffusion model › controllable generation
constrained generation |
0.8 | 1 | 2024 | MILP-FBGen: LP/MILP Instance Generation with Feasibility/Boundedness · ICML 2024 |
Machine learning › Generative modeling
diffusion model |
0.8 | 1 | 2024 | MILP-FBGen: LP/MILP Instance Generation with Feasibility/Boundedness · ICML 2024 |
Digital forensics and information hiding
deepfake detection |
0.8 | 1 | 2024 | Pixel Bleach Network for Detecting Face Forgery Under Compression · IEEE Trans. Multim. 2024 |
Digital forensics and information hiding › forgery detection
face forgery detection |
0.8 | 1 | 2024 | Pixel Bleach Network for Detecting Face Forgery Under Compression · IEEE Trans. Multim. 2024 |
Mathematical optimization
instance generation |
0.8 | 1 | 2024 | MILP-FBGen: LP/MILP Instance Generation with Feasibility/Boundedness · ICML 2024 |
Mathematical optimization
linear programming |
0.8 | 1 | 2024 | MILP-FBGen: LP/MILP Instance Generation with Feasibility/Boundedness · ICML 2024 |
Data mining › text mining › topic modeling
short text topic modeling |
0.7 | 1 | 2023 | Topic Modeling of Short Texts: A Pseudo-Document View With Word Embedding Enhancement · IEEE Trans. Knowl. Data Eng. 2023 |
Data mining › text mining
topic modeling |
0.7 | 1 | 2023 | Topic Modeling of Short Texts: A Pseudo-Document View With Word Embedding Enhancement · IEEE Trans. Knowl. Data Eng. 2023 |
Image and video coding › quality assessment
compression artifacts |
0.2 | 1 | 2024 | Pixel Bleach Network for Detecting Face Forgery Under Compression · IEEE Trans. Multim. 2024 |
Image and video coding
image compression |
0.2 | 1 | 2024 | Pixel Bleach Network for Detecting Face Forgery Under Compression · IEEE Trans. Multim. 2024 |
Machine learning › Representation and self-supervised learning › word representation
word embedding |
0.2 | 1 | 2023 | Topic Modeling of Short Texts: A Pseudo-Document View With Word Embedding Enhancement · IEEE Trans. Knowl. Data Eng. 2023 |
Methods — techniques the papers use, named apart from their topics
pre-training · 1.7multi-view graph transformer · 1.7hybrid self-attention · 1.7contrastive learning · 1.7structure-preserving generation · 1.5generative model · 1.5feature representation learning · 1.5feasibility-constrained sampling · 1.5diffusion model · 1.5word embeddings · 1.3self-aggregation · 1.3probabilistic topic model · 1.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | OPTFM: A Scalable Multi-View Graph Transformer for Hierarchical Pre-Training in Combinatorial OptimizationabstractFoundation Models (FMs) have demonstrated remarkable success in fields like computer vision and natural language processing, yet their application to combinatorial optimization remains underexplored. Optimization problems, often modeled as graphs, pose unique challenges due to their diverse structures, varying distributions, and NP-hard complexity. To address these challenges, we propose OPTFM, the first graph foundation model for general combinatorial optimization. OPTFM introduces a scalable multi-view graph transformer with hybrid self-attention and cross-attention to model large-scale heterogeneous graphs in $O(N)$ time complexity while maintaining semantic consistency throughout the attention computation. A Dual-level pre-training framework integrates node-level graph reconstruction and instance-level contrastive learning, enabling robust and adaptable representations at multiple levels. Experimental results across diverse optimization tasks show that models trained on OPTFM embeddings without fine-tuning consistently outperform task-specific approaches, establishing a new benchmark for solving combinatorial optimization problems. Hao Yuan 0002, Wenli Ouyang, Changwen Zhang, Congrui Li |
NeurIPS | 4 |
| 2024 | MILP-FBGen: LP/MILP Instance Generation with Feasibility/BoundednessabstractMachine learning (ML) has been actively adopted in Linear Programming (LP) and Mixed-Integer Linear Programming (MILP), whose potential is hindered by instance scarcity. Current synthetic instance generation methods often fall short in closely mirroring the distribution of original datasets or ensuring the feasibility and boundedness of the generated data — a critical requirement for obtaining reliable supervised labels in model training. In this paper, we present a diffusion-based LP/MILP instance generative framework called MILP-FBGen. It strikes a balance between structural similarity and novelty while maintaining feasibility/boundedness via a meticulously designed structure-preserving generation module and a feasibility/boundedness-constrained sampling module. Our method shows superiority on two fronts: 1) preservation of key properties (hardness, feasibility, and boundedness) of LP/MILP instances, and 2) enhanced performance on downstream tasks. Extensive studies show two-fold superiority that our method ensures higher distributional similarity and 100% feasibility in both easy and hard datasets, surpassing current state-of-the-art techniques. Yahong Zhang, Congrui Li, Wenli Ouyang, Mingda Zhu, Junchi Yan |
ICML | 4 |
| 2024 | Pixel Bleach Network for Detecting Face Forgery Under CompressionabstractThe existing face forgery algorithms have achieved remarkable progress in how to generate reasonable facial images and can even successfully deceive human beings. Considering public security, face forgery detection is of vital importance, making it essential to design face forgery detection algorithms to detect forgery images over the Internet. Despite the great success achieved by the existing Deepfake detection algorithms, they usually failed to achieve satisfactory Deepfake detection performance when deployed to handle the forgery videos in practice. One significant reason is compression. The videos over the Internet are inevitably compressed considering the transmission efficiency. To address this issue, in this paper, we propose a generic, simple yet effective “bleaching” pre-processing module based on the generative model and the high-level feature representations to produce ableached image, which shares a similar appearance with the compressed images. The bleached images with recovered information can be identified accurately by the optimized Deepfake detection models without retraining. The proposed method has utilized a redesigned feature representation, which serves as a navigator to effectively and sufficiently alter the feature distribution in the high-dimensional space to remedy the difference between real facial images and forgery counterparts. Thus, the proposed method can successfully avoid misclassification. Comprehensive and extensive experiments are carried out on four low-quality Faceforensics++ datasets, demonstrating the effectiveness of our method in recovering the information loss caused by the compression artifacts across various backbones and compression. Congrui Li, Ziqiang Zheng, Yi Bin, Guoqing Wang 0001, Yang Yang 0002, Xuesheng Li, Heng Tao Shen |
IEEE Trans. Multim. | 1 |
| 2023 | Topic Modeling of Short Texts: A Pseudo-Document View With Word Embedding EnhancementabstractRecent years have witnessed the unprecedented growth of online social media, resulting in short texts being the prevalent format of information on the Internet. Given the sparsity of data, however, short-text topic modeling remains a critical yet much-watched challenge in both academia and industry. Research has been devoted to building different types of probabilistic topic models for short texts, among which self-aggregation methods emerged recently to provide informative cross-text word co-occurrences. However, models along this line are still in their infancy and typically yield overfit results and exhibit high computational costs. In this paper, we propose a novel model called Pseudo-document-based Topic Model (PTM), which introduces the concept of pseudo-document to implicitly aggregate short texts against data sparsity. By modeling the topic distributions of latent pseudo-documents rather than short texts, PTM yields excellent performance in accuracy and efficiency. A word embedding-enhanced PTM (WE-PTM) is also proposed to leverage pre-trained word embeddings, which is essential to further alleviating data sparsity. Extensive experiments with self-aggregation or word embedding-based baselines on four real-world datasets including two online media short texts, demonstrate the high-quality topics learned by our models. Robustness to limited training samples and the explainable semantics of topics are also investigated. Yuan Zuo, Congrui Li, Hao Lin 0002, Junjie Wu 0002 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2019 | Bi-Intra Prediction for Versatile Video CodingabstractThis paper presents a novel Bi-Intra Prediction (BIP) algorithm to improve intra coding performance for the next generation of video coding. In the proposed algorithm, a new predictor is generated by combining two existing intra prediction modes as an additional mode, being able to provide more accurate prediction. To remove the number of unnecessary combinations, restrictions on the search candidates and block sizes are carried out based on the statistical analyses. In addition, an efficient mode coding method of syntax elements in BIP is introduced to improve the coding performance. Moreover, a rough mode decision scheme is adopted to avoid high computation complexity in encoder side. Experimental results show that compared with the Versatile Video Coding reference software VTM-1.0, the proposed algorithm achieves 0.71%, 0.46% and 0.43% BD-Rate gains on average for Y, U and V components under all intra configuration, respectively, and 0.32%, 0.39%, 0.49% BD-Rate gains under random access configuration. Congrui Li, Zhenghui Zhao, Xiang Zhang 0004, Siwei Ma 0001 |
DCC | 1 |