EDBT 2026 Demo / reviewers in the wild / expert
Wenye Li 0001
dblp:39/5505
· DBLP profile ↗
40ranked-venue papers
23as first author
13since 2021 · last 2025
0000-0002-5679-9670ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 21 first-author · 11 since 2021Databases, data management, data science and information retrieval · 8 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 3 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | UltraTWD: Optimizing Ultrametric Trees for Tree-Wasserstein DistanceabstractThe Wasserstein distance is a widely used metric for measuring differences between distributions, but its super-cubic time complexity introduces substantial computational burdens. To mitigate this, the tree-Wasserstein distance (TWD) offers a linear-time approximation by leveraging a tree structure; however, existing TWD methods often compromise accuracy due to suboptimal tree structures and edge weights. To address it, we introduce UltraTWD, a novel unsupervised framework that simultaneously optimizes both ultrametric tree structures and edge weights to more faithfully approximate the cost matrix. Specifically, we develop algorithms based on minimum spanning trees, iterative projection, and gradient descent to efficiently learn high-quality ultrametric trees. Empirical results across document retrieval, ranking, and classification tasks demonstrate that UltraTWD achieves superior approximation accuracy and competitive downstream performance. Code is available at: https://github.com/NeXAIS/UltraTWD. Fangchen Yu, Yanzhen Chen, Jiaxing Wei, Jianfeng Mao, Wenye Li 0001, Qiang Sun 0007 |
ICML | 5 |
| 2025 | A Theory-Driven Approach to Inner Product Matrix Estimation for Incomplete Data: An Eigenvalue PerspectiveabstractAddressing the critical challenge of data incompleteness in inner product matrix estimation, we introduce a novel eigenvalue correction method designed to precisely reconstruct true inner product matrices from incomplete data. Utilizing random matrix theory, our method adjusts the eigenvalue distribution of the estimated inner product matrix to align with the ground truth. This approach significantly reduces estimation errors for both inner product matrices and the associated Euclidean distance matrices, thereby enhancing the effectiveness of similarity searches on incomplete data. Our method surpasses traditional data imputation and similarity calibration techniques in both maximum inner product search and nearest neighbor search tasks, demonstrating marked advancements in managing incomplete data. Fangchen Yu, Yicheng Zeng, Jianfeng Mao, Wenye Li 0001 |
WWW | 4 |
| 2024 | DocReal: Robust Document Dewarping of Real-Life Images via Attention-Enhanced Control Point PredictionabstractDocument image dewarping is a crucial task in computer vision with numerous practical applications. The control point method, as a popular image dewarping approach, has attracted attention due to its simplicity and efficiency. However, inaccurate control point prediction due to varying background noises and deformation types can result in unsatisfactory performance. To address these issues, we propose a robust document dewarping approach for real-life images, namely DocReal, which utilizes Enet to effectively remove background noise and an attention-enhanced control point (AECP) module to better capture local deformations. Moreover, we augment the training data by synthesizing 2D images with 3D deformations and additional deformation types. Our proposed method achieves state-of-the-art performance on the DocUNet benchmark and a newly proposed benchmark of 200 Chinese distorted images, exhibiting superior dewarping accuracy, OCR performance, and robustness to various types of image distortion. Fangchen Yu, Yina Xie, Yafei Wen, Guozhi Wang, Shuai Ren 0002, Xiaoxin Chen 0001, Jianfeng Mao, Wenye Li 0001 |
WACV | 9 |
| 2023 | Metric Nearness Made PracticalabstractGiven a square matrix with noisy dissimilarity measures between pairs of data samples, the metric nearness model computes the best approximation of the matrix from a set of valid distance metrics. Despite its wide applications in machine learning and data processing tasks, the model faces non-trivial computational requirements in seeking the solution due to the large number of metric constraints associated with the feasible region. Our work designed a practical approach in two stages to tackle the challenge and improve the model's scalability and applicability. The first stage computes a fast yet high-quality approximate solution from a set of isometrically embeddable metrics, further improved by an effective heuristic. The second stage refines the approximate solution with the Halpern-Lions-Wittmann-Bauschke projection algorithm, which converges quickly to the optimal solution. In empirical evaluations, the proposed approach runs at least an order of magnitude faster than the state-of-the-art solutions, with significantly improved scalability, complete conformity to constraints, less memory consumption, and other desirable features in real applications. Wenye Li 0001, Fangchen Yu, Zichen Ma |
AAAI | 1 |
| 2023 | Highly-Efficient Robinson-Foulds Distance Estimation with Matrix CorrectionabstractPhylogenetic trees are essential in studying evolutionary relationships, and the Robinson-Foulds (RF) distance is a widely used metric to calculate pairwise dissimilarities between phylogenetic trees, with various applications in both the biology and computing communities. However, generating a precise RF distance matrix becomes difficult or even intractable when tree information is partially missing. To address this issue, we introduce a novel distance correction algorithm for estimating the RF distance matrix of incomplete phylogenetic trees. Our method innovatively harnesses the assumption of Euclidean embedding, correcting an approximate distance matrix into a valid distance metric, guaranteed to be closer to the unknown ground-truth. Despite its simplicity, our approach exhibits robust performance, efficiency, and scalability in empirical evaluations, outperforming classical distance correction algorithms and holding potential benefits in downstream applications. Our code is available at https://github.com/CUHKSZ-Yu/EMC. Fangchen Yu, Rui Bao, Jianfeng Mao, Wenye Li 0001 |
ECAI | 4 |
| 2023 | From Incompleteness to Unity: A Framework for Multi-view Clustering with Missing Values
Fangchen Yu, Jianfeng Mao, Wenye Li 0001 |
ICONIP (11) | 5 |
| 2023 | Boosting Spectral Clustering on Incomplete Data via Kernel Correction and Affinity LearningabstractSpectral clustering has gained popularity for clustering non-convex data due to its simplicity and effectiveness. It is essential to construct a similarity graph using a high-quality affinity measure that models the local neighborhood relations among the data samples. However, incomplete data can lead to inaccurate affinity measures, resulting in degraded clustering performance. To address these issues, we propose an imputation-free framework with two novel approaches to improve spectral clustering on incomplete data. Firstly, we introduce a new kernel correction method that enhances the quality of the kernel matrix estimated on incomplete data with a theoretical guarantee, benefiting classical spectral clustering on pre-defined kernels. Secondly, we develop a series of affinity learning methods that equip the self-expressive framework with $\ell_p$-norm to construct an intrinsic affinity matrix with an adaptive extension. Our methods outperform existing data imputation and distance calibration techniques on benchmark datasets, offering a promising solution to spectral clustering on incomplete data in various real-world applications. Fangchen Yu, Jicong Fan 0001, Yicheng Zeng, Jianfeng Mao, Wenye Li 0001 |
NeurIPS | 8 |
| 2023 | Online estimation of similarity matrices with incomplete dataabstractThe similarity matrix measures pairwise similarities between a set of data points and is an essential concept in data processing, routinely used in practical applications. Obtaining a similarity matrix is typically straightforward when data points are completely observed. However, incomplete observations can make it challenging to obtain a high-quality similarity matrix, which becomes even more complex in online data. To address this challenge, we propose matrix correction algorithms that leverage the positive semi-definiteness (PSD) of the similarity matrix to improve similarity estimation in both offline and online scenarios. Our approaches have a solid theoretical guarantee of performance and excellent potential for parallel execution on large-scale data. Empirical evaluations demonstrate their high effectiveness and efficiency with significantly improved results over classical imputation-based methods, benefiting downstream applications with superior performance. Our code is available at \url{https://github.com/CUHKSZ-Yu/OnMC}. Fangchen Yu, Yicheng Zeng, Jianfeng Mao, Wenye Li 0001 |
UAI | 4 |
| 2022 | Asynchronous Personalized Federated Learning with Irregular Clients
Zichen Ma, Yu Lu 0013, Wenye Li 0001, Shuguang Cui |
ACML | 3 |
| 2022 | Calibrating Distance Metrics Under Uncertainty
Wenye Li 0001, Fangchen Yu |
ECML/PKDD (3) | 1 |
| 2022 | Beyond Random Selection: A Perspective from Model Inversion in Personalized Federated Learning
Zichen Ma, Yu Lu 0013, Wenye Li 0001, Shuguang Cui |
ECML/PKDD (4) | 3 |
| 2021 | PFedAtt: Attention-based Personalized Federated Learning on Heterogeneous ClientsabstractIn federated learning, heterogeneity among the clients’ local datasets results in large variations in the number of local updates performed by each client in a communication round. Simply aggregating such local models into a global model will confine the capacity of the system, that is, the single global model will be restricted from delivering good performance on each client’s task. This paper provides a general framework to analyze the convergence of personalized federated learning algorithms. It subsumes previously proposed methods and provides a principled understanding of the computational guarantees. Using insights from this analysis, we propose PFedAtt, a personalized federated learning method that incorporates attention-based grouping to facilitate similar clients’ collaborations. Theoretically, we provide the convergence guarantee for the algorithm, and empirical experiments corroborate the competitive performance of PFedAtt on heterogeneous clients. Zichen Ma, Yu Lu 0013, Wenye Li 0001, Jinfeng Yi, Shuguang Cui |
ACML | 3 |
| 2021 | Learning Sparse Binary Code for Maximum Inner Product SearchabstractMaximum inner product search (MIPS), combined with the hashing method, has become a standard solution to similarity search problems. It often achieves an order of magnitude speedup over nearest neighbor search (NNS) under similar settings. Motivated by the work and achievements along this line, in this paper, we developed a sparse binary hashing method for MIPS to preserve the pairwise similarities with the support of two asymmetric hash functions. We proposed a simple and efficient algorithm that learns two hash functions for the query database and the search database respectively. We conducted experiments to evaluate the proposed method, relying on image retrieval tasks on four benchmark datasets. The empirical results clearly demonstrated the algorithm's promising potential on practical applications in terms of search accuracy and scalability. Changyi Ma, Fangchen Yu, Yueyao Yu, Wenye Li 0001 |
CIKM | 4 |
| 2020 | Scalable Calibration of Affinity Matrices from Incomplete ObservationsabstractEstimating pairwise affinity matrices for given data samples is a basic problem in data processing applications. Accurately determining the affinity becomes impossible when the samples are not fully observed and approximate estimations have to be sought. In this paper, we investigated calibration approaches to improve the quality of an approximate affinity matrix. By projecting the matrix onto a closed and convex subset of matrices that meets specific constraints, the calibrated result is guaranteed to get nearer to the unknown true affinity matrix than the un-calibrated matrix, except in rare cases they are identical. To realize the calibration, we developed two simple, efficient, and yet effective algorithms that scale well. One algorithm applies cyclic updates and the other algorithm applies parallel updates. In a series of evaluations, the empirical results justified the theoretical benefits of the proposed algorithms, and demonstrated their high potential in practical applications. Wenye Li 0001 |
ACML | 1 |
| 2020 | Sparse Lifting of Dense Vectors: A Unified Approach to Word and Sentence Representations
Senyue Hao, Wenye Li 0001 |
ICONIP (4) | 2 |
| 2020 | Modeling Winner-Take-All Competition in Sparse Binary Projections
Wenye Li 0001 |
ECML/PKDD (1) | 1 |
| 2020 | Large-scale Image Retrieval with Sparse Binary ProjectionsabstractInspired by the recent discoveries in neuroscience, the study of the sparse binary projection model started to attract people's attention, shedding new light on image retrieval. Different from the classical work that tries to reduce the dimension of the data for faster retrieval speed, the model projects dense input samples into a higher-dimensional space and outputs sparse binary data representations after winner-take-all competition. Following the work along this line, this paper designed a new algorithm which obtains a high-quality sparse binary projection matrix through unsupervised training. Simple as it is, the algorithm reported significantly improved results over the state-of-the-art methods in both search accuracy and retrieval speed in a series of empirical evaluations on large-scale image retrieval tasks, which exhibited its promising potential in industrial applications. Changyi Ma, Chonglin Gu, Wenye Li 0001, Shuguang Cui |
SIGIR | 3 |
| 2019 | AuxBlocks: Defense Adversarial Examples via Auxiliary BlocksabstractDeep learning models are vulnerable to adversarial examples, which poses an indisputable threat to their applications. However, recent studies have observed that gradient-masking defenses are self-deceiving methods if an attacker can realize this defense. In this paper, we propose a new defense method based on appending information. We introduce the Aux Block model to produce extra outputs as a self-ensemble algorithm and analytically investigate the robustness mechanism of Aux Block. We have empirically studied the efficiency of our method against adversarial examples in two types of white- box attacks, and found that even in the full white-box attack where an adversary can craft malicious examples from defense models, our method has a more robust performance of about 54.6% precision on Cifar10 dataset and 38.7% precision on MiniImagenet dataset. Another advantage of our method is that it is able to maintain the prediction accuracy of the classification model on clean images, and thereby exhibits its high potential in practical applications. Yueyao Yu, Wenye Li 0001 |
IJCNN | 3 |
| 2019 | GreenFlowing: A Green Way of Reducing Electricity Cost for Cloud Data Center Using Heterogeneous ESDsabstractIn this paper, we propose a scheduling scheme called GreenFlowing to reduce the electricity cost for a cloud data center by leveraging heterogeneous ESDs. In our model, the data center can be powered by intermittent green energy like wind and solar, and the electricity with time-varying prices from the power grid. The energy from different sources can also choose to flow into long-term or short-term ESDs for later use. Note that, the former can sustain energy for a long time but with low charging/discharging rate, while the latter can charge/discharge very fast but with high energy leakage that can sustain energy only for a few hours. By combining them together, the energy cost can further be reduced. However, it is hard to decide when and how much energy from different sources should be used to power the data center directly or charged into different types of ESDs. We formulate our scheduling into a large-scale linear programming (LP) problem, which can be solved using CPLEX. Numerical experiments show that our scheduling can significantly reduce the total electricity cost for a cloud data center. Chonglin Gu, Wenye Li 0001, Shuguang Cui |
ISCC | 2 |
| 2018 | Learning Word Vectors with Linear Constraints: A Matrix Factorization ApproachabstractLearning vector space representation of words, or word embedding, has attracted much recent research attention. With the objective of better capturing the semantic and syntactic information inherent in words, we propose two new embedding models based on the singular value decomposition of lexical co-occurrences of words. Different from previous work, our proposed models allow for injecting linear constraints when performing the decomposition, with which the desired semantic and syntactic information will be maintained in word vectors. Conceptually the models are flexible and convenient to encode prior knowledge about words. Computationally they can be easily solved by direct matrix factorization. Surprisingly simple yet effective, the proposed models have reported significantly improved performance in empirical word analogy and sentence classification evaluations, and demonstrated high potentials in practical applications. Wenye Li 0001, Laizhong Cui |
IJCAI | 1 |
| 2018 | Fast Similarity Search via Optimal Sparse LiftingabstractSimilarity search is a fundamental problem in computing science with various applications and has attracted significant research attention, especially in large-scale search with high dimensions. Motivated by the evidence in biological science, our work develops a novel approach for similarity search. Fundamentally different from existing methods that typically reduce the dimension of the data to lessen the computational complexity and speed up the search, our approach projects the data into an even higher-dimensional space while ensuring the sparsity of the data in the output space, with the objective of further improving precision and speed. Specifically, our approach has two key steps. Firstly, it computes the optimal sparse lifting for given input samples and increases the dimension of the data while approximately preserving their pairwise similarity. Secondly, it seeks the optimal lifting operator that best maps input samples to the optimal sparse lifting. Computationally, both steps are modeled as optimization problems that can be efficiently and effectively solved by the Frank-Wolfe algorithm. Simple as it is, our approach has reported significantly improved results in empirical evaluations, and exhibited its high potentials in solving practical problems. Wenye Li 0001, Jingwei Mao, Shuguang Cui |
NeurIPS | 1 |
| 2015 | Estimating Jaccard Index with Missing Observations: A Matrix Calibration ApproachabstractThe Jaccard index is a standard statistics for comparing the pairwise similarity between data samples. This paper investigates the problem of estimating a Jaccard index matrix when there are missing observations in data samples. Starting from a Jaccard index matrix approximated from the incomplete data, our method calibrates the matrix to meet the requirement of positive semi-definiteness and other constraints, through a simple alternating projection algorithm. Compared with conventional approaches that estimate the similarity matrix based on the imputed data, our method has a strong advantage in that the calibrated matrix is guaranteed to be closer to the unknown ground truth in the Frobenius norm than the un-calibrated matrix (except in special cases they are identical). We carried out a series of empirical experiments and the results confirmed our theoretical justification. The evaluation also reported significantly improved results in real learning tasks on benchmarked datasets. Wenye Li 0001 |
NIPS | 1 |
| 2015 | Highlighting data clusters by graph embedding
Wenye Li 0001 |
Neurocomputing | 1 |
| 2015 | Visualizing network communities with a semi-definite programming method
Wenye Li 0001 |
Inf. Sci. | 1 |
| 2014 | A Unified Framework for Privacy Preserving Data Clustering
Wenye Li 0001 |
ICONIP (1) | 1 |
| 2014 | Privacy Preserving Clustering: A k-Means Type Extension
Wenye Li 0001 |
ICONIP (2) | 1 |
| 2014 | Gaussian Process Learning: A Divide-and-Conquer Approach
Wenye Li 0001 |
ISNN | 1 |
| 2014 | Early-Stopping Regularized Least-Squares Classification
Wenye Li 0001 |
ISNN | 1 |
| 2013 | Modularity Embedding
Wenye Li 0001 |
ICONIP (2) | 1 |
| 2013 | Modularity Segmentation
Wenye Li 0001 |
ICONIP (2) | 1 |
| 2013 | Revealing network communities with a nonlinear programming method
Wenye Li 0001 |
Inf. Sci. | 1 |
| 2012 | r-Anonymized Clustering
Wenye Li 0001 |
ICONIP (1) | 1 |
| 2012 | Clustering with Uncertainties: An Affinity Propagation-Based Approach
Wenye Li 0001 |
ICONIP (5) | 1 |
| 2011 | Modular Community Detection in NetworksabstractNetwork community detection—the problem of dividing a network of interest into clusters for intelligent analysis—has recently attracted significant attention in diverse fields of research. To discover intrinsic community structure a quantitative measure called modularity has been widely adopted as an optimization objective. Unfortunately, modularity is inherently NP-hard to optimize and approximate solutions must be sought if tractability is to be ensured. In practice, a spectral relaxation method is most often adopted, after which a community partition is recovered from relaxed fractional values by a rounding process. In this paper, we propose an iterative rounding strategy for identifying the partition decisions that is coupled with a fast constrained power method that sequentially achieves tighter spectral relaxations. Extensive evaluation with this coupled relaxation-rounding method demonstrates consistent and sometimes dramatic improvements in the modularity of the communities discovered. 1 Wenye Li 0001, Dale Schuurmans |
IJCAI | 1 |
| 2009 | Fast normalized cut with linear constraintsabstractNormalized Cut is a widely used technique for solving a variety of problems. Although finding the optimal normalized cut has proven to be NP-hard, spectral relaxations can be applied and the problem of minimizing the normalized cut can be approximately solved using eigen-computations. However, it is a challenge to incorporate prior information in this approach. In this paper, we express prior knowledge by linear constraints on the solution, with the goal of minimizing the normalized cut criterion with respect to these constraints. We develop a fast and effective algorithm that is guaranteed to converge. Convincing results are achieved on image segmentation tasks, where the prior knowledge is given as the grouping information of features. Linli Xu 0002, Wenye Li 0001, Dale Schuurmans |
CVPR | 2 |
| 2008 | Lower integrals and upper integrals with respect to nonadditive set functions
Zhenyuan Wang, Wenye Li 0001, Kin-Hong Lee, Kwong-Sak Leung |
Fuzzy Sets Syst. | 2 |
| 2007 | Large-scale RLSC learning without agonyabstractThe advances in kernel-based learning necessitate the study on solving a large-scale non-sparse positive definite linear system. To provide a deterministic approach, recent researches focus on designing fast matrixvector multiplication techniques coupled with a conjugate gradient method. Instead of using the conjugate gradient method, our paper proposes to use a domain decomposition approach in solving such a linear system. Its convergence property and speed can be understood within von Neumann’s alternating projection framework. We will report significant and consistent improvements in convergence speed over the conjugate gradient method when the approach is applied to recent machine learning problems. 1. Wenye Li 0001, Kin-Hong Lee, Kwong-Sak Leung |
ICML | 1 |
| 2007 | Generalizing the Bias Term of Support Vector Machines
Wenye Li 0001, Kwong-Sak Leung, Kin-Hong Lee |
IJCAI | 1 |
| 2006 | Clustering with a Semantic Criterion Based on Dimensionality Analysis
Wenye Li 0001, Kin-Hong Lee, Kwong-Sak Leung |
ICONIP (2) | 1 |
| 2006 | Generalized Regularized Least-Squares Learning with Predefined Features in a Hilbert SpaceabstractKernel-based regularized learning seeks a model in a hypothesis space by minimizing the empirical error and the model's complexity. Based on the representer theorem, the solution consists of a linear combination of translates of a kernel. This paper investigates a generalized form of representer theorem for kernel-based learning. After mapping predefined features and translates of a kernel simultaneously onto a hypothesis space by a specific way of constructing kernels, we proposed a new algorithm by utilizing a generalized regularizer which leaves part of the space unregularized. Using a squared-loss function in calculating the empirical error, a simple convex solution is obtained which combines predefined features with translates of the kernel. Empirical evaluations have confirmed the effectiveness of the algorithm for supervised learning tasks. Wenye Li 0001, Kin-Hong Lee, Kwong-Sak Leung |
NIPS | 1 |