VLDB 2026 Research / reviewers in the wild / expert
Wu Lin
dblp:70/10338
· DBLP profile ↗
27ranked-venue papers
11as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 9 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Expensive multiobjective immune algorithm using a novel differential evolution in objective space
Wu Lin, Daxin Zhu, Anhui Tan, Ka-Chun Wong, Qiuzhen Lin |
Expert Syst. Appl. | 2 |
| 2026 | Federated surrogate-assisted evolutionary framework for distributed data-driven multiobjective optimization
Qianhui Ding, Wu Lin, Yinglan Feng, Lijia Ma, Ka-Chun Wong, Qiuzhen Lin, Jianqiang Li 0001 |
Inf. Sci. | 2 |
| 2025 | An Enhanced Search Direction-Based Knowledge Transfer for Multiobjective Many-Tasking Evolutionary OptimizationabstractThis paper proposes a multiobjective many-tasking evolutionary algorithm with enhanced search direction-based knowledge transfer (MMaTEA-ESD). Specifically, the search directions for all tasks are first dynamically computed based on the populations obtained during the evolutionary search process. Subsequently, the search direction of the target task is divided into multiple subsearch directions. After that, the most similar subsearch directions from candidate source tasks are identified by measuring their similarity to those of the target task. Finally, the selected subsearch directions are combined to form the enhanced search direction for knowledge transfer, effectively mitigating negative transfer. Numerical results on a widely used multiobjective many-tasking benchmark test suite demonstrate the competitive performance of MMaTEA-ESD over some state-of-the-art algorithms. Wu Lin, Songbai Liu, Qingling Zhu, Qiuzhen Lin |
CEC | 1 |
| 2024 | Can We Remove the Square-Root in Adaptive Gradient Methods? A Second-Order PerspectiveabstractAdaptive gradient optimizers like Adam(W) are the default training algorithms for many deep learning architectures, such as transformers. Their diagonal preconditioner is based on the gradient outer product which is incorporated into the parameter update via a square root. While these methods are often motivated as approximate second-order methods, the square root represents a fundamental difference. In this work, we investigate how the behavior of adaptive methods changes when we remove the root, i.e. strengthen their second-order motivation. Surprisingly, we find that such square-root-free adaptive methods close the generalization gap to SGD on convolutional architectures, while maintaining their root-based counterpart's performance on transformers. The second-order perspective also has practical benefits for the development of non-diagonal adaptive methods through the concept of preconditioner invariance. In contrast to root-based methods like Shampoo, the root-free counterparts do not require numerically unstable matrix root decompositions and inversions, thus work well in half precision. Our findings provide new insights into the development of adaptive methods and raise important questions regarding the currently overlooked role of adaptivity for their success. Wu Lin, Felix Dangel, Runa Eschenhagen, Juhan Bae, Richard E. Turner, Alireza Makhzani |
ICML | 1 |
| 2024 | Structured Inverse-Free Natural Gradient Descent: Memory-Efficient & Numerically-Stable KFACabstractSecond-order methods such as KFAC can be useful for neural net training. However, they are often memory-inefficient since their preconditioning Kronecker factors are dense, and numerically unstable in low precision as they require matrix inversion or decomposition. These limitations render such methods unpopular for modern mixed-precision training. We address them by (i) formulating an inverse-free KFAC update and (ii) imposing structures in the Kronecker factors, resulting in structured inverse-free natural gradient descent (SINGD). On modern neural networks, we show that SINGD is memory-efficient and numerically robust, in contrast to KFAC, and often outperforms AdamW even in half precision. Our work closes a gap between first- and second-order methods in modern low-precision training. Wu Lin, Felix Dangel, Runa Eschenhagen, Kirill Neklyudov, Agustinus Kristiadi, Richard E. Turner, Alireza Makhzani |
ICML | 1 |
| 2024 | Training Data Attribution via Approximate UnrollingabstractMany training data attribution (TDA) methods aim to estimate how a model's behavior would change if one or more data points were removed from the training set. Methods based on implicit differentiation, such as influence functions, can be made computationally efficient, but fail to account for underspecification, the implicit bias of the optimization algorithm, or multi-stage training pipelines. By contrast, methods based on unrolling address these issues but face scalability challenges. In this work, we connect the implicit-differentiation-based and unrolling-based approaches and combine their benefits by introducing Source, an approximate unrolling-based TDA method that is computed using an influence-function-like formula. While being computationally efficient compared to unrolling-based approaches, Source is suitable in cases where implicit-differentiation-based approaches struggle, such as in non-converged models and multi-stage training pipelines. Empirically, Source outperforms existing TDA techniques in counterfactual prediction, especially in settings where implicit-differentiation-based approaches fall short. Juhan Bae, Wu Lin, Jonathan Lorraine, Roger B. Grosse |
NeurIPS | 2 |
| 2024 | Ensemble of Domain Adaptation-Based Knowledge Transfer for Evolutionary MultitaskingabstractRecently, a number of domain adaptation (DA) methods have been proposed for knowledge transfer in evolutionary multitasking (EMT). However, the learned mappings in these methods often have unique biases in representing the connection between source and target tasks. Few studies have paid attention to the complementarity of different mappings in knowledge transfer. To fill this research gap, this article proposes an ensemble method to combine multiple DA methods for knowledge transfer in EMT by considering the efficacy and diversity of these methods. First, a hierarchical clustering method is used to divide the population of each task into multiple clusters. Then, when two parental solutions are selected for knowledge transfer across tasks, the solutions within the same cluster are checked. In particular, if none of these solutions has been transferred before, the efficacy of DA methods is considered first by using roulette wheel selection based on the corresponding performance improvements in the evolutionary optimization process. Otherwise, the diversity of DA methods is emphasized by randomly selecting one of the DA methods for knowledge transfer. The effectiveness of our proposed ensemble method is validated by embedding it into existing state-of-the-art EMT algorithms, and the experimental results show that our algorithm outperforms several recently proposed EMT algorithms on most cases of two multitasking benchmark suites and one practical case. Wu Lin, Qiuzhen Lin, Liang Feng 0001, Kay Chen Tan |
IEEE Trans. Evol. Comput. | 1 |
| 2023 | Simplifying Momentum-based Positive-definite Submanifold Optimization with Applications to Deep LearningabstractRiemannian submanifold optimization with momentum is computationally challenging because, to ensure that the iterates remain on the submanifold, we often need to solve difficult differential equations. Here, we simplify such difficulties for a class of structured symmetric positive-definite matrices with the affine-invariant metric. We do so by proposing a generalized version of the Riemannian normal coordinates that dynamically orthonormalizes the metric and locally converts the problem into an unconstrained problem in the Euclidean space. We use our approach to simplify existing approaches for structured covariances and develop matrix-inverse-free $2^\text{nd}$-order optimizers for deep learning in low precision settings. Wu Lin, Valentin Duruisseaux, Melvin Leok, Frank Nielsen, Mohammad Emtiyaz Khan, Mark Schmidt 0001 |
ICML | 1 |
| 2023 | A Kriging model-based evolutionary algorithm with support vector machine for dynamic multimodal optimization
Xunfeng Wu, Qiuzhen Lin, Wu Lin, Yulong Ye, Qingling Zhu, Victor C. M. Leung |
Eng. Appl. Artif. Intell. | 3 |
| 2022 | From Multitask Gradient Descent to Gradient-Free Evolutionary Multitasking: A Proof of Faster ConvergenceabstractEvolutionary multitasking, which solves multiple optimization tasks simultaneously, has gained increasing research attention in recent years. By utilizing the useful information from related tasks while solving the tasks concurrently, improved performance has been shown in various problems. Despite the success enjoyed by the existing evolutionary multitasking algorithms, still there is a lack of theoretical studies guaranteeing faster convergence compared to the conventional single task case. To analyze the effects of transferred information from related tasks, in this article, we first put forward a novel multitask gradient descent (MTGD) algorithm, which enhances the standard gradient descent updates with a multitask interaction term. The convergence of the resulting MTGD is derived. Furthermore, we present the first proof of faster convergence of MTGD relative to its single task counterpart. Utilizing MTGD, we formulate a gradient-free evolutionary multitasking algorithm called multitask evolution strategies (MTESs). Importantly, the single task evolution strategies (ESs) we utilize are shown to asymptotically approximate gradient descent and, hence, the faster convergence results derived for MTGD extend to the case of MTES as well. Numerical experiments comparing MTES with single task ES on synthetic benchmarks and practical optimization examples serve to substantiate our theoretical claim. Lu Bai 0005, Wu Lin, Abhishek Gupta 0001, Yew-Soon Ong |
IEEE Trans. Cybern. | 2 |
| 2021 | Tractable structured natural-gradient descent using local parameterizationsabstractNatural-gradient descent (NGD) on structured parameter spaces (e.g., low-rank covariances) is computationally challenging due to difficult Fisher-matrix computations. We address this issue by using \emph{local-parameter coordinates} to obtain a flexible and efficient NGD method that works well for a wide-variety of structured parameterizations. We show four applications where our method (1) generalizes the exponential natural evolutionary strategy, (2) recovers existing Newton-like algorithms, (3) yields new structured second-order algorithms, and (4) gives new algorithms to learn covariances of Gaussian and Wishart-based distributions. We show results on a range of problems from deep learning, variational inference, and evolution strategies. Our work opens a new direction for scalable structured geometric methods. Wu Lin, Frank Nielsen, Mohammad Emtiyaz Khan, Mark Schmidt 0001 |
ICML | 1 |
| 2021 | Multimodal Multiobjective Evolutionary Optimization With Dual Clustering in Decision and Objective SpacesabstractThis article suggests a multimodal multiobjective evolutionary algorithm with dual clustering in decision and objective spaces. One clustering is run in decision space to gather nearby solutions, which will classify solutions into multiple local clusters. Nondominated solutions within each local cluster are first selected to maintain local Pareto sets, and the remaining ones with good convergence in objective space are also selected, which will form a temporary population with more than${N}$solutions (${N}$is the population size). After that, a second clustering is run in objective space for this temporary population to get${N}$final clusters with good diversity in objective space. Finally, a pruning process is repeatedly run on the above clusters until each cluster has only one solution, which removes the most crowded solution in decision space from the most crowded cluster in objective space each time. This way, the clustering in decision space can distinguish all Pareto sets and avoid the loss of local Pareto sets, while that in objective space can maintain diversity in objective space. When solving all the benchmark problems from the competition of multimodal multiobjective optimization in the IEEE Congress on Evolutionary Computation 2019, the experiments validate our advantages to maintain diversity in both objective and decision spaces. Qiuzhen Lin, Wu Lin, Zexuan Zhu 0001, Maoguo Gong, Jianqiang Li 0001, Carlos A. Coello Coello |
IEEE Trans. Evol. Comput. | 2 |
| 2020 | A Novel Decomposition-Based Multimodal Multi-objective Evolutionary Algorithm
Wu Lin, Naili Luo |
ICIC (2) | 1 |
| 2020 | Handling the Positive-Definite Constraint in the Bayesian Learning RuleabstractThe Bayesian learning rule is a natural-gradient variational inference method, which not only contains many existing learning algorithms as special cases but also enables the design of new algorithms. Unfortunately, when variational parameters lie in an open constraint set, the rule may not satisfy the constraint and requires line-searches which could slow down the algorithm. In this work, we address this issue for positive-definite constraints by proposing an improved rule that naturally handles the constraints. Our modification is obtained by using Riemannian gradient methods, and is valid when the approximation attains a block-coordinate natural parameterization (e.g., Gaussian distributions and their mixtures). Our method outperforms existing methods without any significant increase in computation. Our work makes it easier to apply the rule in the presence of positive-definite constraints in parameter spaces. Wu Lin, Mark Schmidt 0001, Mohammad Emtiyaz Khan |
ICML | 1 |
| 2019 | Multimodal Multi-objective Optimization Using A Density-based One-by-One Update StrategyabstractFor real-world optimization problems, a uniformly and widely distributed Pareto optimal set (PS) in the decision space can provide more choices for decision makers. However, most of multi-objective evolutionary algorithms (MOEAs) only consider convergence and diversity in the objective space, which rarely pay attention to diversity in the decision space. Especially for multimodal multi-objective optimization problems (MMOPs), there may exist multiple distinct PSs corresponding to the same Pareto front (PF). Thus, we propose a novel multimodal multi-objective evolutionary algorithm using a density-based one-by-one update strategy in this paper, which considers diversity in both the objective and decision spaces. In the proposed algorithm, once an offspring is generated during evolution, the most crowded subregion with the largest niche count in the objective space has to be identified again, helpful to maintain diversity in the objective space. Furthermore, the harmonic average distance approach is used to estimate the global density of solutions in the decision space, trying to maintain the population's diversity in the decision space. Our proposed algorithm is compared with several state-of-the-art algorithms on MMOPs. The experimental results demonstrate that our algorithm is capable of preserving promising solutions with even distribution in both of decision space and objective space and also shows the superiority on solving the adopted MMOPs. Ruizhi Shi, Wu Lin, Qiuzhen Lin, Zexuan Zhu 0001, Jianyong Chen |
CEC | 2 |
| 2019 | Fast and Simple Natural-Gradient Variational Inference with Mixture of Exponential-family ApproximationsabstractNatural-gradient methods enable fast and simple algorithms for variational inference, but due to computational difficulties, their use is mostly limited to minimal exponential-family (EF) approximations. In this paper, we extend their application to estimate structured approximations such as mixtures of EF distributions. Such approximations can fit complex, multimodal posterior distributions and are generally more accurate than unimodal EF approximations. By using a minimal conditional-EF representation of such approximations, we derive simple natural-gradient updates. Our empirical results demonstrate a faster convergence of our natural-gradient method compared to black-box gradient-based methods. Our work expands the scope of natural gradients for Bayesian inference and makes them more widely applicable than before. Wu Lin, Mohammad Emtiyaz Khan, Mark Schmidt 0001 |
ICML | 1 |
| 2018 | Variational Message Passing with Structured Inference Networks
Wu Lin, Nicolas Hubacher, Mohammad Emtiyaz Khan |
ICLR (Poster) | 1 |
| 2018 | Fast and Scalable Bayesian Deep Learning by Weight-Perturbation in AdamabstractUncertainty computation in deep learning is essential to design robust and reliable systems. Variational inference (VI) is a promising approach for such computation, but requires more effort to implement and execute compared to maximum-likelihood methods. In this paper, we propose new natural-gradient algorithms to reduce such efforts for Gaussian mean-field VI. Our algorithms can be implemented within the Adam optimizer by perturbing the network weights during gradient evaluations, and uncertainty estimates can be cheaply obtained by using the vector that adapts the learning rate. This requires lower memory, computation, and implementation effort than existing VI methods, while obtaining uncertainty estimates of comparable quality. Our empirical results confirm this and further suggest that the weight-perturbation in our algorithm could be useful for exploration in reinforcement learning and stochastic optimization. Mohammad Emtiyaz Khan, Didrik Nielsen, Voot Tangkaratt, Wu Lin, Yarin Gal, Akash Srivastava |
ICML | 4 |
| 2017 | Conjugate-Computation Variational Inference: Converting Variational Inference in Non-Conjugate Models to Inferences in Conjugate ModelsabstractVariational inference is computationally challenging in models that contain both conjugate and non-conjugate terms. Methods specifically designed for conjugate models, even though computationally efficient, find it difficult to deal with non-conjugate terms. On the other hand, stochastic-gradient methods can handle the non-conjugate terms but they usually ignore the conjugate structure of the model which might result in slow convergence. In this paper, we propose a new algorithm called Conjugate-computation Variational Inference (CVI) which brings the best of the two worlds together – it uses conjugate computations for the conjugate terms and employs stochastic gradients for the rest. We derive this algorithm by using a stochastic mirror-descent method in the mean-parameter space, and then expressing each gradient step as a variational inference in a conjugate model. We demonstrate our algorithm’s applicability to a large class of models and establish its convergence. Our experimental results show that our method converges much faster than the methods that ignore the conjugate structure of the model. Mohammad Emtiyaz Khan, Wu Lin |
AISTATS | 2 |
| 2016 | SMOS simplified iterative full-pol brightness temperature retrievalabstractSMOS is the acronym for the Soil Moisture and Ocean Salinity mission by the European Space Agency (ESA) [1]. Its single payload, the Microwave Imaging Radiometer using Aperture Synthesis (MIRAS), was successfully launched in November 2009. Along more than six years of operation, SMOS calibration and imaging algorithms have undergone a continuous evolution to further improve the accuracy of the retrieved geophysical parameters for a variety of scientific applications related with soil moisture and ocean salinity [2]. This paper analyzes a simplified iterative image reconstruction method that exclusively takes into account the dominant terms in the full-polarimetric equation [3]. This simplified method, that reduces computation time up to 50% at the cost of small radiometric performance degradation, is useful to the L1 teams to fine tune the calibration of the instrument when processing a large amount of data is required. Israel Durán 0001, Wu Lin, Francesc Torres 0002, Ignasi Corbella, Nuria Duffo, Manuel Martín-Neira |
IGARSS | 2 |
| 2016 | Faster Stochastic Variational Inference using Proximal-Gradient Methods with General Divergence Functions
Mohammad Emtiyaz Khan, Reza Babanezhad 0001, Wu Lin, Mark Schmidt 0001, Masashi Sugiyama |
UAI | 3 |
| 2015 | Mitigation of land-sea contamination in SMOSabstractSince its launch in November 2009, the Soil Moisture and Ocean Salinity (SMOS) mission by the European Space Agency (ESA) has undergone a continuous calibration and imaging procedures improvement to produce valuable geophysical data over land, ice and ocean. However, a problem which does persist is the so-called land-sea contamination (LSC) effect. This paper unveils the nature of this artifact, which is caused by a multiplicative (scene dependent) bias in the retrieved brightness temperature images. The origin of LSC is traced down to SMOS calibration parameters to yield a simple correction scheme which is validated against several geophysical scenarios. This paper shows how autoconsistency rules in interferometric synthesis together with redundant and complementary calibration procedures provide a robust SMOS calibration scheme. Ignasi Corbella, Israel Durán 0001, Wu Lin, Francesc Torres 0002, Nuria Duffo, Ali Khazaal, Manuel Martín-Neira |
IGARSS | 3 |
| 2015 | SMOS floor error impact and migation on ocean imagingabstractSMOS is the acronym for the Soil Moisture and Ocean Salinity mission by the European Space Agency (ESA) [1]. Its single payload, the Microwave Imaging Radiometer using Aperture Synthesis (MIRAS), was successfully launched in November 2009 to provide a continuous flow of fully polarimetric brightness temperature images. After more than five years of operation, SMOS has proven to be highly useful for a variety of scientific applications related to soil moisture over land, ocean salinity and winds over ocean, as well as specific studies over the ice covered surfaces [2][3]. Along these years of operation, SMOS calibration and imaging algorithms have undergone a continuous evolution to further improve the accuracy of the retrieved geophysical parameters. These improvements are particularly useful in Ocean Salinity applications where accuracy and stability requirements are very demanding and spatial and temporal averaging is required to achieve SMOS mission objectives. This paper analyzes the impact on ocean images of the so-called “floor error”, an image reconstruction artifact that, currently, is significant on the 3rdand 4thStokes parameters. Error mitigation techniques are evaluated by computing spatial bias over the ocean and also by providing a 4thStokes global error map over the ocean that takes into account SMOS full spatial and temporal averaging capabilities. Israel Durán 0001, Wu Lin, Ignasi Corbella, Francesc Torres 0002, Nuria Duffo, Manuel Martín-Neira |
IGARSS | 2 |
| 2014 | Residual calibration error impact on SMOS SLL performanceabstractSMOS Side Lobe Level (SLL), which is caused by the limited coverage of the measured visibility samples in the frequency domain, is responsible for non-negligible radiometric errors at pixels placed close to large brightness temperature transitions (e.g. coastline, ice-sea border, RFI, sun reflection, etc.). In order to mitigate this effect the visibility samples are tapered by means of a Blackman window before the image inversion procedure is undertaken. However, the theoretical SLL performance of such a specific taper is degraded by residual visibility calibration errors. Simulations performed in this work show that SLL performance is very sensitive to calibration errors and that the SLL performance of the rectangular window cannot be improved to a large extend, even in the case that calibration errors are constrained by very stringent requirements. Therefore, smarter inversion techniques or data handling procedures must be devised in the future to effectively deal with this issue. Francesc Torres 0002, Ignasi Corbella, Israel Durán 0001, Wu Lin, Nuria Duffo, Manuel Martín-Neira |
IGARSS | 4 |
| 2012 | Spatial decorrelation of radiometric noise in SMOS measurementsabstractThis paper analyzes the degree of correlation of MIRAS/SMOS redundant baselines by analyzing the radiometric noise reduction when N simultaneous redundant visibilities over the Ocean are averaged. This analysis benefits from the high degree of redundancy that is available along each of the three arms in MIRAS “Y” shape array to verify the spatial decorrelation of SMOS visibility samples due to the Earth as an extended source of thermal radiation. Francesc Torres 0002, Wu Lin, Nuria Duffo, Ignasi Corbella, Manuel Martín-Neira |
IGARSS | 2 |
| 2012 | Minimization of Image Distortion in SMOS Brightness Temperature Maps Over the OceanabstractSoil Moisture and Ocean Salinity (SMOS) brightness temperature synthesized images are obtained after a comprehensive error correction procedure that takes into account both on-ground and in-flight calibration measurements. However, the final images are still contaminated by small, although nonnegligible, spatial errors: the so-called pixel bias. Since spatial errors in the 2-D SMOS images are not zero mean along track, these errors produce clearly visible artifacts aligned to this direction. Fortunately, spatial errors have been found to be very stable and can be minimized once the image distortion pattern is properly measured by observing a target at a uniform brightness temperature distribution. This letter describes the procedure to compute a multiplicative mask that largely reduces spatial errors over the ocean. Preliminary results to assess the mask performance are also presented by computing the reduction of the rms spatial error for a number of targets selected to have significant temporal and geographical diversity. Francesc Torres 0002, Ignasi Corbella, Wu Lin, Nuria Duffo, Jérôme Gourrion, Jordi Font, Manuel Martín-Neira |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2011 | Correction of spatial errors in SMOS brightness temperature imagesabstractSystematic spatial errors in SMOS brightness temperature images are successfully estimated by using a statistical analysis of measured data. A constant multiplicative mask is derived by averaging spatial errors of a large number of snapshots over the ocean. The mask has been obtained for the alias- free field of view region without the need of any geophysical model. It can be considered as part of the instrument characterization and is totally independent of the target to measure. When this mask is applied to regular SMOS brightness temperatures, the spatial artifacts are clearly reduced. Wu Lin, Ignasi Corbella, Francesc Torres 0002, Nuria Duffo, Manuel Martín-Neira |
IGARSS | 1 |