VLDB 2026 Research / reviewers in the wild / expert
Zhili Feng
dblp:189/7590
· DBLP profile ↗
17ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0001-6573-7933ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 3 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RankCLIP: Ranking-Consistent Language-Image PretrainingabstractSelf-supervised contrastive learning models, such as CLIP, have set new benchmarks for vision-language models in many downstream tasks. However, their dependency on rigid one-to-one mappings overlooks the complex and often multifaceted relationships between and within texts and images. To this end, we introduce RankCLIP, a novel pre-training method that extends beyond the rigid one-to-one matching framework of CLIP and its variants. By extending the traditional pair-wise loss to list-wise, and leveraging both in-modal and cross-modal ranking consistency, RankCLIP improves the alignment process, enabling it to capture the nuanced many-to-many relationships between and within each modality. Through comprehensive experiments, we demonstrate the effectiveness of RankCLIP in various downstream tasks, notably achieving significant gains in zero-shot classifications over state-of-the-art methods, underscoring the importance of this enhanced learning process. Zhuokai Zhao, Zhaorun Chen, Zhili Feng, Zenghui Ding, Yining Sun |
ICCV | 4 |
| 2025 | Adaptive Data Optimization: Dynamic Sample Selection with Scaling LawsabstractThe composition of pretraining data is a key determinant of foundation models' performance, but there is no standard guideline for allocating a limited computational budget across different data sources. Most current approaches either rely on extensive experiments with smaller models or dynamic data adjustments that also require proxy models, both of which significantly increase the workflow complexity and computational overhead. In this paper, we introduce Adaptive Data Optimization (ADO), an algorithm that optimizes data distributions in an online fashion, concurrent with model training. Unlike existing techniques, ADO does not require external knowledge, proxy models, or modifications to the model update. Instead, ADO uses per-domain scaling laws to estimate the learning potential of each domain during training and adjusts the data mixture accordingly, making it more scalable and easier to integrate. Experiments demonstrate that ADO can achieve comparable or better performance than prior methods while maintaining computational efficiency across different computation scales, offering a practical solution for dynamically adjusting data distribution without sacrificing flexibility or increasing costs. Beyond its practical benefits, ADO also provides a new perspective on data collection strategies via scaling laws. Yiding Jiang, Allan Zhou, Zhili Feng, Sadhika Malladi, J. Zico Kolter |
ICLR | 3 |
| 2025 | Unnatural Languages Are Not Bugs but Features for LLMsabstractLarge Language Models (LLMs) have been observed to process non-human-readable text sequences, such as jailbreak prompts, often viewed as a bug for aligned LLMs. In this work, we present a systematic investigation challenging this perception, demonstrating that unnatural languages - strings that appear incomprehensible to humans but maintain semantic meanings for LLMs - contain latent features usable by models. Notably, unnatural languages possess latent features that can be generalized across different models and tasks during inference. Furthermore, models fine-tuned on unnatural versions of instruction datasets perform on-par with those trained on natural language, achieving $49.71$ win rates in Length-controlled AlpacaEval 2.0 in average across various base models. In addition, through comprehensive analysis, we demonstrate that LLMs process unnatural languages by filtering noise and inferring contextual meaning from filtered words. Our code is publicly available at https://github.com/John-AI-Lab/Unnatural_Language. Keyu Duan, Yiran Zhao 0006, Zhili Feng, Jinjie Ni, Tianyu Pang, Qian Liu 0033, Tianle Cai, Longxu Dou, Kenji Kawaguchi, Anirudh Goyal, J. Zico Kolter, Michael Shieh |
ICML | 3 |
| 2025 | Antidistillation SamplingabstractFrontier models that generate extended reasoning traces inadvertently produce token sequences that can facilitate model distillation. Recognizing this vulnerability, model owners may seek sampling strategies that limit the effectiveness of distillation without compromising model performance. *Antidistillation sampling* provides exactly this capability. By strategically modifying a model's next-token probability distribution, antidistillation sampling poisons reasoning traces, rendering them significantly less effective for distillation while preserving the model's utility. Yash Savani, Asher Trockman, Zhili Feng, Yixuan Even Xu, Avi Schwarzschild, Alexander Robey, Marc Finzi, J. Zico Kolter |
NeurIPS | 3 |
| 2024 | Rethinking LLM Memorization through the Lens of Adversarial CompressionabstractLarge language models (LLMs) trained on web-scale datasets raise substantial concerns regarding permissible data usage.
One major question is whether these models "memorize" all their training data or they integrate many data sources in some way more akin to how a human would learn and synthesize information. The answer hinges, to a large degree, on \emph{how we define memorization.} In this work, we propose the Adversarial Compression Ratio (ACR) as a metric for assessing memorization in LLMs. A given string from the training data is considered memorized if it can be elicited by a prompt (much) shorter than the string itself---in other words, if these strings can be ``compressed'' with the model by computing adversarial prompts of fewer tokens. The ACR overcomes the limitations of existing notions of memorization by (i) offering an adversarial view of measuring memorization, especially for monitoring unlearning and compliance; and (ii) allowing for the flexibility to measure memorization for arbitrary strings at a reasonably low compute. Our definition serves as a practical tool for determining when model owners may be violating terms around data usage, providing a potential legal tool and a critical lens through which to address such scenarios. Avi Schwarzschild, Zhili Feng, Pratyush Maini, Zachary C. Lipton, J. Zico Kolter |
NeurIPS | 2 |
| 2022 | Learning-Augmented $k$-means Clustering
Jon Ergun, Zhili Feng, Sandeep Silwal, David P. Woodruff, Samson Zhou |
ICLR | 2 |
| 2022 | Provable Adaptation across Multiway Domains via Representation Learning
Zhili Feng, Shaobo Han, Simon S. Du |
ICLR | 1 |
| 2022 | UIR-Net: Object Detection in Infrared Imaging of Thermomechanical Processes in Automotive ManufacturingabstractThermomechanical processes (TMPs) such as resistance spot welding (RSW) and hot stamping are widely used in automotive manufacturing. Recent advancement in sensing technology has led to an increasing adoption of thermographic cameras to capture the infrared (IR) radiation of a metal part (or component of a part) during its thermomechanical processing or immediately after the process when the part is still hot. Detecting the object(s) of interest from raw IR images is an essential step in analyzing these data. Deep learning (DL) has been a recent success for object detection (OD), but the application of DL-based OD for industrial IR images in manufacturing is largely lagging behind. The major contribution of this work, which is also the distinction from previous OD studies, is the capability of building the OD model with unlabeled IR images, i.e., imaging data without accurate information indicating the object position. The architecture of Unsupervised IR Image Net (UIR-Net) is designed to accommodate the unique characteristics of IR images from TMPs in manufacturing. This study presents a novel method for OD in unlabeled IR images from TMPs. The proposed method, called UIR-Net, consists of two components: label generation and DL model construction. Two case studies from automotive manufacturing, RSW and hot stamping, are reported to demonstrate the feasibility and effectiveness of the proposed method. Note to Practitioners—This article was motivated by the problem of detecting objects such as weld nugget or metal piece in infrared (IR) imaging of thermomechanical processes (TMPs) in automotive manufacturing. The method is applicable to in situ IR images or videos that contain one or more objects to be detected. It only requires that the data are in image form and come from TMPs. Currently, there is no existing deep learning (DL)-based method for generic object detection (OD) in unlabeled IR images from TMPs. The proposed method takes advantages of the recent advancement in DL. This article suggests a systematic approach to build a DL-based OD model, named Unsupervised IR Image Net (UIR-Net), to extract objects from raw IR images collected for TMPs. A step-by-step procedure is given in this article to guide users through label generation, data quality evaluation, and model training to establish the proposed UIR-Net model. Results from resistance spot welding and hot stamping suggest that this approach is feasible and effective. It is one of the few generic OD works designed for manufacturing applications. Simple implementation, feasibility, and effectiveness make this method a suitable candidate for online data analytics and process monitoring in a wide range of manufacturing applications. Shenghan Guo, Dali Wang, Zhili Feng, Weihong Grace Guo |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2021 | Dimensionality Reduction for the Sum-of-Distances MetricabstractWe give a dimensionality reduction procedure to approximate the sum of distances of a given set of $n$ points in $R^d$ to any “shape” that lies in a $k$-dimensional subspace. Here, by “shape” we mean any set of points in $R^d$. Our algorithm takes an input in the form of an $n \times d$ matrix $A$, where each row of $A$ denotes a data point, and outputs a subspace $P$ of dimension $O(k^{3}/\epsilon^6)$ such that the projections of each of the $n$ points onto the subspace $P$ and the distances of each of the points to the subspace $P$ are sufficient to obtain an $\epsilon$-approximation to the sum of distances to any arbitrary shape that lies in a $k$-dimensional subspace of $R^d$. These include important problems such as $k$-median, $k$-subspace approximation, and $(j,l)$ subspace clustering with $j \cdot l \leq k$. Dimensionality reduction reduces the data storage requirement to $(n+d)k^{3}/\epsilon^6$ from nnz$(A)$. Here nnz$(A)$ could potentially be as large as $nd$. Our algorithm runs in time nnz$(A)/\epsilon^2 + (n+d)$poly$(k/\epsilon)$, up to logarithmic factors. For dense matrices, where nnz$(A) \approx nd$, we give a faster algorithm, that runs in time $nd + (n+d)$poly$(k/\epsilon)$ up to logarithmic factors. Our dimensionality reduction algorithm can also be used to obtain poly$(k/\epsilon)$ size coresets for $k$-median and $(k,1)$-subspace approximation problems in polynomial time. Zhili Feng, Praneeth Kacham, David P. Woodruff |
ICML | 1 |
| 2021 | Non-PSD matrix sketching with applications to regression and optimizationabstractA variety of dimensionality reduction techniques have been applied for computations involving large matrices. The underlying matrix is randomly compressed into a smaller one, while approximately retaining many of its original properties. As a result, much of the expensive computation can be performed on the small matrix. The sketching of positive semidefinite (PSD) matrices is well understood, but there are many applications where the related matrices are not PSD, including Hessian matrices in non-convex optimization and covariance matrices in regression applications involving complex numbers. In this paper, we present novel dimensionality reduction methods for non-PSD matrices, as well as their "square-roots", which involve matrices with complex entries. We show how these techniques can be used for multiple downstream tasks. In particular, we show how to use the proposed matrix sketching techniques for both convex and non-convex optimization, lp-regression for every 1<=p Cite this Paper BibTeX @InProceedings{pmlr-v161-feng21a, title = {Non-PSD matrix sketching with applications to regression and optimization}, author = {Feng, Zhili and Roosta, Fred and Woodruff, David P.}, booktitle = {Proceedings of the Thirty-Seventh Conference on Uncertainty in Artificial Intelligence}, pages = {1841--1851}, year = {2021}, editor = {de Campos, Cassio and Maathuis, Marloes H.}, volume = {161}, series = {Proceedings of Machine Learning Research}, month = {27--30 Jul}, publisher = {PMLR}, pdf = {https://proceedings.mlr.press/v161/feng21a/feng21a.pdf}, url = {https://proceedings.mlr.press/v161/feng21a.html}, abstract = {A variety of dimensionality reduction techniques have been applied for computations involving large matrices. The underlying matrix is randomly compressed into a smaller one, while approximately retaining many of its original properties. As a result, much of the expensive computation can be performed on the small matrix. The sketching of positive semidefinite (PSD) matrices is well understood, but there are many applications where the related matrices are not PSD, including Hessian matrices in non-convex optimization and covariance matrices in regression applications involving complex numbers. In this paper, we present novel dimensionality reduction methods for non-PSD matrices, as well as their "square-roots", which involve matrices with complex entries. We show how these techniques can be used for multiple downstream tasks. In particular, we show how to use the proposed matrix sketching techniques for both convex and non-convex optimization, lp-regression for every 1<=p Copy to Clipboard Download Endnote %0 Conference Paper %T Non-PSD matrix sketching with applications to regression and optimization %A Zhili Feng %A Fred Roosta %A David P. Woodruff %B Proceedings of the Thirty-Seventh Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2021 %E Cassio de Campos %E Marloes H. Maathuis %F pmlr-v161-feng21a %I PMLR %P 1841--1851 %U https://proceedings.mlr.press/v161/feng21a.html %V 161 %X A variety of dimensionality reduction techniques have been applied for computations involving large matrices. The underlying matrix is randomly compressed into a smaller one, while approximately retaining many of its original properties. As a result, much of the expensive computation can be performed on the small matrix. The sketching of positive semidefinite (PSD) matrices is well understood, but there are many applications where the related matrices are not PSD, including Hessian matrices in non-convex optimization and covariance matrices in regression applications involving complex numbers. In this paper, we present novel dimensionality reduction methods for non-PSD matrices, as well as their "square-roots", which involve matrices with complex entries. We show how these techniques can be used for multiple downstream tasks. In particular, we show how to use the proposed matrix sketching techniques for both convex and non-convex optimization, lp-regression for every 1<=p Copy to Clipboard Download APA Feng, Z., Roosta, F. & Woodruff, D.P.. (2021). Non-PSD matrix sketching with applications to regression and optimization. Proceedings of the Thirty-Seventh Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 161:1841-1851 Available from https://proceedings.mlr.press/v161/feng21a.html. Copy to Clipboard Download Related Material Download PDF Supplementary PDF This site last compiled Sun, 05 Jul 2026 14:53:07 +0000 Github Account Copyright © The authors and PMLR 2026. MLResearchPress Zhili Feng, Fred (Farbod) Roosta, David P. Woodruff |
UAI | 1 |
| 2020 | DeepWelding: A Deep Learning Enhanced Approach to GTAW Using Multisource Sensing ImagesabstractDeep learning has great potential to reshape manufacturing industries. In this article, we present DeepWelding, a novel framework that applies deep learning techniques to improve gas tungsten arc welding process monitoring and penetration detection using multisource sensing images. The framework is capable of analyzing multiple types of optical sensing images synchronously and consists of three deep learning enhanced consecutive phases: image preprocessing, image selection, and weld penetration classification. Specifically, we adopted generative adversarial networks (pix2pix) for image denoising and classic convolutional neural networks (AlexNet) for image selection. Both pix2pix and AlexNet delivered satisfactory performance. However, five individual neural networks with heterogeneous architectures demonstrated inconsistent generalization capabilities in the classification phase when holding out multisource images generated with specific experimental settings. Therefore, two ensemble methods combining multiple neural networks are designed to improve the model performance on unseen data collected from different experimental settings. We have also found that the quality of model prediction is heavily influenced by the data stream collection environment. We think these findings are beneficial for the broad intelligent welding community. Yunhe Feng, Zongyao Chen, Dali Wang, Jian Chen 0033, Zhili Feng |
IEEE Trans. Ind. Informatics | 5 |
| 2019 | Does Data Augmentation Lead to Positive Margin?abstractData augmentation (DA) is commonly used during model training, as it significantly improves test error and model robustness. DA artificially expands the training set by applying random noise, rotations, crops, or even adversarial perturbations to the input data. Although DA is widely used, its capacity to provably improve robustness is not fully understood. In this work, we analyze the robustness that DA begets by quantifying the margin that DA enforces on empirical risk minimizers. We first focus on linear separators, and then a class of nonlinear models whose labeling is constant within small convex hulls of data points. We present lower bounds on the number of augmented data points required for non-zero margin, and show that commonly used DA techniques may only introduce significant margin after adding exponentially many points to the data set. Shashank Rajput, Zhili Feng, Zachary Charles, Po-Ling Loh, Dimitris S. Papailiopoulos |
ICML | 2 |
| 2019 | Multi-AGVs path planning based on improved ant colony algorithm
Guohong Yi, Zhili Feng, Tiancan Mei, Pushan Li, Wang Jin |
J. Supercomput. | 2 |
| 2018 | Joint Reasoning for Temporal and Causal RelationsabstractUnderstanding temporal and causal relations between events is a fundamental natural language understanding task.Because a cause must occur earlier than its effect, temporal and causal relations are closely related and one relation often dictates the value of the other.However, limited attention has been paid to studying these two relations jointly.This paper presents a joint inference framework for them using constrained conditional models (CCMs).Specifically, we formulate the joint problem as an integer linear programming (ILP) problem, enforcing constraints that are inherent in the nature of time and causality.We show that the joint inference framework results in statistically significant improvement in the extraction of both temporal and causal relations from text. 1 Qiang Ning, Zhili Feng, Hao Wu 0034, Dan Roth 0001 |
ACL (1) | 2 |
| 2018 | Online Learning with Graph-Structured Feedback against Adaptive AdversariesabstractWe derive upper and lower bounds for the policy regret of T-round online learning problems with graph-structured feedback, where the adversary is nonoblivious but assumed to have a bounded memory. We obtain upper bounds of O~(T2/3) and Õ(T3/4) for strongly-observable and weakly-observable graphs, respectively, based on analyzing a variant of the Exp3 algorithm. When the adversary is allowed a bounded memory of size 1, we show that a matching lower bound of Ω̃(T2/3) is achieved in the case of full-information feedback. We also study the particular loss structure of an oblivious adversary with switching costs, and show that in such a setting, non-revealing strongly-observable feedback graphs achieve a lower bound of Ω̃(T2/3), as well. Zhili Feng, Po-Ling Loh |
ISIT | 1 |
| 2018 | CogCompNLP: Your Swiss Army Knife for NLP
Daniel Khashabi, Mark Sammons, Ben Zhou, Tom Redman, Christos Christodoulopoulos 0001, Vivek Srikumar, Nick Rizzolo, Lev-Arie Ratinov, Guanheng Luo, Quang Do, Chen-Tse Tsai, Subhro Roy, Stephen Mayhew 0001, Zhili Feng, John Wieting, Xiaodong Yu 0003, Yangqiu Song, Shashank Gupta 0007, Shyam Upadhyay, Naveen Arivazhagan, Qiang Ning, Shaoshi Ling, Dan Roth 0001 |
LREC | 14 |
| 2017 | A Structured Learning Approach to Temporal Relation ExtractionabstractIdentifying temporal relations between events is an essential step towards natural language understanding.However, the temporal relation between two events in a story depends on, and is often dictated by, relations among other events.Consequently, effectively identifying temporal relations between events is a challenging problem even for human annotators.This paper suggests that it is important to take these dependencies into account while learning to identify these relations and proposes a structured learning approach to address this challenge.As a byproduct, this provides a new perspective on handling missing relations, a known issue that hurts existing methods.As we show, the proposed approach results in significant improvements on the two commonly used data sets for this problem. Qiang Ning, Zhili Feng, Dan Roth 0001 |
EMNLP | 2 |