EDBT 2026 Demo / reviewers in the wild / expert
Enming Luo
dblp:98/7627
· DBLP profile ↗
20ranked-venue papers
7as first author
8since 2021 · last 2025
0000-0002-6887-1094ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 3 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ROI Scan: LLM-powered Object-level Similarity Search for Google Ads Content Moderation
Enming Luo, Yintao Liu 0002, Dongjin Kwon, Rich Munoz, Wei Qiao 0004, Nic Trieu, Eric Xiao, Jimin Li, Laurel Graham, Ariel Fuxman |
CIKM | 1 |
| 2025 | Google Ads Content Moderation with RAGabstractKeeping ad content policy classifiers up to date while maintaining the high quality bar is a significant challenge, especially with new threats emerging constantly. This paper introduces a new application to apply RAG-inspired in-context learning to accelerate content policy enforcement, especially when mitigating new emerging violations. Our application leverages RAG-based LLM inference for classification tasks and incorporates augmented reasoning information for better performance. We also developed a practical framework to enforce new violation patterns in O(1) days demonstrating improved memorization and generalization capabilities compared to traditional parametric and non-parametric models. Yuan Wang 0049, Wei Qiao 0004, Tiantian Fang, Eric Xiao, Megan Oftelie, Yintao Liu 0002, Jimin Li, Zhongli Ding, Enming Luo |
CIKM | 12 |
| 2025 | A Implies B: Circuit Analysis in LLMs for Propositional Logical ReasoningabstractDue to the size and complexity of modern large language models (LLMs), it has proven challenging to uncover the underlying mechanisms that models use to solve reasoning problems. For instance, is their reasoning for a specific problem localized to certain parts of the network? Do they break down the reasoning problem into modular components that are then executed as sequential steps as we go deeper in the model? To better understand the reasoning capability of LLMs, we study a minimal propositional logic problem that requires combining multiple facts to arrive at a solution. By studying this problem on Mistral and Gemma models, up to 27B parameters, we illuminate the core components the models use to solve such logic problems. From a mechanistic interpretability point of view, we use causal mediation analysis to uncover the pathways and components of the LLMs' reasoning processes. Then, we offer fine-grained insights into the functions of attention heads in different layers. We not only find a sparse circuit that computes the answer, but we decompose it into sub-circuits that have four distinct and modular uses. Finally, we reveal that three distinct models -- Mistral-7B, Gemma-2-9B and Gemma-2-27B -- contain analogous but not identical mechanisms. Guanzhe Hong, Nishanth Dikkala, Enming Luo, Cyrus Rashtchian, Xin Wang 0116, Rina Panigrahy |
NeurIPS | 3 |
| 2025 | Zero-Shot Image Moderation in Google Ads with LLM-Assisted Textual Descriptions and Cross-modal Co-embeddings
Enming Luo, Wei Qiao 0004, Katie Warren, Eric Xiao, Krishna Viswanathan, Yuan Wang 0049, Yintao Liu 0002, Jimin Li, Ariel Fuxman |
WSDM | 1 |
| 2024 | Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language ModelsabstractSolving complex visual tasks such as “Who invented the musical instrument on the right?” involves a composition of skills: understanding space, recognizing instruments, and also retrieving prior knowledge. Recent work shows promise by decomposing such tasks using a large language model (LLM) into an executable program that invokes specialized vision models. However, generated programs are error-prone: they omit necessary steps, include spurious ones, and are unable to recover when the specialized models give incor-rect outputs. Moreover, they require loading multiple models, incurring high latency and computation costs. We propose Visual Program Distillation (VPD), an instruction tuning framework that produces a vision-language model (VLM) ca-pable of solving complex visual tasks with a single forward pass. VPD distills the reasoning ability of LLMs by using them to sample multiple candidate programs, which are then executed and verified to identify a correct one. It translates each correct program into a language description of the reasoning steps, which are then distilled into a VLM. Exten-sive experiments show that VPD improves the VLM's ability to count, understand spatial relations, and reason compositionally. Our VPD-trained PaLI-X outperforms all prior VLMs, achieving state-of-the-art performance across complex vision tasks, including MMBench, OK-VQA, A-OKVQA, TallyQA, POPE, and Hateful Memes. An evaluation with human annotators also confirms that VPD improves model response factuality and consistency. Finally, experiments on content moderation demonstrate that VPD is also helpful for adaptation to real-world applications with limited data. Yushi Hu, Otilia Stretcu, Chun-Ta Lu, Krishnamurthy Viswanathan, Kenji Hata, Enming Luo, Ranjay Krishna, Ariel Fuxman |
CVPR | 6 |
| 2024 | Modeling Collaborator: Enabling Subjective Vision Classification with Minimal Human Effort via LLM Tool-UseabstractFrom content moderation to wildlife conservation, the number of applications that require models to recognize nuanced or subjective visual concepts is growing. Traditionally, developing classifiers for such concepts requires substantial manual effort measured in hours, days, or even months to identify and annotate data needed for training. Even with recently proposed Agile Modeling techniques, which enable rapid bootstrapping of image classifiers, users are still required to spend 30 minutes or more of monotonous, repetitive data labeling just to train a single classifier. Drawing on Fiske's Cognitive Miser theory, we propose a new framework that alleviates manual effort by replacing human labeling with natural language interactions, reducing the total effort required to define a concept by an order of magnitude: from labeling 2,000 images to only 100 plus some natural language interactions. Our framework leverages recent advances in foundation models, both large language models and vision-language models, to carve out the concept space through conversation and by automatically labeling training data points. Most importantly, our framework eliminates the need for crowd-sourced annotations. Moreover, our framework ultimately produces lightweight classification models that are deploy-able in cost-sensitive scenarios. Across 15 subjective concepts and across 2 public image classification datasets, our trained models outperform traditional Agile Modeling as well as state-of-the-art zero-shot classification models like ALIGN, CLIP, CuPL, and large visual question answering models like PaLI-X. Imad Eddine Toubal, Aditya Avinash, Neil Gordon Alldrin, Jan Dlabal, Wenlei Zhou, Enming Luo, Otilia Stretcu, Chun-Ta Lu, Howard Zhou, Ranjay Krishna, Ariel Fuxman, Tom Duerig |
CVPR | 6 |
| 2024 | Scaling Up LLM Reviews for Google Ads Content ModerationabstractLarge language models (LLMs) are powerful tools for content moderation, but their inference costs and latency make them prohibitive for casual use on large datasets, such as the Google Ads repository. This study proposes a method for scaling up LLM reviews for content moderation in Google Ads. First, we use heuristics to select candidates via filtering and duplicate removal, and create clusters of ads for which we select one representative ad per cluster. We then use LLMs to review only the representative ads. Finally, we propagate the LLM decisions for the representative ads back to their clusters. This method reduces the number of reviews by more than 3 orders of magnitude while achieving a 2x recall compared to a baseline non-LLM model. The success of this approach is a strong function of the representations used in clustering and label propagation; we found that cross-modal similarity representations yield better results than uni-modal representations. Wei Qiao 0004, Tushar Dogra, Otilia Stretcu, Yu-Han Lyu, Tiantian Fang, Dongjin Kwon, Chun-Ta Lu, Enming Luo, Yuan Wang 0049, Chih-Chun Chia, Ariel Fuxman, Ranjay Krishna, Mehmet Tek |
WSDM | 8 |
| 2023 | Agile Modeling: From Concept to Classifier in MinutesabstractThe application of computer vision methods to nuanced, subjective concepts is growing. While crowdsourcing has served the vision community well for most objective tasks (such as labeling a "zebra"), it now falters on tasks where there is substantial subjectivity in the concept (such as identifying "gourmet tuna"). However, empowering any user to develop a classifier for their concept is technically difficult: users are neither machine learning experts nor have the patience to label thousands of examples. In reaction, we introduce the problem of Agile Modeling: the process of turning any subjective visual concept into a computer vision model through real-time user-in-the-loop interactions. We instantiate an Agile Modeling prototype for image classification and show through a user study (N=14) that users can create classifiers with minimal effort in under 30 minutes. We compare this user driven process with the traditional crowdsourcing paradigm and find that the crowd’s notion often differs from that of the user’s, especially as the concepts become more subjective. Finally, we scale our experiments with simulations of users training classifiers for ImageNet21k categories to further demonstrate the efficacy of the approach. Otilia Stretcu, Edward Vendrow, Kenji Hata, Krishnamurthy Viswanathan, Vittorio Ferrari, Sasan Tavakkol, Wenlei Zhou, Aditya Avinash, Enming Luo, Neil Gordon Alldrin, Mohammad Hossein Bateni 0001, Gabriel Berger, Andrew Bunner, Chun-Ta Lu, Javier A Rey, Giulia DeSalvo, Ranjay Krishna, Ariel Fuxman |
ICCV | 9 |
| 2020 | NoiseRank: Unsupervised Label Noise Reduction with Dependence Models
Karishma Sharma, Pinar Donmez, Enming Luo, Yan Liu 0002, Ismet Zeki Yalniz |
ECCV (27) | 3 |
| 2018 | Patch Matching for Image Denoising Using Neighborhood-Based Collaborative FilteringabstractWe consider patch matching as a recommendation system problem and introduce a new patch-matching approach using nearest neighbor-based collaborative filtering (NN-CF). Our approach involves recommending similar patches to a query patch with the help of other similar patches in a noisy image or an external database. Using user-oriented and item-oriented formulations of NN-CF, we present two variations of CF-based patch-matching criterion. To demonstrate the superior matches found with our method, we apply the new patch-matching scheme to patch-based image denoising and evaluate its effect on the denoising performance. We test the methods on two data sets with varying background and image complexities and under different levels of noise. The proposed method not only improves robustness to patch matching but also provides a new formulation to seamlessly combine internal and external denoising. Shibin Parameswaran, Enming Luo, Truong Q. Nguyen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Targeted video denoising for decompressed videosabstractThe paradigm of using clean patches from a targeted external database to design optimal denoising filters, called Targeted Image Denoising (TID), has been shown to outperform state-of-the-art denoising algorithms such as BM3D. In this paper, we introduce Targeted Video Denoising algorithm that extends the TID algorithm to denoise decompressed video without adding complexity. Our algorithm leverages the motion vectors generated during compression to establish temporal coherency between patches in consecutive frames. We test our algorithm on three decompressed video sequences with different foregrounds, backgrounds and movement patterns, and different noise level settings. Experimental results show that our approach is effective and performs better than original TID and the state-of-the-art video denoising algorithm. Shibin Parameswaran, Enming Luo, Truong Q. Nguyen |
ICIP | 2 |
| 2016 | Adaptive Image Denoising by Mixture AdaptationabstractWe propose an adaptive learning procedure to learn patch-based image priors for image denoising. The new algorithm, called the expectation-maximization (EM) adaptation, takes a generic prior learned from a generic external database and adapts it to the noisy image to generate a specific prior. Different from existing methods that combine internal and external statistics in ad hoc ways, the proposed algorithm is rigorously derived from a Bayesian hyper-prior perspective. There are two contributions of this paper. First, we provide full derivation of the EM adaptation algorithm and demonstrate methods to improve the computational complexity. Second, in the absence of the latent clean image, we show how EM adaptation can be modified based on pre-filtering. The experimental results show that the proposed adaptation algorithm yields consistently better denoising results than the one without adaptation and is superior to several state-of-the-art algorithms. Enming Luo, Stanley H. Chan, Truong Q. Nguyen |
IEEE Trans. Image Process. | 1 |
| 2015 | Adaptive Image Denoising by Targeted DatabasesabstractWe propose a data-dependent denoising procedure to restore noisy images. Different from existing denoising algorithms which search for patches from either the noisy image or a generic database, the new algorithm finds patches from a database that contains relevant patches. We formulate the denoising problem as an optimal filter design problem and make two contributions. First, we determine the basis function of the denoising filter by solving a group sparsity minimization problem. The optimization formulation generalizes existing denoising algorithms and offers systematic analysis of the performance. Improvement methods are proposed to enhance the patch search process. Second, we determine the spectral coefficients of the denoising filter by considering a localized Bayesian prior. The localized prior leverages the similarity of the targeted database, alleviates the intensive Bayesian computation, and links the new method to the classical linear minimum mean squared error estimation. We demonstrate applications of the proposed method in a variety of scenarios, including text images, multiview images, and face images. Experimental results show the superiority of the new algorithm over existing methods. Enming Luo, Stanley H. Chan, Truong Q. Nguyen |
IEEE Trans. Image Process. | 1 |
| 2014 | Image denoising by targeted external databasesabstractClassical image denoising algorithms based on single noisy images and generic image databases will soon reach their performance limits. In this paper, we propose to denoise images using targeted external image databases. Formulating denoising as an optimal filter design problem, we utilize the targeted databases to (1) determine the basis functions of the optimal filter by means of group sparsity; (2) determine the spectral coefficients of the optimal filter by means of localized priors. For a variety of scenarios such as text images, multiview images, and face images, we demonstrate superior denoising results over existing algorithms. Enming Luo, Stanley H. Chan, Truong Q. Nguyen |
ICASSP | 1 |
| 2013 | Adaptive non-local means for multiview image denoising: Searching for the right patches via a statistical approachabstractWe present an adaptive non-local means (NLM) denoising method for a sequence of images captured by a multiview imaging system, where direct extensions of existing single image NLM methods are incapable of producing good results. Our proposed method consists of three major components: (1) a robust joint-view distance metric to measure the similarity of patches; (2) an adaptive procedure derived from statistical properties of the estimates to determine the optimal number of patches to be used; (3) a new NLM algorithm to denoise using only a set of similar patches. Experimental results show that the proposed method is robust to disparity estimation error, out-performs existing algorithms in multiview settings, and performs competitively in video settings. Enming Luo, Stanley H. Chan, Shengjun Pan, Truong Q. Nguyen |
ICIP | 1 |
| 2011 | Compressing similar image sets using low frequency templateabstractIn advance of the imaging capturing technology, large amount of similar images are created. Instead of compressing each similar image individually, removing the inter-image redundancy would reduce the storage and transmission time. However, only a few set redundancy methods are proposed to deal with the problem. In this paper, a new method was derived from a theoretical model by extracting the low frequency in an image set. For the similar images, the values of their low frequency components are very close to that of their neighboring pixel in the spatial domain. In our model, a low frequency template is created and used as a prediction for each image to compute its residue. This model proves the reduction in the entropy and hence the bit rates. Experiments were conducted and proved there were up to 30% gains over the existing methods. Chi Ho Yeung, Oscar C. Au, Ketan Tang, Zhiding Yu, Enming Luo, Yannan Wu, Shing Fat Tu |
ICME | 5 |
| 2009 | A robust spatial-temporal line-warping based deinterlacing methodabstractIn this paper, a line-warping based deinterlacing method will be introduced. The missing pixels in interlaced videos can be derived from the warping of pixels in horizontal line pairs. In order to increase the accuracy of temporal prediction, multiple temporal-line pairs, selected according to constant velocity model, are used for warping. The stationary pixels can be well-preserved by accuracy stationary detection. A soft switching between spatial-temporal interpolated values and temporal average is introduced in order to prevent unstable switching. Owing to above novelties, the proposed method can yield higher visual quality deinterlaced videos than conventional methods. Moreover, this method can suppress most deinterlaced visual artifacts, such as line-crawling, flickering and ghost-shadow. Shing Fat Tu, Oscar C. Au, Yannan Wu, Enming Luo, Chi Ho Yeung |
ICME | 4 |
| 2009 | A novel deringing method based on MAP image restorationabstractLong has it become a hot topic that reconstructing images of better visual quality from one or a serial of degraded ones. Although there are thousands of different restoration methods, in this paper, we focus on removing the ringing artifact caused by lossy video compression. Being a sort of restoration method, we choose the max-a-posterior (MAP) method to model this optimization problem. Quantification of the ringing artifacts serves as a prior information of the images. So it is also analyzed in this paper. The MAP optimization is further solved using a gradient decent solver. Although this is a quite classical method, there are still lots of problems with it. For settling them, we transform the solver to a filter format operation named iterative optimization filter. Experiments show that, such method could give an averaging 0.3 dB gain and in special regions more than 0.7 dB gain. What is more, as analyzed in this paper, the method is very preferable in terms of hardware implementation in several aspects. Yannan Wu, Oscar C. Au, Enming Luo, Dennis Tu, Leo Yeung |
ICME | 3 |
| 2009 | Image Registration Method based on Local High Order ApproachabstractImage registration based on gradient and least square optimization technique is one of the most edge-cutting registration algorithms. Such method, especially useful for sub-pixel motion, searches for the best motion in an iterative way. This paper solves the same motion registration problem following this direction. And the well-known Gauss-Newton method (GNM) is employed here as the optimization tool. To achieve a speed-up and reduction of the arithmetic calculation, a simplified high order approach(SHoA) used to calculate several parameters for GNM is introduced. Detailed complexity analysis and performance comparison are presented showing that such an approach is a better trade-off between registration error and the number of math operations. Yannan Wu, Oscar C. Au, Enming Luo, Chi Ho Yeung, Shing Fat Tu |
ISCAS | 3 |
| 2009 | Motion vector predictor selection for the enhancement layer in the H.264/AVC extension-spatial SVCabstractScalable Video Coding (SVC) has been approved as the extension of the H.264/AVC video coding standard recently [1]. In current spatial scalability scheme, the technique to examine both inter modes with residual prediction and without residual prediction for enhancement layers can achieve the highest possible coding efficiency, but it's typically one of the most time-consuming parts of a video encoder. In this paper, we propose a method to skip half of motion estimation processes under “inter modes without residual prediction” based on the motion vector predictor selection under “inter modes with residual prediction”. Experimental results show that our proposed scheme can achieve up to 24% time saving for motion estimation in the enhancement layer and meanwhile the coding efficiency can be preserved very well. Enming Luo, Oscar C. Au, Yannan Wu, Shing Fat Tu, Chi Ho Yeung |
PCS | 1 |