Mimi Zhang

dblp:37/2847 · DBLP profile ↗
← Back
19ranked-venue papers
5as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021
YearPublicationVenuePosition
2026 Multi-modal fusion for billboard categorization in video frames: A text and image-based approach
abstract
Abstract Traditional multimedia classification techniques rely on analyzing either its features or the associated annotated textual information. In this paper, we introduce a technique that leverages deep learning methodologies to extract both low-level visual features and high-level semantic information from the billboard images. We propose a multi-modal hybrid fusion model that integrates text and image features to categorize billboards in video frames into food, sports, and miscellaneous categories. Achieving a 5% accuracy improvement over image-based models and 10% over text-based models, our model demonstrates robust generalization across diverse datasets, benefiting advertising, media, and content creation industries.
Sukriti Dhang, Jason Lok, Mimi Zhang, Soumyabrata Dev
Multim. Tools Appl.3
2026 Occlusion-aware advertisement placement in soccer penalty area
abstract
Abstract Advert placement in sports broadcasts is a growing strategy to boost sponsor visibility without disrupting live gameplay. Achieving realism, however, requires careful handling of scene geometry and dynamic occlusions from players and the sports ball. In this work, we propose an occlusion-aware and perspective-consistent framework specifically for virtual advert placement in the soccer penalty area. We introduce an automatic procedure to select a geometrically consistent quadrilateral region inside the penalty area from predicted field coordinates, which is then used for homography-based warping. We integrate instance-level occlusion masks with Laplacian Alpha Blending for dynamic occlusion-aware blending so that the virtual advert is correctly placed behind players and the ball. Quantitative evaluations demonstrate that our occlusion-aware advert placement method preserves high visual fidelity with an average SSIM of 0.97 and PSNR of 31.6 dB, while maintaining temporal consistency with flicker index increase of <7%. Furthermore, we analyze occlusion preservation by comparing advert insertion with and without occlusion handling. A decrease in this metric indicates that objects overlapping the advert region become incorrectly hidden after advert insertion. The results highlight the importance of occlusion-aware blending for maintaining scene integrity and visual realism. By effectively managing occlusions, the proposed framework reduces visual artifacts and improves perceptual quality, producing augmented sports footage that is realistic and visually coherent.
Sukriti Dhang, Fucheng Zheng, Peter Han Joo Chong, Mimi Zhang, Soumyabrata Dev
Multim. Tools Appl.4
2025 Shape-Informed Clustering of Multi-Dimensional Functional Data via Deep Functional Autoencoders
abstract
We introduce FAEclust, a novel functional autoencoder framework for cluster analysis of multi-dimensional functional data, data that are random realizations of vector-valued random functions. Our framework features a universal-approximator encoder that captures complex nonlinear interdependencies among component functions, and a universal-approximator decoder capable of accurately reconstructing both Euclidean and manifold-valued functional data. Stability and robustness are enhanced through innovative regularization strategies applied to functional weights and biases. Additionally, we incorporate a clustering loss into the network's training objective, promoting the learning of latent representations that are conducive to effective clustering. A key innovation is our shape-informed clustering objective, ensuring that the clustering results are resistant to phase variations in the functions. We establish the universal approximation property of our non-linear decoder and validate the effectiveness of our model through extensive experiments.
Samuel V. Singh, Shirley Coyle, Mimi Zhang
NeurIPS3
2025 Parallelizing Adaptive Reliability Analysis Through Penalizing the Learning Function
abstract
Structural reliability analysis is essential for evaluating system failure probabilities under uncertainties, yet it often faces computational efficiency challenges. While surrogate model-based techniques, including Kriging, are known for their high accuracy and efficiency, they typically employ a sequential learning strategy, which limits their potential for parallel computation. This article introduces the Local Penalization Adaptive Learning (LP-AL) method, which facilitates parallel adaptive reliability analysis; LP-AL introduces a penalty function that emulates the process of sequential learning strategies, thereby achieving parallelization. The method also integrates a global error-based stopping criterion and a sample pool reduction strategy to enhance efficiency. We tested LP-AL with five commonly used learning functions across various engineering scenarios. The results demonstrate that LP-AL achieves high accuracy and significantly reduces computational costs, making it a viable approach for diverse structural reliability analysis tasks.
Guangchen Wang, Michael Monaghan, Mimi Zhang
IEEE Trans. Reliab.3
2024 Learning Mixtures of Gaussian Processes through Random Projection
abstract
We propose an ensemble clustering framework to uncover latent cluster labels in functional data generated from a Gaussian process mixture. Our method exploits the fact that the projection coefficients of the functional data onto any given projection function follow a univariate Gaussian mixture model (GMM). By conducting multiple one-dimensional projections and learning a univariate GMM for each, we create an ensemble of GMMs. Each GMM serves as a base clustering, and applying ensemble clustering yields a consensus clustering. Our approach significantly reduces computational complexity compared to state-of-the-art methods, and we provide theoretical guarantees on the identifiability and learnability of Gaussian process mixtures. Extensive experiments on synthetic and real datasets confirm the superiority of our method over existing techniques.
Emmanuel Akeweje, Mimi Zhang
ICML2
2024 Use-Net: Satellite Data-Based Framework for Optimizing Billboard Placement in Urban Areas
abstract
In this study, we address the key objective of identifying the optimal region within an extensive geographical area for strategic billboard placement, ensuring effective communication with diverse audiences. The challenge lies in finding suitable spaces for billboard placement within a confined geographical area. One potential solution involves extracting urban areas from satellite images to enhance the efficiency of identifying appropriate billboard spaces. To address this, we propose a novel framework named USE-NET, leveraging satellite data to comprehensively cover large geographic regions for the precise identification of optimal billboard locations. Our proposed framework integrates an enhanced UNet model with the Squeeze-and-Excitation (SE) attention method to optimize the identification process based on satellite data. Comparative evaluations against alternative deep learning models, utilizing various metrics, underscore the efficacy of our approach. Experimental results demonstrate a testing accuracy of 87.21%, showcasing a notable improvement of at least 4% compared to Link-Net, MaNet, and UNet models. To enhance the reproducibility of this research, the code of this paper is made available at: https://github.com/sukritidhang/USE-Net_optimalbillboardlocation
Sukriti Dhang, Prasanjit Dey, Mimi Zhang, Soumyabrata Dev
IGARSS3
2024 A Theoretical Analysis of Density Peaks Clustering and the Component-Wise Peak-Finding Algorithm
abstract
Density peaks clustering detects modes as points with high density and large distance to points of higher density. Each non-mode point is assigned to the same cluster as its nearest neighbor of higher density. Density peaks clustering has proved capable in applications, yet little work has been done to understand its theoretical properties or the characteristics of the clusterings it produces. Here, we prove that it consistently estimates the modes of the underlying density and correctly clusters the data with high probability. However, noise in the density estimates can lead to erroneous modes and incoherent cluster assignments. A novel clustering algorithm, Component-wise Peak-Finding (CPF), is proposed to remedy these issues. The improvements are twofold: 1) the assignment methodology is improved by applying the density peaks methodology within level sets of the estimated density; 2) the algorithm is not affected by spurious maxima of the density and hence is competent at automatically deciding the correct number of clusters. We present novel theoretical results, proving the consistency of CPF, as well as extensive experimental results demonstrating its exceptional performance. Finally, a semi-supervised version of CPF is presented, integrating clustering constraints to achieve excellent performance for an important problem in computer vision.
Joshua Tobin, Mimi Zhang
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Reinforced EM Algorithm for Clustering with Gaussian Mixture Models
abstract
Methods that employ the EM algorithm for parameter estimation typically face a notorious yet unsolved problem that the initialization input significantly impacts the algorithm output. We here develop a Reinforced Expectation Maximization (REM) algorithm for cluster analysis using Gaussian mixture models. The competence of REM is achieved by introducing two innovative strategies into the EM framework: (1) a mode-finding strategy for initialization that detects non-trivial modes in the data, and (2) a mode-pruning strategy for detecting true modes/mixture components of the population. The pruning strategy is well-justified in the context of mixture modelling, and we present theoretical guarantees on the quality of the initialization. Extensive experimental studies on both synthetic and real datasets show that our approach achieves better performance compared to state-of-the-art methods.
Joshua Tobin, Chin Pang Ho, Mimi Zhang
SDM3
2023 Incorporating ignorance within game theory: An imprecise probability approach
Bernard Fares, Mimi Zhang
Int. J. Approx. Reason.2
2023 Review of Clustering Methods for Functional Data
abstract
Functional data clustering is to identify heterogeneous morphological patterns in the continuous functions underlying the discrete measurements/observations. Application of functional data clustering has appeared in many publications across various fields of sciences, including but not limited to biology, (bio)chemistry, engineering, environmental science, medical science, psychology, social science, and so on. The phenomenal growth of the application of functional data clustering indicates the urgent need for a systematic approach to develop efficient clustering methods and scalable algorithmic implementations. On the other hand, there is abundant literature on the cluster analysis of time series, trajectory data, spatio-temporal data, and so on, which are all related to functional data. Therefore, an overarching structure of existing functional data clustering methods will enable the cross-pollination of ideas across various research fields. We here conduct a comprehensive review of original clustering methods for functional data. We propose a systematic taxonomy that explores the connections and differences among the existing functional data clustering methods and relates them to the conventional multivariate clustering methods. The structure of the taxonomy is built on three main attributes of a functional data clustering method and therefore is more reliable than existing categorizations. The review aims to bridge the gap between the functional data analysis community and the clustering community and to generate new principles for functional data clustering.
Mimi Zhang, Andrew C. Parnell
ACM Trans. Knowl. Discov. Data1
2022 Weighted clustering ensemble: A review
Mimi Zhang
Pattern Recognit.1
2021 DCF: An Efficient and Robust Density-Based Clustering Method
abstract
Density-based clustering methods have been shown to achieve promising results in modern data mining applications. A recent approach, Density Peaks Clustering (DPC), detects modes as points with high density and large distance to points of higher density, and hence often fails to detect low-density clusters in the data. Furthermore, DPC has quadratic complexity. We here develop a new clustering algorithm, aiming at improving the applicability and efficiency of the peak-finding technique. The improvements are threefold: (1) the new algorithm is applicable to large datasets; (2) the algorithm is capable of detecting clusters of varying density; (3) the algorithm is competent at deciding the correct number of clusters, even when the number of clusters is very high. The clustering performance of the algorithm is greatly enhanced by directing the peak-finding technique to discover modal sets, rather than point modes. We present a theoretical analysis of our approach and experimental results to verify that our algorithm works well in practice. We demonstrate a potential application of our work for unsupervised face recognition.
Joshua Tobin, Mimi Zhang
ICDM2
2019 Forward-stagewise clustering: An algorithm for convex clustering
Mimi Zhang
Pattern Recognit. Lett.1
2017 Continuous-Observation Partially Observable Semi-Markov Decision Processes for Machine Maintenance
abstract
Partially observable semi-Markov decision processes (POSMDPs) provide a rich framework for planning under both state transition uncertainty and observation uncertainty. In this paper, we widen the literature on POSMDP by studying discrete-state discrete-action yet continuous-observation POSMDPs. We prove that the resultant α-vector set is continuous and, therefore, propose a point-based value iteration algorithm. This paper also bridges the gap between POSMDP and machine maintenance by incorporating various types of maintenance actions, such as actions changing machine state, actions changing degradation rate, and the temporally extended action “do nothing.” Both finite and infinite planning horizons are reviewed, and the solution methodology for each type of planning horizon is given. We illustrate the maintenance decision process via a real industrial problem and demonstrate that the developed framework can be readily applied to solve relevant maintenance problems.
Mimi Zhang, Matthew Revie
IEEE Trans. Reliab.1
2013 A Bivariate Maintenance Policy for Multi-State Repairable Systems With Monotone Process
abstract
This paper proposes a sequential failure limit maintenance policy for a repairable system. The objective system is assumed to have$k+1$states, including one working state and$k$failure states, and the multiple failure states are classified potentially by features such as failure severity or failure cause. The system deteriorates over time and will be replaced upon the$N$th failure. Corrective maintenance is performed immediately upon each of the first$(N-1)$failures. To avoid the costly failure, preventive maintenance actions will be performed as soon as the system's reliability drops to a critical threshold$R$. Both preventive maintenance and corrective maintenance are assumed to be imperfect. Increasing and decreasing geometric processes are introduced to characterize the efficiency of preventive maintenance and corrective maintenance. The objective is to derive an optimal maintenance policy$(R^{\ast},N^{\ast})$such that the long-run expected cost per unit time is minimized. The analytical expression of the cost rate function is derived, and the corresponding optimal maintenance policy can be determined numerically. A numerical example is given to illustrate the theoretical results and the maintaining procedure. The decision model shows its adaptability to different possible characteristics of the maintained system.
Mimi Zhang, Min Xie 0001, Olivier Gaudoin
IEEE Trans. Reliab.1
2011 Adaptive regulation of CCD camera for real time eye tracking
Ruian Liu, Nailin Wang, Mimi Zhang
Multim. Tools Appl.4
2009 Brand and its effect on user perception of search engine performance
abstract
Abstract In this research we investigate the effect of search engine brand on the evaluation of searching performance. Our research is motivated by the large amount of search traffic directed to a handful of Web search engines, even though many have similar interfaces and performance. We conducted a laboratory experiment with 32 participants using a 42 factorial design confounded in four blocks to measure the effect of four search engine brands (Google, MSN, Yahoo!, and a locally developed search engine) while controlling for the quality and presentation of search engine results. We found brand indeed played a role in the searching process. Brand effect varied in different domains. Users seemed to place a high degree of trust in major search engine brands; however, they were more engaged in the searching process when using lesser‐known search engines. It appears that branding affects overall Web search at four stages: (a) search engine selection, (b) search engine results page evaluation, (c) individual link evaluation, and (d) evaluation of the landing page. We discuss the implications for search engine marketing and the design of empirical studies measuring search engine performance.
Jim Jansen, Mimi Zhang, Carsten D. Schultz
J. Assoc. Inf. Sci. Technol.2
2009 Twitter power: Tweets as electronic word of mouth
abstract
Abstract In this paper we report research results investigating microblogging as a form of electronic word‐of‐mouth for sharing consumer opinions concerning brands. We analyzed more than 150,000 microblog postings containing branding comments, sentiments, and opinions. We investigated the overall structure of these microblog postings, the types of expressions, and the movement in positive or negative sentiment. We compared automated methods of classifying sentiment in these microblogs with manual coding. Using a case study approach, we analyzed the range, frequency, timing, and content of tweets in a corporate account. Our research findings show that 19% of microblogs contain mention of a brand. Of the branding microblogs, nearly 20% contained some expression of brand sentiments. Of these, more than 50% were positive and 33% were critical of the company or product. Our comparison of automated and manual coding showed no significant differences between the two approaches. In analyzing microblogs for structure and composition, the linguistic structure of tweets approximate the linguistic patterns of natural language expressions. We find that microblogging is an online tool for customer word of mouth communications and discuss the implications for corporations using microblogging as part of their overall marketing strategy.
Jim Jansen, Mimi Zhang, Kate Sobel, Abdur Chowdhury
J. Assoc. Inf. Sci. Technol.2
2007 Brand awareness and the evaluation of search results
abstract
We investigate the effect of search engine brand (i.e., the identifying name or logo that distinguishes a product from its competitors) on evaluation of system performance. This research is motivated by the large amount of search traffic directed to a handful of Web search engines, even though most are of equal technical quality with similar interfaces. We conducted a laboratory study with 32 participants to measure the effect of four search engine brands while controlling for the quality of search engine results. There was a 25% difference between the most highly rated search engine and the lowest using average relevance ratings, even though search engine results were identical in both content and presentation. Qualitative analysis suggests branding affects user views of popularity, trust and specialization. We discuss implications for search engine marketing and the design of search engine quality studies.
Jim Jansen, Mimi Zhang
WWW2