Minda Zhao

dblp:234/3897 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0002-8736-272XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Large Language Models Are Bad Dice Players: LLMs Struggle to Generate Random Numbers from Statistical Distributions
abstract
As large language models (LLMs) transition from chat interfaces to integral components of stochastic pipelines and systems approaching general intelligence, the ability to faithfully sample from specified probability distributions has become a functional requirement rather than a theoretical curiosity.We present the first large-scale, statistically powered audit of native probabilistic sampling in frontier LLMs, benchmarking 11 models across 15 distributions.To disentangle failure modes, we employ a dual-protocol design: Batch Generation, where a model produces N =1000 samples within one response, and Independent Requests, comprising N =1000 stateless calls.We observe a sharp protocol asymmetry: batch generation achieves only modest statistical validity, with a 7% median pass rate, while independent requests collapse almost entirely, with 10 of 11 models passing none of the distributions.Beyond this asymmetry, we reveal that sampling fidelity degrades monotonically with distributional complexity and aggravates as the sampling horizon N increases.Finally, we demonstrate how the propagation of these failures into downstream real-world application tasks introduces systematic biases: models fail to enforce uniform answer-position constraints in Multiple Choice Question generation and systematically violate demographic targets in attribute-constrained text-to-image prompt synthesis.These findings indicate that current LLMs lack a functional internal sampler, necessitating external tools for applications requiring statistical guarantees.
Minda Zhao, Yilun Du
ACL (1)1
2025 EasyCraft: A Robust and Efficient Framework for Automatic Avatar Crafting
abstract
Character customization, or ’face crafting,’ is a vital feature in role-playing games (RPGs), enhancing player engagement by enabling the creation of personalized avatars. Existing automated methods often struggle with generalizability across diverse game engines due to their reliance on the intermediate constraints of specific image domain and typically support only one type of input, either text or image. To overcome these challenges, we introduce EasyCraft, an innovative end-to-end feedforward framework that automates character crafting by uniquely supporting both text and image inputs. Our approach employs a translator capable of converting facial images of any style into crafting parameters. We first establish a unified feature distribution in the translator’s image encoder through self-supervised learning on a large-scale dataset, enabling photos of any style to be embedded into a unified feature representation.Subsequently, we map this unified feature distribution to crafting parameters specific to a game engine, a process that can be easily adapted to most game engines and thus enhances EasyCraft’s generalizability. By integrating text-to-image techniques with our translator, EasyCraft also facilitates precise, text-based character crafting. EasyCraft’s ability to integrate diverse inputs significantly enhances the versatility and accuracy of avatar creation. Extensive experiments on two RPG games demonstrate the effectiveness of our method, achieving state-of-the-art results and facilitating adaptability across various avatar engines.
Suzhen Wang 0001, Wei Zhang 0219, Minda Zhao, Lincheng Li, Zhipeng Hu, Xin Yu 0002
CVPR4
2025 ICE: Interactive 3D Game Character Facial Editing via Dialogue
abstract
Most recent popular Role-Playing Games (RPGs) allow players to create in-game characters with hundreds of adjustable parameters, including bone positions and various makeup options. Although text-driven auto-customization systems have been developed to simplify the complex process of adjusting these intricate character parameters, they are limited by their single-round generation and lack the capability for further editing and fine-tuning. In this paper, we propose an Interactive Character Editing framework (ICE) to achieve a multi-round dialogue-based refinement process. In a nutshell, our ICE offers a more user-friendly way to enable players to convey creative ideas iteratively while ensuring that created characters align with the expectations of players. Specifically, we propose an Instruction Parsing Module (IPM) that utilizes large language models (LLMs) to parse multi-round dialogues into clear editing instruction prompts in each round. To reliably and swiftly modify character control parameters at a fine-grained level, we propose a Semantic-guided Low-dimension Parameter Solver (SLPS) that edits character control parameters according to prompts in a zero-shot manner. Our SLPS first localizes the character control parameters related to the fine-grained modification, and then optimizes the corresponding parameters in a low-dimension space to avoid unrealistic results. Extensive experimental results demonstrate the effectiveness of our proposed ICE for in-game character creation and the superior editing performance of ICE. Code:https://github.com/NeteaseFuxi/ICE-Interactive-3D-Game-Character.
Haoqian Wu, Minda Zhao, Zhipeng Hu, Changjie Fan, Lincheng Li, Rui Zhao 0019, Xin Yu 0002
IEEE Trans. Multim.2
2024 EfficientDreamer: High-Fidelity and Stable 3D Creation via Orthogonal-view Diffusion Priors
abstract
While image diffusion models have made significant progress in text-driven 3D content creation, they often fail to accurately capture the intended meaning of text prompts, especially for view information. This limitation leads to the Janus problem, where multi-faced 3D models are generated under the guidance of such diffusion models. In this paper, we propose a robust high-quality 3D content generation pipeline by exploiting orthogonal-view image guidance. First, we introduce a novel 2D diffusion model that generates an image consisting of four orthogonal-view sub-images based on the given text prompt. Then, the 3D content is created using this diffusion model. Notably, the generated orthogonal-view image provides strong geometric structure priors and thus improves 3D consistency. As a result, it effectively resolves the Janus problem and significantly enhances the quality of 3D content creation. Additionally, we present a 3D synthesis fusion network that can further improve the details of the generated 3D contents. Both quantitative and qualitative evaluations demonstrate that our method surpasses previous text-to-3D techniques. Project page: https://efficientdreamer.github.io.
Zhipeng Hu, Minda Zhao, Chaoyi Zhao, Lincheng Li, Zeng Zhao, Changjie Fan, Xiaowei Zhou 0001, Xin Yu 0002
CVPR2
2024 Calligraphy Font Generation via Explicitly Modeling Location-Aware Glyph Component Deformations
abstract
Automatic font generation is a challenging and time-consuming task, particularly in languages that consist of large amounts of characters with complicated structures. Typical component-wise font generation methods decompose the source character into components and search for them from the reference glyph set as candidate components. These candidate components are then utilized to learn the local styles of the target glyph. However, these methods overlook that the same component at different locations may have different profiles. When the candidate components locate differently from their corresponding components in the target glyph, the style of a generated glyph will look inconsistent. It is observed that for arbitrary components at two specific locations, the deformation patterns are similar. Driven by this, we present a location-aware component-deformable font generation method. Specifically, we search for candidate components and their corresponding deformative component pairs from the reference glyph set. Each deformative component pair can accurately depict how to deform the candidate component to the desired profile in the target glyph. Hence, we introduce a location-dependent deformation module to perform component warping. In this way, we significantly improve the component deformation ability. Lastly, we integrate deformed components into target glyphs while enforcing their styles to be consistent with the reference ones. Extensive experiments demonstrate that our method produces target-font consistent glyphs and outperforms the state-of-the-art on both seen and unseen fonts.
Minda Zhao, Xingqun Qi, Zhipeng Hu, Lincheng Li, Yongqiang Zhang 0003, Zi Huang, Xin Yu 0002
IEEE Trans. Multim.1
2023 Towards Unbiased Volume Rendering of Neural Implicit Surfaces with Geometry Priors
abstract
Learning surface by neural implicit rendering has been a promising way for multi-view reconstruction in recent years. Existing neural surface reconstruction methods, such as NeuS [24] and VolSDF [32], can produce reliable meshes from multi-view posed images. Although they build a bridge between volume rendering and Signed Distance Function (SDF), the accuracy is still limited. In this paper, we argue that this limited accuracy is due to the bias of their volume rendering strategies, especially when the viewing direction is close to be tangent to the surface. We revise and provide an additional condition for the unbiased volume rendering. Following this analysis, we propose a new rendering method by scaling the SDF field with the angle between the viewing direction and the surface normal vector. Experiments on simulated data indicate that our rendering method reduces the bias of SDF-based volume rendering. Moreover, there still exists non-negligible bias when the learnable standard deviation of SDF is large at early stage, which means that it is hard to supervise the rendered depth with depth priors. Alternatively we supervise zero-level set with surface points obtained from a pre-trained Multi-View Stereo network. We evaluate our method on the DTU dataset and show that it outperforms the state-of-the-arts neural implicit surface methods without mask supervision.
Yongqiang Zhang 0003, Zhipeng Hu, Haoqian Wu, Minda Zhao, Lincheng Li, Zhengxia Zou, Changjie Fan
CVPR4
2023 Beyond First Impressions: Integrating Joint Multi-modal Cues for Comprehensive 3D Representation
abstract
In recent years, 3D representation learning has turned to 2D vision-language pre-trained models to overcome data scarcity challenges. However, existing methods simply transfer 2D alignment strategies, aligning 3D representations with single-view 2D images and coarse-grained parent category text. These approaches introduce information degradation and insufficient synergy issues, leading to performance loss. Information degradation arises from overlooking the fact that a 3D representation should be equivalent to a series of multi-view images and more fine-grained subcategory text. Insufficient synergy neglects the idea that a robust 3D representation should align with the joint vision-language space, rather than independently aligning with each modality. In this paper, we propose a multi-view joint modality modeling approach, termed JM3D, to obtain a unified representation for point cloud, text, and image. Specifically, a novel Structured Multimodal Organizer (SMO) is proposed to address the information degradation issue, which introduces contiguous multi-view images and hierarchical text to enrich the representation of vision and language modalities. A Joint Multi-modal Alignment (JMA) is designed to tackle the insufficient synergy problem, which models the joint modality by incorporating language knowledge into the visual modality. Extensive experiments on ModelNet40 and ScanObjectNN demonstrate the effectiveness of our proposed method, JM3D, which achieves state-of-the-art performance in zero-shot 3D classification. JM3D outperforms ULIP by approximately 4.3% on PointMLP and achieves an improvement of up to 6.5% accuracy on PointNet++ in top-1 accuracy for zero-shot 3D classification on ModelNet40. The source code and trained models for all our experiments are publicly available at https://github.com/Mr-Neko/JM3D.
Haowei Wang 0001, Jiji Tang, Jiayi Ji, Xiaoshuai Sun, Minda Zhao, Lincheng Li, Zeng Zhao, Tangjie Lv, Rongrong Ji
ACM Multimedia7
2021 Adaptively Meshed Video Stabilization
abstract
Video stabilization is essential for improving the visual quality of shaky videos. Current video stabilization methods usually take feature trajectories in the background to estimate one global transformation matrix or several transformation matrices based on a fixed mesh, and warp shaky frames into their stabilized views. However, these methods may not model the shaky camera motion well in complicated scenes, such as scenes containing large foreground objects or strong parallax, and may result in notable visual artifacts in the stabilized videos. To resolve the above issues, this paper proposes an adaptively meshed method to stabilize a shaky video based on all of its feature trajectories and an adaptive blocking strategy. More specifically, we first extract the feature trajectories of the shaky video and then generate a triangle mesh according to the distribution of the feature trajectories in each frame. Then, the transformations between shaky frames and their stabilized views over all triangular grids of the mesh are calculated to stabilize the shaky video. Since more feature trajectories can usually be extracted from all of the regions, including both the background and foreground regions, a finer mesh will be obtained and provided for camera motion estimation and frame warping. We estimate the mesh-based transformations of each frame by solving a two-stage optimization problem. Moreover, foreground and background feature trajectories are no longer distinguished and both contribute to the estimation of the camera motion in the proposed optimization problem, yielding better estimation performance than previous works, particularly in challenging videos with large foreground objects or strong parallax. To further enhance the robustness of our method, we propose two adaptive weighting mechanisms to improve its spatial and temporal adaptability. Experimental results demonstrate the effectiveness of our method in producing visually pleasing stabilization effects in various challenging videos.
Minda Zhao, Qiang Ling 0001
IEEE Trans. Circuits Syst. Video Technol.1
2020 PWStableNet: Learning Pixel-Wise Warping Maps for Video Stabilization
abstract
As the videos captured by hand-held cameras are often perturbed by high-frequency jitters, stabilization of these videos is an essential task. Many video stabilization methods have been proposed to stabilize shaky videos. However, most methods estimate one global homography or several homographies based on fixed meshes to warp the shaky frames into their stabilized views. Due to the existence of parallax, such single or a few homographies can not well handle the depth variation. In contrast to these traditional methods, we propose a novel video stabilization network, called PWStableNet, which comes up pixel-wise warping maps, i.e., potentially different warping for different pixels, and stabilizes each pixel to its stabilized view. To our best knowledge, this is the first deep learning based pixel-wise video stabilization. The proposed method is built upon a multi-stage cascade encoder-decoder architecture and learns pixel-wise warping maps from consecutive unstable frames. Inter-stage connections are also introduced to add feature maps of a former stage to the corresponding feature maps at a latter stage, which enables the latter stage to learn the residual from the feature maps of former stages. This cascade architecture can produce more precise warping maps at latter stages. To ensure the correct learning of pixel-wise warping maps, we use a well-designed loss function to guide the training procedure of the proposed PWStableNet. The proposed stabilization method achieves comparable performance with traditional methods, but stronger robustness and much faster processing speed. Moreover, the proposed stabilization method outperforms some typical CNN-based stabilization methods, especially in videos with strong parallax. Codes will be provided at https://github.com/mindazhao/pix-pix-warping-video-stabilization.
Minda Zhao, Qiang Ling 0001
IEEE Trans. Image Process.1
2019 An Iterative Feedback-Based Change Detection Algorithm for Flood Mapping in SAR Images
abstract
This letter proposes a novel algorithm for the unsupervised detection of flood mapping in synthetic aperture radar (SAR) images. In the literature, unsupervised change detection of SAR images mainly consists of two steps, i.e., first generating a difference image from two given images and then binarizing the difference image to produce the desired change map. Conventional change detection algorithms usually execute these two steps sequentially and separately. In contrast, our algorithm introduces the feedback of the obtained intermediate change maps into both generation and binarization of the difference image. More specifically, we adjust the weights of neighboring pixels in generating the difference image according to the intermediate change maps. With the fed-back intermediate change maps, we also extend the conventional single binarizing threshold for all pixels of the difference image to threshold maps, i.e., two individual binarizing thresholds are defined for each pixel of the difference image and the threshold maps are adjusted accordingly. Due to such feedback of the intermediate change maps, we may obtain a better difference image and generate more precise change maps, which can surely be fed back again. This iterative execution of the above-mentioned generation and binarization of difference images is terminated when a predefined Markov energy function stops decreasing, i.e., it reaches a local minimum. Experiments with several SAR image data sets with floods show that our algorithm consistently outperforms several state-of-the-art algorithms.
Minda Zhao, Qiang Ling 0001, Feng Li 0042
IEEE Geosci. Remote. Sens. Lett.1
2019 Stabilization of Traffic Videos Based on Both Foreground and Background Feature Trajectories
abstract
This paper considers stabilizing traffic videos, which are recorded by cameras mounted on moving vehicles. Compared with videos captured by hand-held cameras, traffic videos are more difficult to stabilize due to dynamic scenes, higher frequency camera jitter, more moving foreground objects and more serious parallax. The conventional video stabilization methods usually estimate the camera jitter from the background feature trajectories which are mainly determined by the camera motion, and then stabilize videos with the estimated jitter. These methods have to correctly distinguish background and foreground feature trajectories, which is not trivial, and may suffer from performance degradation when large foreground objects exist and no enough number of background feature trajectories can be obtained. To resolve these issues, this paper proposes a novel stabilization method, under which background and foreground feature trajectories are no longer distinguished and work together to yield stabilized trajectories. More specifically, the movement of all feature trajectories is modeled as the summation of the camera motion and the object motion. By solving an optimization problem, we can remove the high frequency components of the camera motion, i.e., the camera jitter, and stabilize videos. Parallax is also treated as object motion and contributes feature trajectories for video stabilization. As our method makes use of both foreground and background feature trajectories, it can outperform the conventional stabilization methods which use only background feature trajectories, especially when there are large foreground objects and the number of extracted background feature trajectories is small. Furthermore, some refinements are proposed to speed up our method and enhance its robustness. Experiments are done and confirm its performance superiority.
Qiang Ling 0001, Minda Zhao
IEEE Trans. Circuits Syst. Video Technol.2