VLDB 2026 Research / reviewers in the wild / expert
Fan Ma
dblp:126/0861
· DBLP profile ↗
38ranked-venue papers
9as first author
27since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 8 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 5 first-author · 20 since 2021Systems, architecture and hardware · 3 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Image Regeneration: Evaluating Text-to-Image Model via Generating Identical Image with Multimodal Large Language ModelsabstractDiffusion models have revitalized the image generation domain, playing crucial roles in both academic research and artistic expression. With the emergence of new diffusion models, assessing the performance of text-to-image models has become increasingly important. Current metrics focus on directly matching the input text with the generated image, but due to cross-modal information asymmetry, this leads to unreliable or incomplete assessment results. Motivated by this, we introduce the Image Regeneration task in this study to assess text-to-image models by tasking the T2I model with generating an image according to the reference image. We use GPT4V to bridge the gap between the reference image and the text input for the T2I model, allowing T2I models to understand image content. This evaluation process is simplified as comparisons between the generated image and the reference image are straightforward. Two regeneration datasets spanning content-diverse and style-diverse evaluation dataset are introduced to evaluate the leading diffusion models currently available. Additionally, we present ImageRepainter framework to enhance the quality of generated images by improving content comprehension via MLLM guided iterative generation and revision. Our comprehensive experiments have showcased the effectiveness of this framework in assessing the generative capabilities of models. By leveraging MLLM, we have demonstrated that a robust T2M can produce images more closely resembling the reference image. Chutian Meng, Fan Ma, Jiaxu Miao, Yi Yang 0001, Yueting Zhuang |
AAAI | 2 |
| 2025 | Autonomous LLM-Enhanced Adversarial Attack for Text-to-MotionabstractHuman motion generative models have enabled promising applications, but the ability of text-to-motion (T2M) models to produce realistic motions raises security concerns if exploited maliciously. Despite growing interest in T2M, limited research focus on safeguarding these models against adversarial attacks, with existing work on text-to-image models proving insufficient for the unique motion domain. In the paper, we propose ALERT-Motion, an autonomous framework that leverages large language models (LLMs) to generate targeted adversarial attacks against black-box T2M models. Unlike prior methods that modify prompts through predefined rules, ALERT-Motion uses the knowledge of LLMs of human motion to autonomously generate subtle yet powerful adversarial text descriptions. It comprises two key modules: an adaptive dispatching module that constructs an LLM-based agent to iteratively refine and search for adversarial prompts; and a multimodal information contrastive module that extracts semantically relevant motion information to guide the agent's search. Through this LLM-driven approach, ALERT-Motion produces adversarial prompts querying victim models to produce outputs closely matching targeted motions, while avoiding obvious perturbations. Evaluations across popular T2M models demonstrate ALERT-Motion's superiority over previous methods, achieving higher attack success rates with stealthier adversarial prompts. This pioneering work on T2M adversarial attacks highlights the urgency of developing defensive measures as motion generation technology advances, urging further research into safe and responsible deployment. Honglei Miao, Fan Ma, Ruijie Quan, Kun Zhan, Yi Yang 0001 |
AAAI | 2 |
| 2025 | BrainGuard: Privacy-Preserving Multisubject Image Reconstructions from Brain ActivitiesabstractReconstructing perceived images from human brain activity forms a crucial link between human and machine learning through Brain-Computer Interfaces. Early methods primarily focused on training separate models for each individual to account for individual variability in brain activity, overlooking valuable cross-subject commonalities. Recent advancements have explored multisubject methods, but these approaches face significant challenges, particularly in data privacy and effectively managing individual variability. To overcome these challenges, we introduce BrainGuard, a privacy-preserving collaborative training framework designed to enhance image reconstruction from multisubject fMRI data while safeguarding individual privacy. BrainGuard employs a collaborative global-local architecture where personalized models are trained on each subject's data and operate in conjunction with a shared commonality model that captures and leverages cross-subject patterns. This architecture eliminates the need to aggregate fMRI data across subjects, thereby ensuring privacy preservation. To tackle the complexity of fMRI data, BrainGuard integrates a hybrid synchronization strategy, enabling individual models to dynamically incorporate parameters from the global model. By establishing a secure and collaborative training environment, BrainGuard not only protects sensitive brain activity data but also improves the accuracy of image reconstructions. Extensive experiments demonstrate that BrainGuard sets a new benchmark in both high-level and low-level metrics, advancing the state-of-the-art in brain decoding through its innovative design. Zhibo Tian, Ruijie Quan, Fan Ma, Kun Zhan, Yi Yang 0001 |
AAAI | 3 |
| 2025 | Imagine and Seek: Improving Composed Image Retrieval with an Imagined ProxyabstractThe Zero-shot Composed Image Retrieval (ZSCIR) requires retrieving images that match the query image and the relative captions. Current methods focus on projecting the query image into the text feature space, subsequently combining them with features of query texts for retrieval. However, retrieving images only with the text features can-not guarantee detailed alignment due to the natural gap between images and text. In this paper, we introduce Imagined Proxy for CIR (IP-CIR), a training-free method that creates a proxy image aligned with the query image and text description, enhancing query representation in the retrieval process. We first leverage the large language model’s generalization capability to generate an image layout, and then apply both the query text and image for conditional generation. The robust query features are enhanced by merging the proxy image, query image, and text semantic perturbation. Our newly proposed balancing metric integrates text-based and proxy retrieval similarities, allowing for more accurate retrieval of the target image while incorporating image-side information into the process. Experiments on three public datasets demonstrate that our method significantly improves retrieval performances. We achieve state-of-the-art (SOTA) results on the CIRR dataset with a Recall@K of 70.07 at K=10. Additionally, we achieved an improvement in Recall@10 on the FashionIQ dataset, rising from 45.11 to 45.74, and improved the baseline performance in CIRCO with a mAPK@10 score, increasing from 32.24 to 34.26. Fan Ma |
CVPR | 2 |
| 2025 | Zero-1-to-A: Zero-Shot One Image to Animatable Head Avatars Using Video DiffusionabstractAnimatable head avatar generation typically requires extensive data for training. To reduce the data requirements, a natural solution is to leverage existing data-free static avatar generation methods, such as pre-trained diffusion models with score distillation sampling (SDS), which align avatars with pseudo ground-truth outputs from the diffusion model. However, directly distilling 4D avatars from video diffusion often leads to over-smooth results due to spatial and temporal inconsistencies in the generated video. To address this issue, we propose Zero-1-to-A, a robust method that synthesizes a spatial and temporal consistency dataset for 4D avatar reconstruction using the video diffusion model. Specifically, Zero-1-to-A iteratively constructs video datasets and optimizes animatable avatars in a progressive manner, ensuring that avatar quality increases smoothly and consistently throughout the learning process. This progressive learning involves two stages: (1) Spatial Consistency Learning fixes expressions and learns from front-to-side views, and (2) Temporal Consistency Learning fixes views and learns from relaxed to exaggerated expressions, generating 4D avatars in a simple-to-complex manner. Extensive experiments demonstrate that Zero-1-to-A improves fidelity, animation quality, and rendering speed compared to existing diffusion-based methods, providing a solution for lifelike avatar creation. Code is publicly available at: https://github.com/ZhenglinZhou/Zero-1-to-A. Zhenglin Zhou, Fan Ma, Hehe Fan, Tat-Seng Chua |
CVPR | 2 |
| 2025 | From Trial to Triumph: Advancing Long Video Understanding via Visual Context Sample Scaling and Self-Reward AlignmentabstractMulti-modal Large language models (MLLMs) show remarkable ability in video understanding. Nevertheless, understanding long videos remains challenging as the models can only process a finite number of frames in a single inference, potentially omitting crucial visual information. To address the challenge, we propose generating multiple predictions through visual context sampling, followed by a scoring mechanism to select the final prediction. Specifically, we devise a bin-wise sampling strategy that enables MLLMs to generate diverse answers based on various combinations of keyframes, thereby enriching the visual context. To determine the final prediction from the sampled answers, we employ a self-reward by linearly combining three scores: (1) a frequency score indicating the prevalence of each option, (2) a marginal confidence score reflecting the inter-intra sample certainty of MLLM predictions, and (3) a reasoning score for different question types, including clue-guided answering for global questions and temporal self-refocusing for local questions. The frequency score ensures robustness through majority correctness, the confidence-aligned score reflects prediction certainty, and the typed-reasoning score addresses cases with sparse key visual information using tailored strategies. Experiments show that this approach covers the correct answer for a high percentage of long video questions, on seven datasets show that our method improves the performance of three MLLMs. Yucheng Suo, Fan Ma, Linchao Zhu, Fengyun Rao, Yi Yang 0001 |
ICCV | 2 |
| 2025 | InfiniDreamer: Arbitrarily Long Human Motion Generation Via Segment Score DistillationabstractWe present InfiniDreamer, a novel framework for arbitrarily long human motion generation. InfiniDreamer addresses the limitations of current motion generation methods, which are typically restricted to short sequences due to the lack of long motion training data. To achieve this, we first generate sub-motions corresponding to each textual description and then assemble them into a coarse, extended sequence using randomly initialized transition segments. We then introduce an optimization-based method called Segment Score Distillation (SSD) to refine the entire long motion sequence. SSD is designed to utilize an existing motion prior, which is trained only on short clips, in a training-free manner. Specifically, SSD iteratively refines overlapping short segments sampled from the coarsely extended long motion sequence, progressively aligning them with the pre-trained motion diffusion prior. This process ensures local coherence within each segment, while the refined transitions between segments maintain global consistency across the entire sequence. Extensive qualitative and quantitative experiments validate the superiority of our framework, showcasing its ability to generate coherent, contextually aware motion sequences of arbitrary length. Wenjie Zhuo, Fan Ma, Hehe Fan |
ICCV | 2 |
| 2025 | Long-horizon Visual Instruction Generation with Logic and Attribute Self-reflectionabstractVisual instructions for long-horizon tasks are crucial as they intuitively clarify complex concepts and enhance retention across extended steps.
Directly generating a series of images using text-to-image models without considering the context of previous steps results in inconsistent images, increasing cognitive load. Additionally, the generated images often miss objects or the attributes such as color, shape, and state of the objects are inaccurate.
To address these challenges, we propose LIGER, the first training-free framework for Long-horizon Instruction GEneration with logic and attribute self-Reflection. LIGER first generates a draft image for each step with the historical prompt and visual memory of previous steps. This step-by-step generation approach maintains consistency between images in long-horizon tasks. Moreover, LIGER utilizes various image editing tools to rectify errors including wrong attributes, logic errors, object redundancy, and identity inconsistency in the draft images. Through this self-reflection mechanism, LIGER improves the logic and object attribute correctness of the images.
To verify whether the generated images assist human understanding, we manually curated a new benchmark consisting of various long-horizon tasks. Human-annotated ground truth expressions reflect the human-defined criteria for how an image should appear to be illustrative.
Experiments demonstrate the visual instructions generated by LIGER are more comprehensive compared with baseline methods. The code and dataset will be available once accepted. Yucheng Suo, Fan Ma, Kaixin Shen, Linchao Zhu, Yi Yang 0001 |
ICLR | 2 |
| 2025 | DreamDPO: Aligning Text-to-3D Generation with Human Preferences via Direct Preference OptimizationabstractText-to-3D generation automates 3D content creation from textual descriptions, which offers transformative potential across various fields. However, existing methods often struggle to align generated content with human preferences, limiting their applicability and flexibility. To address these limitations, in this paper, we propose DreamDPO, an optimization-based framework that integrates human preferences into the 3D generation process, through direct preference optimization. Practically, DreamDPO first constructs pairwise examples, then validates their alignment with human preferences using reward or large multimodal models, and lastly optimizes the 3D representation with a preference-driven loss function. By leveraging relative preferences, DreamDPO reduces reliance on precise quality evaluations while enabling fine-grained controllability through preference-guided optimization. Experiments demonstrate that DreamDPO achieves state-of-the-art results, and provides higher-quality and more controllable 3D content compared to existing methods. The code and models will be open-sourced. Zhenglin Zhou, Xiaobo Xia, Fan Ma, Hehe Fan, Yi Yang 0001, Tat-Seng Chua |
ICML | 3 |
| 2025 | MIGC++: Advanced Multi-Instance Generation Controller for Image SynthesisabstractWe introduce the Multi-Instance Generation (MIG) task, which focuses on generating multiple instances within a single image, each accurately placed at predefined positions with attributes such as category, color, and shape, strictly following user specifications. MIG faces three main challenges: avoiding attribute leakage between instances, supporting diverse instance descriptions, and maintaining consistency in iterative generation. To address attribute leakage, we propose the Multi-Instance Generation Controller (MIGC). MIGC generates multiple instances through a divide-and-conquer strategy, breaking down multi-instance shading into single-instance tasks with singular attributes, later integrated. To provide more types of instance descriptions, we developed MIGC++. MIGC++ allows attribute control through text & images and position control through boxes & masks. Lastly, we introduced the Consistent-MIG algorithm to enhance the iterative MIG ability of MIGC and MIGC++. This algorithm ensures consistency in unmodified regions during the addition, deletion, or modification of instances, and preserves the identity of instances when their attributes are changed. We introduce the COCO-MIG and Multimodal-MIG benchmarks to evaluate these methods. Extensive experiments on these benchmarks, along with the COCO-Position benchmark and DrawBench, demonstrate that our methods substantially outperform existing techniques, maintaining precise control over aspects including position, attribute, and quantity. Project page: https://github.com/limuloo/MIGC. Dewei Zhou, Fan Ma, Zongxin Yang, Yi Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Small Signal Synchronizing Stability of PLL-Based Wind Turbine Converter During Current Injection to Low Voltage Grid FaultabstractIn renewable highly integrated power system, reactive current injection of renewable generating units is important to the security operation of power system during grid fault, which is usually required by LVRT (low voltage ride through) grid code. Keeping stability during LVRT is the precondition for the current injection. However, in high impedance AC grid, the strengthened interaction of renewable generating units and AC grid may deteriorate the stability, resulting in current injection failure. In this paper, the small signal synchronizing stability of PLL (phase locked loop) based WTC (wind turbine converter) during current injection to low voltage grid fault is studied. First, a synchronizing dynamic model is developed, in which PLL is modelled in the form of rotor motion and current injection control adjusts the equivalent driving force similar as governor in SG (synchronous generator). Based on the developed model, two categories of instability issues are identified. One is the nonexistence of equilibrium point related with K-factor, grid impedance and grid voltage sag. The other is the insufficiency of damping. Current injection control may introduce negative damping to PLL’s equivalent motion in some cases, bringing synchronizing oscillation instability. Finally, simulated results are presented to verify the analytical results. Fan Ma |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2025 | VLAB: Enhancing Video Language Pretraining by Feature Adapting and BlendingabstractLarge-scale image-text contrastive pre-training models, such as CLIP, have been demonstrated to effectively learn high-quality multimodal representations. However, there is limited research on learning video-text representations for general video multimodal tasks based on these powerful features. Towards this goal, we propose a novel video-text pre-training method dubbed VLAB:VideoLanguage pre-training by featureAdapting andBlending, which transfers CLIP representations to video pre-training tasks and develops unified video multimodal models for a wide range of video-text tasks. Specifically, VLAB is founded on two key strategies: feature adapting and feature blending. In the former, we introduce a new video adapter module to address CLIP's deficiency in modeling temporal information and extend the model's capability to encompass both contrastive and generative tasks. In the latter, we propose an end-to-end training method that further enhances the model's performance by exploiting the complementarity of image and video features. We validate the effectiveness and versatility of VLAB through extensive experiments on highly competitive video multimodal tasks, including video text retrieval, video captioning, and video question answering. Remarkably, VLAB outperforms competing methods significantly and sets new records in video question answering on MSRVTT, MSVD, and TGIF datasets. It achieves an accuracy of 49.6, 60.9, and 79.0, respectively. Xingjian He, Fan Ma, Zhicheng Huang 0002, Xiaojie Jin 0004, Dongmei Fu, Yi Yang 0001, Jing Liu 0001, Jiashi Feng |
IEEE Trans. Multim. | 3 |
| 2024 | Stitching Segments and Sentences towards Generalization in Video-Text Pre-trainingabstractVideo-language pre-training models have recently achieved remarkable results on various multi-modal downstream tasks. However, most of these models rely on contrastive learning or masking modeling to align global features across modalities, neglecting the local associations between video frames and text tokens. This limits the model’s ability to perform fine-grained matching and generalization, especially for tasks that selecting segments in long videos based on query texts. To address this issue, we propose a novel stitching and matching pre-text task for video-language pre-training that encourages fine-grained interactions between modalities. Our task involves stitching video frames or sentences into longer sequences and predicting the positions of cross-model queries in the stitched sequences. The individual frame and sentence representations are thus aligned via the stitching and matching strategy, encouraging the fine-grained interactions between videos and texts. in the stitched sequences for the cross-modal query. We conduct extensive experiments on various benchmarks covering text-to-video retrieval, video question answering, video captioning, and moment retrieval. Our results demonstrate that the proposed method significantly improves the generalization capacity of the video-text pre-training models. Fan Ma, Xiaojie Jin 0004, Jingjia Huang, Linchao Zhu, Yi Yang 0001 |
AAAI | 1 |
| 2024 | LSK3DNet: Towards Effective and Efficient 3D Perception with Large Sparse KernelsabstractAutonomous systems need to process large-scale, sparse, and irregular point clouds with limited compute resources. Consequently, it is essential to develop LiDAR perception methods that are both efficient and effective. Although naive-ly enlarging 3D kernel size can enhance performance, it will also lead to a cubically-increasing overhead. Therefore, it is crucial to develop streamlined 3D large kernel designs that eliminate redundant weights and work effectively with larger kernels. In this paper, we propose an efficient and effective Large Sparse Kernel 3D Neural Network (LSK3DNet) that leverages dynamic pruning to amplify the 3D kernel size. Our method comprises two core components: Spatial-wise Dynamic Sparsity (SDS) and Channel-wise Weight Selection (CWS). SDS dynamically prunes and regrows volumetric weights from the beginning to learn a large sparse 3D kernel. It not only boosts performance but also significantly reduces model size and computational cost. Moreover, CWS selects the most important channels for 3D convolution during training and subsequently prunes the redundant channels to accelerate inference for 3D vision tasks. We demonstrate the effectiveness of LSK3DNet on three benchmark datasets and five tracks compared with classical models and large kernel designs. Notably, LSK3DNet achieves the state-of-the-art performance on SemanticKITTI (i.e., 75.6% on single-scan and 63.4% on multi-scan), with roughly 40% model size re-duction and 60% computing operations reduction compared to the naive large 3D kernel model. Tuo Feng 0001, Wenguan Wang, Fan Ma, Yi Yang 0001 |
CVPR | 3 |
| 2024 | CapHuman: Capture Your Moments in Parallel UniversesabstractWe concentrate on a novel human-centric image synthesis task, that is, given only one reference facial photograph, it is expected to generate specific individual images with diverse head positions, poses, facial expressions, and illuminations in different contexts. To accomplish this goal, we argue that our generative model should be capable of the following favorable characteristics: (1) a strong visual and semantic understanding of our world and human society for basic object and human image generation. (2) generalizable identity preservation ability. (3) flexible and fine-grained head control. Recently, large pre-trained text-to-image diffusion models have shown remarkable results, serving as a powerful generative foundation. As a basis, we aim to unleash the above two capabilities of the pre-trained model. In this work, we present a new framework named CapHuman. We embrace the “encode then learn to align” paradigm, which enables generalizable identity preservation for new individuals without cumbersome tuning at inference. CapHuman encodes identity features and then learns to align them into the latent space. Moreover, we introduce the 3D facial prior to equip our model with control over the human head in a flexible and 3D-consistent manner. Extensive qualitative and quantitative analyses demonstrate our CapHuman can produce well-identity-preserved, photo-realistic, and high-fidelity portraits with content-rich representations and various head renditions, superior to established baselines. Code and checkpoint will be released at https://github.com/VamosC/CapHuman. Chao Liang 0002, Fan Ma, Linchao Zhu, Yingying Deng, Yi Yang 0001 |
CVPR | 2 |
| 2024 | Vista-llama: Reducing Hallucination in Video Language Models via Equal Distance to Visual TokensabstractRecent advances in large video-language models have displayed promising outcomes in video comprehension. Current approaches straightforwardly convert video into language tokens and employ large language models for multi-modal tasks. However, this method often leads to the generation of irrelevant content, commonly known as “hallucination”, as the length of the text increases and the impact of the video diminishes. To address this problem, we propose Vista-llama, a novel framework that maintains the consistent distance between all visual tokens and any language tokens, irrespective of the generated text length. Vista-llama omits relative position encoding when determining attention weights between visual and text tokens, retaining the position encoding for text and text tokens. This amplifies the effect of visual tokens on text generation, especially when the relative distance is longer between visual and text tokens. The proposed attention mechanism significantly reduces the chance of producing irrelevant text related to the video content. Furthermore, we present a sequential visual projector that projects the current video frame into tokens of language space with the assistance of the previous frame. This approach not only captures the temporal relationship within the video, but also allows less visual tokens to encompass the entire video. Our approach significantly outperforms various previous methods (e.g., Video-ChatGPT, MovieChat) on four challenging open-ended video question answering benchmarks. We reach an accuracy of 60.7 on the zero-shot NExT-QA and 60.5 on the zero-shot MSRVTT-QA, setting a new state-of-the-art performance. This project is available at https://jinxxian.github.iolVista-LLaMA. Fan Ma, Xiaojie Jin 0004, Yuchen Xian, Jiashi Feng, Yi Yang 0001 |
CVPR | 1 |
| 2024 | Clustering for Protein Representation LearningabstractProtein representation learning is a challenging task that aims to capture the structure and function of proteins from their amino acid sequences. Previous methods largely ignored the fact that not all amino acids are equally important for protein folding and activity. In this article, we propose a neural clustering framework that can automatically discover the critical components of a protein by considering both its primary and tertiary structure information. Our framework treats a protein as a graph, where each node represents an amino acid and each edge represents a spatial or sequential connection between amino acids. We then apply an iterative clustering strategy to group the nodes into clusters based on their 1D and 3D positions and assign scores to each cluster. We select the highest-scoring clusters and use their medoid nodes for the next iteration of clustering, until we obtain a hierarchical and informative representation of the protein. We evaluate on four protein-related tasks: protein fold classification, enzyme reaction classification, gene ontology term prediction, and enzyme commission number prediction. Experimental results demonstrate that our method achieves state-of-the-art performance. Ruijie Quan, Wenguan Wang, Fan Ma, Hehe Fan, Yi Yang 0001 |
CVPR | 3 |
| 2024 | Psychometry: An Omnifit Model for Image Reconstruction from Human Brain ActivityabstractReconstructing the viewed images from human brain activity bridges human and computer vision through the Brain-Computer Interface. The inherent variability in brain function between individuals leads existing literature to focus on acquiring separate models for each individual using their respective brain signal data, ignoring commonalities between these data. In this article, we devise Psychometry, an omnifit model for reconstructing images from functional Magnetic Resonance Imaging (fMRI) obtained from different subjects. Psychometry incorporates an omni mixture-of-experts (Omni MoE) module where all the experts work together to capture the inter-subject commonalities, while each expert associated with subject-specific parameters copes with the individual differences. Moreover, Psychometry is equipped with a retrieval-enhanced inference strategy, termed Ecphory, which aims to enhance the learned fMRI representation via retrieving from prestored subject-specific memories. These designs collectively render Psychometry omnifit and efficient, enabling it to capture both inter-subject commonality and individual specificity across subjects. As a result, the enhanced fMRI representations serve as conditional signals to guide a generation model to reconstruct high-quality and realistic images, establishing Psychometry as state-of-the-art in terms of both high-level and low-level metrics. Ruijie Quan, Wenguan Wang, Zhibo Tian, Fan Ma, Yi Yang 0001 |
CVPR | 4 |
| 2024 | Knowledge-Enhanced Dual-Stream Zero-Shot Composed Image RetrievalabstractWe study the zero-shot Composed Image Retrieval (ZS- CIR) task, which is to retrieve the target image given a reference image and a description without training on the triplet datasets. Previous works generate pseudo-word tokens by projecting the reference image features to the text embedding space. However, they focus on the global visual representation, ignoring the representation of detailed attributes, e.g., color, object number and layout. To address this challenge, we propose a Knowledge-Enhanced Dual-stream zero-shot composed image retrieval framework (KEDs). KEDs implicitly models the attributes of the reference images by incorporating a database. The database enriches the pseudo-word tokens by providing relevant images and captions, emphasizing shared attribute information in various aspects. In this way, KEDs recognizes the reference image from diverse perspectives. Moreover, KEDs adopts an extra stream that aligns pseudo-word tokens with textual concepts, leveraging pseudo-triplets mined from image-text pairs. The pseudo-word tokens generated in this stream are explicitly aligned with fine-grained semantics in the text embedding space. Extensive experiments on widely used benchmarks, i.e. ImageNet-R, COCO object, Fashion-IQ and CIRR, show that KEDs outperforms previous zero-shot composed image retrieval methods. Code is available at https://github.com/suoych/KEDs. Yucheng Suo, Fan Ma, Linchao Zhu, Yi Yang 0001 |
CVPR | 2 |
| 2024 | MIGC: Multi-Instance Generation Controller for Text-to-Image SynthesisabstractWe present a Multi-Instance Generation (MIG) task, si-multaneously generating multiple instances with diverse controls in one image. Given a set of predefined coordinates and their corresponding descriptions, the task is to ensure that generated instances are accurately at the designated locations and that all instances' attributes adhere to their corresponding description. This broadens the scope of current research on Single-instance generation, elevating it to a more versatile and practical dimension. Inspired by the idea of divide and conquer, we introduce an innovative approach named Multi-Instance Generation Controller (MIGC) to address the challenges of the MIG task. Ini-tially, we break down the MIG task into several subtasks, each involving the shading of a single instance. To ensure precise shading for each instance, we introduce an instance enhancement attention mechanism. Lastly, we aggregate all the shaded instances to provide the necessary information for accurately generating multiple instances in stable diffusion (SD). To evaluate how well generation models per-form on the MIG task, we provide a COCO-MIG bench-mark along with an evaluation pipeline. Extensive experiments were conducted on the proposed COCO-MIG bench-mark, as well as on various commonly used benchmarks. The evaluation results illustrate the exceptional control ca-pabilities of our model in terms of quantity, position, at-tribute, and interaction. Code and demos will be released at https://migcproject.github.io/. Dewei Zhou, Fan Ma, Yi Yang 0001 |
CVPR | 3 |
| 2024 | HeadStudio: Text to Animatable Head Avatars with 3D Gaussian Splatting
Zhenglin Zhou, Fan Ma, Hehe Fan, Zongxin Yang, Yi Yang 0001 |
ECCV (32) | 2 |
| 2024 | VividDreamer: Invariant Score Distillation for Hyper-Realistic Text-to-3D Generation
Wenjie Zhuo, Fan Ma, Hehe Fan, Yi Yang 0001 |
ECCV (88) | 2 |
| 2024 | FedPAM: Federated Personalized Augmentation Model for Text-to-Image RetrievalabstractCLIP-based models have made significant advancements in text-to-image retrieval tasks. However, these retrieval models are typically trained on public datasets with optimizing all parameters, which limits their ability to generalize and adapt quickly to personalized private datasets. In this paper, we introduce a lightweight personalized federated learning solution, namely Federated Personalized Augmentation Model (FedPAM), to achieve personalized text-to-image retrieval from multiple private database. Specifically, for the query text, we fetch the top-k most similar text-image pairs from the private database. We then use an attention-based module to generate personalized representations for different clients. The updated representation includes client-specific information for text-to-image matching, resolving issues of data heterogeneity. Additionally, we ensure efficient and secure communication by fine-tuning a small portion of network parameters. Our experiments demonstrate the effectiveness of the proposed framework, exhibiting a significant performance improvement over recently proposed methods: +5.36 on IAPR TC-12, +2.86 on CC3M, and +1.72 on Flickr30k. Yueying Feng, Fan Ma, Chang Yao 0001, Jingyuan Chen 0003, Yi Yang 0001 |
ICMR | 2 |
| 2024 | A new adversarial malware detection method based on enhanced lightweight neural network
Caixia Gao, Fan Ma, Qiuyan Lan, Jianying Chen |
Comput. Secur. | 3 |
| 2022 | Unified Transformer Tracker for Object TrackingabstractAs an important area in computer vision, object tracking has formed two separate communities that respectively study Single Object Tracking (SOT) and Multiple Object Tracking (MOT). However, current methods in one tracking scenario are not easily adapted to the other due to the divergent training datasets and tracking objects of both tasks. Although UniTrack [45] demonstrates that a shared appearance model with multiple heads can be used to tackle individual tracking tasks, it fails to exploit the large-scale tracking datasets for training and performs poorly on the single object tracking. In this work, we present the Unified Transformer Tracker (UTT) to address tracking problems in different scenarios with one paradigm. A track transformer is developed in our UTT to track the target in both SOT and MOT where the correlation between the target feature and the tracking frame feature is exploited to localize the target. We demonstrate that both SOT and MOT tasks can be solved within this framework, and the model can be simultaneously end-to-end trained by alternatively optimizing the SOT and MOT objectives on the datasets of individual tasks. Extensive experiments are conducted on several benchmarks with a unified model trained on both SOT and MOT datasets. Fan Ma, Zheng Shou 0001, Linchao Zhu, Haoqi Fan 0001, Yilei Xu, Yi Yang 0001, Zhicheng Yan 0001 |
CVPR | 1 |
| 2022 | Weakly Supervised Moment Localization with Decoupled Consistent Concept Prediction
Fan Ma, Linchao Zhu, Yi Yang 0001 |
Int. J. Comput. Vis. | 1 |
| 2022 | Learning With Noisy Labels via Self-Reweighting From Class CentroidsabstractAlthough deep neural networks have been proved effective in many applications, they are data hungry, and training deep models often requires laboriously labeled data. However, when labeled data contain erroneous labels, they often lead to model performance degradation. A common solution is to assign each sample with a dynamic weight during optimization, and the weight is adjusted in accordance with the loss. However, those weights are usually unreliable since they are measured by the losses of corrupted labels. Thus, this scheme might impede the discriminative ability of neural networks trained on noisy data. To address this issue, we propose a novel reweighting method, dubbed self-reweighting from class centroids (SRCC), by assigning sample weights based on the similarities between the samples and our online learned class centroids. Since we exploit statistical class centers in the image feature space to reweight data samples in learning, our method is robust to noise caused by corrupted labels. In addition, even after reweighting the noisy data, the decision boundaries might still suffer distortions. Thus, we leverage mixed inputs that are generated by linearly interpolating two random images and their labels to further regularize the boundaries. We employ the learned class centroids to evaluate the confidence of our generated mixed data via measuring feature similarities. During the network optimization, the class centroids are updated as more discriminative feature representations of original images are learned. In doing so, SRCC will generate more robust weighting coefficients for noisy and mixed data and facilitates our feature representation learning in return. Extensive experiments on both the synthetic and real image recognition tasks demonstrate that our method SRCC outperforms the state of the art on learning with noisy data. Fan Ma, Yu Wu 0011, Xin Yu 0002, Yi Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2020 | Context Modulated Dynamic Networks for Actor and Action Video Segmentation with Language QueriesabstractActor and action video segmentation with language queries aims to segment out the expression referred objects in the video. This process requires comprehensive language reasoning and fine-grained video understanding. Previous methods mainly leverage dynamic convolutional networks to match visual and semantic representations. However, the dynamic convolution neglects spatial context when processing each region in the frame and is thus challenging to segment similar objects in the complex scenarios. To address such limitation, we construct a context modulated dynamic convolutional network. Specifically, we propose a context modulated dynamic convolutional operation in the proposed framework. The kernels for the specific region are generated from both language sentences and surrounding context features. Moreover, we devise a temporal encoder to incorporate motions into the visual features to further match the query descriptions. Extensive experiments on two benchmark datasets, Actor-Action Dataset Sentences (A2D Sentences) and J-HMDB Sentences, demonstrate that our proposed approach notably outperforms state-of-the-art methods. Hao Wang 0062, Cheng Deng 0002, Fan Ma, Yi Yang 0001 |
AAAI | 3 |
| 2020 | SF-Net: Single-Frame Supervision for Temporal Action Localization
Fan Ma, Linchao Zhu, Yi Yang 0001, Shengxin Zha, Gourab Kundu, Matt Feiszli, Zheng Shou 0001 |
ECCV (4) | 1 |
| 2020 | Self-paced Multi-view Co-trainingabstractCo-training is a well-known semi-supervised learning approach which trains classifiers on two or more different views and exchanges pseudo labels of unlabeled instances in an iterative way. During the co-training process, pseudo labels of unlabeled instances are very likely to be false especially in the initial training, while the standard co-training algorithm adopts a 'draw without replacement' strategy and does not remove these wrongly labeled instances from training stages. Besides, most of the traditional co-training approaches are implemented for two-view cases, and their extensions in multi-view scenarios are not intuitive. These issues not only degenerate their performance as well as available application range but also hamper their fundamental theory. Moreover, there is no optimization model to explain the objective a co-training process manages to optimize. To address these issues, in this study we design a unified self-paced multi-view co-training (SPamCo) framework which draws unlabeled instances with replacement. Two specified co-regularization terms are formulated to develop different strategies for selecting pseudo-labeled instances during training. Both forms share the same optimization strategy which is consistent with the iteration process in co-training and can be naturally extended to multi-view scenarios. A distributed optimization strategy is also introduced to train the classifier of each view in parallel to further improve the efficiency of the algorithm. Furthermore, the SPamCo algorithm is proved to be PAC learnable, supporting its theoretical soundness. Experiments conducted on synthetic, text categorization, person re-identification, image recognition and object detection data sets substantiate the superiority of the proposed method. Fan Ma, Deyu Meng, Xuanyi Dong, Yi Yang 0001 |
J. Mach. Learn. Res. | 1 |
| 2019 | Online Learning to Rank in a Listwise Approach for Information RetrievalabstractA common approach to learning to rank is to minimize the pair-wise loss. However, established analysis shows that pair-wise loss does not necessarily lead to an optimal list-wise ranking measures, e.g., average precision (AP) or area under precision-recall curve (AUPRC). It becomes more difficult in the online learning setting, where the data arrives sequentially and is scanned only once. This paper proposes an online learning-to-rank algorithm by minimizing the list-wise ranking error, which achieves a vanishing gap between the list-wise loss and the ranking measures. Experiments also testify the effectiveness and robustness of the proposed online List-wise algorithm. Fan Ma, Haoyun Yang, Haibing Yin, Xiaofeng Huang, Chenggang Yan 0001 |
ICME | 1 |
| 2019 | Few-Example Object Detection with Model CommunicationabstractIn this paper, we study object detection using a large pool of unlabeled images and only a few labeled images per category, named "few-example object detection". The key challenge consists in generating trustworthy training samples as many as possible from the pool. Using few training examples as seeds, our method iterates between model training and high-confidence sample selection. In training, easy samples are generated first and, then the poorly initialized model undergoes improvement. As the model becomes more discriminative, challenging but reliable samples are selected. After that, another round of model improvement takes place. To further improve the precision and recall of the generated training samples, we embed multiple detection models in our framework, which has proven to outperform the single model baseline and the model ensemble method. Experiments on PASCAL VOC'07, MS COCO'14, and ILSVRC'13 indicate that by using as few as three or four samples selected for each category, our method produces very competitive results when compared to the state-of-the-art weakly-supervised approaches using a large number of image-level labels. Xuanyi Dong, Liang Zheng 0001, Fan Ma, Yi Yang 0001, Deyu Meng |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2017 | Self-Paced Co-trainingabstractCo-training is a well-known semi-supervised learning approach which trains classifiers on two different views and exchanges labels of unlabeled instances in an iterative way. During co-training process, labels of unlabeled instances in the training pool are very likely to be false especially in the initial training rounds, while the standard co-training algorithm utilizes a “draw without replacement” manner and does not remove these false labeled instances from training. This issue not only tends to degenerate its performance but also hampers its fundamental theory. Besides, there is no optimization model to explain what objective a cotraining process optimizes. To these issues, in this study we design a new co-training algorithm named self-paced cotraining (SPaCo) with a “draw with replacement” learning mode. The rationality of SPaCo can be proved under theoretical assumptions utilized in traditional co-training research, and furthermore, the algorithm exactly complies with the alternative optimization process for an optimization model of self-paced curriculum learning, which can be finely explained in robust learning manner. Experimental results substantiate the superiority of the proposed method as compared with current state-of-the-art co-training methods. Fan Ma, Deyu Meng, Qi Xie 0002, Zina Li, Xuanyi Dong |
ICML | 1 |
| 2017 | Directional interlocking overcurrent protection of microgrids powered by inverters injected with characteristic currentsabstractA microgrid powered by inverters has become a typical approach to efficient use of renewable energy source. However, the behavior of constantcurrent of the inverters will make it difficult to protect the network. To solve this problem, this paper offers the directional interlocking overcurrent protection for the microgrid powered by the inverters injected with characteristic current. When a fault occurs, the inverter switches to the constantcurrent mode. Then, the N multiple frequency and small amplitude current are output as a characteristic signal to decide the direction of fault current in outputting the fundamental frequency current. The digital relay protection device is used to protect the network. Finally, with the microgrid powered by four inverters as an example, the proposed protection strategy is verified by means of the PSCAD/EMTDC software. Xiaoliang Hao, Fan Ma |
IECON | 2 |
| 2017 | Emergency control strategy of hybrid power system under sudden load applyingabstractThe hybrid power system consisting of new energy inverters and diesel generators has become a typical case of independent power systems in remote areas. Aiming at the problem of generator overload caused by sudden load applying in such systems, this paper proposes a novel emergency control strategy which increases the inverter output power according to the changing rate of system frequency. This control strategy is based on the constant current control of inverters and voltage and rotate speed droop control of generators. And it effectively solves the generator overload problem and guarantees the continuous power supply for all the loads in the system. The control strategy is validated by time-domain electromagnetic transient simulations. Fan Ma |
IECON | 3 |
| 2017 | A co-training approach to the classification of local climate zones with multi-source dataabstractLocal climate zone (LCZ) classification system provides standard urban morphological classification for urban heat island studies and weather and climate modelling. Based on the definition of the LCZ, various semi-supervised classification approaches have been proposed to generate LCZ maps for different cities using available satellite data. Given that the acquisition of training data is labor intensive, it is practical to develop new models that are suitable for LCZ classification for any cities without the need for training data/samples. In this study, a novel domain-adaptation co-training approach with self-paced learning is designed to generate LCZ maps for new cities with which valid training samples from existing cities are explored and transferred to new target cities for classification. Experimental results show that the proposed approach could derive LCZ maps for the four testing cities, with an overall accuracy of 69.8%, which is over 10% more accurate than conventional approaches. Compared with conventional approaches, the novel approach does not need prior knowledge about the target cities, and it can automatically generate worldwide LCZ maps to support urban-climate studies for cities in the world. Yong Xu 0002, Fan Ma, Deyu Meng, Chao Ren 0004, Yee Leung |
IGARSS | 2 |
| 2017 | A Dual-Network Progressive Approach to Weakly Supervised Object DetectionabstractA major challenge that arises in Weakly Supervised Object Detection (WSOD) is that only image-level labels are available, whereas WSOD trains instance-level object detectors. A typical approach to WSOD is to 1) generate a series of region proposals for each image and assign the image-level label to all the proposals in that image; 2) train a classifier using all the proposals; and 3) use the classifier to select proposals with high confidence scores as the positive instances for another round of training. In this way, the image-level labels are iteratively transferred to instance-level labels. Xuanyi Dong, Deyu Meng, Fan Ma, Yi Yang 0001 |
ACM Multimedia | 3 |
| 2012 | Fast image super resolution via local regression
Shuhang Gu, Nong Sang, Fan Ma |
ICPR | 3 |