Jiawei Ma

dblp:201/7741 · DBLP profile ↗
← Back
39ranked-venue papers
17as first author
37since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 11 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 8 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Stroke Modeling Enables Vectorized Character Generation with Large Vectorized Glyph Model
abstract
Vectorized glyphs are widely used in poster design, network animation, art display, and various other fields due to their scalability and flexibility. In typography, they are often seen as special sequences composed of ordered strokes. This concept extends to the token sequence prediction abilities of large language models (LLMs), enabling vectorized character generation through stroke modeling. In this paper, we propose a novel Large Vectorized Glyph Model (LVGM) designed to generate vectorized Chinese glyphs by predicting the next stroke. Initially, we encode strokes into discrete latent variables called stroke embeddings. Subsequently, we train our LVGM via fine-tuning DeepSeek LLM by predicting the next stroke embedding. With limited strokes given, it can generate complete characters, semantically elegant words, and even unseen verses in vectorized form. Moreover, we release a new large-scale Chinese SVG dataset containing 907, 267 samples based on strokes for dynamically vectorized glyph generation. Experimental results show that our model has scaling behaviors on data scales. Our generated vectorized glyphs have been validated by experts and relevant individuals.
Jiawei Ma, Chen Ye 0002
WACV3
2026 Dual-network optimized backstepping: predefined-time function projective synchronization for hyperchaotic economic systems with external perturbations
Jiawei Ma, Huaguang Zhang, Muxuan Li
Expert Syst. Appl.1
2026 Multi-Modal Few-Shot Object Detection with Meta-Learning-Based Cross-Modal Prompting
Guangxing Han, Long Chen 0016, Jiawei Ma, Shiyuan Huang 0001, Rama Chellappa, Shih-Fu Chang
Int. J. Comput. Vis.3
2026 A novel dynamic signal Lemma for predefined-time stabilization of high-order nonlinear systems with dynamic uncertainties
Jiawei Ma, Huaguang Zhang, Juan Zhang 0002
Neural Networks1
2026 Observer-Based Fuzzy Secure Control for High-Order MASs Against Communication Delays Under Jointly Connected Switching Topology
abstract
This research considers the observer-based adaptive fuzzy secure consensus control issue for high-order nonlinear multiagent systems (MASs) against communication delays and deception attacks under jointly connected switching topology. To address the challenges of time-varying communication delays, deception attacks, and leader-state inaccessibility in MASs with jointly connected switching topology, this article develops a distributed consensus observer with delay-attack resilience. The proposed observer simultaneously compensates for time-varying delays and counteracts deception attacks while reconstructing the leader's state information through a consensus observer. To handle system uncertainties with unknown nonlinear functions, a fuzzy logic system (FLS) is employed to approximate the unknown dynamics. A fuzzy state observer is subsequently constructed to reconstruct the unmeasurable states by utilizing accessible post-attack signals and applying adaptive approximation technology. With the help of the backstepping design scheme and the adding power integral method, an observer-based adaptive fuzzy secure consensus control approach is proposed so that the designed communication-delay-attack-related distributed consensus observer errors converge to zero exponentially. Furthermore, the proposed consensus control protocol ensures the stability of high-order MASs while achieving that tracking errors converge in the neighborhood of the origin. The validation through two simulation scenarios demonstrates the protocol's effectiveness.
Jiawei Ma, Huaguang Zhang, Juan Zhang 0002
IEEE Trans. Cybern.1
2026 Predefined-Time Safe Cooperative Control for Multiagent Systems With Privacy Preservation and Unknown Disturbances
abstract
Most output-constrained methods necessitate reference command within a predefined safe region, without considering cases where the command itself may conflict with safety boundaries. To handle this problem, this article proposes a predefined-time safe cooperative control scheme for multiagent systems under output constraints, privacy preservation and unknown disturbances. At the communication layer, an encryption-decryption mechanism is developed to safeguard information exchange among agents, preventing internal states from being identified by eavesdroppers. At the control layer, to ensure strict adherence to output constraints regardless of whether the original command complies with safety limits, an improved boundary protection method is explored to generate a safety reference trajectory, which is subsequently used in the controller design. Adaptive laws are then formulated to counteract the effects of unknown nonlinearities and disturbances. Finally, by leveraging predefined-time stability theory, a predefined-time safe cooperative controller is designed to ensure error convergence within a user-defined settling time. Theoretical analysis rigorously confirms the closed-loop stability, and simulations verify the effectiveness of the proposed method.
Xiaohui Yue, Huaguang Zhang, Jiawei Ma
IEEE Trans. Cybern.3
2026 A Dual-Network Optimized Control Framework for Predefined-Time Secure Backstepping of Nonlinear Multiagent Systems
abstract
This article investigates a predefined-time optimized consensus secure control problem for nonlinear multiagent systems, where the leader to follower and follower to neighbor agent network communication are subjected to deferred denial-of-service (DoS) attacks. To observe the leader's state and mitigate the adverse effects of DoS attacks on the system, a switching consensus leader observer is designed. Through the synergistic use of backstepping control, predefined-time control theory, and adaptive dynamic programming, a Hamilton–Jacobi–Bellman equation is constructed for each subsystem to ensure the optimal control performance of the overall system. The identifier neural network is incorporated to approximate the unknown uncertainties existing in the system. Through the construction of the critic network, the proposed controller satisfies the Bellman optimality principle, from which the optimal controller of the system is derived. By employing the Lyapunov stability theorem, it is proven that all system signals remain bounded within the predefined time interval, and the followers' outputs ultimately synchronize with the leader's state. Finally, simulation studies are conducted to validate the feasibility and effectiveness of the proposed control scheme.
Jiawei Ma, Huaguang Zhang, Xiyue Guo
IEEE Trans. Ind. Informatics1
2026 Constraint-Based Finite-Time Tracking Control for Intelligent Multivehicle Systems Considering Communication and Operation Restrictions
abstract
This article presents a constraint-based finite-time tracking control scheme for multivehicle systems (MVSs) with communication and operation restrictions, including network attacks and actuator faults. The stable operation of MVSs significantly depends on the effective communication transmission of state information (position and velocity) between vehicles and the ideal operation of the vehicle’s own actuators. However, the actual transmission signals may be tampered with due to malicious attacks, and different faults may occur during the operation of transportation. Therefore, first, a novel coordinate transformation is developed to cope with the difficulty of distorted position and velocity information, which will lead to follower vehicles making incorrect decisions. Then, an adaptive controller is designed with the help of the fuzzy approximation method to solve the problems of actuator fault compensation and unknown nonlinearity of vehicle’s dynamic. Besides, the nonlinear velocity constraint function is designed to avoid collisions and congestion between vehicles so as to ensure a safe and comfortable transportation environment. Not only a safe distance between vehicles is achieved in a finite time, but also the velocity constraint of vehicles is maintained based on the proposed finite-time control scheme. Finally, the simulation example elaborates the effectiveness of the presented method in this article.
Huaguang Zhang, Jiayue Sun, Jiawei Ma
IEEE Trans. Ind. Informatics4
2026 UncNeRF: Uncovering Heavily Occluded Object With Multi-View Clues
abstract
Neural Radiance Fields can achieve photo-realistic rendering results, but the occlusion in front of the target object is a common and extreme scenario in practice that cannot be neglected. The prevailing works attempt to remove the occlusions using external 2D visual priors, which are not constrained to provide 3D-consistent guidance for the specific scenarios. In this paper, we propose UncNeRF, which utilizes multi-view clues from captured defective images to uncover the heavily occluded object. Specifically, we provide additional multi-view complementary optimization supervisions using object-centric forward warping and enhance the target object reconstruction by sampling pseudo-training views and introducing external spatial-relation regularization. To evaluate the reconstruction performance of occluded objects, we present the challenging and diverse Heavy Occlusion Removal (HOR) dataset consisting of synthetic and real-world scenes, whose target objects to be reconstructed are heavily occluded. Experimental results show that our method achieves state-of-the-art performance in heavy occlusion removal compared to other methods.
Jiawei Ma, Yifan Zhao 0002, Jia Li 0003
IEEE Trans. Image Process.1
2025 Explaining Domain Shifts in Language: Concept Erasing for Interpretable Image Classification
abstract
Concept-based models can map black-box representations to human-understandable concepts, which makes the decision-making process more transparent and then allows users to understand the reason behind predictions. However, domain-specific concepts often impact the final predictions, which subsequently undermine the model generalization capabilities, and prevent the model from being used in high-stake applications. In this paper, we propose a novel Language-guided Concept-Erasing (LanCE) framework. In particular, we empirically demonstrate that pre-trained vision-language models (VLMs) can approximate distinct visual domain shifts via domain descriptors while prompting large Language Models (LLMs) can easily simulate a wide range of descriptors of unseen visual domains. Then, we introduce a novel plug-in domain descriptor orthogonality (DDO) regularizer to mitigate the impact of these domain-specific concepts on the final predictions. Notably, the DDO regularizer is agnostic to the design of concept-based models and we integrate it into several prevailing models. Through evaluation of domain generalization on four standard benchmarks and three newly introduced benchmarks, we demonstrate that DDO can significantly improve the out-of-distribution (OOD) generalization over the previous state-of-the-art concept-based models. Our code is available at https://github.com/joeyz0z/LanCE.
Zequn Zeng, Yudi Su, Tiansheng Wen, Hao Zhang 0050, Zhengjue Wang, Bo Chen 0001, Hongwei Liu 0001, Jiawei Ma
CVPR9
2025 Conservative-Radical Complementary Learning for Class-Incremental Medical Image Analysis with Pre-trained Foundation Models
Xinyao Wu, Zhe Xu 0012, Donghuan Lu, Jinghan Sun, Sadia Shakil, Jiawei Ma, Yefeng Zheng 0001, Raymond Kai-Yu Tong
MICCAI (14)7
2025 A blockchain-based one-to-many traceless covert communication model for secure high-capacity information transmission
Wei She, Jiawei Ma, Kebing Xia, Kong Cheng, Wei Liu 0043
Peer Peer Netw. Appl.2
2025 Predefined-Time Control for Multi-Agent Systems With Input Saturation: An Improved Dynamic Surface Control Scheme
abstract
This paper researches the predefined-time fuzzy adaptive dynamics surface consensus control problem for nonlinear multi-agent systems with input saturation. With regard to nonlinear functions existing in the control systems, fuzzy logic system is employed to estimated them. A novel improved predefined-time dynamics surface filter is designed that can avoid the explosion of complexity problem, and the proposed filter satisfies predefined-time stable simultaneously. With the support of piecewise function, the singularity problem that may exist in virtual controller can be dodged greatly. Furthermore, a ameliorative predefined-time auxiliary dynamic system is presented to cope with input saturation. Combining adaptive backstepping control and predefined-time theory, a predefined-time adaptive fuzzy dynamics surface controller is presented that can assure the systems are predefined-time bounded, though the followers exists in the control saturation.Note to Practitioners—Numerous actual physical systems and devices can be modeled as uncertain nonlinear MAS. Furthermore, MAS also can be widely applied to disaster relief, spacecraft and multitudinous fields. On the one hand, the practical systems may exist in input saturation phenomenon due to human factors and in most of the relevant literatures, which can affect the system performance or bring about instability. On the other hand, the initial values of actual systems usually cannot be chosen freely because of the influence of environment and other factors. In addition, it is currently ordinarily supposed that realize the stabilization of controlled systems when time approaches infinity. Consequently, a predefined-time adaptive fuzzy controller is presented for MAS with control saturation, which can achieve the predefined-time stabilization of MAS.
Jiawei Ma, Huaguang Zhang, Juan Zhang 0002, Xiyue Guo
IEEE Trans Autom. Sci. Eng.1
2025 Event-Based Adaptive Fault-Tolerant Control for Nonlinear Cyber-Physical Systems via Intermittent Available Signals
abstract
This research considers the problem of event-based adaptive fault-tolerant control for nonlinear cyber-physical systems with deception attacks and actuator faults via intermittent available signals. Through the application of fuzzy logic systems, the unknown nonlinear functions of the systems are approximated. Then, a novel state observer is designed that can be driven by actuator faults and intermittent available signals arising from triggered attack-state signals, which can realize directly triggering the states after deception attacks and avoid the problem of virtual controller non-differentiability under the backstepping framework. In order to avoid the complexity explosion problem, the dynamics surface control method is introduced. Meanwhile, the dynamics surface filter signals also realize event-triggered control, and it results in substantial savings in communication resources. To verify the validity of the proposed method, two illustrative examples are presented. Note to Practitioners—Cyber-physical systems are frequently applied in modern industrial processes such as smart grids, industrial Internet of Things and so on. Nevertheless, the open network environment makes system components more vulnerable to attacks, posing a significant security threat to the operation of cyber-physical systems. When attackers deliver deceptive information to the sensors, the system state becomes inaccessible, which is a challenging issue based on the signals available after the attacks. Moreover, when actuators are affected by faults, system performance deteriorates, potentially leading to instability. Consequently, the event-triggered control problem of cyber-physical systems, considering both actuator faults and deceptive attacks, presents challenges for controller design. In order to resolve the issues outlined above, a fault-tolerant state observer based on intermittent available signals is developed. This design employs backstepping recursion and adaptive techniques to achieve stability for system under deception attacks and greatly alleviate communication burden, thereby enhancing the practicality of the proposed control strategy.
Jiawei Ma, Huaguang Zhang, Juan Zhang 0002, Xiyue Guo
IEEE Trans Autom. Sci. Eng.1
2025 Predefined-Time Fuzzy Formation Control for High-Order Multiagent Systems via Event-Triggered Schemes
abstract
This research considers the predefined-time adaptive fuzzy formation control issue for high-order nonlinear multi-agent systems. By applying fuzzy logic systems, the systems unknown nonlinear functions can be approximated. To refrain from “explosion of complexity problem”, a novel dynamics surface for high-order nonlinear multi-agent systems is presented. Further, to minimize the communication burden, an event-triggered mechanism suitable for high-order nonlinear multi-agent systems is applied in the control methods. With the help of the backstepping design scheme and the adding power integral method, an adaptive fuzzy predefined-time formation control approach is proposed so that all signals in the considered systems are bounded and realize the desired formation control within predefined time. The illustrative examples are presented to verify the validity of the suggested method.
Jiawei Ma, Huaguang Zhang, Juan Zhang 0002
IEEE Trans. Fuzzy Syst.1
2025 Self-Triggered Optimal Control for Unknown Nonlinear Random Power Systems With Markovian Switching
abstract
This article explores the challenge of triggered optimal control for random differential equations (RDEs) with Markovian switching. We initially address the inherent contradiction between whether to comply with or bypass the event-triggered mechanism. By navigating this challenge, we ensure noise-to-state stability (NSS) for RDEs through event-triggered control (ETC). Furthermore, we establish that random nonlinear systems utilizing self-triggered control (STC) can achieve NSS, by setting a minimum triggering time to prevent Zeno behavior. Lastly, by adopting the adaptive dynamic programming (ADP) strategy, we develop self-triggered optimal control for random systems with Markovian switching, ensuring the uniform ultimate boundedness (UUB) of the signals in all closed-loop systems. This article addresses three key gaps in the field of RDE optimal control, contributing substantially to both theoretical and practical advancements. To demonstrate the method’s feasibility, we include a representative example with simulation results.
Zhongyang Ming, Huaguang Zhang, Shuhang Yu, Jiawei Ma
IEEE Trans. Syst. Man Cybern. Syst.4
2025 Predictor-Based Fixed-Time Neural Dynamics Surface Tracking Control for Nonlinear Systems With Unknown Backlash-Like Hysteresis
abstract
The issue of predictor-based neural fixed-time dynamic surface control for the nonlinear systems with unknown backlash-like hysteresis is the research focus of this article. By applying the predictor-based neural control scheme, the system nonlinear functions can be smoothly estimated. In addition, an improved dynamics surface is proposed to decrease the difficulty of the controller design procedure while ensuring that the dynamic surface compensating signals can satisfy the fixed-time stability. Further, on the basis of fixed-time theorem and backstepping control technology, the designed controller can ensure all signals of the considered closed-loop systems are fixed-time bounded in the presence of unknown backlash-like hysteresis. Eventually, the simulation cases are given to imply the effectiveness of the designed method.
Huaguang Zhang, Jiawei Ma, Juan Zhang 0002
IEEE Trans. Syst. Man Cybern. Syst.2
2024 Towards Automated Chinese Ancient Character Restoration: A Diffusion-Based Method with a New Dataset
abstract
Automated Chinese ancient character restoration (ACACR) remains a challenging task due to its historical significance and aesthetic complexity. Existing methods are constrained by non-professional masks and even overfitting when training on small-scale datasets, which hinder their interdisciplinary application to traditional fields. In this paper, we are proud to introduce the Chinese Ancient Rubbing and Manuscript Character Dataset (ARMCD), which consists of 15,553 real-world ancient single-character images with 42 rubbings and manuscripts, covering the works of over 200 calligraphy artists spanning from 200 to 1,800 AD. We are also dedicated to providing professional synthetic masks by extracting localized erosion from real eroded images. Moreover, we propose DiffACR (Diffusion model for automated Chinese Ancient Character Restoration), a diffusion-based method for the ACACR task. Specifically, we regard the synthesis of eroded images as a special form of cold diffusion on uneroded ones and extract the prior mask directly from the eroded images. Our experiments demonstrate that our method comprehensively outperforms most existing methods on the proposed ARMCD. Dataset and code are available at https://github.com/lhl322001/DiffACR.
Chenghao Du, Ziheng Jiang, Jiawei Ma, Chen Ye 0002
AAAI5
2024 MoDE: CLIP Data Experts via Clustering
abstract
The success of contrastive language-image pretraining (CLIP) relies on the supervision from the pairing between images and captions, which tends to be noisy in web- crawled data. We present Mixture of Data Experts (MoDE) and learn a system of CLIP data experts via clustering. Each data expert is trained on one data cluster, being less sensitive to false negative noises in other clusters. At inference time, we ensemble their outputs by applying weights determined through the correlation between task metadata and cluster conditions. To estimate the correlation pre-cisely, the samples in one cluster should be semantically similar, but the number of data experts should still be rea-sonable for training and inference. As such, we consider the ontology in human language and propose to use fine- grained cluster centers to represent each data expert at a coarse-grained level. Experimental studies show that four CLIP data experts on ViT-B/16 outperform the ViT-L/14 by OpenAI CLIP and OpenCLIP on zero-shot image classification but with less (<35%) training cost. Meanwhile, MoDE can train all data expert asynchronously and can flexibly include new data experts. The code is available here.
Jiawei Ma, Po-Yao Huang 0001, Saining Xie, Shang-Wen Li 0001, Luke Zettlemoyer, Shih-Fu Chang, Scott Yih, Hu Xu 0001
CVPR1
2024 How to Use Diffusion Priors under Sparse Views?
abstract
Novel view synthesis under sparse views has been a long-term important challenge in 3D reconstruction. Existing works mainly rely on introducing external semantic or depth priors to supervise the optimization of 3D representations. However, the diffusion model, as an external prior that can directly provide visual supervision, has always underperformed in sparse-view 3D reconstruction using Score Distillation Sampling (SDS) due to the low information entropy of sparse views compared to text, leading to optimization challenges caused by mode deviation. To this end, we present a thorough analysis of SDS from the mode-seeking perspective and propose Inline Prior Guided Score Matching (IPSM), which leverages visual inline priors provided by pose relationships between viewpoints to rectify the rendered image distribution and decomposes the original optimization objective of SDS, thereby offering effective diffusion visual guidance without any fine-tuning or pre-training. Furthermore, we propose the IPSM-Gaussian pipeline, which adopts 3D Gaussian Splatting as the backbone and supplements depth and geometry consistency regularization based on IPSM to further improve inline priors and rectified distribution. Experimental results on different public datasets show that our method achieves state-of-the-art reconstruction quality. The code is released at https://github.com/iCVTEAM/IPSM.
Yifan Zhao 0002, Jiawei Ma, Jia Li 0003
NeurIPS3
2024 Adaptive fixed-time dynamic surface tracking control for high-order nonstrict-feedback nonlinear switched systems
Huanqing Wang 0001, Zhu Meng, Jiawei Ma, Xudong Zhao 0001
Neurocomputing3
2024 Neuroadaptive event-triggered tracking control for nonlinear systems with dynamic fault feedback and prescribed time convergence
Huaguang Zhang, Juan Zhang 0002, Jiawei Ma
Neurocomputing4
2024 Event-Based Fixed-Time Fuzzy Containment Fault-Tolerant Control for Multiagent Systems With Positive Odd Rational Powers
abstract
In this research, an event-based adaptive fixed-time fuzzy containment fault-tolerant control issue for multi-agent systems (MASs) with positive odd rational powers is discussed. As a result of the presence of unknown nonlinear functions, fuzzy logic systems (FLSs) can be employed to estimate them with the support of FLSs’ approximation capability. The adding a power integrator (API) can be introduced to handle the difficulties encountered in the backstepping design process for high-order nonlinear systems. Further, the event-triggered control ideology was carried out among neighbors of MASs, which is used to minimize the communication burden. In addition, the controlled MASs considered four types of sensor faults, adopting the adaptive control methods to achieve effective estimation of unknown fault parameters. According to the frame of backstepping control, an adaptive fuzzy containment event-triggered controller for MASs with positive odd rational powers was designed that can realize all signals are fixed-time bounded, even if MASs may exist sensor faults. In the last resort, the illustrative example can manifest the validity of the suggested approach.
Jiawei Ma, Huaguang Zhang, Juan Zhang 0002, Zeyi Liu 0003
IEEE Trans. Fuzzy Syst.1
2024 Adaptive Neural Fixed-Time Tracking Control for High-Order Nonlinear Systems
abstract
The problem of adaptive neural fixed-time tracking control for high-order systems is addressed in this article. In order to handle the difficulties from the uncertain nonlinearities within the original systems, the radial basis function neural networks (RBF NNs) are introduced to approximate the unknown nonlinear functions, and the adding a power integrator is applied to overcome the obstacle from high-order terms. It is proven that all signals in the closed-loop system are bounded and the output signal can eventually converge to a small neighborhood of the reference signal. Simulation results further verify the approaches developed.
Jiawei Ma, Huanqing Wang 0001, Junfei Qiao 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 Supervised Masked Knowledge Distillation for Few-Shot Transformers
abstract
Vision Transformers (ViTs) emerge to achieve impressive performance on many data-abundant computer vision tasks by capturing long-range dependencies among local features. However, under few-shot learning (FSL) settings on small datasets with only a few labeled data, ViT tends to overfit and suffers from severe performance degradation due to its absence of CNN-alike inductive bias. Previous works in FSL avoid such problem either through the help of self-supervised auxiliary losses, or through the dextile uses of label information under supervised settings. But the gap between self-supervised and supervised few-shot Transformers is still unfilled. Inspired by recent advances in self-supervised knowledge distillation and masked image modeling (MIM), we propose a novel Supervised Masked Knowledge Distillation model (SMKD) for few-shot Transformers which incorporates label information into self-distillation frameworks. Compared with previous self-supervised methods, we allow intra-class knowledge distillation on both class and patch tokens, and introduce the challenging task of masked patch tokens reconstruction across intra-class images. Experimental results on four few-shot classification benchmark datasets show that our method with simple design outperforms previous methods by a large margin and achieves a new start-of-the-art. Detailed ablation studies confirm the effectiveness of each component of our model. Code for this paper is available here: https://github.com/HL-hanlin/SMKD.
Guangxing Han, Jiawei Ma, Shiyuan Huang 0001, Xudong Lin 0003, Shih-Fu Chang
CVPR3
2023 DiGeo: Discriminative Geometry-Aware Learning for Generalized Few-Shot Object Detection
abstract
Generalized few-shot object detection aims to achieve precise detection on both base classes with abundant annotations and novel classes with limited training data. Existing approaches enhance few-shot generalization with the sacrifice of base-class performance, or maintain high precision in base-class detection with limited improvement in novel-class adaptation. In this paper, we point out the reason is insufficient Discriminative feature learning for all of the classes. As such, we propose a new training framework, DiGeo, to learn Geometry-aware features of interclass separation and intra-class compactness. To guide the separation of feature clusters, we derive an offline simplex equiangular tight frame (ETF) classifier whose weights serve as class centers and are maximally and equally separated. To tighten the cluster for each class, we include adaptive class-specific margins into the classification loss and encourage the features close to the class centers. Experimental studies on two few-shot benchmark datasets (VOC, COCO) and one long-tail dataset (LVIS) demonstrate that, with a single model, our method can effectively improve generalization on novel classes without hurting the detection of base classes. Our code can be found here.
Jiawei Ma, Yulei Niu, Jincheng Xu, Shiyuan Huang 0001, Guangxing Han, Shih-Fu Chang
CVPR1
2023 TempCLR: Temporal Alignment Representation with Contrastive Learning
Yuncong Yang, Jiawei Ma, Shiyuan Huang 0001, Long Chen 0016, Xudong Lin 0003, Guangxing Han, Shih-Fu Chang
ICLR2
2023 Improved finite-time prescribed performance based adaptive neural control for nonlinear systems with sensor faults
Huanqing Wang 0001, Kexin Lu, Fu Zheng, Jiawei Ma, Cungen Liu
Neurocomputing4
2023 Adaptive Fuzzy Fixed-Time Control for High-Order Nonlinear Systems With Sensor and Actuator Faults
abstract
In this article, an adaptive fuzzy fixed-time fault-tolerant tracking control problem for high-order nonlinear systems (HONSs) with sensor and actuator faults is considered. The fuzzy logic systems are introduced to approximate the unknown nonlinear functions of the HONS. In addition, based on backstepping technology and fixed-time theory, an adaptive fuzzy fixed-time fault-tolerant controller is developed to ensure that all the signals of the closed-loop HONS are bounded. Eventually, a numerical example can be shown to prove the rationality of the developed method
Huanqing Wang 0001, Jiawei Ma, Xudong Zhao 0001, Ben Niu 0003, Ming Chen 0020
IEEE Trans. Fuzzy Syst.2
2022 Meta Faster R-CNN: Towards Accurate Few-Shot Object Detection with Attentive Feature Alignment
abstract
Few-shot object detection (FSOD) aims to detect objects using only a few examples. How to adapt state-of-the-art object detectors to the few-shot domain remains challenging. Object proposal is a key ingredient in modern object detectors. However, the quality of proposals generated for few-shot classes using existing methods is far worse than that of many-shot classes, e.g., missing boxes for few-shot classes due to misclassification or inaccurate spatial locations with respect to true objects. To address the noisy proposal problem, we propose a novel meta-learning based FSOD model by jointly optimizing the few-shot proposal generation and fine-grained few-shot proposal classification. To improve proposal generation for few-shot classes, we propose to learn a lightweight metric-learning based prototype matching network, instead of the conventional simple linear object/nonobject classifier, e.g., used in RPN. Our non-linear classifier with the feature fusion network could improve the discriminative prototype matching and the proposal recall for few-shot classes. To improve the fine-grained few-shot proposal classification, we propose a novel attentive feature alignment method to address the spatial misalignment between the noisy proposals and few-shot classes, thus improving the performance of few-shot object detection. Meanwhile we learn a separate Faster R-CNN detection head for many-shot base classes and show strong performance of maintaining base-classes knowledge. Our model achieves state-of-the-art performance on multiple FSOD benchmarks over most of the shots and metrics.
Guangxing Han, Shiyuan Huang 0001, Jiawei Ma, Yicheng He, Shih-Fu Chang
AAAI3
2022 Few-Shot Object Detection with Fully Cross-Transformer
abstract
Few-shot object detection (FSOD), with the aim to detect novel objects using very few training examples, has recently attracted great research interest in the community. Metric-learning based methods have been demonstrated to be effective for this task using a two-branch based siamese network, and calculate the similarity between image regions and few-shot examples for detection. However, in previous works, the interaction between the two branches is only restricted in the detection head, while leaving the remaining hundreds of layers for separate feature extraction. Inspired by the recent work on vision transformers and vision-language transformers, we propose a novel Fully Cross-Transformer based model (FCT) for FSOD by incorporating cross-transformer into both the feature backbone and detection head. The asymmetric-batched cross-attention is proposed to aggregate the key information from the two branches with different batch sizes. Our model can improve the few-shot similarity learning between the two branches by introducing the multi-level interactions. Comprehensive experiments on both PASCAL VOC and MSCOCO FSOD benchmarks demonstrate the effectiveness of our model.
Guangxing Han, Jiawei Ma, Shiyuan Huang 0001, Long Chen 0016, Shih-Fu Chang
CVPR2
2022 Task-Adaptive Negative Envision for Few-Shot Open-Set Recognition
abstract
We study the problem of few-shot open-set recognition (FSOR), which learns a recognition system capable of both fast adaptation to new classes with limited labeled exam-ples and rejection of unknown negative samples. Traditional large-scale open-set methods have been shown in-effective for FSOR problem due to data limitation. Current FSOR methods typically calibrate few-shot closed-set clas-sifiers to be sensitive to negative samples so that they can be rejected via thresholding. However, threshold tuning is a challenging process as different FSOR tasks may require different rejection powers. In this paper, we instead propose task-adaptive negative class envision for FSOR to integrate threshold tuning into the learning process. Specifically, we augment the few-shot closed-set classifier with additional negative prototypes generated from few-shot examples. By incorporating few-shot class correlations in the negative generation process, we are able to learn dynamic rejection boundaries for FSOR tasks. Besides, we extend our method to generalized few-shot open-set recognition (GF-SOR), which requires classification on both many-shot and few-shot classes as well as rejection of negative samples. Extensive experiments on public benchmarks validate our methods on both problems.11Code available at https://github.com/shiyuanh/TANE
Shiyuan Huang 0001, Jiawei Ma, Guangxing Han, Shih-Fu Chang
CVPR2
2022 Few-Shot End-to-End Object Detection via Constantly Concentrated Encoding Across Heads
Jiawei Ma, Guangxing Han, Shiyuan Huang 0001, Yuncong Yang, Shih-Fu Chang
ECCV (26)1
2022 Few-Shot Gaze Estimation with Model Offset Predictors
abstract
Due to the variance of optical properties across different people, the performance of a person-agnostic gaze estimation model may not generalize well on a specific person. Though one may achieve better performance by training a person-specific model, it typically requires a large number of samples which is not available in real-life scenarios. Hence, few-shot gaze estimation method is preferred for the small number of samples from a target person. However, the key question is how to close the performance gap between a "few-shot" model and the "many-shot" model. In this paper, we propose to learn a person-specific offset predictor which outputs the difference between the person-agnostic model and the many-shot person-specific model with as few as one training sample. We adapt the knowledge to a new person by using the average of meta-learned offset predictors parameters as the initialization of the new offset predictor. Experiments show that the proposed few-shot person-specific model is not only closer to the corresponding many-shot person-specific model but also has better accuracy than the SOTA few-shot gaze estimation methods in multiple gaze datasets.
Jiawei Ma, Xu Zhang 0022, Yue Wu 0001, Varsha Hedau, Shih-Fu Chang
ICASSP1
2021 Query Adaptive Few-Shot Object Detection with Heterogeneous Graph Convolutional Networks
abstract
Few-shot object detection (FSOD) aims to detect never-seen objects using few examples. This field sees recent improvement owing to the meta-learning techniques by learning how to match between the query image and few-shot class examples, such that the learned model can generalize to few-shot novel classes. However, currently, most of the meta-learning-based methods perform parwise matching between query image regions (usually proposals) and novel classes separately, therefore failing to take into account multiple relationships among them. In this paper, we propose a novel FSOD model using heterogeneous graph convolutional networks. Through efficient message passing among all the proposal and class nodes with three different types of edges, we could obtain context-aware proposal features and query-adaptive, multiclass-enhanced prototype representations for each class, which could help promote the pairwise matching and improve final FSOD accuracy. Extensive experimental results show that our proposed model, denoted as QA-FewDet, outperforms the current state-of-the-art approaches on the PASCAL VOC and MSCOCO FSOD benchmarks under different shots and evaluation metrics.
Guangxing Han, Yicheng He, Shiyuan Huang 0001, Jiawei Ma, Shih-Fu Chang
ICCV4
2021 Partner-Assisted Learning for Few-Shot Image Classification
abstract
Few-shot Learning has been studied to mimic human visual capabilities and learn effective models without the need of exhaustive human annotation. Even though the idea of meta-learning for adaptation has dominated the few-shot learning methods, how to train a feature extractor is still a challenge. In this paper, we focus on the design of training strategy to obtain an elemental representation such that the prototype of each novel class can be estimated from a few labeled samples. We propose a two-stage training scheme, Partner-Assisted Learning (PAL), which first trains a Partner Encoder to model pair-wise similarities and extract features serving as soft-anchors, and then trains a Main Encoder by aligning its outputs with soft-anchors while attempting to maximize classification performance. Two alignment constraints from logit-level and feature-level are designed individually. For each few-shot task, we perform prototype classification. Our method consistently outperforms the state-of-the-art methods on four benchmarks. Detailed ablation studies of PAL are provided to justify the selection of each component involved in training.
Jiawei Ma, Hanchen Xie, Guangxing Han, Shih-Fu Chang, Aram Galstyan, Wael Abd-Almageed
ICCV1
2021 Class Incremental Learning for Video Action Classification
abstract
Class Incremental Learning (CIL) is a hot topic in machine learning for CNN models to learn new classes incrementally. However, most of the CIL studies are for image classification and object recognition tasks and few CIL studies are available for video action classification. To mitigate this problem, in this paper, we present a new Grow When Required network (GWR) based video CIL framework for action classification. GWR learns knowledge incrementally by modeling the manifold of video frames for each encountered action class in feature space. We also introduce a Knowledge Consolidation (KC) method to separate the feature manifolds of old class and new class and introduce an associative matrix for label prediction. Experimental results on KTH and Weizmann demonstrate the effectiveness of the framework.
Jiawei Ma, Jianxing Ma, Xiaopeng Hong, Yihong Gong
ICIP1
2020 End-to-End Low Cost Compressive Spectral Imaging with Spatial-Spectral Self-Attention
Ziyi Meng 0001, Jiawei Ma, Xin Yuan 0002
ECCV (23)2
2019 Deep Tensor ADMM-Net for Snapshot Compressive Imaging
abstract
Snapshot compressive imaging (SCI) systems have been developed to capture high-dimensional (≥ 3) signals using low-dimensional off-the-shelf sensors, i.e., mapping multiple video frames into a single measurement frame. One key module of a SCI system is an accurate decoder that recovers the original video frames. However, existing model-based decoding algorithms require exhaustive parameter tuning with prior knowledge and cannot support practical applications due to the extremely long running time. In this paper, we propose a deep tensor ADMM-Net for video SCI systems that provides high-quality decoding in seconds. Firstly, we start with a standard tensor ADMM algorithm, unfold its inference iterations into a layer-wise structure, and design a deep neural network based on tensor operations. Secondly, instead of relying on a pre-specified sparse representation domain, the network learns the domain of low-rank tensor through stochastic gradient descent. It is worth noting that the proposed deep tensor ADMM-Net has potentially mathematical interpretations. On public video data, the simulation results show the proposed method achieves average 0.8 ~ 2.5 dB improvement in PSNR and 0.07 ~ 0.1 in SSIM, and 1500× ~ 3600× speedups over the state-of-the-art methods. On real data captured by SCI cameras, the experimental results show comparable visual results with the state-of-the-art methods but in much shorter running time.
Jiawei Ma, Xiao-Yang Liu, Xin Yuan 0002
ICCV1