Jian Cao 0002

dblp:50/2102-2 · DBLP profile ↗
← Back
27ranked-venue papers
1as first author
22since 2021 · last 2026
0000-0002-4724-7065ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 12 since 2021Artificial intelligence and machine learning · 9 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 2 since 2021Systems, architecture and hardware · 4 · 3 since 2021
YearPublicationVenuePosition
2026 Two Heads Are Better than One: Distilling Large Language Model Features into Small Models with Feature Decomposition and Mixture
abstract
Market making (MM) through Reinforcement Learning (RL) has attracted significant attention in financial trading. With the development of Large Language Models (LLMs), more and more attempts are being made to apply LLMs to financial areas. A simple, direct application of LLM as an agent shows significant performance. Such methods are hindered by their slow inference speed, while most of the current research has not studied LLM distillation for this specific task. To address this, we first propose the normalized fluorescent probe to study the mechanism of the LLM’s feature. Based on the observation found by our investigation, we propose Cooperative Market Making (CMM), a novel framework that decouples LLM features across three orthogonal dimensions: layer, task, and data. Various student models collaboratively learn simple LLM features along with different dimensions, with each model responsible for a distinct feature to achieve knowledge distillation. Furthermore, CMM introduces an Hájek-MoE to integrate the output of the student models by investigating the contribution of different models in a kernel function-generated common feature space. Extensive experimental results on four real-world market datasets demonstrate the superiority of CMM over the current distillation method and RL-based market-making strategies.
Tianhao Fu, Xinxin Xu 0006, Weichen Xu 0001, Ruilong Ren, Jian Cao 0002, Xixin Cao
AAAI8
2026 PAICar: a prototype of an embodied neuromorphic intelligent robot platform
Mingkai Liu, Jingyi Zhong, Yi Zhong 0002, Zilin Wang 0001, Chenglong Zou, Xiaoxin Cui, Jian Cao 0002, Yuan Wang 0001
Sci. China Inf. Sci.9
2025 Human-Inspired Situated Question Answering with Large Language Models
abstract
Situated Question Answering (SQA) involves reasoning the question based on the scene context from the questioner's perspective, enabling potential applications in human-centered embodied intelligence. Previous works handle this task by data-driven methods, generating answers in a closed-set or open-ended manner without explicit reasoning. Therefore, these works are often trapped in dataset bias, overfitting, and interpretability challenges. Inspired by the reasoning process of humans, we propose a Human-Inspired SQA model (Hi-QA), the first embodied agent that leverages human-like skills to understand the scene and perform interpretable reasoning. Specifically, Hi-QA leverages a specialized perception tool to gather information from 3D scenes for the reasoning module to reason questions with Large Language Models. The reflection module amends invalid reasoning and builds methodology memory to learn problem-solving methods. The distillation module distills prototypical information about objects, supporting the agent with supplementary knowledge. Experiments on ScanQA and SQA3D datasets validate that Hi-QA achieves state-of-the-art performance and highlight that human-inspired design significantly boosts the model's interpretability and generalizability.
Weichen Xu 0001, Jian Cao 0002, Tianhao Fu, Ruilong Ren, Xing Zhang 0002
ICME3
2025 A survey of language-grounded multimodal 3D scene understanding
Ruilong Ren, Weichen Xu 0001, Jian Cao 0002, Xinxin Xu 0006, Xing Zhang 0002
Knowl. Based Syst.4
2025 Mutual information-driven self-supervised point cloud pre-training
Weichen Xu 0001, Tianhao Fu, Jian Cao 0002, Xinxin Xu 0006, Xixin Cao, Xing Zhang 0002
Knowl. Based Syst.3
2024 J-MAE: Jigsaw Meets Masked Autoencoders in X-Ray Security Inspection
abstract
The X-ray security inspection aims to identify any restricted items to protect public safety. Due to the lack of focus on unsupervised learning in this field, using pre-trained models on natural images leads to suboptimal results in downstream tasks. Previous works would lose the relative positional relationships during the pre-training process, which is detrimental for X-ray images that lack texture and rely on shape. In this paper, we propose the jigsaw style MAE (J-MAE) to preserve the relative position information by shuffling the position encoding of visible patches. This forces the network to perform semantic reasoning to understand the shape and composition of X-ray objects. Meanwhile, we propose the Incremental Shuffling Module (ISM) and Permute Predicting Module (PPM) to make the training process more stable and accelerate convergence. Our proposed method has consistently outperformed other methods on three downstream X-ray security inspection datasets.
Weichen Xu 0001, Jian Cao 0002, Tianhao Fu, Awen Bai, Ruilong Ren, Zicong Hu, Xixin Cao, Xing Zhang 0002
ICASSP2
2024 SweepMM: A High-Quality Multimodal Dataset for Sweeping Robots in Home Scenarios for Vision-Language Model
abstract
Embodied intelligence based on vision-language models aims to learn from interactions and derive general intelligence. However, existing generalized vision-language models cannot understand domain knowledge in home scenarios due to the lack of sweeping robot multimodal datasets. In this paper, we propose the first multimodal dataset for sweeping robots, called SweepMM. We create textual data such as room type, scene descriptions, and moving recommendations using various approaches including rule-based, manual-based, and off-the-shelf model-based methods. Based on this dataset, we fine-tune the first generative pretrained model for sweeping robots, called SweepGPM. This model enables human-robot dialogue and surpasses previous state-of-the-art methods by 0.8% in room type recognition, 0.4% in obstacle detection, and 8.0% in lost item search, demonstrating the potential of embodied intelligence in sweeping robots.
Weichen Xu 0001, Xinxin Xu 0006, Tianhao Fu, Jian Cao 0002, Yuetian Huang, Xixin Cao, Xing Zhang 0002
ICASSP4
2024 A Facial Expression Transfer Method Based on 3DMM and Diffusion Models
abstract
Due to the complex geometric features of facial expressions, achieving realistic facial expression transfer is a challenging task. This paper proposes a two-stage facial expression transfer method based on 3D Morphable Model (3DMM) and diffusion models. In the first stage, 3DMM-based DECA model is employed to map the source identity face and the target expression face 2D image into 3D facial parameters. Subsequently, the expression parameters of the source identity face are substituted with the expression parameters obtained from the target expression face. This procedure reconstructs a 3D flame model of the face with the transferred expression. Finally, the 3D flame model, in conjunction with the source identity face rendering-related parameters, is fed into a differentiable renderer to generate a 2D image, completing a rough expression transfer. However, due to the inability of the 3DMM to reconstruct the hair and teeth in the face, the rough expression transfer 2D images obtained in the first stage have obvious artificial traces. To address the above issues, this paper uses the diffusion model in the second stage to repair artificial traces in 2D images, aiming to achieve a close resemblance to the source identity face after expression transfer.
Xinyan Cao, Siqi Liu 0008, Jinming Che, Jian Cao 0002, Jinlong Lin
ICASSP6
2024 Empirical Research On Quantization For 3D Multi-Modal Vit Models
abstract
Model quantization finds success in simplifying model inference in practical applications. However, it predominantly focuses on CNNs and 2D ViT models, with limited attention given to quantizing 3D models. We pensively explore 3D model quantization challenges and discover similar numerical distributions of Softmax and LayerNorm between 3D and 2D models. Consequently, we apply the quantization algorithms FQ-ViT and I-ViT designed for 2D ViT models to 3D model quantization to address performance issues caused by uneven numerical distributions in Softmax and LayerNorm. Our research includes extensive experiments using transformer architectures and establishes benchmarks, demonstrating successful quantization of 3D multimodal model UNITR. Notably, our approach experiences a slight decrease compared to FP32 while outperforming other state-of-the-art models. For example, in the 3D object detection task on the nuScenes dataset, the 8 -bit UNITR (FQ-ViT) achieves impressive NDS and mAP scores of $73.0 \%$ and $70.0 \%$, surpassing the full precision BEVFusion model.
Zicong Hu, Jian Cao 0002, Weichen Xu 0001, Ruilong Ren, Tianhao Fu, Xinxin Xu 0006, Xing Zhang 0002
ICIP2
2024 Boosting 3D Visual Grounding by Object-Centric Referring Network
abstract
3D visual grounding is tasked with locating a specific object within a 3D scene, as described by a given textual reference. This task is challenging because it requires (1) the accurate recognition of various objects in a 3D scene and (2) the understanding of spatial relations in the description. However, current studies encounter difficulties in situations where multiple similar objects are present or when the descriptions involve intricate and abstract relations. In this paper, a novel, simple, and efficient Object-Centric Referring network, namely 3D-OCR, is presented to take high-quality semantic representation and deep relation modeling into account. Specifically, an offline Fine-grained Semantic Enhancement (FSE) module is designed to reinforce the object-centric semantic awareness with fine-grained high-quality object semantic representations. To achieve superior object-centric relation awareness, we propose a Deep Relation Modeling (DRM) module with the explicit and implicit relation self-attention module, enriching object features with relational context. Moreover, we utilize a vision-language contrastive loss to further improve the matching process between point cloud and language. Comprehensive experiments conducted on the challenging ScanRefer and Nr3D datasets corroborate the exceptional performance of our method, with an increase of +1.47% on ScanRefer and +1.2% on Nr3D.
Ruilong Ren, Jian Cao 0002, Weichen Xu 0001, Tianhao Fu, Yilei Dong, Xinxin Xu 0006, Zicong Hu, Xing Zhang 0002
IROS2
2024 An End-to-End SoC for Brain-Inspired CNN-SNN Hybrid Applications
abstract
Inspired by the brain, Spiking Neural Network (SNN) applies temporally sparse spiking communication to gain more bio-mimetic and highly energy efficient computing. The current mainstream platforms for SNN applications are typically the combination of Host+FPGA+Chip Array, which requires an efficient host to preprocess and encode data. It’s not suitable for end-to-end tasks in edge due to its high system power consumption of host and non-negligible high latency of protocol conversion on FPGA. In addition, Convolutional Neural Network (CNN), exhibits strong feature extraction capabilities. Like the brain's visual system, a hierarchical CNN-SNN hybrid network, in which SNN can make use of CNN’s feature extraction capabilities during encoding, can achieve better performance. In this study, we design a 64Neural-Core Array and integrate it with a CNN encoder and a low-power RISC-V CPU within a System-on-Chip (SoC) to enable comprehensive end-to-end hybrid network application support. The proposed heterogeneous SoC is implemented on a Virtex UltraScale+ XCVU9P FPGA, featuring 32.8K neurons, 37.7M synapses and 578GOPS/s peak performance. It processes MNIST classification with a peak throughput of 2022 images per second at frequency of 250MHz. This design gains a balance between high throughput and recognition accuracy simultaneously.
Zhaotong Zhang, Yi Zhong 0002, Yingying Cui, Yawei Ding, Yukun Xue, Qibin Li, Ruining Yang, Jian Cao 0002, Yuan Wang 0001
ISCAS8
2024 Point Cloud Reconstruction Is Insufficient to Learn 3D Representations
abstract
This paper revisits the development of generative self-supervised learning in 2D images and 3D point clouds in autonomous driving. In 2D images, the pretext task has evolved from low-level to high-level features. Inspired by this, through explore model analysis, we find that the gap in weight distribution between self-supervised learning and supervised learning is substantial when employing only low-level features as the pretext task in 3D point clouds. Low-level features represented by PoInt Cloud reconsTruction are insUfficient to learn 3D REpresentations (dubbed PICTURE). To advance the development of pretext tasks, we propose a unified generative self-supervised framework. Firstly, high-level features are demonstrated to exhibit semantic consistency with downstream tasks. We utilize the high-level features as an additional pretext task to enhance the understanding of semantic information during the pre-training. Next, we propose inter-class and intra-class discrimination-guided masking (I2Mask) based on the attributes of the high-level features, adaptively setting the masking ratio for each superclass. On Waymo and nuScenes datasets, we achieve 75.13% mAP and 72.69% mAPH for 3D object detection, 79.4% mIoU for 3D semantic segmentation, and 18.4% mIoU for occupancy prediction. Extensive experiments have demonstrated the effectiveness and necessity of high-level features.
Weichen Xu 0001, Jian Cao 0002, Tianhao Fu, Ruilong Ren, Zicong Hu, Xixin Cao, Xing Zhang 0002
ACM Multimedia2
2024 NeuroREC: A 28-nm Efficient Neuromorphic Processor for Radar Emitter Classification
abstract
Radar emitter classification (REC) plays an important role in modern warfare. Traditional REC methods have difficulty identifying complex radar signals in the present day. Inspired by biology, spiking neural networks (SNNs) have gradually gained widespread attention due to their low power characteristics. Compared with convolutional neural networks (CNNs), SNNs are more suitable for application in the field of REC. The reason is that SNN can not only maintain higher accuracy in the presence of noise interference, but also reduce the power consumption of mobile devices. However, it is challenging to make full use of the input sparsity of radar emitter signals and the weight sparsity of pruned SNN models. In this paper, a 28-nm neuromorphic processor for REC named NeuroREC is proposed. It uses matrix compression algorithms to store sparse weights on chip, and designs corresponding spike detection circuits for this purpose. As a single-core design, we propose a ping-pong running mechanism to alleviate the imbalance between IO throughput and peak performance. Two SNN models for classifying RadioML2016.b and RadioML2018.a datasets are deployed on the chip, achieving competitive accuracy with only 8 timesteps, and demonstrating better robustness than CNN. Fabricated in 28-nm CMOS process, NeuroREC runs at frequencies ranging from 22.5MHz to 744MHz. Under specific sparsity conditions, it can reach an energy efficiency of 7.22TSOP/W for 8-bit weight.
Zilin Wang 0001, Zehong Ou, Yi Zhong 0002, Youming Yang 0002, Li Lun, Hufei Li, Jian Cao 0002, Xiaoxin Cui, Song Jia, Yuan Wang 0001
IEEE Trans. Circuits Syst. I Regul. Pap.7
2024 Leveraging Permuted Image Restoration for Improved Interpretation of Remote Sensing Images
abstract
In this study, we introduce a novel self-supervised learning adapter based on permutated image restoration (PIR) for effectively transferring pretrained weights from natural images to remote sensing object detection tasks. The adapter’s unique methodology encompasses a three-phase process: segmenting and permuting image blocks, estimating permutation matrices for sequence reconstruction, and applying specialized loss functions for accurate block positioning. The use of our approach results in the maintenance of fidelity in both absolute and relative block positions as demonstrated by the evaluation of block similarities. The empirical results indicate significant performance enhancements for diverse datasets spanning optical and SAR data types, including HRSC2016, SODA-A, and RSDD, while effectively avoiding overfitting.
Awen Bai, Jie Chen 0009, Wei Yang 0004, Zhirong Men, Hongcheng Zeng 0001, Weichen Xu 0001, Jian Cao 0002
IEEE Trans. Geosci. Remote. Sens.8
2023 Razor SNN: Efficient Spiking Neural Network with Temporal Embeddings
Yuan Zhang 0020, Jian Cao 0002, Wenyu Sun, Yuan Wang 0001
ICANN (5)2
2023 Channel Pruning Via Attention Module And Memory Curve
abstract
As an effective pruning method, dynamic pruning introduces gate modules that allow different input data to choose different channels. This shows that the choice of channels strongly depends on data. However, this creates an additional computational burden because of the gate modules. In this paper, we propose a simple, efficient and transferable channel pruning method via attention module and memory curve, dubbed as CPAM, which not only takes advantage of the strong correlation between data and channels, but also does not impose any additional computational burden on the model. Inspired by the memory curve, we use a progressive method without any sparse operation. Moreover, our method has been demonstrated effective for many advanced CNN architectures. Notably, on CIFAR-10, CPAM reduces 50% FLOPs on ResNet-56 with 0.31% relative accuracy improvement, which has advanced the state-of-the-art.
Hufei Li, Jian Cao 0002, Xiangcheng Liu, Jingjie Shang, Yuan Wang 0001
ICIP2
2023 Masked Distillation with Receptive Tokens
Tao Huang 0020, Yuan Zhang 0020, Shan You, Fei Wang 0032, Chen Qian 0006, Jian Cao 0002, Chang Xu 0002
ICLR6
2023 Design of a Multimodal Short Video Classification Model
Xinyan Cao, Changjin Li, Jinming Che, Jinlong Lin, Jian Cao 0002
ICONIP (9)8
2023 A New ANN-SNN Conversion Method with High Accuracy, Low Latency and Good Robustness
abstract
Due to the advantages of low energy consumption, high robustness and fast inference speed, Spiking Neural Networks (SNNs), with good biological interpretability and the potential to be applied on neuromorphic hardware, are regarded as the third generation of Artificial Neural Networks (ANNs). Despite having so many advantages, the biggest challenge encountered by spiking neural networks is training difficulty caused by the non-differentiability of spike signals. ANN-SNN conversion is an effective method that solves the training difficulty by converting parameters in ANNs to those in SNNs through a specific algorithm. However, the ANN-SNN conversion method also suffers from accuracy degradation and long inference time. In this paper, we reanalyzed the relationship between Integrate-and-Fire (IF) neuron model and ReLU activation function, proposed a StepReLU activation function more suitable for SNNs under membrane potential encoding, and used it to train ANNs. Then we converted the ANNs to SNNs with extremely small conversion error and introduced leakage mechanism to the SNNs and get the final models, which have high accuracy, low latency and good robustness, and have achieved the state-of-the-art performance on various datasets such as CIFAR and ImageNet.
Bingsen Wang, Jian Cao 0002, Yuan Wang 0001
IJCAI2
2023 Avatar Knowledge Distillation: Self-ensemble Teacher Paradigm with Uncertainty
abstract
Knowledge distillation is an effective paradigm for boosting the performance of pocket-size model, especially when multiple teacher models are available, the student would break the upper limit again. However, it is not economical to train diverse teacher models for the disposable distillation. In this paper, we introduce a new concept dubbed Avatars for distillation, which are the inference ensemble models derived from the teacher. Concretely, (1) For each iteration of distillation training, various Avatars are generated by a perturbation transformation. We validate that Avatars own higher upper limit of working capacity and teaching ability, aiding the student model in learning diverse and receptive knowledge perspectives from the teacher model. (2) During the distillation, we propose an uncertainty-aware factor from the variance of statistical differences between the vanilla teacher and Avatars, to adjust Avatars' contribution on knowledge transfer adaptively. Avatar Knowledge Distillation (AKD) is fundamentally different from existing methods and refines with the innovative view of unequal training. Comprehensive experiments demonstrate the effectiveness of our Avatars mechanism, which polishes up the state-of-the-art distillation methods for dense prediction without more extra computational cost. The AKD brings at most 0.7 AP gains on COCO 2017 for Object Detection and 1.83 mIoU gains on Cityscapes for Semantic Segmentation, respectively.
Yuan Zhang 0020, Tao Huang 0020, Xiuyu Sun, Jian Cao 0002
ACM Multimedia6
2022 Boosting Dense Long-Tailed Object Detection from Data-Centric View
Weichen Xu 0001, Jian Cao 0002, Tianhao Fu, Hongyi Yao, Yuan Wang 0001
ACCV (3)2
2022 A free lunch from ViT: adaptive attention multi-scale fusion Transformer for fine-grained visual recognition
abstract
Learning subtle representation about object parts plays a vital role in fine-grained visual recognition (FGVR) field. The vision transformer (ViT) achieves promising results on computer vision due to its attention mechanism. Nonetheless, with the fixed size of patches in ViT, the class token in deep layer focuses on the global receptive field and cannot generate multi-granularity features for FGVR. To capture region attention without box annotations and compensate for ViT shortcomings in FGVR, we propose a novel method named Adaptive attention multi-scale Fusion Transformer (AFTrans). The Selective Attention Collection Module (SACM) in our approach leverages attention weights in ViT and filters them adaptively to correspond with the relative importance of input patches. The multiple scales (global and local) pipeline is supervised by our weights sharing encoder and can be easily trained end-to-end. Comprehensive experiments demonstrate that AFTrans can achieve SOTA performance on three published fine-grained benchmarks: CUB-200-2011, Stanford Dogs and iNat2017.
Yuan Zhang 0020, Jian Cao 0002, Xiangcheng Liu, Weiqian Chen
ICASSP2
2016 A novel low-leakage power-rail ESD clamp circuit with adjustable triggering voltage and superior false-triggering immunity for nanoscale applications
abstract
This work presents a novel power-rail electrostatic discharge (ESD) clamp circuit for nanoscale applications. By skillfully incorporating transient and static ESD detection mechanisms into its detection circuit, the proposed circuit achieves a wide range of adjustable triggering voltage (Ft1) while maintaining low standby leakage current (Ileak). Besides, the proposed circuit achieves significantly-improved false-triggering immunity compared with the transient-triggered circuit. All investigated circuits are fabricated in a 65-nm CMOS process. Simulation and test results have both confirmed the superiority of the proposed circuit. In addition, the proposed circuit achieves similar triggering behaviors in both transmission line pulsing (TLP) and very fast TLP (VF-TLP) tests.
Guangyi Lu, Yuan Wang 0001, Jian Cao 0002, Song Jia, Xing Zhang 0002
ISCAS3
2016 A compact SCR model using advanced BJT models and standard SPICE elements
Jian Cao 0002, Jingya Xu, Yuan Wang 0001, Guangyi Lu, Xing Zhang 0002
Sci. China Inf. Sci.1
2016 Design of a novel static-triggered power-rail ESD clamp circuit in a 65-nm CMOS process
Guangyi Lu, Yuan Wang 0001, Lizhong Zhang, Jian Cao 0002, Xing Zhang 0002
Sci. China Inf. Sci.4
2016 Area-efficient transient power-rail electrostatic discharge clamp circuit with mis-triggering immunity in a 65-nm CMOS process
Yuan Wang 0001, Guangyi Lu, Haibing Guo, Jian Cao 0002, Song Jia, Xing Zhang 0002
Sci. China Inf. Sci.4
2015 Investigation on the layout strategy of ggNMOS ESD protection devices for uniform conduction behavior and optimal width scaling
Guangyi Lu, Yuan Wang 0001, Lizhong Zhang, Jian Cao 0002, Song Jia, Xing Zhang 0002
Sci. China Inf. Sci.4