Jianxing Liao

dblp:211/9831 · DBLP profile ↗
← Back
19ranked-venue papers
3as first author
17since 2021 · last 2026
0009-0002-5474-4163ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing
abstract
Large language models are extensively utilized in creative writing applications. Creative writing requires a balance between subjective writing quality (e.g., literariness and emotional expression) and objective constraint following (e.g., format requirements and word limits). Existing reinforcement learning methods struggle to balance these two aspects: single reward strategies fail to improve both abilities simultaneously, while fixed-weight mixed-reward methods lack the ability to adapt to different writing scenarios. To address this problem, we propose Reinforcement Learning with Mixed Rewards (RLMR), utilizing a dynamically mixed reward system from a writing reward model evaluating subjective writing quality and a constraint verification model assessing objective constraint following. The constraint following reward weight is adjusted dynamically according to the writing quality within sampled groups, ensuring that samples violating constraints get negative advantage in GRPO and thus penalized during training, which is the key innovation of this proposed method. We conduct automated and manual evaluations across diverse model families from 8B to 72B parameters. Additionally, we construct a real-world writing benchmark named WriteEval for comprehensive evaluation. Results illustrate that our method achieves consistent improvements in both instruction following (IFEval from 83.36% to 86.65%) and writing quality (72.75% win rate in manual expert pairwise evaluations on WriteEval). To the best of our knowledge, RLMR is the first work to combine subjective preferences with objective verification in online RL training, providing an effective solution for multi-dimensional creative writing optimization.
Jianxing Liao, Yusong Zhang, Haorui Wang, Bosi Wen, Ziying Wang, Runzhi Shi
AAAI1
2026 TDSS: Task Dynamic-Synergistic Skill Adaptation for Boosting Efficient and Scalable Multi-Task Learning in Dense Visual Prediction
abstract
The transfer of knowledge from large-scale pre-trained models to diverse downstream tasks has achieved remarkable success. Beyond the traditional full fine-tuning paradigm, Parameter-Efficient Fine-Tuning (PEFT) has emerged as a more efficient model adaptation approach. However, applying existing PEFT methods to adapt dense vision models, particularly in multi-task settings, remains inadequately explored due to their low efficiency, limited task scalability, and neglect of cross-task fine-tuning interactions. To address these challenges, we propose the Task Dynamic-Synergistic Skill Adaptation, termed TDSS, an efficient and scalable multi-task model adaptation framework for dense visual predictions. TDSS comprises two key components: Task-Dynamic Skill Adapters (TDSA) and Task-Synergistic Adaptation Interaction (TSAI). Specifically, TDSA are inserted in parallel into pre-trained vision models to extract task-specific adapted features through the construction of skill representation experts and task dynamic gating. TSAI is developed to enhance cross-task adaptation interaction by bridging global generic and task-specific adapted features. Extensive experiments on multi-task dense visual predictions demonstrate that TDSS surpasses existing state-of-the-art parameter-efficient fine-tuning methods, while exhibiting remarkable efficiency and scalability in parameters and computational complexity.
Haiming Yao, Qiyu Chen 0002, Jianxing Liao
AAAI5
2026 Parameter-, Memory-, Time-Efficient Multi-Task Dense Vision Adaptation
abstract
While adapting pretrained vision models to downstream dense prediction tasks is widely used, current methods often overlook adaptation efficiency, especially in the context of multi-task learning (MTL). Although parameter-efficient fine-tuning (PEFT) methods can enhance parameter efficiency, broader aspects such as GPU memory and training time efficiency remain underexplored. In this paper, we propose a new paradigm that simultaneously achieves efficiency in Parameters, GPU Memory, and Training Time for Multi-Task Dense Vision Adaptation. Specifically, we propose a dual-branch framework, in which a frozen pretrained backbone serves as the generic main branch, and the proposed Bi-Directional Task Adaptation (BDTA) modules are integrated in parallel to form a task bypass branch that extracts adaptation features required by multiple specific tasks. This adaptation module is lightweight, efficient, and does not require backpropagation through the large pre-trained backbone, thus avoiding resource-intensive gradient computations. Moreover, a Mixture of Task Experts mechanism (MoTE) is further proposed to integrate adaptation features across tasks and scales, thereby obtaining more robust representations tailored for dense prediction tasks. On the PASCAL-Context benchmark, our method achieves over 2× relative performance improvement compared to the best prior multi-task PEFT method, while using only ~30% of the parameters, ~50% of the memory, and ~60% of the training time, demonstrating superior overall adaptation efficiency.
Haiming Yao, Qiyu Chen 0002, Jianxing Liao
AAAI4
2025 Automated CAD Modeling Sequence Generation from Text Descriptions via Transformer-Based Large Language Models
abstract
Designing complex computer-aided design (CAD) models is often time-consuming due to challenges such as computational inefficiency and the difficulty of generating precise models. We propose a novel language-guided framework for industrial design automation to address these issues, integrating large language models (LLMs) with computer-automated design (CAutoD).Through this framework, CAD models are automatically generated from parameters and appearance descriptions, supporting the automation of design tasks during the detailed CAD design phase. Our approach introduces three key innovations: (1) a semi-automated data annotation pipeline that leverages LLMs and vision-language large models (VLLMs) to generate high-quality parameters and appearance descriptions; (2) a Transformer-based CAD generator (TCADGen) that predicts modeling sequences via dual-channel feature aggregation; (3) an enhanced CAD modeling generation model, called CADLLM, that is designed to refine the generated sequences by incorporating the confidence scores from TCADGen. Experimental results demonstrate that the proposed approach outperforms traditional methods in both accuracy and efficiency, providing a powerful tool for automating industrial workflows and generating complex CAD models from textual prompts.The code is available at https://jianxliao.github.io/cadllm-page/
Jianxing Liao, Junyan Xu, Yatao Sun, Maowen Tang, Jingxian Liao, Shui Yu 0002, Yun Li 0002, Xiaohong Guan
ACL (1)1
2025 Scaling Spike-Driven Transformer With Efficient Spike Firing Approximation Training
abstract
The ambition of brain-inspired Spiking Neural Networks (SNNs) is to become a low-power alternative to traditional Artificial Neural Networks (ANNs). This work addresses two major challenges in realizing this vision: the performance gap between SNNs and ANNs, and the high training costs of SNNs. We identify intrinsic flaws in spiking neurons caused by binary firing mechanisms and propose a Spike Firing Approximation (SFA) method using integer training and spike-driven inference. This optimizes the spike firing pattern of spiking neurons, enhancing efficient training, reducing power consumption, improving performance, enabling easier scaling, and better utilizing neuromorphic chips. We also develop an efficient spike-driven Transformer architecture and a spike-masked autoencoder to prevent performance degradation during SNN scaling. On ImageNet-1k, we achieve state-of-the-art top-1 accuracy of 78.5%, 79.8%, 84.0%, and 86.2% with models containing 10 M, 19 M, 83 M, and 173 M parameters, respectively. For instance, the 10 M model outperforms the best existing SNN by 7.2% on ImageNet, with training time acceleration and inference energy efficiency improved by 4.5× and 3.9×, respectively. We validate the effectiveness and efficiency of the proposed method across various tasks, including object detection, semantic segmentation, and neuromorphic vision tasks. This work enables SNNs to match ANN performance while maintaining the low-power advantage, marking a significant step towards SNNs as a general visual backbone.
Man Yao, Xuerui Qiu, Tianxiang Hu, Yuhong Chou, Keyu Tian, Jianxing Liao, Luziwei Leng, Bo Xu 0002, Guoqi Li 0002
IEEE Trans. Pattern Anal. Mach. Intell.7
2025 Efficient Deep Spiking Multilayer Perceptrons With Multiplication-Free Inference
abstract
Advancements in adapting deep convolution architectures for spiking neural networks (SNNs) have significantly enhanced image classification performance and reduced computational burdens. However, the inability of multiplication-free inference (MFI) to align with attention and transformer mechanisms, which are critical to superior performance on high-resolution vision tasks, imposes limitations on these gains. To address this, our research explores a new pathway, drawing inspiration from the progress made in multilayer perceptrons (MLPs). We propose an innovative spiking MLP architecture that uses batch normalization (BN) to retain MFI compatibility and introduce a spiking patch encoding (SPE) layer to enhance local feature extraction capabilities. As a result, we establish an efficient multistage spiking MLP network that blends effectively global receptive fields with local feature extraction for comprehensive spike-based computation. Without relying on pretraining or sophisticated SNN training techniques, our network secures a top-one accuracy of 66.39% on the ImageNet-1K dataset, surpassing the directly trained spiking ResNet-34 by 2.67%. Furthermore, we curtail computational costs, model parameters, and simulation steps. An expanded version of our network compares with the performance of the spiking VGG-16 network with a 71.64% top-one accuracy, all while operating with a model capacity 2.1 times smaller. Our findings highlight the potential of our deep SNN architecture in effectively integrating global and local learning abilities. Interestingly, the trained receptive field in our network mirrors the activity patterns of cortical cells.
Boyan Li 0001, Luziwei Leng, Shuaijie Shen, Jianguo Zhang 0001, Jianxing Liao, Ran Cheng 0004
IEEE Trans. Neural Networks Learn. Syst.6
2025 Accurate and Efficient Event-Based Semantic Segmentation Using Adaptive Spiking Encoder-Decoder Network
abstract
Spiking neural networks (SNNs), known for their low-power, event-driven computation, and intrinsic temporal dynamics, are emerging as promising solutions for processing dynamic, asynchronous signals from event-based sensors. Despite their potential, SNNs face challenges in training and architectural design, resulting in limited performance in challenging event-based dense prediction tasks compared with artificial neural networks (ANNs). In this work, we develop an efficient spiking encoder-decoder network (SpikingEDN) for large-scale event-based semantic segmentation (EbSS) tasks. To enhance the learning efficiency from dynamic event streams, we harness the adaptive threshold which improves network accuracy, sparsity, and robustness in streaming inference. Moreover, we develop a dual-path spiking spatially adaptive modulation (SSAM) module, which is specifically tailored to enhance the representation of sparse events and multimodal inputs, thereby considerably improving network performance. Our SpikingEDN attains a mean intersection over union (MIoU) of 72.57% on the DDD17 dataset and 58.32% on the larger DSEC-Semantic dataset, showing competitive results to the state-of-the-art ANNs while requiring substantially fewer computational resources. Our results shed light on the untapped potential of SNNs in event-based vision applications. The source codes are publicly available at https://github.com/EMI-Group/spikingedn.
Luziwei Leng, Kaiwei Che, Qinghai Guo, Jianxing Liao, Ran Cheng 0004
IEEE Trans. Neural Networks Learn. Syst.7
2024 AutoForma: A Large Language Model-Based Multi-Agent for Computer-Automated Design
abstract
With the proliferation of artificial intelligence, Computer-Aided Design (CAD) is being transformed into Computer-Automated Design (CAutoD). In this paper, the advent of Large Language Models (LLMs) introduces new opportunities for CAutoD. This study develops AutoForma, an LLM-based multi-agent system, for automatic conversion from natural language descriptions to 3D models. By harnessing the comprehension capabilities of LLMs, AutoForma streamlines the CAutoD workflow by efficiently translating design intents into precise models in CAD. Through a comprehensive set of evaluations, AutoForma is seen to offer automation performance across various design tasks, particularly in generating non-standard parts that meet specific requirements, with higher efficiency and accuracy than using just an LLM like GPT-4.
Jianxing Liao, Junyan Xu, Zeke Chen, Shui Yu 0002, Yun Li 0002
SMC1
2024 GraDiNet: Implicit Self-Distillation of Graph Structural Knowledge
abstract
Graph Knowledge Distillation (GKD) in artificial intelligence typically employs a teacher-student model, which faces challenges such as rigidity, time-consumption, and teacher training. To improve, this paper develops a Graph self-Distillation Network (GraDiNet), a framework that operates without the need for a teacher model or graph neural network (GNN) during training and inferencing phases. GraDiNet uniquely utilizes multi-layer perceptrons (MLPs) to harness both the structural knowledge of graphs and the semantic information of nodes, thus facilitating hierarchical self-distillation between a target node and its neighbors. Additionally, the GraDiNet approach incorporates a novel similarity-based difference enhancement technique and a penalty factor within the training loss to further delineate the distinction between positive and negative samples. This allows GraDiNet not only to bypass the necessity for a GNN teacher in learning graph structure knowledge but also to predict node classification efficiently. Extensive evaluations show that standard MLPs can significantly boost their performance through this implicit hierarchical self-distillation and the similarity difference enhancement. GraDiNet thus achieves an average improvement of 15% over conventional MLPs and outperforms leading state-of-the-art GKD methods across three real-world datasets.
Junyan Xu, Jianxing Liao, Rucong Xu
SMC2
2024 A novel multi-step ahead prediction method for landslide displacement based on autoregressive integrated moving average and intelligent algorithm
Peng Shao, Guangyu Long, Jianxing Liao, Fei Gan, Yuhang Teng
Eng. Appl. Artif. Intell.4
2023 Test-Time Training-Free Domain Adaptation
abstract
Deploying deep learning models to new environments is very challenging. Domain adaptation (DA) is a promising paradigm to solve the problem by collecting and adapting to unlabeled data in new environments. Though research efforts have led to steady performance improvement over the past decade, DA algorithms are still hard to deploy, as training on unlabeled new data makes tuning difficult and not feasible for inference-only devices. To make DA practical, in this paper we study a new problem named Test-time Training-Free Domain Adaptation (TTDA), where trained models must adapt to a single input (mimicking the test-time scenario) without training. By exploiting spatial activation that was previously overlooked and simply averaged out, we propose a simple method based on Feature Statistics Transformation (FST) on-the-fly for each test example. The proposed algorithm is tested in the TTDA setting on two standard DA benchmarks. Surprisingly, it surpasses or performs on par with state-of-the-art DA methods, even though they require additional training. We envision that this training-free paradigm has the potential to bring DA to embedded devices and would be of interest to audience of community.
Yongxiang Feng, Weihua He, Kaichao You, Yaoyuan Wang, Yihang Lou, Jianxing Liao
ICASSP11
2023 Weakly-Supervised Action Localization by Hierarchically-structured Latent Attention Modeling
abstract
Weakly-supervised action localization aims to recognize and localize action instancese in untrimmed videos with only video-level labels. Most existing models rely on multiple instance learning(MIL), where the predictions of unlabeled instances are supervised by classifying labeled bags. The MIL-based methods are relatively well studied with cogent performance achieved on classification but not on localization. Generally, they locate temporal regions by the video-level classification but overlook the temporal variations of feature semantics. To address this problem, we propose a novel attention-based hierarchically-structured latent model to learn the temporal variations of feature semantics. Specifically, our model entails two components, the first is an unsupervised change-points detection module that detects change-points by learning the latent representations of video features in a temporal hierarchy based on their rates of change, and the second is an attention-based classification model that selects the change-points of the foreground as the boundaries. To evaluate the effectiveness of our model, we conduct extensive experiments on two benchmark datasets, THUMOS-14 and ActivityNet-v1.3. The experiments show that our method outperforms current state-of-the-art methods, and even achieves comparable performance with fully-supervised methods.
Guiqin Wang, Peng Zhao 0001, Cong Zhao 0001, Shusen Yang, Luziwei Leng, Jianxing Liao, Qinghai Guo
ICCV7
2022 Replay-Oriented Gradient Projection Memory for Continual Learning in Medical Scenarios
abstract
Despite the tremendous progress recently achieved by deep learning (DL) in medical image analysis, most DL models only concentrate on single data distribution, which follows the independent and identically distributed (i.i.d) assumption. However, in practice, image data distribution changes with clinical conditions, such as different scanner manufacturers, imaging settings, and statistics regions. Although one can further train the model on new data samples, updating a model with data from an unknown distribution will always result in the model’s performance degradation on the learned data, a notorious phenomenon called catastrophic forgetting. Therefore affects the applicability of DL algorithms in continuously changing clinical scenarios. In this study, we have proposed a new method to address the impact of changing distributions in continual learning scenarios and alleviate catastrophic forgetting. A gradient regularization approach is used to suppress forgetting, and a replay-oriented consistency calculation method combined with a subspace weighting strategy is proposed to improve the model plasticity further. The proposed replay-oriented gradient projection memory (RO-GPM) is evaluated on multiple fundus disease diagnosis datasets including a real-world application and a continual learning benchmark. The quantitative and visualization results demonstrate that the proposed RO-GPM achieves superior performance to state-of-the-art algorithms by a large margin.1
Kuang Shu, Heng Li 0010, Qinghai Guo, Luziwei Leng, Jianxing Liao, Jiang Liu 0001
BIBM6
2022 TimeReplayer: Unlocking the Potential of Event Cameras for Video Interpolation
abstract
Recording fast motion in a high FPS (frame-per-second) requires expensive high-speed cameras. As an alternative, interpolating low-FPS videos from commodity cameras has attracted significant attention. If only low-FPS videos are available, motion assumptions (linear or quadratic) are necessary to infer intermediate frames, which fail to model complex motions. Event camera, a new camera with pixels producing events of brightness change at the temporal resolution of μs (10–6second), is a game-changing device to enable video interpolation at the presence of arbitrarily complex motion. Since event camera is a novel sensor, its potential has not been fulfilled due to the lack of processing algorithms. The pioneering work Time Lens introduced event cameras to video interpolation by designing optical devices to collect a large amount of paired training data of high-speed frames and events, which is too costly to scale. To fully unlock the potential of event cameras, this paper proposes a novel TimeReplayer algorithm to interpolate videos captured by commodity cameras with events. It is trained in an unsupervised cycleconsistent style, canceling the necessity of high-speed training data and bringing the additional ability of video extrapolation. Its state-of-the-art results and demo videos in supplementary reveal the promising future of event-based vision.
Weihua He, Kaichao You, Zhendong Qiao, Xu Jia 0012, Wenhui Wang 0001, Huchuan Lu, Yaoyuan Wang, Jianxing Liao
CVPR9
2022 Video Interpolation by Event-Driven Anisotropic Adjustment of Optical Flow
Kaichao You, Weihua He, Yaoyuan Wang, Jianxing Liao
ECCV (7)8
2022 Meta Talk: Learning To Data-Efficiently Generate Audio-Driven Lip-Synchronized Talking Face With High Definition
abstract
Audio-driven talking face, driving talking face by audio, has received considerable attention in multi-modal learning due to its widespread use in virtual reality. However, long-time recording of target high-quality video is needed by most existing audio-driven talking face studies, which significantly increases customization costs. This paper proposes a novel data-efficient audio-driven talking face generation method, which uses just a short target video to produce both lip-synchronized and high-definition face video driven by arbitrary audio in the wild. Current methods suffer from many problems, such as low definition, asynchronization of lip movement and voice, and intense demands for videos for training. In this work, the original target character’s face images are decomposed into 3D face model parameters including expression, geometry, illumination, etc. Then, low-definition pseudo video generated by an adapted target face video bridges the powerful pre-trained audio-driven model to our audio-to-expression transformation network and help to transfer the ability of audio-identity disentanglement. The expression is replaced via an audio and then combined with other face parameters to render a synthetic face. Finally, a neural rendering network translates the synthetic face into talking face without loss of definition. Experimental results show that the proposed method has the best performance in high-definition image quality, and comparable performance in lip synchronization compared with the existing state-of-the-art methods.
Yuhan Zhang 0006, Weihua He, Yaoyuan Wang, Jianxing Liao
ICASSP8
2022 Differentiable hierarchical and surrogate gradient search for spiking neural networks
abstract
Spiking neural network (SNN) has been viewed as a potential candidate for the next generation of artificial intelligence with appealing characteristics such as sparse computation and inherent temporal dynamics. By adopting architectures of deep artificial neural networks (ANNs), SNNs are achieving competitive performances in benchmark tasks such as image classification. However, successful architectures of ANNs are not necessary ideal for SNN and when tasks become more diverse effective architectural variations could be critical. To this end, we develop a spike-based differentiable hierarchical search (SpikeDHS) framework, where spike-based computation is realized on both the cell and the layer level search space. Based on this framework, we find effective SNN architectures under limited computation cost. During the training of SNN, a suboptimal surrogate gradient function could lead to poor approximations of true gradients, making the network enter certain local minima. To address this problem, we extend the differential approach to surrogate gradient search where the SG function is efficiently optimized locally. Our models achieve state-of-the-art performances on classification of CIFAR10/100 and ImageNet with accuracy of 95.50%, 76.25% and 68.64%. On event-based deep stereo, our method finds optimal layer variation and surpasses the accuracy of specially designed ANNs meanwhile with 26$\times$ lower energy cost ($6.7\mathrm{mJ}$), demonstrating the advantage of SNN in processing highly sparse and dynamic signals. Codes are available at \url{https://github.com/Huawei-BIC/SpikeDHS}.
Kaiwei Che, Luziwei Leng, Jianguo Zhang 0001, Qinghu Meng, Qinghai Guo, Jianxing Liao
NeurIPS8
2019 Reasoning mechanism: An effective data reduction algorithm for on-line point cloud selective sampling of sculptured surfaces
Jianxing Liao
Comput. Aided Des.4
2018 Benchmarking the GPU memory at the warp level
Minquan Fang, Jianbin Fang, Haifang Zhou, Jianxing Liao, Yuangang Wang
Parallel Comput.5