Jianxiong Yin

dblp:132/8514 · also Jianxiong (Terry) Yin · DBLP profile ↗
← Back
22ranked-venue papers
2as first author
14since 2021 · last 2025
0000-0003-4686-2768ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 since 2021Computer networks · 3 · 2 first-authorSystems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Masked Sensory-Temporal Attention for Sensor Generalization in Quadruped Locomotion
abstract
With the rising focus on quadrupeds, a generalized policy capable of handling different robot models and sensor inputs becomes highly beneficial. Although several methods have been proposed to address different morphologies, it remains a challenge for learning-based policies to manage various combinations of proprioceptive information. This paper presents Masked Sensory-Temporal Attention (MSTA), a novel transformer-based mechanism with masking for quadruped locomotion. It employs direct sensor-level attention to enhance the sensory-temporal understanding and handle different combinations of sensor data, serving as a foundation for incorporating unseen information. MSTA can effectively understand its states even with a large portion of missing information, and is flexible enough to be deployed on physical systems despite the long input sequence.
Dikai Liu, Tianwei Zhang 0004, Jianxiong Yin, Simon See
ICRA3
2025 Unified Locomotion Transformer with Simultaneous Sim-to-Real Transfer for Quadrupeds
abstract
Quadrupeds have gained rapid advancement in their capability of traversing across complex terrains. The adoption of deep Reinforcement Learning (RL), transformers and various knowledge transfer techniques can greatly reduce the sim-to-real gap. However, the classical teacher-student framework commonly used in existing locomotion policies requires a pre-trained teacher and leverages the privilege information to guide the student policy. With the implementation of large-scale models in robotics controllers, especially transformers-based ones, this knowledge distillation technique starts to show its weakness in efficiency, due to the requirement of multiple supervised stages. In this paper, we propose Unified Locomotion Transformer (ULT), a new transformer-based framework to unify the processes of knowledge transfer and policy optimization in a single network while still taking advantage of privilege information. The policies are optimized with reinforcement learning, next state-action prediction, and action imitation, all in just one training stage, to achieve zero-shot deployment. Evaluation results demonstrate that with ULT, optimal teacher and student policies can be obtained at the same time, greatly easing the difficulty in knowledge transfer, even with complex transformer-based models.
Dikai Liu, Tianwei Zhang 0004, Jianxiong Yin, Simon See
IROS3
2025 Replay Master: Automatic Sample Selection and Effective Memory Utilization for Continual Semantic Segmentation
abstract
Continual Semantic Segmentation (CSS) extends static semantic segmentation by incrementally introducing new classes for training. To alleviate the catastrophic forgetting issue in this task, replay methods can be adopted, constructing a memory buffer that stores a small number of samples from previous classes for future replay. However, existing replay approaches in CSS often lack a thorough exploration of two critical issues: how to find the most suitable memory samples and how to utilize them for replay more effectively. Common strategies either randomly select samples or rely on hand-crafted, single-factor-driven methods that are hard to be optimal, and often employ conventional training techniques for replay that do not account for class imbalance problem resulting from limited memory capacity. In this work, we tackle these challenges by introducing a novel memory sample selection method that leverages a reinforcement learning framework with innovative state representations and a dual-stage action scheme to automatically learn a selection policy. Additionally, we propose an expert mechanism and a dual-phase training method to address the class imbalance issue, thereby enhancing the effectiveness of replay training by making better use of memory samples. Incorporating the proposed automatic sample selection and effective memory utilization methods, we develop a novel and effective replay-based pipeline for CSS. Our extensive experiments on Pascal VOC 2012 and ADE20 K datasets demonstrate the effectiveness of our approach, which achieves state-of-the-art (SOTA) performance and outperforms previous advanced methods significantly.
Lanyun Zhu, Tianrun Chen, Jianxiong Yin, Simon See, De Wen Soh, Jun Liu 0036
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Addressing Background Context Bias in Few-Shot Segmentation Through Iterative Modulation
abstract
Existing few-shot segmentation methods usually extract foreground prototypes from support images to guide query image segmentation. However, different background contexts of support and query images can cause their foreground features to be misaligned. This phenomenon, known as background context bias, can hinder the effectiveness of support prototypes in guiding query image segmentation. In this work, we propose a novel framework with an it-erative structure to address this problem. In each iteration of the framework, we first generate a query prediction based on a support foreground feature. Next, we extract background context from the query image to modulate the support foreground feature, thus eliminating the foreground feature misalignment caused by the different backgrounds. After that, we design a confidence-biased attention to eliminate noise and cleanse information. By integrating these components through an iterative structure, we create a novel network that can leverage the synergies between different modules to improve their performance in a mutually reinforcing manner. Through these carefully designed components and structures, our network can effectively elimi-nate background context bias in few-shot segmentation, thus achieving outstanding performance. We conduct extensive experiments on the PASCAL-5iand COCO-20idatasets and achieve state-of-the-art (SOTA) results, which demonstrate the effectiveness of our approach.
Lanyun Zhu, Tianrun Chen, Jianxiong Yin, Simon See, Jun Liu 0036
CVPR3
2024 Enhancing Modality Representation and Alignment for Multimodal Cold-start Active Learning
abstract
Training multimodal models requires a large amount of labeled data. Active learning (AL) aim to reduce labeling costs. Most AL methods employ warm-start approaches, which rely on sufficient labeled data to train a well-calibrated model that can assess the uncertainty and diversity of unlabeled data. However, when assembling a dataset, labeled data are often scarce initially, leading to a cold-start problem. Additionally, most AL methods seldom address multimodal data, highlighting a research gap in this field. Our research addresses these issues by developing a two-stage method for Multi-Modal Cold-Start Active Learning (MMCSAL). Firstly, we observe the modality gap, a significant distance between the centroids of representations from different modalities, when only using cross-modal pairing information as self-supervision signals. This modality gap affects data selection process, as we calculate both uni-modal and cross-modal distances. To address this, we introduce uni-modal prototypes to bridge the modality gap. Secondly, conventional AL methods often falter in multimodal scenarios where alignment between modalities is overlooked. Therefore, we propose enhancing cross-modal alignment through regularization, thereby improving the quality of selected multimodal data pairs in AL. Finally, our experiments demonstrate MMCSAL's efficacy in selecting multimodal data pairs across three multimodal datasets.
Meng Shen 0002, Yake Wei, Jianxiong Yin, Deepu Rajan, Di Hu 0001, Simon See
MMAsia3
2024 Going Deeper into Recognizing Actions in Dark Environments: A Comprehensive Benchmark Study
Yuecong Xu, Haozhi Cao, Jianxiong Yin, Zhenghua Chen, Xiaoli Li 0001, Zhengguo Li, Qianwen Xu 0001, Jianfei Yang 0001
Int. J. Comput. Vis.3
2024 Self-Supervised Video Representation Learning by Video Incoherence Detection
abstract
This article introduces a novel self-supervised method that leverages incoherence detection for video representation learning. It stems from the observation that the visual system of human beings can easily identify video incoherence based on their comprehensive understanding of videos. Specifically, we construct the incoherent clip by multiple subclips hierarchically sampled from the same raw video with various lengths of incoherence. The network is trained to learn the high-level representation by predicting the location and length of incoherence given the incoherent clip as input. Additionally, we introduce intravideo contrastive learning to maximize the mutual information between incoherent clips from the same raw video. We evaluate our proposed method through extensive experiments on action recognition and video retrieval using various backbone networks. Experiments show that our proposed method achieves remarkable performance across different backbone networks and different datasets compared to previous coherence-based methods.
Haozhi Cao, Yuecong Xu, Kezhi Mao, Lihua Xie 0001, Jianxiong Yin, Simon See, Qianwen Xu 0001, Jianfei Yang 0001
IEEE Trans. Cybern.5
2023 Continual Semantic Segmentation with Automatic Memory Sample Selection
abstract
Continual Semantic Segmentation (CSS) extends static semantic segmentation by incrementally introducing new classes for training. To alleviate the catastrophic forgetting issue in CSS, a memory buffer that stores a small number of samples from the previous classes is constructed for replay. However, existing methods select the memory samples either randomly or based on a single-factor-driven handcrafted strategy, which has no guarantee to be optimal. In this work, we propose a novel memory sample selection mechanism that selects informative samples for effective replay in a fully automatic way by considering comprehensive factors including sample diversity and class performance. Our mechanism regards the selection operation as a decision-making process and learns an optimal selection policy that directly maximizes the validation performance on a reward set. To facilitate the selection decision, we design a novel state representation and a dual-stage action space. Our extensive experiments on Pascal-VOC 2012 and ADE 20K datasets demonstrate the effectiveness of our approach with state-of-the-art (SOTA) performance achieved, outperforming the second-place one by 12.54% for the 6-stage setting on Pascal-VOC 2012.
Lanyun Zhu, Tianrun Chen, Jianxiong Yin, Simon See, Jun Liu 0036
CVPR3
2023 Learning Gabor Texture Features for Fine-Grained Recognition
abstract
Extracting and using class-discriminative features is critical for fine-grained recognition. Existing works have demonstrated the possibility of applying deep CNNs to exploit features that distinguish similar classes. However, CNNs suffer from problems including frequency bias and loss of detailed local information, which restricts the performance of recognizing fine-grained categories. To address the challenge, we propose a novel texture branch as complimentary to the CNN branch for feature extraction. We innovatively utilize Gabor filters as a powerful extractor to exploit texture features, motivated by the capability of Gabor filters in effectively capturing multi-frequency features and detailed local information. We implement several designs to enhance the effectiveness of Gabor filters, including imposing constraints on parameter values and developing a learning method to determine the optimal parameters. Moreover, we introduce a statistical feature extractor to utilize informative statistical information from the signals captured by Gabor filters, and a gate selection mechanism to enable efficient computation by only considering qualified regions as input for texture extraction. Through the integration of features from the Gabor-filter-based texture branch and CNN-based semantic branch, we achieve comprehensive information extraction. We demonstrate the efficacy of our method on multiple datasets, including CUB-200-2011, NA-bird, Stanford Dogs, and GTOS-mobile. State-of-the-art performance is achieved using our approach.
Lanyun Zhu, Tianrun Chen, Jianxiong Yin, Simon See, Jun Liu 0036
ICCV3
2023 Towards Balanced Active Learning for Multimodal Classification
abstract
Training multimodal networks requires a vast amount of data due to their larger parameter space compared to unimodal networks. Active learning is a widely used technique for reducing data annotation costs by selecting only those samples that could contribute to improving model performance. However, current active learning strategies are mostly designed for unimodal tasks, and when applied to multimodal data, they often result in biased sample selection from the dominant modality. This unfairness hinders balanced multimodal learning, which is crucial for achieving optimal performance. To address this issue, we propose three guidelines for designing a more balanced multimodal active learning strategy. Following these guidelines, a novel approach is proposed to achieve more fair data selection by modulating the gradient embedding with the dominance degree among modalities. Our studies demonstrate that the proposed method achieves more balanced multimodal learning by avoiding greedy sample selection from the dominant modality. Our approach outperforms existing active learning strategies on a variety of multimodal classification tasks. Overall, our work highlights the importance of balancing sample selection in multimodal active learning and provides a practical solution for achieving more balanced active learning for multimodal classification.
Meng Shen 0002, Yizheng Huang 0001, Jianxiong Yin, Heqing Zou, Deepu Rajan, Simon See
ACM Multimedia3
2021 ACT: an Attentive Convolutional Transformer for Efficient Text Classification
abstract
Recently, Transformer has been demonstrating promising performance in many NLP tasks and showing a trend of replacing Recurrent Neural Network (RNN). Meanwhile, less attention is drawn to Convolutional Neural Network (CNN) due to its weak ability in capturing sequential and long-distance dependencies, although it has excellent local feature extraction capability. In this paper, we introduce an Attentive Convolutional Transformer (ACT) that takes the advantages of both Transformer and CNN for efficient text classification. Specifically, we propose a novel attentive convolution mechanism that utilizes the semantic meaning of convolutional filters attentively to transform text from complex word space to a more informative convolutional filter space where important n-grams are captured. ACT is able to capture both local and global dependencies effectively while preserving sequential information. Experiments on various text classification tasks and detailed analyses show that ACT is a lightweight, fast, and effective universal text classifier, outperforming CNNs, RNNs, and attentive models including Transformer.
Peixiang Zhong, Kezhi Mao, Dongzhe Wang, Xuefeng Yang, Jianxiong Yin, Simon See
AAAI7
2021 Exploiting inter-frame regional correlation for efficient action recognition
Yuecong Xu, Jianfei Yang 0001, Kezhi Mao, Jianxiong Yin, Simon See
Expert Syst. Appl.4
2021 PNL: Efficient long-range dependencies extraction with pyramid non-local module for action recognition
Yuecong Xu, Haozhi Cao, Jianfei Yang 0001, Kezhi Mao, Jianxiong Yin, Simon See
Neurocomputing5
2021 Effective action recognition with embedded key point shifts
Haozhi Cao, Yuecong Xu, Jianfei Yang 0001, Kezhi Mao, Jianxiong Yin, Simon See
Pattern Recognit.5
2020 MLModelCI: An Automatic Cloud Platform for Efficient MLaaS
abstract
MLModelCI provides multimedia researchers and developers with a one-stop platform for efficient machine learning (ML) services. The system leverages DevOps techniques to optimize, test, and manage models. It also containerizes and deploys these optimized and validated models as cloud services (MLaaS). In its essence, MLModelCI serves as a housekeeper to help users publish models. The models are first automatically converted to optimized formats for production purpose and then profiled under different settings (e.g., batch size and hardware). The profiling information can be used as guidelines for balancing the trade-off between performance and cost of MLaaS. Finally, the system dockerizes the models for ease of deployment to cloud environments. A key feature of MLModelCI is the implementation of a controller, which allows elastic evaluation which only utilizes idle workers while maintaining online service quality. Our system bridges the gap between current ML training and serving systems and thus free developers from manual and tedious work often associated with service deployment. We release the platform as an open-source project on GitHub under Apache 2.0 license, with the aim that it will facilitate and streamline more large-scale ML applications and research projects.
Huaizheng Zhang, Yuanming Li, Yizheng Huang 0001, Yonggang Wen 0001, Jianxiong Yin, Kyle Guan
ACM Multimedia5
2019 DeepHunter: a coverage-guided fuzz testing framework for deep neural networks
abstract
The past decade has seen the great potential of applying deep neural network (DNN) based software to safety-critical scenarios, such as autonomous driving. Similar to traditional software, DNNs could exhibit incorrect behaviors, caused by hidden defects, leading to severe accidents and losses. In this paper, we propose DeepHunter, a coverage-guided fuzz testing framework for detecting potential defects of general-purpose DNNs. To this end, we first propose a metamorphic mutation strategy to generate new semantically preserved tests, and leverage multiple extensible coverage criteria as feedback to guide the test generation. We further propose a seed selection strategy that combines both diversity-based and recency-based seed selection. We implement and incorporate 5 existing testing criteria and 4 seed selection strategies in DeepHunter. Large-scale experiments demonstrate that (1) our metamorphic mutation strategy is useful to generate new valid tests with the same semantics as the original seed, by up to a 98% validity ratio; (2) the diversity-based seed selection generally weighs more than recency-based seed selection in boosting the coverage and in detecting defects; (3) DeepHunter outperforms the state of the arts by coverage as well as the quantity and diversity of defects identified; (4) guided by corner-region based criteria, DeepHunter is useful to capture defects during the DNN quantization for platform migration.
Xiaofei Xie, Lei Ma 0003, Felix Juefei-Xu, Minhui Xue 0001, Hongxu Chen 0001, Yang Liu 0003, Jianjun Zhao 0001, Bo Li 0026, Jianxiong Yin, Simon See
ISSTA9
2019 Improving Deep Lesion Detection Using 3D Contextual and Spatial Attention
Qingyi Tao, ZongYuan Ge, Jianfei Cai 0001, Jianxiong Yin, Simon See
MICCAI (6)4
2018 Stochastic Downsampling for Cost-Adjustable Inference and Improved Regularization in Convolutional Networks
abstract
It is desirable to train convolutional networks (CNNs) to run more efficiently during inference. In many cases however, the computational budget that the system has for inference cannot be known beforehand during training, or the inference budget is dependent on the changing real-time resource availability. Thus, it is inadequate to train just inference-efficient CNNs, whose inference costs are not adjustable and cannot adapt to varied inference budgets. We propose a novel approach for cost-adjustable inference in CNNs - Stochastic Downsampling Point (SDPoint). During training, SDPoint applies feature map downsampling to a random point in the layer hierarchy, with a random downsampling ratio. The different stochastic downsampling configurations known as SDPoint instances (of the same model) have computational costs different from each other, while being trained to minimize the same prediction loss. Sharing network parameters across different instances provides significant regularization boost. During inference, one may handpick a SDPoint instance that best fits the inference budget. The effectiveness of SDPoint, as both a cost-adjustable inference approach and a regularizer, is validated through extensive experiments on image classification.
Jason Kuen, Xiangfei Kong, Zhe Lin 0001, Gang Wang 0012, Jianxiong Yin, Simon See, Yap-Peng Tan
CVPR5
2018 Fast MPEG-CDVS Encoder With GPU-CPU Hybrid Computing
abstract
The compact descriptors for visual search (CDVS) standard from ISO/IEC moving pictures experts group has succeeded in enabling the interoperability for efficient and effective image retrieval by standardizing the bitstream syntax of compact feature descriptors. However, the intensive computation of a CDVS encoder unfortunately hinders its widely deployment in industry for large-scale visual search. In this paper, we revisit the merits of low complexity design of CDVS core techniques and present a very fast CDVS encoder by leveraging the massive parallel execution resources of graphics processing unit (GPU). We elegantly shift the computation-intensive and parallel-friendly modules to the state-of-the-arts GPU platforms, in which the thread block allocation as well as the memory access mechanism are jointly optimized to eliminate performance loss. In addition, those operations with heavy data dependence are allocated to CPU for resolving the extra but non-necessary computation burden for GPU. Furthermore, we have demonstrated the proposed fast CDVS encoder can work well with those convolution neural network approaches which enables to leverage the advantages of GPU platforms harmoniously, and yield significant performance improvements. Comprehensive experimental results over benchmarks are evaluated, which has shown that the fast CDVS encoder using GPU-CPU hybrid computing is promising for scalable visual search.
Ling-Yu Duan, Wei Sun 0029, Xinfeng Zhang 0001, Shiqi Wang 0001, Jie Chen 0006, Jianxiong Yin, Simon See, Tiejun Huang 0001, Alex Chichung Kot, Wen Gao 0001
IEEE Trans. Image Process.6
2015 Transient data delivery using fine-grained mobility data in spontaneous smartphone networks
abstract
Abstract The commercial success of smartphones increases the feasibility of mobile ad hoc networking in daily life; we define such networks as spontaneous smartphone networks (SSNs). Efficient data delivery in SSNs is challenging because of the low node density, ambiguous contact opportunities, and short message lifetime. The existing schemes attempt to select optimal relays via various cumulative metrics (e.g., encounter history, social centrality, or contact distribution), whose effectiveness is ambiguous and suboptimal. In this paper, we introduce a Markov predictor‐based transient delivery scheme that quantifies the regularity of small time scale movement for forwarding decisions. Unlike previous works, we utilized fine‐grained mobility data to reduce errors of estimating contact opportunities and contact duration. On the basis of this forwarding strategy, we developed a multi‐copy routing scheme. The evaluation using real traces indicates that the proposed approach outperforms compared alternatives in terms of delivery rate and cost. Copyright © 2013 John Wiley & Sons, Ltd.
Jianxiong Yin, Yohan Chon, Elmurod Talipov, Hojung Cha
Wirel. Commun. Mob. Comput.1
2014 A context-rich and extensible framework for spontaneous smartphone networking
Elmurod Talipov, Jianxiong Yin, Yohan Chon, Hojung Cha
Comput. Commun.2
2013 Cloud3DView: an interactive tool for cloud data center operations
abstract
The emergence of cloud computing has promoted growing demand and rapid deployment of data centers. However, data center operations require a set of sophisticated skills (e.g., command-line-interface), resulting in a high operational cost. In this demo, to reduce the data center operational cost, we design and build a novel cloud data center management system, based on the concept of 3D gamification. In particular, we apply data visualization techniques to overlay operational status upon a data center 3D model, allowing the operators to monitor the real-time situation and control the data center from a friendly user interface. This demo highlights: (1)a data center 3D view from a First Person Shooter (FPS) camera, (2)a run-time presentation of visualized infrastructures information. Moreover, to improve the user experience, we employ cutting-edge HCI technologies from multi-touch, for remote access to Cloud3DView.
Jianxiong Yin, Peng Sun 0006, Yonggang Wen 0001, Hai-gang Gong, Ming Liu 0002, Xuelong Li 0001, Haipeng You, Jinqi Gao, Cynthia Lin
SIGCOMM1