Liyang Liu

dblp:92/9944 · DBLP profile ↗
← Back
22ranked-venue papers
6as first author
17since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 8 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Design and Comparison of Modular Phase-Unit Yokeless and Segmented Armature Machines
abstract
Axial flux machines (AFMs) have attracted significant attention for their exceptional torque density, while existing AFM topologies lack fault tolerance capability. This paper proposes a series of phase-unit AFMs with yokeless and segmented armature structure (PU-YASA) for direct drive applications. All phases of the machines are physically separated and distributed along the axial direction with a specific mechanical angle between the adjacent phases, meaning that the machine can be easily converted to the multi-phase structure by adjusting the mechanical angle between the adjacent phases. The isolated phase structure avoids the possibility of single-phase short-circuit fault escalating to interphase short-circuit fault. Moreover, YASA structure can perform higher power density due to the large cooling area. The study involves two models, with the key difference being that the angular offset is positioned on the rotor PMs in model I and on the stator side in model II. A simplification method is presented to increase simulation efficiency and reduce the required computational resources. The simulation results of different models are compared to validate the proposed approach.
Zhijun Ou, Wenyuan Mi, Yuanfeng Qu, Liyang Liu, Hang Zhao 0010
IECON5
2024 Amodal Scene Analysis via Holistic Occlusion Relation Inference and Generative Mask Completion
abstract
Amodal scene analysis entails interpreting the occlusion relationship among scene elements and inferring the possible shapes of the invisible parts. Existing methods typically frame this task as an extended instance segmentation or a pair-wise object de-occlusion problem. In this work, we propose a new framework, which comprises a Holistic Occlusion Relation Inference (HORI) module followed by an instance-level Generative Mask Completion (GMC) module. Unlike previous approaches, which rely on mask completion results for occlusion reasoning, our HORI module directly predicts an occlusion relation matrix in a single pass. This approach is much more efficient than the pair-wise de-occlusion process and it naturally handles mutual occlusion, a common but often neglected situation. Moreover, we formulate the mask completion task as a generative process and use a diffusion-based GMC module for instance-level mask completion. This improves mask completion quality and provides multiple plausible solutions. We further introduce a large-scale amodal segmentation dataset with high-quality human annotations, including mutual occlusions. Experiments on our dataset and two public benchmarks demonstrate the advantages of our method. code public available at https://github.com/zbwxp/Amodal-AAAI.
Bowen Zhang 0009, Qing Liu 0017, Jianming Zhang 0001, Yilin Wang 0002, Liyang Liu, Zhe Lin 0001, Yifan Liu 0001
AAAI5
2024 Decomposing Disease Descriptions for Enhanced Pathology Detection: A Multi-Aspect Vision-Language Pre-Training Framework
abstract
Medical vision language pre-training (VLP) has emerged as a frontier of research, enabling zero-shot pathological recognition by comparing the query image with the textual descriptions for each disease. Due to the complex semantics of biomedical texts, current methods struggle to align medical images with key pathological findings in un-structured reports. This leads to the misalignment with the target disease's textual representation. In this paper, we introduce a novel VLP framework designed to dissect disease descriptions into their fundamental aspects, leveraging prior knowledge about the visual manifestations of pathologies. This is achieved by consulting a large language model and medical experts. Integrating a Transformer module, our approach aligns an input image with the diverse elements of a disease, generating aspect-centric image representations. By consolidating the matches from each aspect, we improve the compatibility between an image and its associated disease. Additionally, capitalizing on the aspect-oriented representations, we present a dual-head Transformer tailored to process known and unknown diseases, optimizing the comprehensive detection efficacy. Conducting experiments on seven downstream datasets, ours improves the accuracy of recent methods by up to 8.56% and 17.26% for seen and unseen categories, respectively. Our code is released at https://github.com/HieuPhan33/MAVL.
Vu Minh Hieu Phan, Yutong Xie 0001, Yuankai Qi, Lingqiao Liu, Liyang Liu, Bowen Zhang 0009, Zhibin Liao, Qi Wu 0001, Minh-Son To, Johan Verjans
CVPR5
2024 Design and Optimization of a Dual-Rotor and Flat-Type-Stator Transverse Flux Machine
abstract
The development and application of transverse flux machines (TFMs) with permanent magnet (PM) excitation have received more and more attention in recent years because of the high torque density at low rotating speed applications. This paper proposes a novel topology structure of TFM with dual-rotor and flat-type-stator (DRFTS-TFM). One of the advantages of the proposed structure is that the sandwiched design greatly reduces the manufacturing and assembling difficulties. Another advantage is that the number of the armature phase can be easily changed by converting the mechanical angle of the stator. To investigate the best electromagnetic torque performance of the machine, the multi-objective particle swarm optimization (MOPSO) algorithm is adapted and validated by the finite element method (FEM). The corresponding electromagnetic performances in different conditions are analyzed.
Zhijun Ou, Hang Zhao 0010, Liyang Liu, Xiangdong Su, Hui Wang 0147
IECON3
2024 BPKD: Boundary Privileged Knowledge Distillation For Semantic Segmentation
abstract
Current knowledge distillation approaches in semantic segmentation tend to adopt a holistic approach that treats all spatial locations equally. However, for dense prediction, students’ predictions on edge regions are highly uncertain due to contextual information leakage, requiring higher spatial sensitivity knowledge than the body regions. To address this challenge, this paper proposes a novel approach called boundary-privileged knowledge distillation (BPKD). BPKD distills the knowledge of the teacher model’s body and edges separately to the compact student model. Specifically, we employ two distinct loss functions: (i) edge loss, which aims to distinguish between ambiguous classes at the pixel level in edge regions; (ii) body loss, which utilizes shape constraints and selectively attends to the inner-semantic regions. Our experiments demonstrate that the proposed BPKD method provides extensive refinements and aggregation for edge and body regions. Additionally, the method achieves state-of-the-art distillation performance for semantic segmentation on three popular benchmark datasets, highlighting its effectiveness and generalization ability. BPKD shows consistent improvements across a diverse array of lightweight segmentation structures, including both CNNs and transformers, underscoring its architecture-agnostic adaptability. The code is available at https://github.com/AkideLiu/BPKD.
Liyang Liu, Minh Hieu Phan, Bowen Zhang 0009, Jinchao Ge, Yifan Liu 0001
WACV1
2024 FAKD: Feature Augmented Knowledge Distillation for Semantic Segmentation
abstract
In this work, we explore data augmentations for knowledge distillation on semantic segmentation. Due the capacity gap, small-sized student networks struggle to discover the discriminative feature space learned by a powerful teacher. Image-level augmentations allow the student to better imitate the teacher by providing extra outputs. However, existing distillation frameworks only augment a limited number of samples, which restricts the learning of a student. Inspired by the recent progress on semantic directions on feature space, this work proposes a feature-level augmented knowledge distillation (FAKD) which infinitely augments features along a semantic direction for optimal knowledge transfer. Furthermore, we introduce novel surrogate loss functions to distill the teacher’s knowledge from an infinite number of samples. The surrogate loss is an upper bound of the expected distillation loss over infinite augmented samples. Extensive experiments on four semantic segmentation benchmarks demonstrate that the proposed method boosts the performance of current knowledge distillation methods without any significant overhead. The code will be released at FAKD.
Jianlong Yuan, Minh Hieu Phan, Liyang Liu
WACV3
2024 SegViT v2: Exploring Efficient and Continual Semantic Segmentation with Plain Vision Transformers
abstract
Abstract This paper investigates the capability of plain Vision Transformers (ViTs) for semantic segmentation using the encoder–decoder framework and introduce SegViTv2 . In this study, we introduce a novel Attention-to-Mask (ATM) module to design a lightweight decoder effective for plain ViT. The proposed ATM converts the global attention map into semantic masks for high-quality segmentation results. Our decoder outperforms popular decoder UPerNet using various ViT backbones while consuming only about $$5\%$$ 5 % of the computational cost. For the encoder, we address the concern of the relatively high computational cost in the ViT-based encoders and propose a Shrunk ++ structure that incorporates edge-aware query-based down-sampling (EQD) and query-based up-sampling (QU) modules. The Shrunk++ structure reduces the computational cost of the encoder by up to $$50\%$$ 50 % while maintaining competitive performance. Furthermore, we propose to adapt SegViT for continual semantic segmentation, demonstrating nearly zero forgetting of previously learned knowledge. Experiments show that our proposed SegViTv2 surpasses recent segmentation methods on three popular benchmarks including ADE20k, COCO-Stuff-10k and PASCAL-Context datasets. The code is available through the following link: https://github.com/zbwxp/SegVit .
Bowen Zhang 0009, Liyang Liu, Minh Hieu Phan, Zhi Tian, Chunhua Shen, Yifan Liu 0001
Int. J. Comput. Vis.2
2022 Group R-CNN for Weakly Semi-supervised Object Detection with Points
abstract
We study the problem of weakly semi-supervised object detection with points (WSSOD-P), where the training data is combined by a small set of fully annotated images with bounding boxes and a large set of weakly-labeled images with only a single point annotated for each instance. The core of this task is to train a point-to-box regressor on well-labeled images that can be used to predict credible bounding boxes for each point annotation. We challenge the prior belief that existing CNN-based detectors are not compatible with this task. Based on the classic R-CNN architecture, we propose an effective point-to-box regressor: Group R-CNN. Group R-CNN first uses instance-level proposal grouping to generate a group of proposals for each point annotation and thus can obtain a high recall rate. To better distinguish different instances and improve precision, we propose instance-level proposal assignment to replace the vanilla assignment strategy adopted in original R-CNN methods. As naive instance-level assignment brings converging difficulty, we propose instance aware representation learning which consists of instance aware feature enhancement and instance-aware parameter generation to overcome this issue. Comprehensive experiments on the MS-COCO benchmark demonstrate the effectiveness of our method. Specifically, Group R-CNN significantly outperforms the prior method Point DETR by 3.9 mAP with 5% well-labeled images, which is the most challenging scenario. The source code can be found at https://github.com/jshilong/GroupRCNN.
Zhuoran Yu, Liyang Liu, Xinjiang Wang, Aojun Zhou, Kai Chen 0026
CVPR3
2022 A Framework to Address the Challenges of Surface Mining through Appropriate Sensing and Perception
abstract
The majority of sensing systems in current mining procedures are not optimized to capture the high-fidelity inputs for perception systems. This is due to the lack of accuracy, resolution, update rate, and other shortcomings in the quality and quantity of captured data. High-fidelity input from the multi-modal sensing system is a key component in the development of superior perception technologies. These technologies address critical mining concerns such as performance analysis, progress monitoring, and policy verification for safe operating practices. This paper proposes the non-trivial procedure of proper sensor selection and appropriate deployment of the sensing system to ensure high fidelity multi-modal sensing to serve the above-ground mining applications. Furthermore, the development of perception systems that incorporate this sensing system and machine learning tools is discussed. The mentioned perception systems are mainly used for in-pit visualization, analysis, and decision support applications. Nonetheless, we hope the proposed techniques could be extended to cover complex control applications too. Additionally, we demonstrate the framework for the incorporation of a representative simulated environment, along with the simulated sensors that were chosen through a precise sensor selection analysis, to produce training data for machine learning models. The models are made to be used for processing the real inputs produced by the physical Mobile Sensing Trailer (MST), which is installed in the mining environment. To the best of our knowledge, such simulation environment and its components did not exist or were not accessible before.
Mehala Balamurali, Andrew John Hill, Javier Martinez, Rami N. Khushaba, Liyang Liu, Najmeh Kamyabpour, Ehsan Mihankhah
ICARCV5
2022 Object-Aware Self-Supervised Multi-Label Learning
abstract
Multi-label Learning on image data has been widely exploited with deep learning models. However, supervised training on deep CNN models often cannot discover sufficient discriminative features for classification. As a result, numerous self-supervision methods are proposed to learn more robust image representations. However, most self-supervised approaches focus on single-instance single-label data and fall short on more complex images with multiple objects. Therefore, we propose an Object-Aware Self-Supervision (OASS) method to obtain more fine-grained representations for multi-label learning, dynamically generating auxiliary tasks based on object locations. Secondly, the robust representation learned by OASS can be leveraged to efficiently generate Class-Specific Instances (CSI) in a proposal-free fashion to better guide multi-label supervision signal transfer to instances. Extensive experiments on the VOC2012 dataset for multi-label classification demonstrate the effectiveness of the proposed method against the state-of-the-art counterparts.
Kaixin Xu, Liyang Liu, Ziyuan Zhao, Zeng Zeng, Bharadwaj Veeravalli
ICIP2
2022 GenDet: Meta Learning to Generate Detectors From Few Shots
abstract
Object detection has made enormous progress and has been widely used in many applications. However, it performs poorly when only limited training data is available for novel classes that the model has never seen before. Most existing approaches solve few-shot detection tasks implicitly without directly modeling the detectors for novel classes. In this article, we propose GenDet, a new meta-learning-based framework that can effectively generate object detectors for novel classes from few shots and, thus, conducts few-shot detection tasks explicitly. The detector generator is trained by numerous few-shot detection tasks sampled from base classes each with sufficient samples, and thus, it is expected to generalize well on novel classes. An adaptive pooling module is further introduced to suppress distracting samples and aggregate the detectors generated from multiple shots. Moreover, we propose to train a reference detector for each base class in the conventional way, with which to guide the training of the detector generator. The reference detectors and the detector generator can be trained simultaneously. Finally, the generated detectors of different classes are encouraged to be orthogonal to each other for better generalization. The proposed approach is extensively evaluated on the ImageNet, VOC, and COCO data sets under various few-shot detection settings, and it achieves new state-of-the-art results.
Liyang Liu, Bochao Wang, Zhanghui Kuang, Jing-Hao Xue, Wenming Yang, Qingmin Liao, Wayne Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2021 Pseudo-mask Matters in Weakly-supervised Semantic Segmentation
abstract
Most weakly supervised semantic segmentation (WSSS) methods follow the pipeline that generates pseudo-masks initially and trains the segmentation model with the pseudo-masks in fully supervised manner after. However, we find some matters related to the pseudo-masks, including high quality pseudo-masks generation from class activation maps (CAMs), and training with noisy pseudo-mask supervision. For these matters, we propose the following designs to push the performance to new state-of-art: (i) Coefficient of Variation Smoothing to smooth the CAMs adaptively; (ii) Proportional Pseudo-mask Generation to project the expanded CAMs to pseudo-mask based on a new metric indicating the importance of each class on each location, instead of the scores trained from binary classifiers. (iii) Pretended Under-Fitting strategy to suppress the influence of noise in pseudo-mask; (iv) Cyclic Pseudo-mask to boost the pseudo-masks during training of fully supervised semantic segmentation (FSSS). Experiments based on our methods achieve new state-of-art results on two changeling weakly supervised semantic segmentation datasets, pushing the mIoU to 70.0% and 40.2% on PAS-CAL VOC 2012 and MS COCO 2014 respectively. Codes including segmentation framework are released at https://github.com/Eli-YiLi/PMM
Yi Li 0050, Zhanghui Kuang, Liyang Liu, Wayne Zhang 0001
ICCV3
2021 Towards Impartial Multi-task Learning
Liyang Liu, Yi Li 0050, Zhanghui Kuang, Jing-Hao Xue, Wenming Yang, Qingmin Liao, Wayne Zhang 0001
ICLR1
2021 Group Fisher Pruning for Practical Network Compression
abstract
Network compression has been widely studied since it is able to reduce the memory and computation cost during inference. However, previous methods seldom deal with complicated structures like residual connections, group/depth-wise convolution and feature pyramid network, where channels of multiple layers are coupled and need to be pruned simultaneously. In this paper, we present a general channel pruning approach that can be applied to various complicated structures. Particularly, we propose a layer grouping algorithm to find coupled channels automatically. Then we derive a unified metric based on Fisher information to evaluate the importance of a single channel and coupled channels. Moreover, we find that inference speedup on GPUs is more correlated with the reduction of memory rather than FLOPs, and thus we employ the memory reduction of each channel to normalize the importance. Our method can be used to prune any structures including those with coupled channels. We conduct extensive experiments on various backbones, including the classic ResNet and ResNeXt, mobile-friendly MobileNetV2, and the NAS-based RegNet, both on image classification and object detection which is under-explored. Experimental results validate that our method can effectively prune sophisticated networks, boosting inference speed without sacrificing accuracy.
Liyang Liu, Zhanghui Kuang, Aojun Zhou, Jing-Hao Xue, Xinjiang Wang, Wenming Yang, Qingmin Liao, Wayne Zhang 0001
ICML1
2021 Seasonal Variability of GPP and Phenology in Remote Sensed Observations and Land Surface Models
abstract
Surface carbon fluxes associated with terrestrial vegetation play a key role in the global carbon cycle. Remote sensing (RS) and land surface models (LSM) have demonstrated to be valuable tools in assessing the gross primary production (GPP). Yet, the seasonal variability of this flux, and timing of the seasonal cycle remain challenging to observe and simulate accurately. Here, the ability of four RS products and two LSM to simulate GPP and its variability was assessed. It was found that the mean seasonal GPP was simulated accurately with RS-based models, but that it failed to capture the seasonal variability and timing of the seasonal cycle. In contrast, the LSMs demonstrated their ability to simulate seasonal anomalies in GPP, but had trouble to simulate the associated phenology.
Jan De Pue, Sebastian Wieneke, José Miguel Barrios, Liyang Liu, Maral Maleki, Philippe Ciais, Alirio Arboleda, Rafiq Hamdi, Ana Bastos, Ivan A. Janssens, Fabienne Maignan, Françoise Gellens-Meulenberghs, Manuela Balzarolo
IGARSS4
2021 Diverse Message Passing for Attribute with Heterophily
abstract
Most of the existing GNNs can be modeled via the Uniform Message Passing framework. This framework considers all the attributes of each node in its entirety, shares the uniform propagation weights along each edge, and focuses on the uniform weight learning. The design of this framework possesses two prerequisites, the simplification of homophily and heterophily to the node-level property and the ignorance of attribute differences. Unfortunately, different attributes possess diverse characteristics. In this paper, the network homophily rate defined with respect to the node labels is extended to attribute homophily rate by taking the attributes as weak labels. Based on this attribute homophily rate, we propose a Diverse Message Passing (DMP) framework, which specifies every attribute propagation weight on each edge. Besides, we propose two specific strategies to significantly reduce the computational complexity of DMP to prevent the overfitting issue. By investigating the spectral characteristics, existing spectral GNNs are actually equivalent to a degenerated version of DMP. From the perspective of numerical optimization, we provide a theoretical analysis to demonstrate DMP's powerful representation ability and the ability of alleviating the over-smoothing issue. Evaluations on various real networks demonstrate the superiority of our DMP on handling the networks with heterophily and alleviating the over-smoothing issue, compared to the existing state-of-the-arts.
Liang Yang 0002, Mengzhe Li, Liyang Liu, Bingxin Niu, Chuan Wang 0002, Xiaochun Cao, Yuanfang Guo
NeurIPS3
2021 IncDet: In Defense of Elastic Weight Consolidation for Incremental Object Detection
abstract
Elastic weight consolidation (EWC) has been successfully applied for general incremental learning to overcome the catastrophic forgetting issue. It adaptively constrains each parameter of the new model not to deviate much from its counterpart in the old model during fine-tuning on new class data sets, according to its importance weight for old tasks. However, the previous study demonstrates that it still suffers from catastrophic forgetting when directly used in object detection. In this article, we show EWC is effective for incremental object detection if with critical adaptations. First, we conduct controlled experiments to identify two core issues why EWC fails if trivially applied to incremental detection: 1) the absence of old class annotations in new class images makes EWC misclassify objects of old classes in these images as background and 2) the quadratic regularization loss in EWC easily leads to gradient explosion when balancing old and new classes. Then, based on the abovementioned findings, we propose the corresponding solutions to tackle these issues: 1) utilize pseudobounding box annotations of old classes on new data sets to compensate for the absence of old class annotations and 2) adopt a novel Huber regularization instead of the original quadratic loss to prevent from unstable training. Finally, we propose a general EWC-based incremental object detection framework and implement it under both Fast R-CNN and Faster R-CNN, showing its flexibility and versatility. In terms of either the final performance or the performance drop with respect to the upper bound of joint training on all seen classes, evaluations on the PASCAL VOC and COCO data sets show that our method achieves a new state of the art.
Liyang Liu, Zhanghui Kuang, Jing-Hao Xue, Wenming Yang, Wayne Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2018 Parallax Bundle Adjustment on Manifold with Improved Global Initialization
Liyang Liu, Teng Zhang 0003, Brenton Leighton, Liang Zhao 0003, Shoudong Huang, Gamini Dissanayake
WAFR1
2018 Distributed-microphones based in-vehicle speech enhancement via sparse and low-rank spectrogram decomposition
Xuliang Li 0002, Liyang Liu, Weifeng Li 0001
Speech Commun.3
2016 A novel approach to steel rivet detection in poorly illuminated steel structural environments
abstract
It is becoming increasingly achievable for steel bridge structures, which are normally both inaccessible and hazardous for humans, to be inspected and maintained by autonomous robots. Steel bridges have been traditionally constructed by securing plate members together with rivets. However, rivets present a challenge for robots both in terms of cleaning and surface traversal. This paper presents a novel approach to RGB-D image and point cloud analysis that enables rivets to be rapidly and robustly located using low cost, non-contact sensing devices that can be easily affixed to a robot. The approach performs classification based on: (a) high-intensity blobs in color images, (b) the non-linear perturbations in depth images, and (c) surface normal clusters in 3D point clouds. The predicted rivet locations from the three classifiers are combined using a probabilistic occupancy mapping technique. Experiments are conducted in several different lab and real-world steel bridge environments, where there is no external lighting infrastructure, and the sensors are attached to a mobile platform, i.e. a climbing inspection robot. The location of rivets within 2m of the robot can be robustly located within 10mm of their correct location. The state of voxels can be predicted with above 95% accuracy, in approximately 1 second per frame.
Gavin Paul, Liyang Liu, Dikai Liu
ICARCV2
2015 Energy-efficient node selection and power control in cooperative spectrum sensing
abstract
In cooperative spectrum sensing, wireless reporting channels (from local sensing nodes to fusion center) may suffer severe unreliability, which would make a correct local sensing result incorrect while received by fusion center. In this study, the reliability of the sensing result transmission from local sensor to fusion center is considered, and to make spectrum sensing more energy efficient, appropriate nodes were selected and activated to participate in spectrum sensing while others remain idle. To further reduce energy consumption, transmission power control for sensing nodes was optimized. Total energy consumption minimization was modeled as a mixed discrete and continuous variable optimization problem and the binary particle swarm optimization with power control (BPSO-PC) was proposed. BPSO-PC adjusted the sensing node transmission power and properly selects nodes for cooperative spectrum sensing. Simulation results showed that total energy consumption was significantly reduced compared with BPSO and three other sensing nodes selection algorithms.
Liyang Liu, Zhaowei Qu, Sixing Yin
PIMRC2
2014 Smart hoist: An assistive robot to aid carers
abstract
Assistive Robotics(AR) is a rapidly expanding field, implementing advanced intelligent machines capable of working collaboratively with a range of human users; as assistants, tools and as companions. These AR devices can provide assistance to stretched carers when transferring non-ambulatory patients safely. This paper presents the preliminary outcomes of the design, development and implementation of a patient lifting AR device, Smart Hoist. This device, an enhanced conventional patient lifter (standard hoist), is fitted with several sensors capable of interacting with the device operator and its environment, and a set of powered wheels. The assisted manoeuvring functionality of the Smart Hoist may help reduce prevailing lower back injuries among the carers while improving the safety of carers and patients. Results collected from an evaluation of the preliminary version of the Smart Hoist conducted at the premises of IRT Woonona residential care facility confirms the system is easy to use and it reduces the effort of the operator, which may help in reducing lower back injuries.
Ravindra Ranasinghe, Lakshitha Dantanarayana, Antony Tran, Stefan Lie, Michael Behrens, Liyang Liu
ICARCV6