Boxiao Liu

dblp:188/2274 · DBLP profile ↗
← Back
20ranked-venue papers
2as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 9 since 2021Systems, architecture and hardware · 6 · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 EPLoN: Exploiting Efficient Parallelism with Selective Rematerialization for Lightning Attention on Ascend NPU
abstract
The quadratic computational complexity of softmax attention presents a fundamental bottleneck to scaling modern language models to long sequences. While the proposed Lightning Attention mechanism offers a linear-complexity alternative, its state-of-the-art implementations remain predominantly optimized for GPU architectures and fail to fully leverage the capabilities of alternative accelerators such as Ascend NPUs. To bridge this gap, we propose EPLoN (Exploiting Efficient Parallelism with Selective Rematerialization for Lightning Attention on NPU). EPLoN presents a high-performance implementation of Lightning Attention optimized for heterogeneous Ascend NPUs. EPLoN reformulates the algorithm, introducing an efficient parallelism scheme with a rematerialization strategy based on inter- and intra-core that maximizes the utilization of the NPU architecture. In a cross-architectural comparison against the state-of-the-art FlashLinearAttention (FLA) on an Nvidia GPU of comparable computational capacity, our evaluation achieves a speedup of up to 3.39 × and a geometric mean speedup of 1.73 ×, while reducing peak memory consumption by approximately 33%.
Zhenfeng Su, Alexander Setyaev, Stanislav Kamenev, Alexander Gneushev, Junmin Xiao, Anastasiya Bistrigova, Sergey Buzykanov, Evgeny Tetin, Guangming Tan, Boxiao Liu, Xueyi Zou, Zhenhua Dong, Constantine Korikov, Xianzhi Yu, Zhongzhe Hu
ICS13
2026 A 10-MS/s Sub-0.1-mW Power Scalable SAR ADC with 10/12-bit Reconfigurable Resolution for Portable Ultrasound Systems
Jieming Ni, Yuhao Mao, Aitong Gong, Boxiao Liu
ISCAS6
2025 See Further When Clear: Curriculum Consistency Model
abstract
Significant advances have been made in the sampling efficiency of diffusion and flow matching models, driven by Consistency Distillation (CD), which trains a student model to mimic the output of a teacher model at a later timestep. However, we found that the knowledge discrepancy between student and teacher varies significantly across different timesteps, leading to suboptimal performance in CD. To address this issue, we propose the Curriculum Consistency Model (CCM), which stabilizes and balances the knowledge discrepancy across timesteps. Specifically, we regard the distillation process at each timestep as a curriculum and introduce a metric based on the Peak Signal-to-Noise Ratio (PSNR) to quantify the knowledge discrepancy of this curriculum, then ensure that the curriculum maintains consistent knowledge discrepancy across different timesteps by having the teacher model iterate more steps when the noise intensity is low. Our method achieves competitive single-step sampling Fréchet Inception Distance (FID) scores of 1.64 on CIFAR-10 and 2.18 on ImageNet 64x64. Moreover, we have extended our method to large-scale text-to-image models and confirmed that it generalizes well to both diffusion models (Stable Diffusion XL) and flow matching models (Stable Diffusion 3). The generated samples demonstrate improved image-text alignment and semantic structure since CCM enlarges the distillation step at large timesteps and reduces the accumulated error.
Boxiao Liu, Yi Zhang 0108, Xingzhong Hou, Guanglu Song, Yu Liu 0015, Haihang You
CVPR2
2025 MM-instruct: Generated visual instructions for large multimodal model alignment
abstract
This paper presents MM-Instruct, an automated pipeline for generating diverse and high-quality visual instruction data to better align large multimodal models (LMMs) with real-world use cases. While previous works have focused on question-answering data, their generated instruction datasets pose challenges for broader application scenarios. Additionally, manually collecting diverse instruction data at scale from users is prohibitively costly. To mitigate these issues, MM-Instruct leverages ChatGPT to automatically generate diverse instructions from a limited set of seed instructions through augmentation and summarization. It then uses an open-sourced large language model (LLM) to construct instruction-following answers and builds a large-scale visual instruction dataset. Evaluating the LLaVA-Instruct models trained with the generated data shows significant improvements in instruction-following capabilities compared to LLaVA-1.5 models. MM-Instruct releases its synthetic dataset containing diverse instructions and high-quality instruction-answer pairs to support training LMMs for real-world applications.
Xin Huang 0027, Jihao Liu, Jinliang Zheng, Boxiao Liu, Jia Wang 0025, Yu Liu 0015, Hongsheng Li 0001, Osamu Yoshie
Neurocomputing4
2024 EasyDrag: Efficient Point-Based Manipulation on Diffusion Models
abstract
Generative models are gaining increasing popularity, and the demand for precisely generating images is on the rise. However, generating an image that perfectly aligns with users' expectations is extremely challenging. The shapes of objects, the poses of animals, the structures of landscapes, and more may not match the user's desires, and this applies to real images as well. This is where point-based image editing becomes essential. An excellent image editing method needs to meet the following criteria: user-friendly interaction, high performance, and good generalization capability. Due to the limitations of StyleGAN, DragGAN exhibits limited robustness across diverse scenarios, while DragDiffusion lacks user-friendliness due to the necessity of LoRA fine-tuning and masks. In this paper, we introduce a novel interactive point-based image editing framework, called EasyDrag, that leverages pretrained diffusion models to achieve high-quality editing outcomes and user-friendship. Extensive experimentation demonstrates that our approach surpasses DragDiffusion in terms of both image quality and editing precision for point-based image manipulation tasks. The code will be available on https://github.com/Ace-Pegasus/EasyDrag.
Xingzhong Hou, Boxiao Liu, Yi Zhang 0108, Jihao Liu, Yu Liu 0015, Haihang You
CVPR2
2024 Fast Subgraph Matching by Dynamic Graph Editing
abstract
Subgraph matching is a challenging NP-complete problem that involves finding identical subgraphs of a query graph$q$in a larger data graph$G$. It has numerous applications in diverse fields, including social and biological networks. However, existing subgraph matching algorithms assume that the graph structure is fixed, which limits their performance in solving more difficult matching cases. To address this issue, we propose a novel approach called Dynamic Graph Editing (DGE), which dynamically edits the query graph to optimize the subgraph matching algorithm. Based on this approach, we introduce an efficient enumeration method called Dynamic Graph Editing Enumeration, which significantly improves the performance of the algorithm. Our experimental results show that DGE outperforms current state-of-the-art algorithms in terms of computational efficiency and ability to solve more complex subgraph matching cases.
Zite Jiang, Shuai Zhang 0040, Boxiao Liu, Xingzhong Hou, Mengting Yuan 0001, Haihang You
IEEE Trans. Serv. Comput.3
2023 UniKD: Universal Knowledge Distillation for Mimicking Homogeneous or Heterogeneous Object Detectors
abstract
Knowledge distillation (KD) has become a standard method to boost the performance of lightweight object detectors. Most previous works are feature-based, where students mimic the features of homogeneous teacher detectors. However, distilling the knowledge from the heterogeneous teacher fails in this manner due to the serious semantic gap, which greatly limits the flexibility of KD in practical applications. Bridging this semantic gap now requires case-by-case algorithm design which is time-consuming and heavily relies on experienced adjustment. To alleviate this problem, we propose Universal Knowledge Distillation (UniKD), introducing additional decoder heads with deformable cross-attention called Adaptive Knowledge Extractor (AKE). In UniKD, AKEs are first pretrained on the teacher’s output to infuse the teacher’s content and positional knowledge into a fixed-number set of knowledge embeddings. The fixed AKEs are then attached to the student’s backbone to encourage the student to absorb the teacher’s knowledge in these knowledge embeddings. In this query-based distillation paradigm, detection-relevant information can be dynamically aggregated into a knowledge embedding set and transferred between different detectors. When the teacher model is too large for online inference, its output can be stored on disk in advance to save the computation overhead, which is more storage efficient than feature-based methods. Extensive experiments demonstrate that our UniKD can plug and play in any homogeneous or heterogeneous teacher-student pairs and significantly outperforms conventional feature-based KD.
Shanshan Lao, Guanglu Song, Boxiao Liu, Yu Liu 0015, Yujiu Yang 0001
ICCV3
2023 Masked Autoencoders Are Stronger Knowledge Distillers
abstract
Knowledge distillation (KD) has shown great success in improving student’s performance by mimicking the intermediate output of the high-capacity teacher in fine-grained visual tasks, e.g. object detection. This paper proposes a technique called Masked Knowledge Distillation (MKD) that enhances this process using a masked autoencoding scheme. In MKD, random patches of the input image are masked, and the corresponding missing feature is recovered by forcing it to imitate the output of the teacher. MKD is based on two core designs. First, using the student as the encoder, we develop an adaptive decoder architecture, which includes a spatial alignment module that operates on the multi-scale features in the feature pyramid network (FPN) [20], a simple decoder, and a spatial recovery module that mimics the teacher’s output from the latent representation and mask tokens. Second, we introduce the masked convolution in each convolution block to keep the masked patches unaffected by others. By coupling these two designs, we can further improve the completeness and effectiveness of teacher knowledge learning. We conduct extensive experiments on different architectures with object detection and semantic segmentation. The results show that all the students can achieve further improvements compared to the conventional KD. Notably, we establish the new state-of-the-art results by boosting RetinaNet ResNet-18, and ResNet-50 from 33.4 to 37.5 mAP, and 37.4 to 41.5 mAP, respectively.
Shanshan Lao, Guanglu Song, Boxiao Liu, Yu Liu 0015, Yujiu Yang 0001
ICCV3
2023 GeoMIM: Towards Better 3D Knowledge Transfer via Masked Image Modeling for Multi-view 3D Understanding
abstract
Multi-view camera-based 3D detection is a challenging problem in computer vision. Recent works leverage a pretrained LiDAR detection model to transfer knowledge to a camera-based student network. However, we argue that there is a major domain gap between the LiDAR BEV features and the camera-based BEV features, as they have different characteristics and are derived from different sources. In this paper, we propose Geometry Enhanced Masked Image Modeling (GeoMIM) to transfer the knowledge of the LiDAR model in a pretrain-finetune paradigm for improving the multi-view camera-based 3D detection. GeoMIM is a multi-camera vision transformer with Cross-View Attention (CVA) blocks that uses LiDAR BEV features encoded by the pretrained BEV model as learning targets. During pretraining, GeoMIM’s decoder has a semantic branch completing dense perspective-view features and the other geometry branch reconstructing dense perspective-view depth maps. The depth branch is designed to be camera-aware by inputting the camera’s parameters for better transfer capability. Extensive results demonstrate that GeoMIM outperforms existing methods on nuScenes benchmark, achieving state-of-the-art performance for camera-based 3D object detection and 3D segmentation.
Jihao Liu, Boxiao Liu, Qihang Zhang, Yu Liu 0015, Hongsheng Li 0001
ICCV3
2023 A High-Gain and Low-Noise Mixer with Hybrid $G_{m}$-Boosting for 5G FR2 Applications
abstract
This paper presents a hybrid transconductance ($g_{m}$) boosting technique exploiting both transformer coupling and cross-coupled PMOS pair to improve the conversion gain (CG) and noise figure (NF) of mm-wave mixers. To demonstrate the effectiveness of the proposed$g_{m}$-boosting technique, a high-gain and low-noise mixer for 5G FR2 frequency band applications is developed in a 40 nm CMOS process. Transformer-based pole splitting and derivative superposition are employed to enhance the mixer bandwidth and improve linearity. The mixer achieves a peak CG of 20.9 dB, a 30% fractional bandwidth ($f_{BW}$), a minimum NF of 7.7 dB, and an input referred 1-dB compression point of −14 dBm, leading to an excellent figure of merit (FOM) of 11.05. The mixer consumes 15.6 mW of power and occupies a die area of 0.31 mm2.
Sijie Fu, Boxiao Liu, Chunqi Shi, Leilei Huang, Jinghong Chen, Runxi Zhang
ISCAS3
2023 A 3.84 GHz 32 fs RMS Jitter Over-Sampling PLL with High-Gain Cross-Switching Phase Detector
abstract
A 32 fs RMS jitter oversampling phase-locked loop (OSPLL) exploiting a high-gain cross-switching phase detector (CSPD) is proposed. The over-sampling PLL increases sam-pling frequency by 4x, reducing the in-band phase noise and overcoming the loop bandwidth limitation due to the reference frequency. Leveraging the increased loop bandwidth, the noise contribution of the voltage-controlled oscillator (VCO) is sig-nificantly suppressed. The high-gain CSPD adopts a common-mode sampling technique with time interleaving switches to ensure that the reference clock is sampled only at the maximum slew rate. The CSPD with a higher gain facilitates reducing the noise contribution from the phase detector (PD) and the transconductance cell. Additionally, an RC poly-phase filter (PPF) is employed to generate quadrature clocks, avoiding the deterioration of the PLL's low offset frequency phase noise. The PLL is implemented in a 40-nm CMOS process. Simulation results show that the PLL achieves a 32 fs RMS jitter integrated from 10 kHz to 100 MHz and a power consumption of 6.5 mW, resulting in an$FoM_{jitter}$of -261 dB. At 3.84 GHz frequency, the in-band phase noise is -136.8 dBc/Hz at 100 kHz offset.
Xuhong Lil, Jianghu Hong, Chunqi Shi, Leilei Huang, Boxiao Liu, Hao Deng 0003, Jinghong Chen, Runxi Zhang
ISCAS5
2023 A 88%-Peak-Efficiency 10-mV-Voltage-Ripple Dual-Mode Switched-Capacitor DC-DC Converter for Ultra-Low-Power Battery Management
abstract
This paper proposes a high-efficiency low-ripple dual-mode switched-capacitor (SC) DC-DC converter for low-power IoT and wearable device applications. A hybrid self-biased current scheme (HSBC) is developed to achieve low output voltage ripple and fast response. Two supply voltage domains of HSBC and clock drive controller circuit are introduced to reduce the power loss of the control circuit. An on-chip ultra-low-power bias circuit is also designed to minimize the power loss of the bias generator. The proposed DC-DC converter is implemented in a 40 nm CMOS process. Post-layout simulation results show that the converter realizes 1.6-1.8 V to 0.4 V conversion. The peak efficiency is up to 88% at$5\ \mu\mathrm{A}$, and the voltage ripple is less than 10 mV over a load range of 10 nA-$10\ \mu\mathrm{A}$. The response time of the converter is less than$15\ \mu\mathrm{s}$.
Xiaoyuan Wu, Leilei Huang, Boxiao Liu, Chunqi Shi, Jinghong Chen, Runxi Zhang
ISCAS5
2023 RAPHAEL: Text-to-Image Generation via Large Mixture of Diffusion Paths
abstract
Text-to-image generation has recently witnessed remarkable achievements. We introduce a text-conditional image diffusion model, termed RAPHAEL, to generate highly artistic images, which accurately portray the text prompts, encompassing multiple nouns, adjectives, and verbs. This is achieved by stacking tens of mixture-of-experts (MoEs) layers, i.e., space-MoE and time-MoE layers, enabling billions of diffusion paths (routes) from the network input to the output. Each path intuitively functions as a "painter" for depicting a particular textual concept onto a specified image region at a diffusion timestep. Comprehensive experiments reveal that RAPHAEL outperforms recent cutting-edge models, such as Stable Diffusion, ERNIE-ViLG 2.0, DeepFloyd, and DALL-E 2, in terms of both image quality and aesthetic appeal. Firstly, RAPHAEL exhibits superior performance in switching images across diverse styles, such as Japanese comics, realism, cyberpunk, and ink illustration. Secondly, a single model with three billion parameters, trained on 1,000 A100 GPUs for two months, achieves a state-of-the-art zero-shot FID score of 6.61 on the COCO dataset. Furthermore, RAPHAEL significantly surpasses its counterparts in human evaluation on the ViLG-300 benchmark. We believe that RAPHAEL holds the potential to propel the frontiers of image generation research in both academia and industry, paving the way for future breakthroughs in this rapidly evolving field. More details can be found on a webpage: https://raphael-painter.github.io/.
Zeyue Xue, Guanglu Song, Qiushan Guo, Boxiao Liu, Zhuofan Zong, Yu Liu 0015, Ping Luo 0002
NeurIPS4
2023 A hybrid-order local search algorithm for set k-cover problem in wireless sensor networks
Boxiao Liu, Mengting Yuan 0001, Haihang You
Frontiers Comput. Sci.1
2022 TokenMix: Rethinking Image Mixing for Data Augmentation in Vision Transformers
Jihao Liu, Boxiao Liu, Hang Zhou 0009, Hongsheng Li 0001, Yu Liu 0015
ECCV (26)2
2022 Rethinking Robust Representation Learning Under Fine-Grained Noisy Faces
Bingqi Ma, Guanglu Song, Boxiao Liu, Yu Liu 0015
ECCV (12)3
2022 Dynamic Weighted Semantic Correspondence for Few-Shot Image Generative Adaptation
abstract
Few-shot image generative adaptation, which finetunes well-trained generative models on limited examples, is of practical importance. The main challenge is that the few-shot model easily becomes overfitting. It can be attributed to two aspects: the lack of sample diversity for the generator and the failure of fidelity discrimination for the discriminator. In this paper, we introduce two novel methods to solve the diversity and fidelity respectively. Concretely, we propose dynamic weighted semantic correspondence to keep the diversity for the generator, which benefits from the richness of samples generated by source models. To prevent discriminator overfitting, we propose coupled training paradigm across the source and target domains to keep the feature extraction capability of the discriminator backbone. Extensive experiments show that our method outperforms previous methods both on image quality and diversity significantly.
Xingzhong Hou, Boxiao Liu, Shuai Zhang 0040, Lulin Shi, Zite Jiang, Haihang You
ACM Multimedia2
2021 Switchable K-class Hyperplanes for Noise-Robust Representation Learning
abstract
Optimizing the K-class hyperplanes in the latent space has become the standard paradigm for efficient representation learning. However, it’s almost impossible to find an optimal K-class hyperplane to accurately describe the latent space of massive noisy data. For this potential problem, we constructively propose a new method, named Switchable K-class Hyperplanes (SKH), to sufficiently describe the latent space by the mixture of K-class hyperplanes. It can directly replace the conventional single K-class hyperplane optimization as the new paradigm for noise-robust representation learning. When collaborated with the popular ArcFace on million-level data representation learning, we found that the switchable manner in SKH can effectively eliminate the gradient conflict generated by real-world label noise on a single K-class hyperplane. Moreover, combined with the margin-based loss functions (e.g. ArcFace), we propose a simple Posterior Data Clean strategy to reduce the model optimization deviation on clean dataset caused by the reduction of valid categories in each K-class hyperplane. Extensive experiments demonstrate that the proposed SKH easily achieves new state-of-the-art on IJB-B and IJB-C by encouraging noise-robust representation learning. Our code will be available at https://github.com/liubx07/SKH.git.
Boxiao Liu, Guanglu Song, Manyuan Zhang, Haihang You, Yu Liu 0015
ICCV1
2019 C-MIDN: Coupled Multiple Instance Detection Network With Segmentation Guidance for Weakly Supervised Object Detection
abstract
Weakly supervised object detection (WSOD) that only needs image-level annotations has obtained much attention recently. By combining convolutional neural network with multiple instance learning method, Multiple Instance Detection Network (MIDN) has become the most popular method to address the WSOD problem and been adopted as the initial model in many works. We argue that MIDN inclines to converge to the most discriminative object parts, which limits the performance of methods based on it. In this paper, we propose a novel Coupled Multiple Instance Detection Network (C-MIDN) to address this problem. Specifically, we use a pair of MIDNs, which work in a complementary manner with proposal removal. The localization information of the MIDNs is further coupled to obtain tighter bounding boxes and localize multiple objects. We also introduce a Segmentation Guided Proposal Removal (SGPR) algorithm to guarantee the MIL constraint after the removal and ensure the robustness of C-MIDN. Through a simple implementation of the C-MIDN with online detector refinement, we obtain 53.6% and 50.3% mAP on the challenging PASCAL VOC 2007 and 2012 benchmarks respectively, which significantly outperform the previous state-of-the-arts.
Gao Yan, Boxiao Liu, Nan Guo 0003, Xiaochun Ye, Fang Wan 0001, Haihang You, Dongrui Fan
ICCV2
2019 Live Demonstration: A Pulmonary Conditions Monitor Based on Electrical Impedance Tomography Measurement
abstract
In this demonstration, we present a non-invasive, real-time lung imaging system based on the electrical impedance tomography (EIT) technique. EIT is a medical imaging technique based on the electrical properties, i.e. resistivity and permittivity of tissues and organs. This demo is based on an EIT system-on-chip that utilizes frequency division multiplexing scheme to improve the throughput by 10, allowing clinicians to identify and prevent mechanical pulmonary injury during lung ventilation.
Boxiao Liu, Yongfu Li 0002, Guoxing Wang, Yong Lian 0001, Chun-Huat Heng
ISCAS2