Lijie Liu

dblp:63/1773 · DBLP profile ↗
← Back
17ranked-venue papers
12as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 8 first-author · 2 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 2 since 2021Systems, architecture and hardware · 5 · 4 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
abstract
Human-Centric Video Generation (HCVG) methods seek to synthesize human videos from multimodal inputs, including text, images, and audio. Existing methods struggle to effectively coordinate these heterogeneous modalities due to two challenges: the scarcity of modality-complete data and the difficulty of jointly modeling triplet conditions without performance degradation. In this work, we present HuMo, a unified HCVG framework for collaborative multimodal control. For the first challenge, we construct an incomplete-yet-complementary dataset for improved data utilization efficiency and training scalability. For the second challenge, we propose a two-stage progressive multimodal training paradigm with task-specific strategies at each stage. In the first stage, to balance the text-following and subject-preservation abilities, we adopt the minimal-invasive image injection strategy. In the second stage, to enhance audio-visual sync, we propose a focus-by-predicting strategy that implicitly guides the model to associate audio with facial regions. For joint learning of controllabilities across multi-modal inputs, we progressively incorporate the audio-visual sync task, building on previously acquired capabilities. During inference, for flexible and fine-grained multimodal control, we design a stage-adaptive Classifier-Free Guidance strategy that dynamically adjusts guidance weights across denoising steps. Extensive experimental results demonstrate that HuMo surpasses specialized state-of-the-art methods in sub-tasks, establishing a unified framework for collaborative multimodal-conditioned HCVG.
Liyang Chen, Tianxiang Ma, Bingchuan Li, Zhuowei Chen, Lijie Liu, Zhiyong Wu 0001
AAAI6
2025 Phantom: Subject-Consistent Video Generation via Cross-Modal Alignment
abstract
The continuous development of foundational models for video generation is evolving into various applications, with subject-consistent video generation still in the exploratory stage. We refer to this as Subject-to-Video, which extracts subject elements from reference images and generates subject-consistent videos following textual instructions. We believe that the essence of subject-to-video lies in balancing the dual-modal prompts of text and image, thereby deeply and simultaneously aligning both text and visual content. To this end, we propose Phantom, a unified video generation framework for both single- and multi-subject references. Building on existing text-to-video and image-to-video architectures, we redesign the joint text-image injection model and drive it to learn cross-modal alignment via text-image-video triplet data. The proposed method achieves high-fidelity subject-consistent video generation while addressing issues of image content leakage and multi-subject confusion. Evaluation results indicate that our method outperforms other state-of-the-art closed-source commercial solutions. In particular, we emphasize subject consistency in human generation, covering existing ID-preserving video generation while offering enhanced advantages.
Lijie Liu, Tianxiang Ma, Bingchuan Li, Zhuowei Chen
ICCV1
2025 Universal Content-Agnostic Backscatter for OFDM WiFi
abstract
Ambient backscatter is one of the most promising solutions for the widespread deployment of low-power IoT. We present CAB, a content-agnostic backscatter system that can demodulate both tag and ambient data from ambient backscattered WiFi alone. In contrast to prior ambient backscatter systems that use ambient data (content) to demodulate tag data, we focus on zero-subcarriers, which are invariant and independent for any ambient OFDM WiFi. The idea of using zero-subcarriers to convey tag data is simple and elegant. Not only does it for the first time remove the dependency of tag-data demodulation on ambient data, but it also significantly improves the practicality of ambient backscatter. CAB is also universal for various OFDM-WiFi signals because the zero-subcarrier we use is virtual and is essentially the phase estimation of the pilot, which is not affected by different pilot patterns. In order to verify the universality, we prototype CAB using off-the-shelf FPGAs and SDRs. Extensive experiments show CAB is universal as it can work with multi-band, multi-stream, and multi-user ambient traffic, including WiFi 3/4/5/6. In addition, CAB can work with excitations at different data rates over various commercial NICs. As the first content-agnostic backscatter system with 340.9 Mbps aggregate throughput, we believe CAB takes a crucial step forward on ubiquitous battery-free IoTs.
Wei Gong 0001, Lijie Liu
IEEE Trans. Netw.2
2023 Energy-Efficient WiFi Backscatter Communication for Green IoTs
abstract
The boom of the Internet of Things has revolutionized people's lives, but it has also resulted in massive resource consumption and environmental pollution. Recently, Green IoT (GIoT) has become a worldwide consensus to address this issue. In this paper, we propose EEWScatter, an energy-efficient WiFi backscatter communication system to pursue the goal of GIoT. Unlike previous backscatter systems that solely focus on tags, our approach offers a comprehensive system-wide view on energy conservation. Specifically, we reuse ambient signals as carriers and utilize an ultra-low-power and battery-free design for tag nodes by backscatter. Further, we design a new CRC-based algorithm that enables the demodulation of both ambient and tag data by only a single receiver while using ambient carriers. Such a design eliminates system reliance on redundant transceivers with high power consumption. Results demonstrate that EEWScatter achieves the lowest overall system power consumption and saves at least half of the energy. What's more, the power consumption of our tag is only 1/1000 of that of active radio.
Yimeng Huang, Lijie Liu, Jihong Yu, Yuguang Fang, Wei Gong 0001
GLOBECOM2
2023 A Flexible Power Allocation Strategy for Dual-DC-Port Inverter-Connected PV-Battery Hybrid Systems
abstract
Dual-dc-port inverter can directly connect photo-voltaic (PV) and battery to ac-grid without dc-dc converter, providing a low-cost, small-volume, and high-efficiency solution for PV-battery hybrid systems. However, it is challenging to design the modulation scheme for the dual-dc-port inverter due to the unbalanced dc-link and coupled port power control feature. To realize flexible power allocation and achieve satisfactory current quality under unbalanced dc-link voltage, this article proposes a hybrid modulation-based flexible power allocation strategy for the dual-dc-port inverter. In the proposed strategy, flexible active power allocation for each port is achieved by proportionally splitting the desired voltage vector, and the split voltage vector for each port is implemented by the hybrid modulation. With the hybrid modulation, challenges for modulation scheme design under unbalanced dc port voltage are also avoided. Experimental tests are conducted to verify the effectiveness of the proposed power allocation strategy. And experimental results indicate that the proposed strategy has beneficial steady-state and dynamic performance.
Lijie Liu, Dehong Zhou, Jianxiao Zou
IECON1
2023 High-Efficiency Quasi-Single-Stage Battery-Supercapacitor Hybrid Energy Storage System
abstract
The battery-supercapacitor hybrid energy storage system (HESS) integrates high-energy-density units and high-power-density units together, which has been widely used in microgrids (MGs) applications. However, the conventional HESS has the disadvantages of low efficiency and large volume. To improve the system power density, the quasi single-stage con-verter is an attractive solution for HESS due to it offering direct power flows from dc-side to ac-side. To maximize the system efficiency, this article proposes a novel space vector modulation with its basic idea and implementation process. Finally, the effectiveness of the proposed modulation is verified by quasi-single-stage converter-based virtual synchronous generator (VSG) experimental tests.
Lijie Liu, Dehong Zhou, Jianxiao Zou
IECON1
2022 A Hybrid Si/GaN-Based Quasi-Single-Stage Converter for Microgrid Applications with Simplified Space-Vector Modulation
abstract
In low-voltage energy storage microgird systems, the quasi-single-stage architecture is a promising alternative to improve the efficiency, which contains a direct power flow path from the battery to the inverter, resulting in reduced power losses by dc/dc converter. However, the unbalanced dc-link voltage of two ports certainly generates asymmetric space vectors. Thus, how to design the modulation strategy under this scenario is the major challenge. Meanwhile, the power losses are not only determined by the topology, but the device material such as gallium nitride (GaN) also has a significant impact on it. Therefore, a novel hybrid Si/GaN-based quasi-single-stage converter (HSG-QSSC) is proposed in this paper. Furthermore, a simplified space-vector modulation (SVM) scheme is presented to concentrate all the high-frequency switching events on the GaN HEMTs while the Si IGBTs operate with low frequency and avoid complicated triangle functions. As a result, the total power losses are reduced due to the decoupled frequency switching, and the high efficiency of calculation is achieved. Islanded microgrid experimental results with a hybrid Si/GaN active-neutral-point converter (ANPC) prototype are provided to verify the feasibility and effectiveness of the presented modulation scheme.
Dehong Zhou, Jianxiao Zou, Zewei Shen, Lijie Liu, Xiaoming Fu 0005
IECON5
2020 Reinforced Axial Refinement Network for Monocular 3D Object Detection
Lijie Liu, Chufan Wu, Jiwen Lu, Lingxi Xie, Jie Zhou 0001, Qi Tian 0001
ECCV (17)1
2019 Deep Fitting Degree Scoring Network for Monocular 3D Object Detection
abstract
In this paper, we propose to learn a deep fitting degree scoring network for monocular 3D object detection, which aims to score fitting degree between proposals and object conclusively. Different from most existing monocular frameworks which use tight constraint to get 3D location, our approach achieves high-precision localization through measuring the visual fitting degree between the projected 3D proposals and the object. We first regress the dimension and orientation of the object using an anchor-based method so that a suitable 3D proposal can be constructed. We propose FQNet, which can infer the 3D IoU between the 3D proposals and the object solely based on 2D cues. Therefore, during the detection process, we sample a large number of candidates in the 3D space and project these 3D bounding boxes on 2D image individually. The best candidate can be picked out by simply exploring the spatial overlap between proposals and the object, in the form of the output 3D IoU score of FQNet. Experiments on the KITTI dataset demonstrate the effectiveness of our framework.
Lijie Liu, Jiwen Lu, Chunjing Xu, Qi Tian 0001, Jie Zhou 0001
CVPR1
2018 Adversarial Transfer Networks for Visual Tracking
abstract
Visual tracking plays an important role in unmanned systems. In many cases, the system needs to keep track of targets it has never seen before, and the only training sample available is the specified object in the initial frame. In this paper, we propose a deep architecture called adversarial transfer networks (ATNet), which aims to make well use of offline video training data and solve the problem of lacking training samples in visual tracking. Different from most existing trackers which neglect significant differences between videos and gulp the training data all together, our method utilizes the special nature of tracking problem and concentrates on transferring domain-specific information across similar tracking tasks. We first propose an efficient way to select a training video that is most similar to online tracking task and regard it as source domain. With the labeled data in the selected source domain, we apply adversarial transfer learning to make the feature distribution of source-domain samples and target-domain samples as similar as possible. Therefore, the transferred source-domain samples can provide various possible appearance of tracked target for training and boost the tracking performance. Experimental results on three OTB tracking benchmarks show that our method outperforms the state-of-the-art trackers in both accuracy and robustness.
Lijie Liu, Jiwen Lu, Jie Zhou 0001
IROS1
2008 Lifting-based Laplacian Pyramid reconstruction schemes
abstract
Laplacian Pyramid (LP) provides a redundant signal representation and can be characterized as an oversampled filter bank (FB). In this paper, a generic lifting-based parameterization reconstruction algorithm is proposed to characterize all LP synthesis banks that can satisfy the perfect reconstruction property. Two typical lifting-based LP reconstruction schemes are then derived from this general representation. The first scheme presents the dual frame LP reconstruction and its closed-form solutions for any LP filters. The second LP reconstruction scheme leads to an efficient FB, which demonstrates improvements over the usual LP reconstruction in the presence of noise.
Lijie Liu, Lu Gan 0002, Trac D. Tran
ICIP1
2007 An 8×8 IEEE-Compliant Lifting-Based Multiplierless IDCT Structure and Algorithm
abstract
In this paper we propose a lifting-based 8times8 IDCT structure and its EEEE-1180 compliant approximation solution. Derived from an efficient Loeffler's 11-multiply IDCT structure, the proposed scheme comprises of butterflies and dyadic-rational lifting steps that can be implemented using only shift and add operations. Our approach also allows the computational scalability with different accuracy-versus-complexity trade-offs. Furthermore, the lifting construction allows a simple construction of the corresponding multiplierless forward DCT, providing bit-exact reconstruction if pairing with our proposed IDCT Our high-accuracy solution provides a very close approximation of the floating-point IDCT. The experiments in MPEG-2 and MPEG-4 video coders under the worst-case assumptions show almost drifting-free reconstructions.
Lijie Liu, Trac D. Tran
ICASSP (1)1
2006 JPEG-compliant image coding with adaptive pre-/post-filtering
abstract
In this paper we propose an image coding scheme with adaptive pre-/post-filtering which produces a fully compliant JPEG bitstream. The basic idea is to introduce pre-filtering to improve the coding performance and post-filtering to reduce JPEG blocking artifacts. The adaptivity of the pre-/post-filters is achieved by varying their filter supports based on two criteria: rate-distortion optimization (RD-opt) and over-/under-flow. Experiments show that despite keeping intact JPEG baseline coding, our proposed coding scheme with these two criteria can improve not only the objective quality (0.3-1.5 dB PSNR gain), but also yield superior visual quality by preserving edge details and mitigating blocking artifacts. Our proposed algorithm is competitive with state-of-the-art deblocking algorithms.
Lijie Liu, Trac D. Tran
ISCAS1
2005 Adaptive Block-Based Image Coding with Pre-/Post-Filtering
abstract
This paper presents an adaptive block-based image coding method, which combines the advantages of variable block size transform and adaptive pre-/post-filtering scheme. Our approach partitions an image into blocks with different sizes, which are best suitable for the characteristics of the underlying data in the rate-distortion (RD) sense. The adaptive block decomposition mitigates the ringing artifacts by adopting a small block size transform in nonstationary regions, and improves the coding efficiency by using a large block size transform in homogenous regions. Moreover, pre-/post-filtering is adaptively applied along the block boundaries to improve coding efficiency and minimize blocking artifacts. Simulation results show that the proposed coder can achieve competitive objective performance as well as yield superior reconstruction visual quality, compared with the RD-optimized JPEG2000 and H.264/AVC I-frame coder.
Lijie Liu, Trac D. Tran
DCC2
2005 Combined key-frame extraction and object-based video segmentation
abstract
Video segmentation has been an important and challenging issue for many video applications. Usually there are two different video segmentation approaches, i.e., shot-based segmentation that uses a set of key-frames to represent a video shot and object-based segmentation that partitions a video shot into objects and background. Representing a video shot at different semantic levels, two segmentation processes are usually implemented separately or independently for video analysis. In this paper, we propose a new approach to combine two video segmentation techniques together. Specifically, a combined key-frame extraction and object-based segmentation method is developed based state-of-the-art video segmentation algorithms and statistical clustering approaches. On the one hand, shot-based segmentation can dramatically facilitate and enhance object-based segmentation by using key-frame extraction to select a few key-frames for statistical model training. On the other hand, object-based segmentation can be used to improve shot-based segmentation results by using model-based key-frame refinement. The proposed approach is able to integrate advantages of these two segmentation methods and provide a new combined shot-based and object-based framework for a variety of advanced video analysis tasks. Experimental results validate effectiveness and flexibility of the proposed video segmentation algorithm.
Lijie Liu
IEEE Trans. Circuits Syst. Video Technol.1
2003 An entropy based segmentation algorithm for computer-generated document images
abstract
This paper presents an efficient compression-oriented segmentation algorithm for computer-generated document images. In this algorithm, a document image is represented in a block-based multiscale pyramid. Then, image blocks will be characterized based on their entropy values of the intensity histogram, and the entropy distribution are assumed to be Gaussian priors in this work. We will discuss two methods, i.e., off-line and online training, to estimate model parameters. We use the multiscale Bayesian estimation to refine the classification results and generate the final segmentation result, where image blocks are classified into four classes, i.e., background, text, graphic and picture. It is expected that the proposed entropy-based segmentation will be suitable for compound document compression and two training approaches apply to different applications.
Lijie Liu, Xiaomu Song
ICIP (1)1
2003 A new JPEG2000 region-of-interest image coding method: partial significant bitplanes shift
abstract
We propose a new region-of-interest (ROI) coding method called partial significant bitplanes shift (PSBShift) that combines the advantages of the two standard ROI coding methods defined in JPEG2000. The PSBShift method not only supports arbitrarily shaped ROI coding without coding the shape, but also enables the flexible adjustment of compression quality in ROI and background. Additionally, the new method can efficiently code multiple ROIs with different degrees of interest in an image.
Lijie Liu
IEEE Signal Process. Lett.1