EDBT 2026 Demo / reviewers in the wild / expert
Jie Liu 0040
dblp:03/2134-40
· DBLP profile ↗
24ranked-venue papers
6as first author
19since 2021 · last 2026
0000-0002-9297-7729ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 11 since 2021Computer networks · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RASR: Retrieval-Augmented Super Resolution for Practical Reference-based Image Restoration
Shuning Xu, Xiangyu Chen 0006, Dell Zhang, Jiantao Zhou 0001, Jie Tang 0006, Gangshan Wu, Jie Liu 0040 |
ISCAS | 8 |
| 2026 | STAR-RIS-Assisted Computation offloading and resource allocation optimization for mobile IoV with NOMA-MEC
Linbo Zhai, Zhiquan Liu 0001, Linfeng Wei, Xiaochuan Li 0001, Jie Liu 0040 |
Comput. Networks | 6 |
| 2026 | Access selection and service placement in mobile edge computing networks
Linbo Zhai, Zhiquan Liu 0001, Linfeng Wei, Xiaochuan Li 0001, Jie Liu 0040 |
Comput. Networks | 6 |
| 2026 | GLAD: Generative Language-Assisted Visual Tracking for Low-Semantic Templates
Xingyu Luo, Yidong Cai, Jie Liu 0040, Jie Tang 0006, Gangshan Wu, Limin Wang 0002 |
Int. J. Comput. Vis. | 3 |
| 2025 | In-the-wild Audio Spatialization with Flexible Text-guided LocalizationabstractBinaural audio enriches immersive experiences by enabling the perception of the spatial locations of sounding objects in AR, VR, and embodied AI applications. While existing audio spatialization methods can generally map any available monaural audio to binaural audio signals, they often lack the flexible and interactive control needed in complex multi-object user-interactive environments. To address this, we propose a Text-guided Audio Spatialization (TAS) framework that utilizes diverse text prompts and evaluates our model from unified generation and comprehension perspectives. Due to the limited availability of high-quality, large-scale stereo data, we construct the SpatialTAS dataset, which encompasses 376,000 simulated binaural audio samples to facilitate the training of our model. Our model learns binaural differences guided by 3D spatial location and relative position prompts, enhanced with flipped-channel audio. Experimental results show that our model can generate high quality binaural audios for various audio types on both simulated and real-recorded datasets. Besides, we establish an assessment model based on Llama-3.1-8B, which evaluates the semantic accuracy of spatial locations through a spatial reasoning task. Results demonstrate that by utilizing text prompts for flexible and interactive control, we can generate binaural audio with both high quality and semantic consistency in spatial locations. Tianrui Pan, Jie Liu 0040, Zewen Huang, Jie Tang 0006, Gangshan Wu |
ACL (1) | 2 |
| 2025 | CATANet: Efficient Content-Aware Token Aggregation for Lightweight Image Super-ResolutionabstractTransformer-based methods have demonstrated impressive performance in low-level visual tasks such as Image Super-Resolution (SR). However, its computational complexity grows quadratically with the spatial resolution. A series of works attempt to alleviate this problem by dividing Low-Resolution images into local windows, axial stripes, or dilated windows. SR typically leverages the redundancy of images for reconstruction, and this redundancy appears not only in local regions but also in long-range regions. However, these methods limit attention computation to content-agnostic local regions, limiting directly the ability of attention to capture long-range dependency. To address these issues, we propose a lightweight Content-Aware Token Aggregation Network (CATANet). Specifically, we propose an efficient Content-Aware Token Aggregation module for aggregating long-range content-similar tokens, which shares token centers across all image tokens and updates them only during the training phase. Then we utilize intra-group self-attention to enable long-range information interaction. Moreover, we design an inter-group cross-attention to further enhance global information interaction. The experimental results show that, compared with the state-of-the-art cluster-based method SPIN, our method achieves superior performance, with a maximum PSNR improvement of 0.33dB and nearly double the inference speed. Xin Liu 0012, Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
CVPR | 2 |
| 2025 | AutoLUT: LUT-Based Image Super-Resolution with Automatic Sampling and Adaptive Residual LearningabstractIn recent years, the increasing popularity of Hi-DPI screens has driven a rising demand for high-resolution images. However, the limited computational power of edge devices poses a challenge in deploying complex super-resolution neural networks, highlighting the need for efficient methods. While prior works have made significant progress, they have not fully exploited pixel-level information. Moreover, their reliance on fixed sampling patterns limits both accuracy and the ability to capture fine details in low-resolution images. To address these challenges, we introduce two plug-and-play modules designed to capture and leverage pixel information effectively in Look-Up Table (LUT) based super-resolution networks. Our method introduces Automatic Sampling (AutoSample), a flexible LUT sampling approach where sampling weights are automatically learned during training to adapt to pixel variations and expand the receptive field without added inference cost. We also incorporate Adaptive Residual Learning (AdaRL) to enhance inter-layer connections, enabling detailed information flow and improving the network’s ability to reconstruct fine details. Our method achieves significant performance improvements on both MuLUT and SPF-LUT while maintaining similar storage sizes. Specifically, for MuLUT, we achieve a PSNR improvement of approximately +0.20 dB improvement on average across five datasets. For SPF-LUT, with more than a 50% reduction in storage space and about a 2/3 reduction in inference time, our method still maintains performance comparable to the original. The code is available at https://github.com/SuperKenVery/AutoLUT. Yuheng Xu, Xin Liu 0012, Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
CVPR | 4 |
| 2025 | Towards Practical Real-Time Low-Latency Music Source SeparationabstractIn recent years, significant progress has been made in the field of deep learning for music demixing. However, there has been limited attention on real-time, low-latency music demixing, which holds potential for various applications, such as hearing aids, audio stream remixing, and live performances. Additionally, a notable tendency has emerged towards the development of larger models, limiting their applicability in certain scenarios. In this paper, we introduce a lightweight real-time low-latency model called Real-Time Single-Path TFC-TDF UNET (RT-STT), which is based on the Dual-Path TFC-TDF UNET (DTTNet). In RT-STT, we propose a feature fusion technique based on channel expansion. We also demonstrate the superiority of single-path modeling over dual-path modeling in real-time models. Moreover, we investigate the method of quantization to further reduce inference time. RT-STT exhibits superior performance with significantly fewer parameters and shorter inference times compared to state-of-the-art models. Junyu Wu, Jie Liu 0040, Tianrui Pan, Jie Tang 0006, Gangshan Wu |
ICME | 2 |
| 2025 | Computation bits maximization in multi-UAV-assisted-multi-vehicle edge computing system
Linbo Zhai, Meiyu Jin, Jiande Sun 0001, Chuanfen Feng, Zhiquan Liu 0001, Linfeng Wei, Xiaochuan Li 0001, Youlei Zhang, Jie Liu 0040 |
J. Netw. Comput. Appl. | 10 |
| 2025 | FSDM: An efficient video super-resolution method based on Frames-Shift Diffusion Model
Chao Chen 0026, Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
Neural Networks | 3 |
| 2024 | Sketch and Refine: Towards Fast and Accurate Lane DetectionabstractLane detection is to determine the precise location and shape of lanes on the road. Despite efforts made by current methods, it remains a challenging task due to the complexity of real-world scenarios. Existing approaches, whether proposal-based or keypoint-based, suffer from depicting lanes effectively and efficiently. Proposal-based methods detect lanes by distinguishing and regressing a collection of proposals in a streamlined top-down way, yet lack sufficient flexibility in lane representation. Keypoint-based methods, on the other hand, construct lanes flexibly from local descriptors, which typically entail complicated post-processing. In this paper, we present a “Sketch-and-Refine” paradigm that utilizes the merits of both keypoint-based and proposal-based methods. The motivation is that local directions of lanes are semantically simple and clear. At the “Sketch” stage, local directions of keypoints can be easily estimated by fast convolutional layers. Then we can build a set of lane proposals accordingly with moderate accuracy. At the “Refine” stage, we further optimize these proposals via a novel Lane Segment Association Module (LSAM), which allows adaptive lane segment adjustment. Last but not least, we propose multi-level feature integration to enrich lane feature representations more efficiently. Based on the proposed “Sketch-and-Refine” paradigm, we propose a fast yet effective lane detector dubbed “SRLane”. Experiments show that our SRLane can run at a fast speed (i.e., 278 FPS) while yielding an F1 score of 78.9%. The source code is available at: https://github.com/passerer/SRLane. Chao Chen 0026, Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
AAAI | 2 |
| 2024 | GTPT: Group-Based Token Pruning Transformer for Efficient Human Pose Estimation
Jie Liu 0040, Jie Tang 0006, Gangshan Wu, Yanbing Chou |
ECCV (69) | 2 |
| 2024 | RAVSS: Robust Audio-Visual Speech Separation in Multi-Speaker Scenarios with Missing Visual CuesabstractWhile existing Audio-Visual Speech Separation (AVSS) methods primarily concentrate on the audio-visual fusion strategy for two-speaker separation, they demonstrate a severe performance drop in the multi-speaker separation scenarios. Typically, AVSS methods employ guiding videos to sequentially isolate individual speakers from the given audio mixture, resulting in notable missing and noisy parts across various segments of the separated speech. In this study, we propose a simultaneous multi-speaker separation framework that can facilitate the concurrent separation of multiple speakers within a singular process. We introduce speaker-wise interactions to establish distinctions and correlations among speakers. Experimental results on the VoxCeleb2 and LRS3 datasets demonstrate that our method achieves state-of-the-art performance in separating mixtures with 2, 3, 4, and 5 speakers, respectively. Additionally, our model can utilize speakers with complete audio-visual information to mitigate other visual-deficient speakers, thereby enhancing its resilience to missing visual cues. We also conduct experiments where visual information for specific speakers is entirely absent or visual frames are partially missing. The results demonstrate that our model consistently outperforms others, exhibiting the smallest performance drop across all settings involving 2, 3, 4, and 5 speakers. Tianrui Pan, Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
ACM Multimedia | 2 |
| 2023 | From Coarse to Fine: Hierarchical Pixel Integration for Lightweight Image Super-resolutionabstractImage super-resolution (SR) serves as a fundamental tool for the processing and transmission of multimedia data. Recently, Transformer-based models have achieved competitive performances in image SR. They divide images into fixed-size patches and apply self-attention on these patches to model long-range dependencies among pixels. However, this architecture design is originated for high-level vision tasks, which lacks design guideline from SR knowledge. In this paper, we aim to design a new attention block whose insights are from the interpretation of Local Attribution Map (LAM) for SR networks. Specifically, LAM presents a hierarchical importance map where the most important pixels are located in a fine area of a patch and some less important pixels are spread in a coarse area of the whole image. To access pixels in the coarse area, instead of using a very large patch size, we propose a lightweight Global Pixel Access (GPA) module that applies cross-attention with the most similar patch in an image. In the fine area, we use an Intra-Patch Self-Attention (IPSA) module to model long-range pixel dependencies in a local patch, and then a spatial convolution is applied to process the finest details. In addition, a Cascaded Patch Division (CPD) strategy is proposed to enhance perceptual quality of recovered images. Extensive experiments suggest that our method outperforms state-of-the-art lightweight SR methods by a large margin. Code is available at https://github.com/passerer/HPINet. Jie Liu 0040, Chao Chen 0026, Jie Tang 0006, Gangshan Wu |
AAAI | 1 |
| 2023 | Robust Object Modeling for Visual TrackingabstractObject modeling has become a core part of recent tracking frameworks. Current popular tackers use Transformer attention to extract the template feature separately or interactively with the search region. However, separate template learning lacks communication between the template and search regions, which brings difficulty in extracting discriminative target-oriented features. On the other hand, interactive template learning produces hybrid template features, which may introduce potential distractors to the template via the cluttered search regions. To enjoy the merits of both methods, we propose a robust object modeling framework for visual tracking (ROMTrack), which simultaneously models the inherent template and the hybrid template features. As a result, harmful distractors can be suppressed by combining the inherent features of target objects with search regions’ guidance. Target-related features can also be extracted using the hybrid template, thus resulting in a more robust object modeling framework. To further enhance robustness, we present novel variation tokens to depict the ever-changing appearance of target objects. Variation tokens are adaptable to object deformation and appearance variations, which can boost overall performance with negligible computation. Experiments show that our ROMTrack sets a new state-of-the-art on multiple benchmarks. Yidong Cai, Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
ICCV | 2 |
| 2023 | Video Frame Interpolation with Densely Queried Bilateral CorrelationabstractVideo Frame Interpolation (VFI) aims to synthesize non-existent intermediate frames between existent frames. Flow-based VFI algorithms estimate intermediate motion fields to warp the existent frames. Real-world motions' complexity and the reference frame's absence make motion estimation challenging. Many state-of-the-art approaches explicitly model the correlations between two neighboring frames for more accurate motion estimation. In common approaches, the receptive field of correlation modeling at higher resolution depends on the motion fields estimated beforehand. Such receptive field dependency makes common motion estimation approaches poor at coping with small and fast-moving objects. To better model correlations and to produce more accurate motion fields, we propose the Densely Queried Bilateral Correlation (DQBC) that gets rid of the receptive field dependency problem and thus is more friendly to small and fast-moving objects. The motion fields generated with the help of DQBC are further refined and up-sampled with context features. After the motion fields are fixed, a CNN-based SynthNet synthesizes the final interpolated frame. Experiments show that our approach enjoys higher accuracy and less inference time than the state-of-the-art. Source code is available at https://github.com/kinoud/DQBC. Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
IJCAI | 2 |
| 2023 | Lightweight Super-Resolution Head for Human Pose EstimationabstractHeatmap-based methods have become the mainstream method for pose estimation due to their superior performance. However, heatmap-based approaches suffer from significant quantization errors with downscale heatmaps, which result in limited performance and the detrimental effects of intermediate supervision. Previous heatmap-based methods relied heavily on additional post-processing to mitigate quantization errors. Some heatmap-based approaches improve the resolution of feature maps by using multiple costly upsampling layers to improve localization precision. To solve the above issues, we creatively view the backbone network as a degradation process and thus reformulate the heatmap prediction as a Super-Resolution (SR) task. We first propose the SR head, which predicts heatmaps with a spatial resolution higher than the input feature maps (or even consistent with the input image) by super-resolution, to effectively reduce the quantization error and the dependence on further post-processing. Besides, we propose SRPose to gradually recover the HR heatmaps from LR heatmaps and degraded features in a coarse-to-fine manner. To reduce the training difficulty of HR heatmaps, SRPose applies SR heads to supervise the intermediate features in each stage. In addition, the SR head is a lightweight and generic head that applies to top-down and bottom-up methods. Extensive experiments on the COCO, MPII, and CrowdPose datasets show that SRPose outperforms the corresponding heatmap-based approaches. Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
ACM Multimedia | 2 |
| 2022 | Two Strategies Toward Lightweight Image Super-ResolutionabstractRecent convolution neural networks (CNNs) have achieved remarkable success in lightweight image super-resolution (LISR). The goal of LISR is to restore more accurate details with less model capacity. However, we observe two phenomena in current micro-architectures, one is the lack of consistent learning ability of high-frequency components, the other is large residual problem which does harm to the stability of residual learning. To tackle the two issues, we propose two strategies, namely global-guided attention strategy (GGAS) and channel-wise scaling strategy (CWSS), which can significantly improve the performance of the state-of-the-arts with negligible overheads. Zongcai Du, Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
ICASSP | 2 |
| 2022 | IAA-VSR: An iterative alignment algorithm for video super-resolution
Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
Appl. Intell. | 1 |
| 2020 | Residual Feature Aggregation Network for Image Super-ResolutionabstractRecently, very deep convolutional neural networks (CNNs) have shown great power in single image super-resolution (SISR) and achieved significant improvements against traditional methods. Among these CNN-based methods, the residual connections play a critical role in boosting the network performance. As the network depth grows, the residual features gradually focused on different aspects of the input image, which is very useful for reconstructing the spatial details. However, existing methods neglect to fully utilize the hierarchical features on the residual branches. To address this issue, we propose a novel residual feature aggregation (RFA) framework for more efficient feature extraction. The RFA framework groups several residual modules together and directly forwards the features on each local residual branch by adding skip connections. Therefore, the RFA framework is capable of aggregating these informative residual features to produce more representative features. To maximize the power of the RFA framework, we further propose an enhanced spatial attention (ESA) block to make the residual features to be more focused on critical spatial contents. The ESA block is designed to be lightweight and efficient. Our final RFANet is constructed by applying the proposed RFA framework with the ESA blocks. Comprehensive experiments demonstrate the necessity of our RFA framework and the superiority of our RFANet over state-of-the-art SISR methods. Jie Liu 0040, Wenjie Zhang 0006, Yuting Tang, Jie Tang 0006, Gangshan Wu |
CVPR | 1 |
| 2020 | Belief Map Enhancement Network for Accurate Human Pose Estimation
Jie Liu 0040, Yishun Dou, Wenjie Zhang 0006, Jie Tang 0006, Gangshan Wu |
ECAI | 1 |
| 2020 | Memory Recursive Network for Single Image Super-ResolutionabstractRecently, extensive works based on convolutional neural network (CNN) have shown great success in single image super-resolution (SISR). In order to improve the SISR performance while reducing the number of model parameters, some methods adopt multiple recursive layers to enhance the intermediate features. However, in the recursive process, these methods only use the output features of current stage as the input of the next stage and neglect the output features of historical stages, which degrades the performance of the recursive blocks. The long-term dependencies can only be learned implicitly during the recursive processes. To address these issues, we propose the memory recursive network (MRNet) to make full use of the output features at each stage. The proposed MRNet utilizes a memory recursive module (MRM) to generate features for each recursive stage, and then these features are fused by our proposed ShuffleConv block. Specifically, MRM adopts a memory updater block to explicitly model the long-term dependencies between the output features of historical recursive stages. The output features from the memory updater will be used as the input of the next recursive stage and will be continuously updated during the recursions. To reduce the number of parameters and ease the training difficulty, we introduce a ShuffleConv module to fuse the features from different recursive stages, which is much more effective than using plain convolutional combinations. Comprehensive experiments demonstrate that the proposed MRNet achieves state-of-the-art SISR performance while using much fewer parameters. Jie Liu 0040, Minqiang Zou, Jie Tang 0006, Gangshan Wu |
ACM Multimedia | 1 |
| 2017 | Deep convolutional neural networks for pedestrian detection with skip poolingabstractWith the big success of deep convolutional neural networks (CNN) in image classification task, many proposal based networks are proposed to detect given objects in an image. Faster R-CNN is such a network that uses a region proposal network (RPN) to generate nearly cost-free region proposals, which has shown excellent performance in ILSVRC and MS COCO datasets. However, Faster R-CNN does not behave so well for the task of pedestrian detection since the images in popular pedestrian detection datasets have more complicated background and contain a lot of small foreground objects. In this work, we leverage the RPN architecture of Faster R-CNN and extend it to a multi-layer version combined with skip pooling to tackle the pedestrian detection problem. Skip pooling is a kind of network connection that combines multiple ROI pooling results from lower layers to form a single input to a higher layer while bypassing intermediate layers. We comprehensively evaluate our network, referred to as SP-CNN, on the Caltech pedestrian detection benchmark and KITTI object detection benchmark. Our method achieves state-of-the-art accuracy on Caltech dataset and presents a comparable result on KITTI dataset while maintaining a good speed. Jie Liu 0040, Xingkun Gao, Nianyuan Bao, Jie Tang 0006, Gangshan Wu |
IJCNN | 1 |
| 2016 | Scalable Single-Source SimRank Computation for Large GraphsabstractSimRank is an effective similarity measure between vertices in a graph, which has become a fundamental technique in graph analytics. Despite its popularity, computation of SimRank is often costly in both space and time, especially with the ever growing scale of graph data nowadays. In this paper, we focus on the computation of Single-Source SimRank: given a query vertex, return the similarities between this vertex and any other vertices in the graph. The traditional centralized SimRank algorithms are not efficient for this problem. To fully utilize the computing power of modern distributed systems, we propose sssSimRank, an efficient distributed algorithm based on the random walk model. Our algorithm achieves scalability via minimizing the total number, the space cost, and the matching time of random walks. We implement our approach on the popular distributed processing platform Spark. Experimental results demonstrate the effectiveness, efficiency and scalability of our method. Xingkun Gao, Nianyuan Bao, Jie Liu 0040, Jie Tang 0006, Gangshan Wu |
ICPADS | 3 |