Yu Liu 0004

dblp:97/2274-4 · DBLP profile ↗
← Back
50ranked-venue papers
6as first author
45since 2021 · last 2027
0000-0002-5949-6587ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 31 · 6 first-author · 26 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Computer networks · 7 · 1 first-author · 7 since 2021Systems, architecture and hardware · 5 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2027 CProPNet: Chrono-progressive dietary synergy network with pathology decoupling for individualized HbA1c prediction
Jing Liu 0002, Peiguang Jing, Yu Liu 0004
Expert Syst. Appl.6
2026 Causality-Aligned Semantic Recovery for Incomplete Cross-Modal Retrieval
abstract
Incomplete cross-modal retrieval (ICMR) requires models to recover missing modalities and robustly align heterogeneous ones for effective retrieval. Existing methods, however, fall short in both aspects. They often rely on limited semantic cues, such as single samples or coarse category prototypes, which compromises reconstruction quality. Moreover, these approaches are vulnerable to learning spurious cross-modal correlations, thereby impairing accurate alignment and hindering retrieval performance. To address these challenges, we propose Causality-Aligned Semantic Recovery (CASR), a novel method designed to both comprehensively restore missing modalities and mitigate spurious associations between vision and language. Our CASR involves two essential components: i) the Missing Modality Imagination (MMI) module, which combines category semantic priors with relevant contextual information to achieve high-quality semantic reconstruction; ii) the Explicit Causal Alignment (ECA) module, which explicitly learns environment-invariant attention, effectively eliminating the interference of spurious correlations and improving retrieval performance. Furthermore, we extend CASR to the challenging task of Partially Aligned Cross-Modal Retrieval, where we treat unlabeled unpaired data as a form of incomplete data. By leveraging MMI and ECA modules, we are able to learn robust representations in this setting. Extensive experiments on benchmark datasets under various missing rates demonstrate that CASR achieves superior robustness and retrieval performance.
Haipeng Chen 0002, Yu Liu 0004, Xun Yang 0001, Yuheng Liang, Yingda Lyu
AAAI2
2026 Pedestrian-Centric Discriminative and Fine-grained Semantic Mining for Text-based Person Retrieval
abstract
Text-based Person Retrieval (TPR) aims to retrieve specific pedestrian images from a gallery based on the given textual descriptions, serving as a fine-grained instance of cross-modal retrieval on the Web. Current mainstream approaches primarily leverage pre-trained models and attention mechanisms to enhance multi-modal representations. Despite notable progress, they still struggle with two major challenges: 1) Intra-instance semantic asymmetry, which mainly derives from the partial semantic relevance conveyed by each image-text pair; and 2) Inter-instance semantic ambiguity, which arises from the high similarity of image-text pairs with different identities. These issues result in suboptimal semantic alignment and degraded retrieval accuracy. To this end, we propose a novel Pedestrian-Centric Discriminative and Fine-grained Semantic Mining (DFSM) framework for TPR. Specifically, our DFSM method comprises two essential components: 1) Text-aware Visual Refinement (TVR), which mitigates visual redundancy by selecting semantically relevant patches under textual guidance, and refines them via adaptive clustering and merging; 2) Token-level Semantic Alignment (TSA), which formulates the matching relationship between image regions and text words as a conditional transport (CT) problem, effectively mining fine-grained semantic differences and enhancing instance discrimination. Extensive experiments on four benchmarks validate the advantages of DFSM in terms of retrieval accuracy and visual interpretability.
Yuheng Liang, Haipeng Chen 0002, Yu Liu 0004, Yingda Lyu
WWW3
2026 IoT-Oriented EEG-Based Memory State Recognition With Music-Facilitated Hierarchical Temporal-Frequency Fusion
abstract
Music dynamically modulates neural activity in memory-related brain regions, enabling more precise retention of details, facilitating easier recall of information, and enhancing working memory capacity. However, most current electroencephalography (EEG)-based memory state recognition approaches neither directly embed music stimuli as inputs nor sufficiently quantify the influence of musical genres. To address this gap, we propose a hierarchical temporal–frequency fusion (HTFF) network that integrates EEG data with musical cues for memory state recognition. Specifically, we propose a music-facilitated EEG temporal–frequency fusion module that progressively fuses low-level perceptual cues into high-level memory representations via cascaded resonant EEG fusion reinforcers (REFRs). These REFR units allocate musical cues into the temporal and frequency domains, thereby enhancing EEG representations through a cross-domain coupling mechanism. We also propose a mixed-constraints mechanism that creates more compact intra-class sample distributions and sharpens decision boundaries for musical genres. The experimental results demonstrate that the HTFF network outperforms state-of-the-art methods in memory state recognition. Moreover, its compatibility with the Internet of Things (IoT) and wearable devices highlights its potential for real-time monitoring and personalized cognitive interventions, paving the way for practical applications in cognitive health.
Zhuang Miao, Yu Liu 0004, Peiguang Jing, Zhiju Huang
IEEE Internet Things J.2
2026 MSSFN: Multi-stimulus stereo spatiotemporal fusion network with pattern disentanglement for Alzheimer's disease diagnosis
Peiguang Jing, Yu Liu 0004, Sun-Yuan Kung
Inf. Process. Manag.5
2026 WUSRVN: Wavelet U-shaped network and sensitivity refinement-based variational network for accelerated magnetic resonance imaging
Haiyuan Li, Jizhong Duan, Haibo Tao, Yu Liu 0004
Signal Process.6
2026 HAOT: Heterogeneous Hardware-Aware Tensor Computation Optimization Framework via Transformer
abstract
Efficient execution of tensor computations on different hardware, such as CPUs, GPUs, and spatial accelerators, poses substantial challenges due to divergent memory hierarchies, compute models, and architectural constraints. However, existing auto-tuning approaches typically fail to integrate hardware-aware formulations, leading to inefficient search processes and suboptimal utilization of hardware resources. To address these challenges, we introduce HAOT, a novel hardware-aware tensor computation optimization framework that combines a learning-based scheduling policy with a hardware analytical model. Unlike traditional methods, HAOT not only adapts the learning-based scheduling strategy but also dynamically constrains the search space based on the availability of hardware resources, ensuring both the feasibility and optimal use of resources. Experiments across a range of hardware platforms show that HAOT achieves up to 1.4× average speedup on individual tensor operations and a 1.8× improvement in end-to-end deep learning workloads, outperforming state-of-the-art baselines. Ablation studies demonstrate the complementary roles of the Transformer-based policy model and the hardware analysis model, each contributing to generating high-quality schedule decisions. Furthermore, HAOT exhibits superior sample efficiency, requiring significantly fewer search trials to converge to the optimal schedule.
Han Wang 0039, Yu Liu 0004
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2026 CSVSUF: A Deep Unfolding Framework for Compressive Spectral Video Sensing
abstract
Spectral videos (SVs) capture spatio-temporal-spectral information from dynamic scenes, but their acquisition traditionally requires expensive and complex systems, motivating the development of compressive spectral video sensing (CSVS). It typically employs the coded aperture snapshot spectral imager (CASSI) to acquire compressed measurements, from which SVs are reconstructed via model-driven or learning-based algorithms. However, two major limitations remain in current CASSI-based reconstruction methods: 1) conventional model-driven algorithms rely on iterative optimization, which limits their representational capacity in complex scenes and results in slow reconstruction; 2) existing deep learning-based approaches overlook the joint modeling of spatial, temporal, and spectral correlations, failing to fully exploit the multi-dimensional dependencies. Hence, we propose a principled compressive spectral video sensing unfolding framework (CSVSUF) in a CASSI system for spectral video reconstruction. Moreover, we develop a novel spatio-temporal-spectral prior-learning Transformer (STS-PLT) to capture the multi-dimensional correlations within each unfolding stage. By treating STS-PLT as a Gaussian denoiser for the prior term in CSVSUF, we establish a deep unfolding-based method for CSVS. Extensive experiments demonstrate that our method consistently outperforms existing approaches in both reconstruction accuracy and visual quality, validating the benefit of combining physics-guided modeling with deep prior learning in CSVS. Code is available at https://github.com/zli1024/CSVSUF.
Han Wang 0039, Jizhong Duan, Baihua Li, Yu Liu 0004
IEEE Trans. Image Process.5
2026 Discriminative Representation Learning for Remote Sensing Visual Question Answering
abstract
Recently, Remote Sensing Visual Question Answering (RSVQA) has attracted increasing attention from both academia and industry, which is the basis for understanding the underlying correspondence between remote sensing imagery and text descriptions. However, current methods are still insufficient in learning discriminative visual and textual representations for answer reasoning, mainly due to two reasons: (1) the remote sensing image environment is complex and changeable, and the target scales vary significantly, making it difficult to extract discriminative visual features; and (2) there is a lack of effective guidance from remote sensing domain knowledge to learn discriminative features. To this end, we propose a D iscriminative R epresentation L earning (DRL) method that includes two key strategies: visual feature enhancement and prior knowledge guidance. Specifically, we employ the Fourier transform to simulate the diverse visual environment and force the model to mine discriminative visual representations by imposing consistency constraints with the original features. In addition, we leverage the Remote Sensing Multimodal Large Language Model (RSMLLM) to generate captions rich in remote sensing domain-specific prior knowledge. These captions, derived from RSMLLM’s powerful knowledge integration and summarization capabilities, can then be compared and fused with visual representations to generate more discriminative representations, which are ultimately used for answer reasoning. Finally, recognizing that most existing RSVQA methods rely solely on static remote sensing images, we introduce RSVideoQA, a novel satellite video question answering dataset. This dataset is designed to facilitate the exploration of the rich spatio-temporal dynamics inherent in video sequences. Experimental results across three distinct datasets validate the effectiveness of our proposed method. Our dataset and code will be released at https://github.com/chill-han/DRL .
Yingda Lyu, Yu Liu 0004, Haipeng Chen 0002
ACM Trans. Multim. Comput. Commun. Appl.3
2025 MPBR: Multimodal Progressive Bidirectional Reasoning for Open-Set Fine-Grained Recognition
Junfu Tan, Peiguang Jing, Yu Liu 0004
ICCV4
2025 Enhancing Semantic Clarity: Discriminative and Fine-grained Information Mining for Remote Sensing Image-Text Retrieval
abstract
Remote sensing image-text retrieval is a fundamental task in remote sensing multimodal analysis, promoting the alignment of visual and language representations. The mainstream approaches commonly focus on capturing shared semantic representations between visual and textual modalities. However, the inherent characteristics of remote sensing image-text pairs lead to a semantic confusion problem, stemming from redundant visual representations and high inter-class similarity. To tackle this problem, we propose a novel Discriminative and Fine-grained Information Mining (DFIM) model, which aims to enhance semantic clarity by reducing visual redundancy and increasing the semantic gap between different classes. Specifically, the Dynamic Visual Enhancement (DVE) module adaptively enhances the visual discriminative features under the guidance of multimodal fusion information. Meanwhile, the Fine-grained Semantic Matching (FSM) module cleverly models the matching relationship between image regions and text words as an optimal transport problem, thereby refining intra-instance matching. Extensive experiments on two benchmark datasets justify the superiority of DFIM in terms of retrieval accuracy and visual interpretability over the leading methods.
Yu Liu 0004, Haipeng Chen 0002, Yuheng Liang, Xun Yang 0001, Yingda Lyu
IJCAI1
2025 FCCQA: A High-Performance FPGA-based Accelerator for Approximate Nearest Neighbor Search
abstract
Data search plays a key role in information retrieval and machine learning. Currently, the Field Programmable Gate Array (FPGA) has emerged as a popular hardware accelerator for such studies. In order to maintain a better balance between search time and accuracy, we introduce FCCQA, an exhaustive and scalable data search architecture, utilizing Approximate Nearest Neighbor (ANN) for rapid interrogation of high-dimensional datasets on FPGA. Our architecture adopts the technique of compositing quantization to achieve a high data compression ratio, reducing memory usage to about 1.67%, while efficiently transforming the intricate computation into a simplified table lookup operation in high-dimensional datasets. Compared with GPU and CPU, our approach establishes a superior balance among accuracy, energy efficiency and speed.
Tianle Miao, Yongxiang Lyu, Mang I Vai, Yu Liu 0004
ISCAS6
2025 UDNet: Unified Deep Network based on Transformer and Multi-stage Fusion for brain tumor classification from undersampled MRI
Jizhong Duan, Yunshuang Xie, Yu Liu 0004
Neurocomputing4
2025 Illumination guided domain adaptation object detection in thermal imagery
Yuan Zhou 0006, Yu Liu 0004, Sun-Yuan Kung
Neurocomputing4
2025 Healthcare IoT-Enabled PMAFNet: A Progressive Multimodal Adaptive Fusion Network for Understanding HbA1c Fluctuations
abstract
The Internet of Things (IoT) has transformed healthcare via Healthcare-IoT (HIoT) by enabling continuous patient monitoring and efficient data analytics, offering promising solutions for managing the escalating global diabetes burden. The glycated hemoglobin (HbA1c) is an important biomarker for diabetes evaluation and control. Accurate prediction of HbA1c fluctuations remains a critical challenge, as existing models often overlook the interplay between stable patient profiles (demographics, lifestyle, clinical indicators) and dynamic dietary nutrient intake, limiting their ability to model glycemic responses. This study constructs a comprehensive multi-source dataset comprising detailed personal information, lifestyle, medical tests, and multi-day dynamic nutrient intake through the HIoT system. Utilizing this dataset, we propose a novel progressive multimodal adaptive fusion network (PMAFNet) to decode the complex determinants of HbA1c variability. PMAFNet employs two core modules: the multimodal graph feature-level enhancement (MGFE) module processes static and dynamic features to capture structural dependencies and temporal patterns, and the adaptive modality-aware progressive fusion (AMPF) module integrates these representations via the hierarchical attention mechanism. Experimental results demonstrate that PMAFNet achieves 95.24% accuracy. Ablation analysis confirms nutrient intake is a critical driver, with accuracy declining by 19.04% upon its exclusion. Furthermore, the SHapley Additive exPlanations (SHAP) was employed to interpret the PMAFNet and identify key features influencing HbA1c. This study highlights the significance of incorporating multimodal data to unravel the complexity of glycemic regulation and reveals the pivotal influence of nutrient intake in diabetes management.
Huaiyan Jiang, Jing Liu 0002, Haoyu Gu, Yu Liu 0004
IEEE Internet Things J.5
2025 SubMap: A Partial Mapping Strategy for CGRA Based on sub-CGRA Exploration
abstract
Coarse-grained reconfigurable array (CGRA) is a quality hardware for compute-intensive loop kernels, with its excellent balance of performance, energy efficiency, and reconfigurability. However, the efficiency of CGRA depends heavily on how the compiler maps the data flow graph (DFG) extracted from application kernels onto the target architecture. Most existing CGRA compilers encounter the challenge of long compilation times due to excessive exploration space. To reduce the exploration space and compilation time, we propose SubMap, which adaptively explores a suitable sub-CGRA for different DFGs in a target CGRA and efficiently performs the mapping process. The experimental results show that SubMap greatly reduces the compilation time compared to the latest methods while maintaining the mapping quality. On HyCube$4\times 4$, SubMap has an average performance improvement of$9.47 \times $and$11.67 \times $, respectively, compared with Morpher (Pathfinder) and Morpher (SA). As the scale of the target CGRA increases, the performance improvement of SubMap becomes more pronounced.
Peiguang Jing, Sio-Hang Pun, Yu Liu 0004
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2025 Sequential Spectral-Spatial Feature Convolution Network With Self-Attention for Remote Sensing Hyperspectral Image Classification
abstract
The rich spatial and spectral information in hyperspectral images (HSIs) makes spectral-spatial relationships essential for HSI classification (HSIC). Recent advancements indicate convolutional neural networks (CNNs) excel in HSIC but often struggle with precise spectral feature extraction. Moreover, the abundance of spectral information presents challenges in efficient feature representation and minimizing cross-domain interference. To address these limitations, we propose an efficient sequential spectral-spatial feature convolution network (S3FCN), employing successive subnetworks for spectral and spatial feature extraction with depthwise separable convolution. This approach balances the preservation of deep spectral and spatial features while significantly reducing network parameters, enhancing both performance and computational efficiency. We also introduce a sequential spectral-spatial attention module (S3AM) to integrate cross-domain correlations. This module utilizes spectral features from the preceding subnetwork and multilevel residual layers for in-depth exploration of spatial features, enabling deep integration for improved classification performance. The proposed architecture’s effectiveness is verified on five benchmark HSI datasets, including Pavia University, Salinas Valley, Kennedy Space Center, Indian Pines, and Houston 2013. Experimental results demonstrate that the sequential spectral-spatial connection in the feature extraction and attention mechanism integrated with depthwise separable convolution collectively surpasses current state-of-the-art (SOTA) techniques in classification accuracy with overall accuracies of 98.28%, 97.63%, 99.31%, 96.72%, and 95.38% across different datasets, while limiting the computation overhead, ensuring balanced network efficiency.
Jiqing Liu, Han Wang 0039, Renhe Liu, Shaochu Wang, Yu Liu 0004
IEEE Trans. Geosci. Remote. Sens.5
2025 Bias Mitigation and Representation Optimization for Noise-Robust Cross-Modal Retrieval
abstract
The remarkable progress in cross-modal retrieval relies on accurately annotated multimedia datasets. In practice, most existing datasets used for training cross-modal retrieval models are automatically collected from the Internet to reduce data collection costs. However, it inevitably contains mismatched pairs, i.e., noisy correspondences, thus degrading the model performance. Recent advances utilize the predicted similarity distribution of individual samples for noise validation and correction, which easily faces two challenging dilemmas: (1) confirmation bias and (2) unstable performance with increasing noise. In light of the above, we propose a generalized Bias Mitigation and Representation Optimization (BMRO) framework. Specifically, we propose a Bias Estimator (BE) to estimate the unbiased confidence factor of a sample by contrasting it against its nearest neighbors. Unbiased confidence factor can precisely adjust sample contribution and enhance accurate sample division. This facilitates the Adaptive Representation Optimizer (ARO) in providing tailored optimization strategies for clean and noisy samples. ARO performs contrastive learning between clean samples and generated hard samples, thus promoting the generalizability and robustness of the representation. Besides, it utilizes complementary learning to reduce incorrect guidance from noisy samples. Extensive experiments on five visual-text benchmarks verify that our BMRO can significantly improve the matching accuracy and performance stability against noisy correspondences.
Yu Liu 0004, Haipeng Chen 0002, Guihe Qin, Jincai Song, Xun Yang 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2024 Causality-Inspired Invariant Representation Learning for Text-Based Person Retrieval
abstract
Text-based Person Retrieval (TPR) aims to retrieve relevant images of specific pedestrians based on the given textual query. The mainstream approaches primarily leverage pretrained deep neural networks to learn the mapping of visual and textual modalities into a common latent space for cross-modality matching. Despite their remarkable achievements, existing efforts mainly focus on learning the statistical cross-modality correlation found in training data, other than the intrinsic causal correlation. As a result, they often struggle to retrieve accurately in the face of environmental changes such as illumination, pose, and occlusion, or when encountering images with similar attributes. In this regard, we pioneer the observation of TPR from a causal view. Specifically, we assume that each image is composed of a mixture of causal factors (which are semantically consistent with text descriptions) and non-causal factors (retrieval-irrelevant, e.g., background), and only the former can lead to reliable retrieval judgments. Our goal is to extract text-critical robust visual representation (i.e., causal factors) and establish domain invariant cross-modality correlations for accurate and reliable retrieval. However, causal/non-causal factors are unobserved, so we emphasize that ideal causal factors that can simulate causal scenes should satisfy two basic principles:1) Independence: being independent of non-causal factors, and 2)Sufficiency: being causally sufficient for TPR across different environments. Building on that, we propose an Invariant Representation Learning method for TPR (IRLT), that enforces the visual representations to satisfy the two aforementioned critical properties. Extensive experiments on three datasets clearly demonstrate the advantages of IRLT over leading baselines in terms of accuracy and generalization.
Yu Liu 0004, Guihe Qin, Haipeng Chen 0002, Zhiyong Cheng 0001, Xun Yang 0001
AAAI1
2024 A Novel Medical Image Fusion Framework Integrating Multi-scale Encoder-Decoder with Discrete Wavelet Decomposition
abstract
In recent years, many fusion algorithms based on multi-scale transform or neural networks have been proposed to improve medical image fusion (MIF) performance. However, there is still enormous potential to explore the combination of different fusion theories. In this paper, we propose a novel MIF framework to integrate powerful feature representation abilities of the deep learning model and accurate frequency decomposition characteristics of discrete wavelet transform (DWT). Firstly, a multi-scale encoder-decoder network is well-trained to extract feature information in different scales and achieve efficient image reconstruction. In particular, DWT is introduced into each scale to decompose the extracted features into high- and low-frequency sub-bands for information preservation during down-sampling. An elaborate feature fusion process is designed to achieve multi-scale fusion while merging different frequency sub-bands. Experiment results on benchmark datasets demonstrate that the proposed fusion framework outperforms current state-of-the-art methods with comparable time complexity in both objective and subjective evaluation.
Renhe Liu, Yu Liu 0004, Shan Du 0001
ICASSP2
2024 Asymmetric Neural Image Compression with High-Preserving Information
abstract
Recently, neural image compression has made significant progress in reducing rate-distortion and has received widespread attention. However, existing methods focus more on perfecting entropy models yet overlook the ability of their encoder networks to extract non-linear features of images, which can promote compression performance. In this paper, we design a learning-based asymmetric image compression network to enhance the feature representation capability for improved compression quality. Firstly, we propose a high-preserving information block (HPIB) consisting of a high-frequency filtering module (HFM) and a feature modulation module (FMM) to fully utilize the different frequency information in images. Secondly, we progressively use the HPIB layer to design a high-performance encoder network for high-fidelity feature extraction. Results from extensive experiments demonstrate that our network performs superior to the prior art in terms of both PSNR and MS-SSIM metrics and achieves 3.91% and 8.88 % BD-rate over VVC on the Kodak and CLIC datasets, respectively.
Yu Liu 0004, Renhe Liu, Shenghui Song 0001
ISCAS2
2024 DDIN: Enhancing Food Ingredient Recognition with Region and Category Discovery Modules
abstract
Food ingredient recognition has received numerous attention for its importance for health-related applications. But there are still some challenges due to food dishes complexity such as detecting ingredients of varying sizes and identifying multiple ingredients within a single image. Addressing this challenge, we propose a novel ingredient recognition network named Dual Discovery Integration Network (DDIN) which consists of two modules: the Region Discovery (RD) module using deconvolution to get probability distribution map for find-grained region discover, the Category Discovery (CD) module using an ingredient dictionary to capture multiple ingredient category. Finally, the output from RD and CD modules are fused to obtain the final prediction results. The experimental results demonstrate that our model achieves state-of-the-art results in ingredient recognition on the Chinese Food dataset Vireo Food-172, and it also outperform existing methods for less parameters and lower computational complexity. Further visualization of discovered ingredient regions also shows the superiority of our method.
Yiheng Ru, Huaiyan Jiang, Hang Song 0001, Bo Wei 0001, Yu Liu 0004
VCIP5
2024 MSCFormer: Multi-Scale Circular Transformer for Image Deblurring
abstract
Currently, with the extensive application of digital cameras in dynamic capturing, implications such as camera jitter, out-of-focus, and target motion induce various types and degrees of image blurring. Deep learning (DL) is a powerful technique that offers data-adaptive recovery without prior characterization of deblurring filter kernels. However, end-to-end networks can still be improved to restore regions with severe localized blurring. Therefore, we propose a multi-scale circular transformer (MSC-Former) employing averaged neighborhood attention (AvgNA) to solve this problem. It computes the local attention of each feature pixel by learning the correlation between the center and the surrounding windowed neighborhood, then produces integrated attention with direct averaging. We employ a multi-scale circular strategy (MSCS) to compute attention at different spatial scales to expand the receptive field while maintaining a low parameter count. It uses concentric circular regions with varying radii to define neighborhoods at different scales, which expands the receptive field during attention computation while capturing spatial continuity across larger neighborhoods. Experimental results demonstrate that the proposed method surpasses the recent state-of-the-art deblurring techniques on the benchmark dataset.
Renhe Liu, Bo Wei 0001, Yu Liu 0004
VCIP6
2024 Dual Preference Perception Network for Fashion Recommendation in Social Internet of Things
abstract
Nowadays, with the continuous development of information technology, the application scenarios of the Internet of Things (IoT) are progressively expanding to the social field, engendering widespread attention to the Social IoT (SIoT). Personalized fashion recommendation that possesses the potential to establish social relationships between clothing and humans has substantially broadened the scope of the SIoT, particularly with the flourishing fashion industry and the ascent of smart home. Compared to conventional recommendations, fashion recommendation generally suggests a collection of items rather than individual pieces for users. Additionally, considering the public acceptance alongside the user-specific preference is reasonable for fashion recommendation, however, current methods often overlook the former. To comprehensively capture the public acceptance and the user-specific preference, we propose a dual preference perception network (DP2Net) for fashion recommendation. First, a fashion corpus is constructed to facilitate the condensation of general taste, wherein adversarial learning and determinantal point process are leveraged to ensure representativeness and diversity of the corpus. Second, a user-general preference perception module is built based on a bottleneck transformer structure to generate aggregated representations for the corpus. Third, a user-specific preference perception module is constructed to acquire collaborative representations of users and outfits by employing an attentive heterogeneous graph embedding. The final loss functions of two preference perception modules are constructed by combining the representations of users, outfits, and the corpus. Experiments on large-scale real-world data sets demonstrate the effectiveness of the proposed method. To facilitate reproducible research, we have made our code publicly available athttps://github.com/KaiZhang1228/DP2Net.
Peiguang Jing, Xianyi Liu, Yun Li 0006, Yu Liu 0004, Yuting Su 0001
IEEE Internet Things J.5
2024 Multimodal deep hierarchical semantic-aligned matrix factorization method for micro-video multi-label classification
Fugui Fan, Yuting Su 0001, Yun Liu 0009, Peiguang Jing, Kaihua Qu, Yu Liu 0004
Inf. Process. Manag.6
2024 HiEval: A scheduling performance estimation approach for spatial accelerators via hierarchical abstraction
Yue Hu 0004, Wei Lu 0026, Yu Liu 0004
J. Syst. Archit.5
2024 A deep low-rank semantic factorization method for micro-video multi-label classification
Fugui Fan, Yuting Su 0001, Yu Liu 0004, Peiguang Jing, Kaihua Qu
Multim. Syst.3
2024 Reinforced visual interaction fusion radiology report generation
Haipeng Chen 0002, Yu Liu 0004, Yingda Lyu
Multim. Syst.3
2024 Weakly Supervised Video Re-Localization Through Multi-Agent-Reinforced Switchable Network
abstract
The objective of video re-localization (VRL) is to localize a successive sequence of frames, namely, the target moment, from untrimmed reference videos that semantically correspond to a given query video. During training, the weakly supervised setting of VRL provides only coarse-grained video-level rather than frame-level annotations. For the weakly supervised VRL (WS-VRL) task, obtaining effective video feature representations that can be used to evaluate the relevance between videos and localizing the accurate temporal boundaries of the target moment remain challenging. In this paper, a novel multi-agent-reinforced switchable network (MARS) is proposed to address these challenges. MARS can adaptively guide video feature encoding and moment localization using multiple learned agents. Specifically, an agent-controlled switchable encoder is used to obtain effective video feature representations, and an agent-reinforced boundary localizer is used to determine accurate localized moments through progressive refinement. Furthermore, a relevance-oriented reward generator was designed to estimate the relevance of the localized moment to the query video and assign a reward to multiple agents. The effectiveness of the proposed MARS model was verified through extensive experiments on the ActivityNet-VRL dataset.
Yuan Zhou 0006, Axin Guo, Shuwei Huo, Yu Liu 0004, Sun-Yuan Kung
IEEE Trans. Circuits Syst. Video Technol.4
2024 Regular Constrained Multimodal Fusion for Image Captioning
abstract
More diverse and closer to human-like captions are of paramount importance in image captioning. Recent research has achieved significant advancements, with the majority adopting end-to-end encoder-decoder architectures that integrate specific feature-text processing. However, the homogeneity of their model structures, the simplicity or complexity of feature-text fusion, and the uniformity of training objectives have all to some extent affected the diversity and effectiveness of caption generation, thus limiting the potential applications of this task. Therefore, in this paper, we propose the Regular Constrained Multimodal Fusion (RCMF) method for image captioning to better integrate information across and within modalities, while also approaching human-like fine-grained semantic perception and relationship reasoning capabilities. Initially, our RCMF preprocesses images using a Swin-Transformer and then an extended encoder with a new intra-modal fusion module, utilizing window-focused linear attention to capture features and leveraging refined grid and global visual features. By combining text features, RCMF employs a cross-modal fusion module and decoder to deeply model the interaction between text and image. Additionally, RCMF first introduces a new additional regulatory modal fusion reasoning (MFR) branch, which surpasses the above architectures. Its MFR loss combined with cross-entropy loss forms a new training objective strategy, effectively mining fine-grained relationships between images and text, perceiving the semantic information of images and their corresponding captions, thereby regulating the generated captions to be more diverse and human-like. Experimental results based on the MS COCO 2014 dataset, particularly under the same experimental conditions, demonstrate the outstanding performance of our method, especially in terms of METEOR, ROUGE-L, CIDEr, and SPICE metrics. Visualization results further intuitively confirm the effectiveness of our RCMF method. Source code inhttps://github.com/200084/RCMF-for-image-caption.
Haipeng Chen 0002, Yu Liu 0004, Yingda Lyu
IEEE Trans. Circuits Syst. Video Technol.3
2024 Deep Learning-Based Eye-Tracking Analysis for Diagnosis of Alzheimer's Disease Using 3D Comprehensive Visual Stimuli
abstract
Alzheimer's Disease (AD) is a neurodegenerative disorder that causes a continuous decline in cognitive functions and eventually results in death. An early AD diagnosis is important for taking active measures to slow its deterioration. Traditional diagnoses are usually based on clinical experience, which is limited by several realistic factors. In this paper, we focus on exploiting deep learning techniques to diagnose AD based on eye-tracking behaviors. Visual attention, as a typical eye-tracking behavior, is of great clinical value in detecting cognitive abnormalities in AD patients. To better analyze the differences in visual attention between AD patients and normals, we first conducted a 3D comprehensive visual task on a noninvasive eye-tracking system to collect visual attention heatmaps. Then a multilayered comparison convolutional neural network (MC-CNN) is proposed to distinguish the visual attention differences between AD patients and normals. In MC-CNN, the multilayered feature representations of heatmaps were obtained by hierarchical residual blocks to better encode eye movement behaviors, which were further integrated into a distance vector to benefit the comprehensive visual task. From evaluation, MC-CNN can distinguish AD patients from normals with 0.84 accuracy, 0.86 recall, 0.82 precision, 0.83 F1-score and 0.90 area under the curve (AUC). The above results demonstrate the effectiveness of the proposed MC-CNN in AD diagnosis based on the comprehensive 3D visual task.
Fangyu Zuo, Peiguang Jing, Jinglin Sun, Jizhong Duan, Yu Liu 0004
IEEE J. Biomed. Health Informatics6
2024 Dual-Domain Aligned Deep Hierarchical Matrix Factorization Method for Micro-Video Multi-Label Classification
abstract
Recently, with the growing popularity of micro-videos, multi-label learning has attracted increasing attention due to its potential commercial value in different scenarios. However, existing methods place more emphasis on the alignment between explicit semantics and visual features, while neglecting the exploration of interactions at fine-grained semantic levels. To address this problem, we propose a novel dual-domain aligned deep hierarchical matrix factorization (DADHMF) method for micro-video multi-label classification. Specifically, we construct a dual-stream deep matrix factorization framework to explore implicit hierarchical semantics and corresponding intrinsic feature representations in top-down and bottom-up ways, respectively. On this basis, we leverage the intralayer alignment strategy to narrow the semantic gap between label and instance domains by introducing adaptive semantic-aware embeddings. Moreover, we further utilize the inverse covariance estimation module to automatically capture latent semantic correlations, and project the structural information into the semantic-aware embeddings to ensure the stability of the intralayer alignment. Extensive experiments on two available micro-video multi-label datasets demonstrate that our proposed method outperforms the state-of-the-art methods.
Fugui Fan, Yuting Su 0001, Liqiang Nie, Peiguang Jing, Daozheng Hong, Yu Liu 0004
IEEE Trans. Multim.6
2024 VMemNet: A Deep Collaborative Spatial-Temporal Network With Attention Representation for Video Memorability Prediction
abstract
Video memorability measures the degree to which a video is remembered by different viewers and has shown great potential in various contexts, including advertising, education, and health care. While extensive research has been conducted on image memorability, the study of video memorability is still in its early stages. Existing methods in this field primarily focus on coarse-grained spatial feature representation and decision fusion strategies, overlooking the crucial interactions between spatial and temporal domains. Therefore, we propose an end-to-end collaborative spatial-temporal network called VMemNet, which incorporates targeted attention mechanisms and intermediation fusion strategies. This enables VMemNet to capture the intricate relationships between spatial and temporal information and uncover more elements of memorability within video visual features. VMemNet integrates spatially and semantically guided attention modules into a dual-stream network architecture, allowing it to simultaneously capture static local cues and dynamic global cues in videos. Specifically, the spatial attention module is used to aggregate more memorable elements from spatial locations, and the semantically guided attention module is used to achieve semantic alignment and intermediate fusion of the local and global cues. In addition, two types of loss functions with complementary decision rules are associated with the corresponding attention modules to guide the training process of the proposed network. Experimental results obtained on a publicly available dataset verify that the proposed VMemNet approach outperforms all current single- and multi-modal methods in terms of video memorability prediction.
Wei Lu 0026, Jiaze Han, Peiguang Jing, Yu Liu 0004, Yuting Su 0001
IEEE Trans. Multim.5
2024 Multimodal Attentive Representation Learning for Micro-video Multi-label Classification
abstract
As one of the representative types of user-generated contents (UGCs) in social platforms, micro-videos have been becoming popular in our daily life. Although micro-videos naturally exhibit multimodal features that are rich enough to support representation learning, the complex correlations across modalities render valuable information difficult to integrate. In this paper, we introduced a multimodal attentive representation network (MARNET) to learn complete and robust representations to benefit micro-video multi-label classification. To address the commonly missing modality issue, we presented a multimodal information aggregation mechanism module to integrate multimodal information, where latent common representations are obtained by modeling the complementarity and consistency in terms of visual-centered modality groupings instead of single modalities. For the label correlation issue, we designed an attentive graph neural network module to adaptively learn the correlation matrix and representations of labels for better compatibility with training data. In addition, a cross-modal multi-head attention module is developed to make the learned common representations label-aware for multi-label classification. Experiments conducted on two micro-video datasets demonstrate the superior performance of MARNET compared with state-of-the-art methods.
Peiguang Jing, Xianyi Liu, Yun Li 0006, Yu Liu 0004, Yuting Su 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2023 A Visible and Infrared Image Fusion Framework Based on Dual-Path Encoder-Decoder and Multi-Scale Discrete Wavelet Transform
abstract
In recent years, extensive research has been conducted on visible and infrared image fusion (VIF) task using traditional multi-scale transform-based and deep learning model-based methods. However, there is still a need to explore the combination of neural networks and multi-scale transform. This paper proposes a novel fusion framework based on a dual-path encoder-decoder and multi-scale transform. A dual-path encoder is trained to extract rich features at different depths from source images, while a shared decoder is trained to efficiently reconstruct images from the extracted feature space. We apply the discrete wavelet transform (DWT) to generate various frequency components from the extracted features. A fusion module is utilized to achieve fusion for low and high-frequency sub-bands, respectively, which is constrained by a gradient-based fusion loss function and an absolute values maximum-selection strategy. Our proposed method is superior to current state-of-the-art fusion methods, as demonstrated through quantitative and qualitative comparisons of publicly available datasets.
Renhe Liu, Shan Du 0001, Yu Liu 0004
ICIP4
2023 Cross-domain Prototype Contrastive loss for Few-shot 2D Image-Based 3D Model Retrieval
abstract
2D image-based 3D model retrieval (IBMR) usually relies on abundant explicit supervision on 2D images, together with unlabeled 3D models to learn domain-aligned yet class-discriminative features for the retrieval task. However, collecting large-scale 2D labels is cost-effective and time-consuming. Therefore, we explore a challenging IBMR task, where only few-shot labeled 2D images are available while the rest of the 2D and 3D samples remain unlabeled. Limited annotation of 2D images further increases the difficulty of domain-aligned yet discriminative feature learning. Therefore, we propose cross-domain prototype contrastive loss (CPCL) for the few-shot IBMR task. Specifically, we capture semantic information to learn class-discriminative features in each domain by minimizing intra-domain prototype contrastive loss. Besides, we perform inter-domain transferable contrastive learning to align the features between instances and prototypes of the same class across domains. Comprehensive experiments on popular benchmarks, MI3DOR and MI3DOR-2, validate the superiority of CPCL.
Yaqian Zhou 0002, Yu Liu 0004, Dan Song 0006, Jiayu Li 0004, Xuanya Li, Anan Liu
ICME2
2023 Non-local low-rank constraint-based self-consistent PMRI reconstruction using eigenvector maps
abstract
Abstract Eigenvector‐based SPIRiT (ESPIRiT) can estimate multiple sets of coil sensitivity maps from the calibration matrix constructed from the auto‐calibration data. Recently, the L1 norm and total variation were combined with the ESPIRiT model to improve the reconstruction quality of magnetic resonance (MR) images. To further improve the reconstruction performance, the non‐local low‐rank regularisation term is incorporated into the ESPIRiT model (NLR‐ESPIRiT) is proposed. The proposed NLR‐ESPIRiT model takes full advantage of the non‐local self‐similarity features of MR images. The resulting optimisation problem can be transformed into a gradient problem and a denoising problem with low‐rank constraints using the operator splitting technique. The weighted nuclear norm (WNN) is applied as a surrogate of the rank. Then the denoising subproblem with the WNN can be effectively solved by using the alternating direction method of multipliers technique. For practical applications, a parameter‐selecting method is proposed to obtain almost optimal parameters for the same kind of MR images. Simulation experiments on in vivo data sets demonstrate that the proposed NLR‐ESPIRiT outperforms all competing traditional model‐based algorithms in terms of three objective metrics and visual comparison.
Jizhong Duan, Yu Liu 0004
IET Signal Process.3
2023 Internet of Things for Diagnosis of Alzheimer's Disease: A Multimodal Machine Learning Approach Based on Eye Movement Features
abstract
Alzheimer’s disease (AD) is a degenerative neurological disease that occurs in the elderly with typical symptoms of decline in cognition, manifested by eye movement behaviors. The key to AD treatment requires early detection of cognitive impairment, which relies on frequent medical screening. This article proposes an Internet of Things (IoT) architecture constructed with eye-tracker (ET) nodes and cloud-based diagnosis enabled by machine learning (ML), which can provide convenient screening of oculomotor abnormalities and automatic identification of early-stage AD. The bespoke ET nodes collect 3-D oculomotor responses from diverse stereo video stimulation trials and transmit data into a dedicated multimodal ML (MMML) algorithm in the cloud. The algorithm incorporates multimodal features extracted from diverse oculomotor types to enhance the classification accuracy and optimized data dimension reduction in feature fusion to improve the classifier’s performance. From evaluation, the proposed method can distinguish AD patients from the control group with 86% accuracy (ACC), 78% true positive rate (TPR), and 90% positive predictive value (PPV). The results confirm the effectiveness of our MMML algorithm in AD diagnosis with the fusion of multimodal oculomotor features and prove the feasibility of our IoT-powered eye-tracking solution for AD screening.
Yunpeng Yin, Han Wang 0039, Jinglin Sun, Peiguang Jing, Yu Liu 0004
IEEE Internet Things J.6
2023 Unsupervised self-training correction learning for 2D image-based 3D model retrieval
Yaqian Zhou 0002, Yu Liu 0004, Jun Xiao 0001, Min Liu 0008, Xuanya Li, Anan Liu
Inf. Process. Manag.2
2023 M-AResNet: a novel multi-scale attention residual network for melting curve image classification
Pengxiang Su, Xuanjing Shen, Haipeng Chen 0002, Di Gai, Yu Liu 0004
Multim. Tools Appl.5
2022 A Novel Lightweight Network for Fast Monocular Depth Estimation
abstract
Depth estimation is of growing interest in many sectors, from robotics to wearable augmented reality gears. Monocular depth estimation attracts more attention due to its cost efficiency and low complexity. Most recent research has developed very large and resource intensive networks which are not suitable for small systems with limited resources. In this paper, we propose a lightweight network which leverages the advantages of dimension-wise convolutions and depthwise separable convolutions to reduce complexity in the architecture. In particular, the proposed depth estimation architecture utilizes a novel DICE unit-based encoder, optimized for a lightweight encoder-decoder structure. Furthermore, we propose a DICE unit-based decoder structure as well as an optimized depthwise separable convolution-based decoder. Both decoders follow a similar five-layer architecture. In the experiments, we have demonstrated the effectiveness of the proposed architecture as well as the comparison between the two proposed decoders. Our novel lightweight network has a significant decrease in both size and complexity at a marginal cost to accuracy when compared to other state-of-the-art lightweight networks.
Tim Heydrich, Yimin Yang 0001, Yu Liu 0004, Shan Du 0001
ICASSP4
2022 A Novel Predictor with Optimized Sampling Method for Hardware-aware NAS
abstract
Designing suitable neural networks in resource-constrained scenarios is a very challenging problem. It means the network architecture must not only have excellent performance on the target task but also meet the requirements of the deployment platform. In previous works, hardware-aware neural architecture search is proposed to automatically search for such architectures by integrating hardware attributes into the neural architecture search (NAS) process. However, hardware-aware NAS methods suffer from the difficulty of architecture evaluation and multi-objective optimization. Using a predictor to directly output the defined evaluation score of the architecture is a good solution, which is both efficient and accurate.In this work, we design a novel multi-objective predictor framework, which better reflects the differences in multi-objective optimization and is more suitable for hardware-aware NAS. We transfer the knowledge of single-objective predictors from different perspectives to the designed multi-objective predictor structure and define the decoding rule with the specific multi-objective score. In addition, higher-performance architecture sampling (HPAS) is proposed to selectively collect training samples from the search space instead of simple random sampling. Because for NAS, the status of architectures is unequal, and high-performance architectures are more important than low-performance architectures. By increasing the proportion of the former in the training dataset, we can obtain better performance under the same actual measurements. Experimental results in multiple scenarios based on NAS-BENCH-201 and HW-NAS-BENCH prove the effectiveness of our methods. This work can also be applied to other hardware platforms, such as control chips in specific fields.
Yue Hu 0004, Chongfei Shen, Yu Liu 0004
ICPR5
2022 Learning Transferable and Discriminative Representations for 2D Image-Based 3D Model Retrieval
abstract
Existing research on the 2D image-based 3D model retrieval task focuses on learning transferable representations directly to narrow the domain discrepancy. However, it is not easy to achieve in practice due to the significant variations across two domains. In addition, some methods design a domain discriminator to distinguish the feature arising from source or target domains for transferable feature representations learning, which will lead to an unexpected deterioration of the feature discriminability. To settle these problems, we propose jointly learning transferable and discriminative representations for 2D image-based 3D model retrieval. Specifically, we extract features from the 2D images and 3D models (described as multiple views) by CNN. Considering the difficulty of directly narrowing the discrepancy of two domains, we are prone to connect 2D image and 3D model domains to an intermediate domain, where the domain gap aims to be eliminated. However, the feature transferability does not denote well discriminability. Based on the batch spectral penalization (BSP) theory, the feature transferability is dominated by feature vectors with higher singular values, while the feature discriminability depends on more eigenvectors with lower singular values to convey rich discriminative structures. Therefore, we penalize the largest singular values so that the feature vectors with lower singular values are appropriately enhanced, thereby strengthening feature discriminability. A series of experiments on two challenging datasets, MI3DOR and MI3DOR-2, indicate that our method can significantly improve performance.
Yaqian Zhou 0002, Yu Liu 0004, Heyu Zhou, Zhiyong Cheng 0001, Xuanya Li, Anan Liu
IEEE Trans. Circuits Syst. Video Technol.2
2021 Illumination-based adaptive saliency detection network through fusion of multi-source features
Chunxu Jiang, Yu Liu 0004, Jinglin Sun, Jichang Guo, Wei Lu 0026
J. Vis. Commun. Image Represent.2
2021 Wasserstein distance feature alignment learning for 2D image-based 3D model retrieval
Yaqian Zhou 0002, Yu Liu 0004, Heyu Zhou, Wenhui Li 0001
J. Vis. Commun. Image Represent.2
2019 A Lightweight Neural Network For Crowd Analysis Of Images With Congested Scenes
abstract
For images with congested scenes, the task of crowd analysis, including crowd counting and crowd distribution prediction, becomes very difficult. To address these issues, various CNN-based approaches have been proposed. However, those methods usually have a large number of parameters and require huge computing resources. In this paper, we focus on low-complexity approaches and propose a lightweight endto-end network for crowd analysis. Our method utilizes an effective scale-aware module to extract multi-scale features and then regresses these features to density maps. The proposed network is consisted by three parts: multi-scale feature extraction, density map estimation and density map correction, and the network which only contains 0.86 M parameters (Lightweight). According to our experiments, our proposal obtain a better result than other existing methods on several testing sequences.
Shan Du 0001, Yu Liu 0004
ICIP3
2019 Fast Palette Mode Decision Methods for Coding Game Videos With HEVC-SCC
abstract
The live broadcasting of game playing is becoming more and more popular. The game video is one of the typical video data which should be coded with high efficiency video coding-screen content coding (HEVC-SCC). In HEVC-SCC, the palette mode is an important technique which can effectively enhance the coding performance. However, this technique also demands high-coding computation. In this paper, we propose a fast palette mode pre-decision method based on the analysis of the color complexity of the video data. Considering only the coding computation of palette mode of HEVC-SCC, our algorithm saves about 74.24% computation in the experiments of game videos. Whereas, the BDBR only increases by about 0.36% and the BDPSNR decreases by about 0.03 dB on average.
Yu Liu 0004, Jinglin Sun, Xiangdong Huang 0002
IEEE Trans. Circuits Syst. Video Technol.1
2015 Interface MB-Based Video Content Editing Transcoding
abstract
In practical multimedia systems, the content of coded video streams often needs to be re-edited at the nodes of transmitting networks. For example, logo insertion is always required for copyright protection at different local transmitting nodes. This kind of video stream editing is denoted as video content editing transcoding (VCET) in this paper. Though some techniques have been suggested for VCET, these methods cannot meet the requirement of dynamic transmitting bandwidth in practical applications. In this paper, we proposed an interface macroblock-based transcoding scheme for VCET, which can reuse the variable length codes of the original video streams as much as possible to achieve the best VCET quality. In order to ensure that the edited video streams can be transmitted by the original bandwidth, we also proposed a rate control algorithm for VCET, which can accurately control the bitrate of edited video streams according to the frame level coding bits of the original video streams. Experimental results showed that the proposed scheme achieved substantially better results in bitrate accuracy, computational complexities, and video quality than many other existing schemes.
Yu Liu 0004, Jizhong Duan, Shaochu Wang, Shenghui Song 0001
IEEE Trans. Circuits Syst. Video Technol.1
2014 A Simple and Efficient Re-Scrambling Scheme for DTV Programs
abstract
In order to guarantee pay-TV services, data of digital television (DTV) programs are scrambled by conditional access systems (CAS). In practical applications, some DTV transmitting nodes need descramble the scrambled DTV programs for editing purposes. After editing, how to re-scramble these edited programs is a challenging task since building CAS at transmitting nodes is expensive and insecure. In this paper, we proposed a novel scheme to solve this problem. Together with the recommended scheme, we also proposed techniques regarding video key data selection and extraction, synchronization of scrambled and descrambled transport stream (TS) packets, and scrambled/unscrambled interleaving multiplexer. The proposed scheme requires much lower complexity than existing methods while maintaining enough security for practical applications. Neither professional CAS equipment nor real-time common key (CK) transmitting is required by our re-scrambling scheme. Various experimental results demonstrated that the proposed re-scrambling scheme achieved superior performance in practical DTV systems, and it also obtained good compatibility with different CAS algorithms.
Yu Liu 0004, Jizhong Duan, Qiang Tang 0006, Yongdong Zhang 0001
IEEE Trans. Multim.1
2013 Bregman Iteration Based Efficient Algorithm for MR Image Reconstruction From Undersampled K-Space Data
abstract
It is difficult to solve Magnetic Resonance (MR) image reconstruction problems with linear combinations of total variation and ℓ1norm regularization terms. In order to solve these compound regularization problems, we propose an efficient algorithm in this letter. The proposed algorithm adopts the Bregman iteration technique to convert the original constrained problem to a sequence of unconstrained problems, which are then solved using operator splitting and variable splitting techniques. Simulation experiments demonstrate that significant improvement of the quality of reconstructed images is achieved by the proposed algorithm when compared to previous methods.
Jizhong Duan, Yu Liu 0004
IEEE Signal Process. Lett.2