EDBT 2026 Demo / reviewers in the wild / expert
Jing Liu 0002
dblp:72/2590-2
· DBLP profile ↗
81ranked-venue papers
27as first author
47since 2021 · last 2027
0000-0003-4690-1886ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 66 · 21 first-author · 40 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 5 since 2021Systems, architecture and hardware · 4 · 4 first-author · 1 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | CProPNet: Chrono-progressive dietary synergy network with pathology decoupling for individualized HbA1c prediction
Jing Liu 0002, Peiguang Jing, Yu Liu 0004 |
Expert Syst. Appl. | 2 |
| 2026 | GeoX-Bench: Benchmarking Cross-View Geo-Localization and Pose Estimation Capabilities of Large Multimodal ModelsabstractLarge multimodal models (LMMs) have demonstrated remarkable capabilities across a wide range of tasks, however their knowledge and abilities in the cross-view geo-localization and pose estimation domains remain unexplored, despite potential benefits for navigation, autonomous driving, outdoor robotics, etc. To bridge this gap, we introduce GeoX-Bench, a comprehensive Benchmark designed to explore and evaluate the capabilities of LMMs in cross-view Geo-localization and pose estimation. Specifically, GeoX-Bench contains 10,859 panoramic-satellite image pairs spanning 128 cities in 49 countries, along with corresponding 755,976 question-answering (QA) pairs. Among these, 42,900 QA pairs are designated for benchmarking, while the remaining are intended to enhance the capabilities of LMMs. Based on GeoX-Bench, we evaluate the capabilities of 25 state-of-the-art LMMs on cross-view geo-localization and pose estimation tasks, and further explore the empowered capabilities of instruction-tuning. Our benchmark demonstrate that while current LMMs achieve impressive performance in geo-localization tasks, their effectiveness declines significantly on the more complex pose estimation tasks, highlighting a critical area for future improvement, and instruction-tuning LMMs on the training data of GeoX-Bench can significantly improve the cross-view geo-sense abilities. Yushuo Zheng, Jiangyong Ying, Huiyu Duan, Chunyi Li 0001, Jing Liu 0002, Xiaohong Liu 0001, Guangtao Zhai |
AAAI | 6 |
| 2026 | Multi-stage Superpixel-guided Mamba-based Network for Change Detection
Jing Liu 0002, Yuting Su 0001, Peiguang Jing |
ISCAS | 1 |
| 2026 | Quality Assessment and Distortion-Aware Saliency Prediction for AI-Generated Omnidirectional ImagesabstractWith the rapid advancement of Artificial Intelligence Generated Content (AIGC) techniques, AI generated images (AIGIs) have attracted widespread attention, among which AI generated omnidirectional images (AIGODIs) hold significant potential for Virtual Reality (VR) and Augmented Reality (AR) applications. AI generated omnidirectional images exhibit unique quality issues, however, research on the quality assessment and optimization of AI-generated omnidirectional images is still lacking. To this end, this work first studies the quality assessment and distortion-aware saliency prediction problems for AIGODIs, and further presents a corresponding optimization process. Specifically, we first establish a comprehensive database to reflecthumanfeedback for AI-generatedomnidirectionals, termed OHF2024, which includes both subjective quality ratings evaluated from three perspectives and distortion-aware salient regions. Based on the constructed OHF2024 database, we propose two models with shared encoders based on the BLIP-2 model to evaluate the human visual experience and predict distortion-aware saliency for AI-generated omnidirectional images, which are named as BLIP2OIQA and BLIP2OISal, respectively. Finally, based on the proposed models, we present an automatic optimization process that utilizes the predicted visual experience scores and distortion regions to further enhance the visual quality of an AI-generated omnidirectional image. Extensive experiments show that our BLIP2OIQA model and BLIP2OISal model achieve state-of-the-art (SOTA) results in the human visual experience evaluation task and the distortion-aware saliency prediction task for AI generated omnidirectional images, and can be effectively used in the optimization process. The database and codes will be released on https://github.com/IntMeGroup/AIGCOIQA to facilitate future research. Huiyu Duan, Jing Liu 0002, Menghan Hu, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Multi-Dimensional Quality Assessment for Single-Image-to-3D Contents: Dataset and ModelabstractThe rapid advancement of AI generation technologies has led to the widespread use of AI-generated multimedia content, including images, videos, and 3D contents, across various applications. While significant progress has been made in quality evaluation for 2D content, evaluating the quality of 3D content synthesized from single image remains an underexplored problem. To bridge this gap, we introduce the first comprehensive subjective evaluation database tailored for assessing the quality of 3D content generated from single image. Our database, named AIGC-SI23DCQA, includes three distinct categories of input images, i.e., realistic images, AI-generated images, and computer graphic (CG) images, with 100 images in each category. Using five representative single-image-to-3D algorithms, we produce 1,500 3D contents and collect 94,500 annotations across three quality dimensions, including texture fidelity, shape accuracy, and overall quality. Based on the constructed database, we first benchmark and evaluate the performance of existing quality assessment methods revealing their limitations in addressing this novel task. Thus, we further propose a novel objective quality assessment method, termed I3DQA, for effective single-image-to-3D content quality assessment. Specifically, I3DQA first extracts the reference features from the source image, and the multi-modal features from the generated 3D content, including the projected video, patches, and large-multimodal model (LMM) features. These features are integrated through symmetric transformer blocks, enabling effective quality-related feature fusion and score prediction. Extensive experiments demonstrate the superior performance of our method and validate the effectiveness of its components. This work provides a foundational resource and a robust framework for advancing research in this emerging field, and our database and model are released at https://github.com/ZedFu/SI23DCQA. Huiyu Duan, Jing Liu 0002, Yun Liu 0009, Xiaohong Liu 0001, Jia Wang 0004, Xiongkuo Min, Patrick Le Callet, Guangtao Zhai |
IEEE Trans. Image Process. | 4 |
| 2025 | Healthcare IoT-Enabled PMAFNet: A Progressive Multimodal Adaptive Fusion Network for Understanding HbA1c FluctuationsabstractThe Internet of Things (IoT) has transformed healthcare via Healthcare-IoT (HIoT) by enabling continuous patient monitoring and efficient data analytics, offering promising solutions for managing the escalating global diabetes burden. The glycated hemoglobin (HbA1c) is an important biomarker for diabetes evaluation and control. Accurate prediction of HbA1c fluctuations remains a critical challenge, as existing models often overlook the interplay between stable patient profiles (demographics, lifestyle, clinical indicators) and dynamic dietary nutrient intake, limiting their ability to model glycemic responses. This study constructs a comprehensive multi-source dataset comprising detailed personal information, lifestyle, medical tests, and multi-day dynamic nutrient intake through the HIoT system. Utilizing this dataset, we propose a novel progressive multimodal adaptive fusion network (PMAFNet) to decode the complex determinants of HbA1c variability. PMAFNet employs two core modules: the multimodal graph feature-level enhancement (MGFE) module processes static and dynamic features to capture structural dependencies and temporal patterns, and the adaptive modality-aware progressive fusion (AMPF) module integrates these representations via the hierarchical attention mechanism. Experimental results demonstrate that PMAFNet achieves 95.24% accuracy. Ablation analysis confirms nutrient intake is a critical driver, with accuracy declining by 19.04% upon its exclusion. Furthermore, the SHapley Additive exPlanations (SHAP) was employed to interpret the PMAFNet and identify key features influencing HbA1c. This study highlights the significance of incorporating multimodal data to unravel the complexity of glycemic regulation and reveals the pivotal influence of nutrient intake in diabetes management. Huaiyan Jiang, Jing Liu 0002, Haoyu Gu, Yu Liu 0004 |
IEEE Internet Things J. | 2 |
| 2025 | What and Where: Semantic Grasping and Contextual Scanning for Moment Retrieval and Highlight DetectionabstractThe current surge in video content highlights the tasks of moment retrieval (MR) and highlight detection (HD), which involve localizing video segments of events and predicting clip-wise saliency scores based on text queries. The recent methods, while effective, may overlook two aspects: 1) Multimodal features often show weak alignment from frozen encoders, hindering thorough semantic exploration of video clips through fine-grained cross-modal interaction. 2) Due to the absence of significant distinction between adjacent video clips, it is challenging for clip-level context modeling to accurately locate query-relevant content. To mitigate these gaps and inspired by the human routine in understanding visual events, we propose a progressive framework dubbed “what and where” to initially grasp the aligned semantics of each video clip, and then proceed to scan moment-level contextual features temporally to identify events matching the query. In the ‘what’ stage, to enable explicit alignment of modal features and achieve a thorough semantic understanding, we firstly devise the Initial Semantic Projection (ISP) loss to bring closer different modal features with similar semantics. Additionally, we develop a Clip Semantic Mining module to deeply mine the relevance of these identified semantics to the specific query (at both word- and sentence-level). In the ‘where’ stage, to enhance feature distinctiveness, we design a Multi-Context Perception module that models moment-level context. It includes an Event Context (EC) branch and a Chronological Context (CC) branch, focusing on possible query-relevant event moments and temporal moments of various lengths. Finally, extensive experiments validate the state-of-the-art performance of our W2W model on three benchmark datasets without additional pre-training. Codes are available athttps://github.com/TJUMMG/W2W. Jing Liu 0002, Zhuo He, Weizhi Nie, Zongbing Zhang, Yuting Su 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | KA-MIN: Knowledge-Aware Multimodal Interaction Network for Emotion Recognition in ConversationabstractEmotion recognition in conversations (ERC) has garnered significant attention for its critical role in human-computer interaction systems. ERC benefits from multimodal data, which offers diverse perspectives on emotional states, and commonsense knowledge (CSK), which enriches the context by incorporating real-world understanding of human behavior. However, existing ERC studies have not fully exploited the potential of multimodal-CSK interactions for complementary information learning from these sources. To address this, we innovatively propose a Knowledge-Aware Multimodal Interaction Network (KA-MIN). KA-MIN is designed to capture complementary emotional information from CSK-multimodal interactions, thereby facilitating the ERC task. To achieve this, KA-MIN begins by combining six relation types of CSK, leveraging their differences between multimodal emotional information. The fused CSK features are then refined to incorporate context and emotional information using multimodal contextual guidance. Subsequently, we construct a novel knowledge-aware multimodal graph structure that allows the CSK information to interact with multimodal information, leading to more comprehensive multimodal and context modeling. During the graph learning process, the CSK-multimodal interactions capture the complementary emotional information between CSK and multimodal features. Finally, we dynamically fuse the multimodal emotional information with the informative CSK and textual guidance to obtain the final utterance representations, which encompass effective emotional information from both multimodal and CSK features. Extensive experiments on two popular multimodal ERC datasets demonstrate the superiority and effectiveness of the proposed KA-MIN framework. Minjie Ren, Xiangdong Huang 0002, Jing Liu 0002, Zan Gao 0001, Yuting Su 0001, Anan Liu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Subjective and Objective Audio-Visual Quality Assessment for Omnidirectional VideosabstractVirtual Reality (VR) has attracted widespread attention in recent years due to its capability to create immersive experiences by presenting multi-modal information to users. Omnidirectional videos (ODVs), as a prominent component of VR content, are essential across diverse applications. This necessitates service providers to monitor and optimize the quality of ODVs throughout the filming, encoding, decoding, and transmission stages to ensure a high-quality viewing experience. However, most existing Quality of Experience (QoE) studies for ODVs only focus on the visual quality, while overlooking the impact of the audio modality on perceptual quality. This paper presents a comprehensive study of omnidirectional audio-visual quality assessment (OD-AVQA) from both subjective and objective perspectives. Specifically, we first establish a large-scale audio-visual quality assessment database for ODVs named OAVQAD+, which includes 625 distorted omnidirectional audio-visual sequences derived from 25 pristine ODVs, and the corresponding collected mean opinion scores (MOSs) for the QoE of these ODVs. This contributes to the largest database for assessing the audio-visual quality of ODVs. To advance the fields of objective OD-AVQA, we construct a benchmark that includes three types of benchmark models. Type I and Type II models integrate well-known video quality assessment (VQA) and audio quality assessment (AQA) methods using support vector regression (SVR) and multi-layer perceptron (MLP), respectively, while Type III consists of AVQA models specifically designed for traditional 2D audio-visual sequences. We also propose a novel Omnidirectional Audio-Visual quality assessment Network (OmniAVNet) that integrates quality-aware audio, visual, and motion features to predict overall audio-visual quality for ODVs effectively, which supports both full-reference (FR) and no-reference (NR) assessment. Extensive experimental results demonstrate that OmniAVNet outperforms the aforementioned benchmark OD-AVQA models on two OD-AVQA databases, and shows great performance on one omnidirectional VQA database. The database and code are available at https://github.com/IntMeGroup/OmniAVNet. Xilei Zhu, Huiyu Duan, Yuqin Cao, Yucheng Zhu, Jing Liu 0002, Xiongkuo Min, Guangtao Zhai, Patrick Le Callet |
IEEE Trans. Image Process. | 6 |
| 2025 | DSDP: Real-Time Asymmetric Dual-Stream Instance Segmentation Embedding Depth-Predictive Architecture for Enhanced Scene UnderstandingabstractInstance segmentation can help vehicles or robots enhance their understanding of a scene through the pixel-level segmentation of different objects. However, occlusion and boundary blur, especially in cases with similar colors or textures, are still challenges encountered in real-time robust segmentation tasks. To segment a complete instance boundary, the existing 2D approaches fuse local and abstract semantic features derived from the color domain, which leads to homogeneous semantic information, and efficiently separating different objects is difficult in some cases. To address these complicated scenes, inspired by a human prediction processing strategy, where “the brain fills in missing information in advance to help make better decisions”, this study proposes a real-time asymmetric dual-stream instance segmentation algorithm embedding a depth-predictive architecture that provides the covisible depth information of objects. Furthermore, a cross-domain data fusion method and an enhancement-decoupling loss are designed to complement RGB data by utilizing the rich foreground and boundary details of the predicted depth map. In addition, our model can be fine-tuned to integrate it with real depth domain data provided by different input devices. Extensive experiments conducted on the COCO, OCHuman and CityScapes datasets demonstrate the effectiveness of our method. We further deployed our DSDP method on a UAV platform for validation purposes and qualitatively confirmed its validity. Qiang Li 0048, Weizhi Nie, Jing Liu 0002, Jingjing Geng, Yongtao Ma |
IEEE Trans. Multim. | 4 |
| 2025 | Learning to Generate Realistic Images for Bit-Depth Enhancement via Camera Imaging ProcessingabstractWith the prevalence of advanced displays devices, many attempts have been successfully made in bit-depth enhancement (BDE) to restore the low bit-depth (LBD) images to visually pleasant high bit-depth (HBD) images. However, most methods are still far from satisfactory when addressing real-world LBD images owing to their heavy dependence on LBD-HBD data pairs through direct pixel quantization. Therefore, in this paper, we propose a novel network dubbed RealGAN to generate real-world LBD images by simulating the complex quantization procedure in camera imaging process. Particularly, we design a two-mode differentiable quantization block embedded in the synthesis network facilitating adaptively simulation of the complicated quantization distortions. Furthermore, a simple residual group network is proposed in order to learn the distribution of degradation and non-linear processing in the Image Signal Processing (ISP) pipeline. In the absence of paired HBD and LBD data, the synthesis model is trained end-to-end within the generative adversarial framework using non-paired LBD and HBD images. Finally, we demonstrate that a series of BDE models can benefit from the proposed synthetic dataset and exhibit improved visual quality with sharper edges and finer textures on real-world scenes compared with the original versions trained on directly quantized LBD-HBD pairs. Jing Liu 0002, Huiyu Duan, Yuting Su 0001, Guangtao Zhai |
IEEE Trans. Multim. | 1 |
| 2025 | Aggregate and Discriminate: Pseudo Clips-Guided Boundary Perception for Video Moment RetrievalabstractVideo moment retrieval (VMR) aims to localize a video segment in an untrimmed video that is semantically relevant to a language query. The challenge of this task lies in effectively aligning the intricate and information-dense video modality with the succinctly summarized textual modality, and further localizing the starting and ending timestamps of the target moments. Previous works have attempted to achieve multi-granularity alignment of video and query in a coarse-to-fine manner, yet these efforts still fall short in addressing the inherent disparities in representation and information density between videos and queries, leading to modal misalignments. In this paper, we propose a progressive video moment retrieval framework, initially retrieving the most relevant and irrelevant video clips to the query as semantic guidance, thereby bridging the semantic gap between video modality and language modality. Futhermore, we introduce a pseudo clips guided aggregation module to aggregate densely relevant moment clips closer together and propose a discriminative boundary-enhanced decoder with the guidance of pseudo clips to push the semantically confusing proposals away. Extensive experiments on the Charades-STA, ActivityNet Captions and TACoS datasets demonstrate that our method outperforms existing methods. Jing Liu 0002, Zongbing Zhang, Yuting Su 0001, Bing Yang 0003, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Multim. | 1 |
| 2025 | Weakly Supervised Referring Video Object Segmentation With Object-Centric Pseudo-GuidanceabstractReferring video object segmentation (RVOS) is an emerging task for multimodal video comprehension while the expensive annotating process of object masks restricts the scalability and diversity of RVOS datasets. To relax the dependency on expensive mask annotations and take advantage from large-scale partially annotated data, in this paper, we explore a novel extended RVOS task, namely weakly supervised referring video object segmentation (WRVOS), which employs multiple weak supervision sources, including object points and bounding boxes. Correspondingly, we propose a unified WRVOS framework. Specifically, an object-centric pseudo mask generation method is introduced to provide effective shape priors for the pseudo guidance of spatial object location. Then, a pseudo-guided optimization strategy is proposed to effectively optimize the object outlines in terms of spatial location and projection density with a multi-stage online learning strategy. Furthermore, a multimodal cross-frame level set evolution method is proposed to iteratively refine the object boundaries considering both temporal consistency and cross-modal interactions. Extensive experiments are conducted on four publicly available RVOS datasets, including A2D Sentences, J-HMDB Sentences, Ref-DAVIS, and Ref-YoutubeVOS. Performance comparison shows that the proposed method achieves state-of-the-art performance in both point-supervised and box-supervised settings. Weikang Wang 0002, Yuting Su 0001, Jing Liu 0002, Wei Sun 0029, Guangtao Zhai |
IEEE Trans. Multim. | 3 |
| 2025 | Disentangled Denoising and Counterfactual Balance for Multimodal RecommendationabstractRecently, graph convolutional network-based dual-view multimodal recommendation methods have achieved great success. They extract multimodal and behavior features based on item-item and user-item graphs, respectively. However, they still have two- fold limitations. First, the relevance between multimodal semantics and user preferences is ignored, resulting in the propagation and coupling of preference-irrelevant noise. Second, the direct use of uneven factual user-item graphs is suboptimal, as both redundant noisy edges and missing positive interaction edges impair recommendations. To solve the above issues, we propose aDisentAngled deNoising andCounterfactual balancEmethod for multimodal recommendation, dubbed asDANCE. Specifically, for multimodal features, we explicitly disentangle them into preference-relevant and preference-irrelevant representations, to absorb and discard irrelevant noise via the latter. An orthogonal regularization and a contrastive learning task on preference relevance score prediction are proposed as the dual safeguard to prevent preference-relevant representations from encoding irrelevant noise. For behavior feature extraction, we construct a balanced user-item graph by integrating factual and counterfactual graphs. In this process, we pre-train a behavior simulator to build the counterfactual graph with full interactions. Top-$K$sampling is adopted to omit noisy edges and add missing edges in the graph. The final recommendation is performed upon the fused representation of preference-relevant multimodal and behavior representations. Extensive experiments on three public datasets verify the power of our DANCE. Xin Wen 0017, Weizhi Nie, Jing Liu 0002, Yuting Su 0001, Anan Liu |
IEEE Trans. Multim. | 3 |
| 2025 | Model Can Be Subtle: Two Important Mechanisms for Social Media Popularity PredictionabstractSocial media popularity prediction is an important channel to explore content sharing and communication on social networks. It aims to capture informative cues by analyzing multi-type data (such as user profile, image, and text) to decide the popularity of a specified post. In this article, we divide social network users into two categories (i.e., active and inactive users) and find a dilemma in existing models: If an active user publishes the low-popularity post, the model will habitually predict the high score. On the contrary, if an inactive user provides the high-popularity post, the model still gives the low score incorrectly. Therefore, how to make the model more subtle to users is important. Comparing to existing methods that directly leverage multi-modal features for regression training, this article stresses more on two novel mechanisms. The first method aims to prevent the over-fitting on user IDs. We propose the attribute-sensitive interactive mechanism (M1) by incorporating explicit user-attribute and post-attribute interaction. It can analyze which type of features a user cares the most and weaken the model’s dependence on user IDs. The second method aims to strengthen the influence of post content. We propose the knowledge embedding mechanism (M2) to revise the popularity scores in existing models by fusing the statistical frequency over multi-type data. Note that both mechanisms are model-agnostic, which can be applicable in any popularity prediction model. Extensive experiments conducted on the Social Media Prediction Dataset further validate the effectiveness. Ning Xu 0003, Jing Liu 0002, Lanjun Wang, Xuanya Li, Mengxiao Zhu 0001, Yongdong Zhang 0001, Anan Liu |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Graph Disentangled Contrastive Learning with Personalized Transfer for Cross-Domain RecommendationabstractCross-Domain Recommendation (CDR) has been proven to effectively alleviate the data sparsity problem in Recommender System (RS). Recent CDR methods often disentangle user features into domain-invariant and domain-specific features for efficient cross-domain knowledge transfer. Despite showcasing robust performance, three crucial aspects remain unexplored for existing disentangled CDR approaches: i) The significance nuances of the interaction behaviors are ignored in generating disentangled features; ii) The user features are disentangled irrelevant to the individual items to be recommended; iii) The general knowledge transfer overlooks the user's personality when interacting with diverse items. To this end, we propose a Graph Disentangled Contrastive framework for CDR (GDCCDR) with personalized transfer by meta-networks. An adaptive parameter-free filter is proposed to gauge the significance of diverse interactions, thereby facilitating more refined disentangled representations. In sight of the success of Contrastive Learning (CL) in RS, we propose two CL-based constraints for item-aware disentanglement. Proximate CL ensures the coherence of domain-invariant features between domains, while eliminatory CL strives to disentangle features within each domains using mutual information between users and items. Finally, for domain-invariant features, we adopt meta-networks to achieve personalized transfer. Experimental results on four real-world datasets demonstrate the superiority of GDCCDR over state-of-the-art methods. Jing Liu 0002, Lele Sun, Weizhi Nie, Peiguang Jing, Yuting Su 0001 |
AAAI | 1 |
| 2024 | Beyond Users: Denoising Behavior-based Contrastive Learning for Disentangled Cross-Domain Recommendation
Lele Sun, Jing Liu 0002, Shenyuan Zhang, Weizhi Nie, Anan Liu, Yuting Su 0001 |
DASFAA (2) | 2 |
| 2024 | Adaptive proposal network based on generative adversarial learning for weakly supervised temporal sentence grounding
Weikang Wang 0002, Yuting Su 0001, Jing Liu 0002, Peiguang Jing |
Pattern Recognit. Lett. | 3 |
| 2024 | Collaborative spatial-temporal video salient object detection with cross attention transformer
Yuting Su 0001, Weikang Wang 0002, Jing Liu 0002, Peiguang Jing |
Signal Process. | 3 |
| 2024 | BAND-2k: Banding Artifact Noticeable Database for Banding Detection and Quality AssessmentabstractBanding, also known as staircase-like contours, frequently occurs in flat areas of images/videos processed by compression or quantization algorithms. As undesirable artifacts, banding destroys the original image structure, thus inevitably degrading users’ quality of experience (QoE). In this paper, we systematically investigate the banding image quality assessment (IQA) problem, aiming to detect the image banding artifacts and evaluate their perceptual visual quality. Considering that the existing image banding databases only contain limited content sources and banding generation methods, and lack perceptual quality labels (i.e. mean opinion scores), we first build the largest banding IQA database so far, namedBanding Artifact Noticeable Database (BAND-2k), which consists of 2,000 banding images generated by 15 compression and quantization schemes. A total of 23 workers participated in the subjective IQA experiment, yielding over 214,000 patch-level banding class labels and 44,371 reliable image-level quality rating scores. Subsequently, we develop an effective no-reference (NR) banding evaluator for banding detection and quality assessment by leveraging frequency characteristics of banding artifacts. To be more specific, a dual convolutional neural network (CNN) is employed to concurrently learn the feature representation from the high-frequency and low-frequency maps, thereby enhancing the ability to discern banding artifacts. The quality score of a banding image is generated by pooling the banding detection maps masked by the spatial frequency filters. The experimental results demonstrate that our banding evaluator achieves remarkably high accuracy in banding detection and also exhibits high SRCC and PLCC results with the perceptual quality labels, even without directly learning a regression model for banding quality evaluation. These findings unveil the strong correlations between the intensity of banding artifacts and the perceptual visual quality, thus validating the necessity of banding quality assessment. The BAND-2k database and the proposed banding evaluator are available at https://github.com/zijianchen98/BAND-2k. Zijian Chen 0001, Wei Sun 0029, Jun Jia, Fangfang Lu, Jing Liu 0002, Ru Huang 0002, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Multi-Stage Spatio-Temporal Fusion Network for Fast and Accurate Video Bit-Depth EnhancementabstractFor video bit-depth enhancement (VBDE) tasks, inter-frame information is critical for removing false contours and recovering the details in low bit-depth (LBD) videos. However, due to different structural distortions and complex motions in the neighboring frames, it is difficult to effectively utilized inter-frame information. Most algorithms rely on alignment operations to provide information of neighboring frames, suffering from slow inference speed due to the complex alignment module design. Meanwhile, most existing methods sequentially perform the intra-frame feature extractions and inter-frame information fusions, but fail to efficiently fuse spatio-temporal information. Therefore, in this paper, we propose a two-stage progressive group (TSPG) network to find complementary information related to the target frame without adopting an alignment operation. To simultaneously achieve intra-frame feature extractions and inter-frame feature fusions, we propose a parallel spatio-temporal fusion (PSTF) module with a dual-branch spatial-temporal residual (DSTR) block to focus on more useful temporal information while ensuring a faster inference speeds. Extensive experiments on public datasets demonstrate that our proposed multi-stage spatio-temporal fusion network (named MSTFN) can quickly and effectively eliminate false contours and recover high quality target frames. Furthermore, our method outperforms the state-of-the-art methods in terms of both PSNR and SSIM, and can reach faster inference speeds. Jing Liu 0002, Ziwen Yang, Yuting Su 0001, Xiaokang Yang 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | Counterfactual Visual Dialog: Robust Commonsense Knowledge Learning From Unbiased TrainingabstractVisual Dialog (VD) requires an agent to answer the current question by engaging in a conversation with humans referring to an image. Despite the recent progress, it is beneficial to introduce external commonsense knowledge to fully understand the given image and dialog history. However, the existing knowledge-based VD models are inclined to rely on severe learning bias brought by commonsense, e.g., the retrieved$< {\mathtt{{bus}}}, {\mathtt{capable\;of}}, {\mathtt{transport\;people}}>$,$< {\mathtt{{bus}}}, {\mathtt{is\;a}}, {\mathtt{public\;transport}}>$, and$< {\mathtt{{bus}}}, {\mathtt{is\;a}}, {\mathtt{car}}>$can induce a spurious correlation between the question “What is the bus used for?” and the false answer “City bus”. There are two challenges to make commonsense learning more robust against spurious correlations: 1) how to disentangle the true effect of “good” commonsense knowledge from the whole, and 2) how to estimate and remove the effect of “bad” commonsense bias on answers. In this article, we propose a novel CounterFactual Commonsense learning scheme for the Visual Dialog task (CFC-VD). First, comparing with the causal graph of existing VD models, we add one new commonsense node and one new link to multi-modal information from history, question, and image. Since the retrieved knowledge prior is subtle and uncontrollable, we consider it as an unobserved confounder in the commonsense node, which leads to spurious correlations for the answer inference. Then, to remove the effect of the confounder, we formulate it as the direct causal effect of commonsense on answers and remove the direct language effect by subtracting it from the total causal effect via counterfactual reasoning. Experimental results certify the effectiveness of our method on the prevailing Visdial v0.9 and Visdial v1.0 datasets. Anan Liu, Ning Xu 0003, Hongshuo Tian, Jing Liu 0002, Yongdong Zhang 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | Pixel-Learnable 3DLUT With Saturation-Aware Compensation for Image EnhancementabstractThe 3D Lookup Table (3DLUT)-based methods are gaining popularity due to their satisfactory and stable performance in achieving automatic and adaptive real time image enhancement. In this paper, we present a new solution to the intractability in handling continuous color transformations of 3DLUT due to the lookup via three independent color channel coordinates in RGB space. Inspired by the inherent merits of the HSV color space, we separately enhance image intensity and color composition. The Transformer-based Pixel-Learnable 3D Lookup Table is proposed to undermine contouring artifacts, which enhances images in a pixel-wise manner with non-local information to emphasize the diverse spatially variant context. In addition, noticing the underestimation of composition color component, we develop the Saturation-Aware Compensation (SAC) module to enhance the under-saturated region determined by an adaptive SA map with Saturation-Interaction block, achieving well balance between preserving details and color rendition. Our approach can be applied to image retouching and tone mapping tasks with fairly good generality, especially in restoring localized regions with weak visibility. The performance in both theoretical analysis and comparative experiments manifests that the proposed solution is effective and robust. Jing Liu 0002, Xiongkuo Min, Yuting Su 0001, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | Inter- and Intra-Domain Potential User Preferences for Cross-Domain RecommendationabstractData sparsity poses a persistent challenge in Recommender Systems (RS), driving the emergence of Cross-Domain Recommendation (CDR) as a potential remedy. However, most existing CDR methods often struggle to circumvent the transfer of domain-specific information, which are perceived as noise in the target domain. Additionally, they primarily concentrate on inter-domain information transfer, disregarding the comprehensive exploration of data within intra-domains. To address these limitations, we propose SUCCDR (SeparatingUser features withCompound samples), a novel approach that tackles data sparsity by leveraging both cross-domain knowledge transfer and comprehensive intra-domain analysis. Specifically, to ensure the exclusion of noisy domain-specific features during the transfer process, user preferences are separated into domain-invariant and domain-specific features through three efficient constraints. Furthermore, the unobserved items are leveraged to generate compound samples that intelligently merge observed and unobserved potential user-item interaction, utilizing a simple yet efficient attention mechanism to enable a comprehensive and unbiased representation of user preferences. We evaluate the performance of SUCCDR on two real-world datasets, Douban and Amazon, and compare it with state-of-the-art single-domain and cross-domain recommendation methods. The experimental results demonstrate that SUCCDR outperforms existing approaches, highlighting its ability to effectively alleviate data sparsity problem. Jing Liu 0002, Lele Sun, Weizhi Nie, Yuting Su 0001, Yongdong Zhang 0001, Anan Liu |
IEEE Trans. Multim. | 1 |
| 2024 | Knowledge-Enhanced Causal Reinforcement Learning Model for Interactive RecommendationabstractOwing to its inherently dynamic nature and economical training cost, offline reinforcement learning (RL) is typically employed to implement an interactive recommender system (IRS). A crucial challenge in offline RL-based IRSs is the data sparsity issue, i.e., it is hard to mine user preferences well from the limited number of user-item interactions. In this article, we propose a knowledge-enhanced causal reinforcement learning model (KCRL) to mitigate data sparsity in IRSs. We make technical extensions to the offline RL framework in terms of the reward function and state representation. Specifically, we first propose a group preference-injected causal user model (GCUM) to learn user satisfaction (i.e., reward) estimation. We introduce beneficial group preference information, namely, the group effect, via causal inference to compensate for incomplete user interests extracted from sparse data. Then, we learn the RL recommendation policy with the reward given by the GCUM. We propose a knowledge-enhanced state encoder (KSE) to generate knowledge-enriched user state representations at each time step, which is assisted by a self-constructed user-item knowledge graph. Extensive experimental results on real-world datasets demonstrate that our model significantly outperforms the baselines. Weizhi Nie, Xin Wen 0017, Jing Liu 0002, Jiawei Chen 0007, Jiancan Wu, Guoqing Jin, Anan Liu |
IEEE Trans. Multim. | 3 |
| 2024 | CDCM: ChatGPT-Aided Diversity-Aware Causal Model for Interactive RecommendationabstractIn recent years, interactive recommender systems (IRSs) have attracted extensive interest. Existing IRSs are typically implemented with offline reinforcement learning (RL). They are devoted to improving recommendation accuracy by optimizing the extraction of users' inherent preferences. However, there hasn't been much attention on recommendation diversity, which could result in the monotony effect,i.e., categories of recommended items are consistently fixed and unchanging. In this paper, we center on category diversification in IRSs while largely preserving or even boosting recommendation accuracy. To this end, we propose a ChatGPT-aided diversity-aware causal model (CDCM) to enhance the offline RL framework with causal inference and ChatGPT. Specifically, we first propose a diversity-aware causal user model (DCUM) to estimate user satisfaction. This model disentangles the causal effect of users' inherent preferences and the monotony effect to obtain user satisfaction with both accuracy and diversity. Then, DCUM is used to assist the RL agent in recommendation policy learning. A ChatGPT-aided state encoder (CSE) is proposed to provide user state representation for each time step of policy learning. With the help of ChatGPT, CSE incorporates multi-category information in line with users' potential preferences to promote diverse and relevant category recommendations. Extensive experiment results on two real-world datasets validate the superiority of our CDCM regarding both accuracy and diversity. Xin Wen 0017, Weizhi Nie, Jing Liu 0002, Yuting Su 0001, Yongdong Zhang 0001, Anan Liu |
IEEE Trans. Multim. | 3 |
| 2024 | Privacy-preserving Multi-source Cross-domain Recommendation Based on Knowledge GraphabstractThe cross-domain recommender systems aim to alleviate the data sparsity problem in the target domain by transferring knowledge from the auxiliary domain. However, existing works ignore the fact that the data sparsity problem may also exist in the single auxiliary domain, and sharing user behavior data is restricted by the privacy policy. In addition, their cross-domain models lack interpretability. To address these concerns, we propose a novel multi-source cross-domain model based on knowledge graph. Specifically, to avoid the insufficiency of single auxiliary domain, we construct a knowledge graph comprehensively leveraging items from multiple auxiliary domains. To avoid the leakage of user privacy when user information is transferred to multiple domains, we construct graph for information transfer between items to effectively avoid the propagation of users’ private information between different domains. We implicitly integrate the user–item interaction by transferring the learned item embeddings. To improve the interpretability of cross-domain knowledge transfer, we propose a knowledge graph-based retrieval and fusion method to transfer knowledge derived from multiple auxiliary domains. An attention-based fusion network is designed to enhance the representation of the targeted user and items with the transferred item embedding. We perform extensive experiments on three real-world datasets, demonstrating that our model outperforms the states of the art. Jing Liu 0002, Litao Shang, Yuting Su 0001, Weizhi Nie, Xin Wen 0017, Anan Liu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Subjective and Objective Quality Assessment for in-the-Wild Computer Graphics ImagesabstractComputer graphics images (CGIs) are artificially generated by means of computer programs and are widely perceived under various scenarios, such as games, streaming media, etc. In practice, the quality of CGIs consistently suffers from poor rendering during production, inevitable compression artifacts during the transmission of multimedia applications, and low aesthetic quality resulting from poor composition and design. However, few works have been dedicated to dealing with the challenge of computer graphics image quality assessment (CGIQA). Most image quality assessment (IQA) metrics are developed for natural scene images (NSIs) and validated on databases consisting of NSIs with synthetic distortions, which are not suitable for in-the-wild CGIs. To bridge the gap between evaluating the quality of NSIs and CGIs, we construct a large-scale in-the-wild CGIQA database consisting of 6,000 CGIs (CGIQA-6k) and carry out the subjective experiment in a well-controlled laboratory environment to obtain the accurate perceptual ratings of the CGIs. Then, we propose an effective deep learning–based no-reference (NR) IQA model by utilizing both distortion and aesthetic quality representation. Experimental results show that the proposed method outperforms all other state-of-the-art NR IQA methods on the constructed CGIQA-6k database and other CGIQA-related databases. The database is released at https://github.com/zzc-1998/CGIQA6K . Wei Sun 0029, Yingjie Zhou 0003, Jun Jia, Jing Liu 0002, Xiongkuo Min, Guangtao Zhai |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2023 | Efficient Spatio-Temporal Video Grounding with Semantic-Guided Feature DecompositionabstractSpatio-temporal video grounding (STVG) aims to localize the spatio-temporal object tube in a video according to a given text query. Current approaches address the STVG task with end-to-end frameworks while suffering from heavy computational complexity and insufficient spatio-temporal interactions. To overcome these limitations, we propose a novel Semantic-Guided Feature Decomposition based Network (SGFDN). A semantic-guided mapping operation is proposed to decompose the 3D spatio-temporal feature into 2D motions and 1D object embedding without losing much object-related semantic information. Thus, the computational complexity in computationally expensive operations such as attention mechanisms can be effectively reduced by replacing the input spatio-temporal feature with the decomposed features. Furthermore, based on this decomposition strategy, a pyramid relevance filtering based attention is proposed to capture the cross-modal interactions at multiple spatio-temporal scales. In addition, a decomposition-based grounding head is proposed to locate the queried objects with less computational complexity. Extensive experiments on two widely-used STVG datasets (VidSTG and HC-STVG) demonstrate that our method enjoys state-of-the-art performance as well as less computational complexity. The code has been available at https://github.com/TJUMMG/SGFDN. Weikang Wang 0002, Jing Liu 0002, Yuting Su 0001, Weizhi Nie |
ACM Multimedia | 2 |
| 2023 | Edge-aware object pixel-level representation tracking
Peiguang Jing, Zijian Huang 0010, Jing Liu 0002, Jiexiao Yu |
J. Vis. Commun. Image Represent. | 3 |
| 2023 | Multi-loop graph convolutional network for multimodal conversational emotion recognition
Minjie Ren, Xiangdong Huang 0002, Wenhui Li 0001, Jing Liu 0002 |
J. Vis. Commun. Image Represent. | 4 |
| 2023 | SMPC: boosting social media popularity prediction with caption
Anan Liu, Ning Xu 0003, Jing Liu 0002, Yuting Su 0001, Shenyuan Zhang, Yejun Tang, Junbo Guo, Guoqing Jin, Xuanya Li |
Multim. Syst. | 4 |
| 2023 | A comprehensive survey on deep-learning-based visual captioning
Bowen Xin, Ning Xu 0003, Yingchen Zhai, Zimu Lu, Jing Liu 0002, Weizhi Nie, Xuanya Li, Anan Liu |
Multim. Syst. | 6 |
| 2023 | Group attention retention network for co-salient object detection
Jing Liu 0002, Weikang Wang 0002, Jiexiao Yu |
Mach. Vis. Appl. | 1 |
| 2023 | MALN: Multimodal Adversarial Learning Network for Conversational Emotion RecognitionabstractMultimodal emotion recognition in conversations (ERC) aims to identify the emotional state of constituent utterances expressed by multiple speakers in dialogue from multimodal data. Existing multimodal ERC approaches focus on modeling the global context of the dialogue and neglect to mine the characteristic information from the corresponding utterances expressed by the same speaker. Additionally, information from different modalities exhibits commonality and diversity for emotional expression. The commonality and diversity of multimodal information are compensated for each other but not effectively exploited in previous multimodal ERC works. To tackle these issues, we propose a novel Multimodal Adversarial Learning Network (MALN). MALN first mines the speaker’s characteristics from context sequences and then incorporate them with the unimodal features. Afterward, we design a novel adversarial module AMDM to exploit both commonality and diversity from the unimodal features. Finally, AMDM fuses different modalities to generate refined utterance representations for emotion classification. Extensive experiments are conducted on two public multimodal ERC datasets, IEMOCAP and MELD. Through the experiments, MALN shows its superiority over the state-of-the-art methods. Minjie Ren, Xiangdong Huang 0002, Jing Liu 0002, Ming Liu 0004, Xuanya Li, Anan Liu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Trajectory Guided Robust Visual Object Tracking With Selective RemedyabstractSiamese trackers have received a lot of attentions due to their promising performance and real-time high speed. However, the robustness is generally limited especially in challenging conditions, such as occlusion. Motivated by that tracking failure does not always occur and the target movement tends to follow certain patterns, in the paper, we propose a generic, fast and flexible approach to improve the robustness of Siamese trackers with two light-load novel modules: Trajectory Guidance Module (TGM) and Selective Refinement Module (SRM). Specifically, TGM encourages to pay a soft attention on possible target location based on short-term historical trajectory. SRM selectively remedies the tracking results at the risk of failure with little impact on the speed. The proposed algorithm can be easily establish upon state-of-the-art Siamese trackers and obtains better performance on seven benchmarks with high real-time tracking speed. The code is available athttps://github.com/TJUMMG/TGSR. Han Wang 0034, Jing Liu 0002, Yuting Su 0001, Xiaokang Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | MRFT: Multiscale Recurrent Fusion Transformer Based Prior Knowledge for Bit-Depth EnhancementabstractBit-depth enhancement (BDE) plays an important role in providing high bit-depth data support for high-dynamic range (HDR) display. Although convolutional neural network (CNN) based BDE methods have achieved top performance, multiscale feature extraction and fusion still suffer from some inherent architectural flaws. Moreover, the training-data-scarce scene has not been effectively explored. To this end, this paper proposes an innovative multiscale recurrent fusion transformer (MRFT) framework, which contains three key components, i.e. multiscale transformer feature encoder, recurrent feature fusion module, and prior knowledge injection. Specifically, the multiscale transformer feature encoder consists of a prior-injected context encoder (PICE) and a multiscale local feature encoder (MLFE). PICE leverages the vanilla self-attention mechanism to extract the global context correlating spatially-distant contents for distinguishing long-distance false contours. MLFE exploits the local self-attention mechanism with varied window sizes to capture different-scale detail features. Then, a hierarchical recurrent decoder (HRD) is proposed as the recurrent feature fusion module to fuse multiscale visual information with global guidance. Via the circular query-key mechanism, global-to-local information is progressively fused. Furthermore, we propose a two-stage alternating optimization strategy for prior knowledge injection. By pre-parameterizing the global auxiliary priors, the training dilemma on the data-scarce domain is significantly alleviated. Extensive analyses on multiple benchmark datasets demonstrate the superiority of our MRFT in terms of quantitative measures and aesthetic effects. Xin Wen 0017, Weizhi Nie, Jing Liu 0002, Yuting Su 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Sequence as a Whole: A Unified Framework for Video Action Localization With Long-Range Text QueryabstractComprehensive understanding of video content requires both spatial and temporal localization. However, there lacks a unified video action localization framework, which hinders the coordinated development of this field. Existing 3D CNN methods take fixed and limited input length at the cost of ignoring temporally long-range cross-modal interaction. On the other hand, despite having large temporal context, existing sequential methods often avoid dense cross-modal interactions for complexity reasons. To address this issue, in this paper, we propose a unified framework which handles the whole video in sequential manner with long-range and dense visual-linguistic interaction in an end-to-end manner. Specifically, a lightweight relevance filtering based transformer (Ref-Transformer) is designed, which is composed of relevance filtering based attention and temporally expanded MLP. The text-relevant spatial regions and temporal clips in video can be efficiently highlighted through the relevance filtering and then propagated among the whole video sequence with the temporally expanded MLP. Extensive experiments on three sub-tasks of referring video action localization, i.e., referring video segmentation, temporal sentence grounding, and spatiotemporal video grounding, show that the proposed framework achieves the state-of-the-art performance in all referring video action localization tasks. The code has been available at https://github.com/TJUMMG/SAW. Yuting Su 0001, Weikang Wang 0002, Jing Liu 0002, Xiaokang Yang 0001 |
IEEE Trans. Image Process. | 3 |
| 2022 | Tracking by dynamic template: Dual update mechanism
Jing Liu 0002, Xiangdong Huang 0002, Yuting Su 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2022 | 3DFP-FCGAN: Face completion generative adversarial network with 3D facial prior
Jing Liu 0002, Weikang Wang 0002, Jiexiao Yu, Chunping Zhang, Yuting Su 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2022 | Iterative Residual Feature Refinement Network for Bit-Depth EnhancementabstractBit-depth enhancement (BDE) restores high bit-depth (HBD) images from low bit-depth (LBD) ones, which has important applications. Recently, residual-optimized BDE algorithms based on convolutional neural networks (CNNs) have achieved top performance. However, they fail to use a single model to accurately recover all frequency information encoded by missing significant bits at one time on challenging large bit-depth recovery tasks. In this paper, we redefine BDE residual recovery from the perspective of image frequency characteristics. On this basis, we propose an iterative residual feature optimization strategy, which provides an implicit error correction mechanism and improves training and inference efficiency. Furthermore, we design a simple but effective iterative residual feature refinement network (IRFRN). By linking model complexity with the recovery of different frequency information, IRFRN enables a single model to simultaneously recover the missing low and high frequency information. Extensive experiments indicate that our method achieves the state-of-the-art quantitative and qualitative performance on large bit-depth recovery tasks. Weizhi Nie, Xin Wen 0017, Jing Liu 0002, Yuting Su 0001 |
IEEE Signal Process. Lett. | 3 |
| 2022 | Residual-Guided Multiscale Fusion Network for Bit-Depth EnhancementabstractBit-depth enhancement (BDE) is a challenging task due to stubborn false contour artifacts and disappeared detailed information. Given the mixture of structural distortions and real edges in low bit-depth (LBD) images, both large and small receptive fields (RFs) are critical for BDE tasks. However, even powerful state-of-the-art CNN-based methods can hardly capture sufficient LBD features under multiple RFs. This paper proposes a residual-guided multiscale fusion network (RMFNet) to explore multiscale features in a residual manner. We find that the shuffling operation provides desired multiscale inputs for effectively distinguishing false contours from real edges without any loss of information. Therefore, we shuffle LBD images to multiple scales and then fully extract residual features under different RFs with corresponding subnets. To facilitate interscale guidance from the global context to the local context, we progressively transfer the encoded residual features between adjacent subnets from top to bottom. We further propose a dual-branch depthwise group fusion (DDGF) module to fully capture inter- and inner correlations of multiscale features with fewer parameters. Finally, extensive experiments show that our algorithm achieves excellent performance improvement both quantitatively and qualitatively, verifying its effectiveness. Jing Liu 0002, Xin Wen 0017, Weizhi Nie, Yuting Su 0001, Peiguang Jing, Xiaokang Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Tripartite Graph Regularized Latent Low-Rank Representation for Fashion Compatibility PredictionabstractIn recent years, an increasing online shopping demand has greatly promoted the innovation and development of the fashion industry. Visual fashion analysis has become a prospective research topic in computer vision and multimedia fields. Among these studies, fashion compatibility analysis is required in many real applications, such as fashion recommendation, matching, and retrieval. However, learning fashion compatibility is nontrivial, not only due to the uncertain and sparse dependencies among fashion items but also the latent and mutual associations among multiple factors such as color, texture, style, and functionality. To better predict fashion compatibility, in this paper, we proposed a tripartite graph regularized latent low-rank representation method, named TGRLLR, for fashion compatibility prediction. In TGRLLR, to learn more low-dimensional and effective representations, we considered the latent low-rank representation by decomposing the original feature matrix in both the column and row directions to tackle the problem of insufficient observations. On this basis, we simultaneously exploited different regularization strategies to encode the structured correlations among features, the high-order relationships among items, and the geometrical structures of outfits for more informative representations. Extensive experiments conducted on a real-world dataset demonstrate the effectiveness of our proposed method compared with state-of-the-art methods. Peiguang Jing, Jing Zhang 0038, Liqiang Nie, Jing Liu 0002, Yuting Su 0001 |
IEEE Trans. Multim. | 5 |
| 2022 | TANet: Target Attention Network for Video Bit-Depth EnhancementabstractVideo bit-depth enhancement (VBDE) reconstructs high-bit-depth (HBD) frames from a low-bit-depth (LBD) video sequence. As neighboring frames contain a considerable amount of complementary information related to the center frame, it is vital for VBDE to exploit neighboring frames as much as possible. Conventional VBDE algorithms with explicit alignment across frames attempt to warp each neighboring frame to the center frame with estimated optical flow, taking into account only pairwise correlation. Most spatiotemporal fusion approaches involve direct concatenation or 3D convolution and treat all features equally, failing to focus on information related to the center frame. Therefore, in this paper, we introduce an improved nonlocal block as a global attentive alignment (GAA) module, which takes the whole input video sequence into consideration to capture features that are globally correlated, to perform implicit alignment. Furthermore, given the bulk of features extracted from the center and neighboring frames, we propose target-guided attention (TGA). TGA can exploit more center-frame-related details and facilitate feature fusion. The proposed network (dubbed TANet) is capable of effectively eliminating false contours and recovering the center frame in high quality, as demonstrated by the experimental results. TANet outperforms state-of-the-art models in terms of both PSNR and SSIM with low time consumption. Jing Liu 0002, Ziwen Yang, Yuting Su 0001, Xiaokang Yang 0001 |
IEEE Trans. Multim. | 1 |
| 2021 | Key Facial Components Guided Micro-Expression Recognition Based on First & Second-Order MotionabstractAlthough there have been many successful attempts in the field of micro-expression recognition, plenty of challenges remain due to the subtle spatio-temporal changes and high locality of micro-expressions. In this paper, to tackle such issues, we propose a novel key facial components guided micro-expression recognition approach (KFC-MER). Face semantic segmentation probability maps involving several key components provide a guidance for feature learning. With the Components-Aware Attention (CAA) module, expression-related areas are highlighted and the relationship between components will also be learned. To cope with the limited size of micro-expression datasets, we design a parallel shallow residual network as the MER network. Both the first- and second-order motion are exploited as the input data, for capturing motive information and non-rigid deformation, respectively. Extensive experiments on three benchmark datasets demonstrate that our method outperforms previous works and achieves state-of-the-art performance. The code is publicly available on GitHub: https://github.com/TJUMMG/KFC-MER. Yuting Su 0001, Jing Liu 0002, Guangtao Zhai |
ICME | 3 |
| 2021 | Spatio-Temporal Mitosis Detection in Time-Lapse Phase-Contrast Microscopy Image Sequences: A BenchmarkabstractIn this paper, we report the results of the first international contest on mitosis detection in phase-contrast microscopy image sequences (https://www.iti-tju.org/mitosisdetection), which was held at the workshop of computer vision for microscopy image analysis (CVMI) in CVPR 2019. This contest aims to promote research on spatiotemporal mitosis detection under microscopy images. In this contest, we released a large-scale time-lapse phase-contrast microscopy image dataset (C2C12-16) for the mitosis detection task. Compared with the previous popular datasets (e.g., C2C12, C3H10), C2C12-16 contains more annotated mitotic events and more diverse cell culture environments. A total of ten different mitosis detection methods were submitted in the contest and evaluated on the test sets of four different cell culture environments in C2C12-16. In this benchmark, we describe all methods and conduct a thorough analysis based on their performances and discuss a feasible direction for mitosis detection. To the best of our knowledge, this is the first benchmark for the mitosis detection problem using a time-lapse phase-contrast microscopy spatiotemporal image sequence model. Yuting Su 0001, Yao Lu 0005, Jing Liu 0002, Anan Liu |
IEEE Trans. Medical Imaging | 3 |
| 2021 | Learning Low-Rank Sparse Representations With Robust Relationship Inference for Image Memorability PredictionabstractImage memorability prediction aims to estimate the degree to which an image will be remembered by observers. Generally, the core problem in image memorability prediction is how to obtain effective representations to characterize the visual content of an image. In contrast to existing methods, which focus more on exploring the factors that make images memorable, in this paper, we first propose a general framework for learning joint low-rank and sparse principal feature representations, called the LSPFR framework, to obtain the lowest-rank intrinsic representation for image memorability prediction. By considering the joint optimization of the nuclear and$\ell _{1}$-norms, the global low-rank structure and the local patterns embedded in data can be exploited to make the learned features more robust and informative. To improve our framework based on the exploitation of sample relationship structure information, we present an extended version of LSPFR, named E-LSPFR, in which the underlying relationship structure matrix is inferred through a negative log-likelihood term with a sparsity constraint. The results of experiments conducted on four publicly available datasets confirm the superior performance of our proposed approaches. Peiguang Jing, Yuechen Shang, Liqiang Nie, Yuting Su 0001, Jing Liu 0002, Meng Wang 0001 |
IEEE Trans. Multim. | 5 |
| 2020 | An End-to-End Perceptual Quality Assessment Method via Score Distribution Prediction
Jing Liu 0002, Jingting Wang, Weizhi Nie, Yuting Su 0001, Anan Liu |
Neural Process. Lett. | 1 |
| 2020 | ABSNet: Aesthetics-Based Saliency Network Using Multi-Task Convolutional NetworkabstractAs a smart visual attention mechanism to analyze visual scenes, visual saliency has been shown to closely correlate with semantic information such as faces. Although many semantic-information-guided saliency models have been proposed, to the best of our knowledge, no semantic information in affective domain has been employed for saliency detection. Aesthetic, the affective perceptual quality that integrates factors like scene composition and contrast, can certainly benefit visual attention that highly depends on these visual factors. In this letter, we propose an end-to-end multi-task framework called aesthetics-based saliency network (ABSNet). We use three commonly-used shared backbones and design two distinct branches for each task. Mean square error (MSE) loss and Earth Mover's Distance (EMD) loss are jointly adopted to alternately train the shared network and individual branch for different tasks, facilitating the proposed model to extract more effective features for visual perception. Moreover, our model is resolution-friendly to predict saliency for images of arbitrary size. It has been shown that the proposed multi-task method is superior over single-task version and outperforms state-of-the-art saliency methods. Jing Liu 0002, Jincheng Lv, Jing Zhang 0038, Yuting Su 0001 |
IEEE Signal Process. Lett. | 1 |
| 2020 | Low-Rank Regularized Multi-Representation Learning for Fashion Compatibility PredictionabstractThe currently flourishing fashion-oriented community websites and the continuous pursuit of fashion have attracted the increased research interest of the fashion analysis community. Many studies show that predicting the compatibility of fashion outfits is a nontrivial task due to the difficulty in capturing the implicit patterns affecting fashion compatibility prediction and the complex relationships presented by raw data. To address these problems, in this paper, we propose a transductive low-rank hypergraph regularizer multiple-representation learning framework (LHMRL), whereby we formulate the processes of feature representation and fashion compatibility prediction in a joint framework. Specifically, we first introduce a low-rank regularized multiple-representation learning framework, in which the lowest-rank multiple representations of samples can be learned to characterize samples from different perspectives. In this framework, we maximize the total difference among multiple representations based on Grassmann manifold theory and incorporate a common hypergraph regularizer to naturally encode the complex relationships between fashion items and an outfit. To enhance the representation ability of our model, we then develop a supervised learning term by exploiting two types of supervision information from labeled data. Experiments on a publicly available large-scale dataset demonstrate the effectiveness of our proposed model over the state-of-the-art methods. Peiguang Jing, Liqiang Nie, Jing Liu 0002, Yuting Su 0001 |
IEEE Trans. Multim. | 4 |
| 2019 | Photo-realistic image bit-depth enhancement via residual transposed convolutional neural network
Yuting Su 0001, Wanning Sun, Jing Liu 0002, Guangtao Zhai, Peiguang Jing |
Neurocomputing | 3 |
| 2019 | Scene graph captioner: Image captioning based on structural visual representation
Ning Xu 0003, Anan Liu, Jing Liu 0002, Weizhi Nie, Yuting Su 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2019 | Comprehensive image quality assessment via predicting the distribution of opinion score
Anan Liu, Jingting Wang, Jing Liu 0002, Yuting Su 0001 |
Multim. Tools Appl. | 3 |
| 2019 | A structure-transfer-driven temporal subspace clustering for video summarization
Jing Zhang 0038, Peiguang Jing, Jing Liu 0002, Yuting Su 0001 |
Multim. Tools Appl. | 4 |
| 2019 | Low-rank regularized tensor discriminant representation for image set classification
Peiguang Jing, Yuting Su 0001, Zhengnan Li, Jing Liu 0002, Liqiang Nie |
Signal Process. | 4 |
| 2019 | A Framework of Joint Low-Rank and Sparse Regression for Image Memorability PredictionabstractImage memorability is to measure the degree to which an image is remembered. Generally image memorability prediction involves two steps: feature representation and prediction. Most previous work just focused on addressing the first step by investigating the factors of making an image memorable. They not only lack the use of a learning mechanism in feature representation, but also often neglect the second step. In this paper, we first propose a joint low-rank and sparse regression (JLRSR) framework to address this problem. JLRSR aims to jointly learn: 1) a low-rank projection matrix that enables us to decompose the original data into a component part and an error part and 2) a sparse regression coefficient vector for image memorability prediction. The projection matrix and the regression coefficients are bound by a sparse constraint to make our approach invariant to training samples. Moreover, a graph regularizer is constructed to improve the generalization performance and prevent overfitting. We then extend JLRSR to a multi-view version called Mv-JLRSR by imposing the block-wise constraint to ensure the group effect and the view correlation constraint to eliminate the heterogeneity among views. Experiment results validate the effectiveness of our proposed approaches. Peiguang Jing, Yuting Su 0001, Liqiang Nie, Huimin Gu, Jing Liu 0002, Meng Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2019 | BE-CALF: Bit-Depth Enhancement by Concatenating All Level Features of DNNabstractThere is a growing demand for monitors to provide high-quality visualization with more bits representing each rendered pixel. However, since most existing images and videos are of low bit-depth (LBD), transforming LBD images to visually pleasant high bit-depth (HBD) versions is of significant value. Most existing bit-depth enhancement methods generate unsatisfactory HBD images with annoying false contour artifacts or blurry details, and some algorithms are also time-consuming. To overcome these drawbacks, we propose a bit-depth enhancement framework via concatenating all level features of deep neural networks (DNNs). A novel deep learning network is proposed based on the deep convolutional variational auto-encoders (VAEs), and skip connections that concatenate every two layers are applied to pass low-level and high-level features to consequent layers, easing the gradient vanishing problem. Meanwhile, the proposed network is optimized to generate the residual between original images and its quantized ones, which performs better than recovering HBD images directly. The experimental results show that the proposed algorithm can eliminate false contour artifacts of the recovered HBD images with low time consumption, and can achieve dramatic restoration performance gains compared with state-of-the-art methods both subjectively and objectively. Jing Liu 0002, Wanning Sun, Yuting Su 0001, Peiguang Jing, Xiaokang Yang 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Spatiotemporal Symmetric Convolutional Neural Network for Video Bit-Depth EnhancementabstractIn contrast to the high sensitivity of human eyes and rapid development of modern display devices in terms of dynamic range, mainstream multimedia sources are generally at relatively lower bit depths (BDs). Therefore, BD enhancement (BDE), which attempts to transform low-BD multimedia sources into high-BD sources, is considered of significant research value. Current BDE algorithms are based on images rather than videos. However, for massive numbers of videos, temporal continuity among frames should be considered. Thus, in this paper, we propose a spatiotemporal symmetric BDE network for videos based on an encoder-decoder network. Consecutive frames are input into five subnets in the encoder, where the convolutional filters in the temporal symmetric subnets share the same weights to achieve lower model complexity. In addition, symmetric skip connections are introduced between the symmetric convolutional/deconvolutional layers of the encoder/decoder to pass features and alleviate the gradient diffusion problem. The experimental results show that our model can efficiently eliminate false contours and chroma distortions. The model significantly outperforms state-of-the-art image BDE algorithms and single-frame baseline models in terms of PSNR and SSIM. Jing Liu 0002, Pingping Liu, Yuting Su 0001, Peiguang Jing, Xiaokang Yang 0001 |
IEEE Trans. Multim. | 1 |
| 2019 | Multi-Channel Decomposition in Tandem With Free-Energy Principle for Reduced-Reference Image Quality AssessmentabstractThe visual quality of perceptions is highly correlated with the mechanisms of the human brain and visual system. Recently, the free-energy principle, which has been widely researched in brain theory and neuroscience, is introduced to quantize the perception, action, and learning in human brain. In the field of image quality assessment (IQA), on one hand, the free-energy principle can resort to the internal generative model to simulate the visual stimulus of the human beings. On the other hand, abundant psychological and neurobiological studies reveal that different frequency and orientation components of one visual stimulus arouse different neurons in the striate cortex, and the striate cortex processes visual information in the cerebral cortex. Motivated by these two aspects, a novel reduce-reference IQA metric called the multi-channel free-energy based reduced-reference quality metric is proposed in this paper. First, a two-level discrete Haar wavelet transform is used to decompose the input reference and distorted images. Next, to simulate the generative model in the human brain, the sparse representation is leveraged to extract the free-energy-based features in subband images. Finally, the overall quality metric is obtained through the support vector regressor. Extensive experimental comparisons on four benchmark image quality databases (LIVE, CSIQ, TID2008, and TID2013) demonstrate that the proposed method is highly competitive with the representative reduced-reference and classical full-reference models. Wenhan Zhu, Guangtao Zhai, Xiongkuo Min, Menghan Hu, Jing Liu 0002, Guodong Guo, Xiaokang Yang 0001 |
IEEE Trans. Multim. | 5 |
| 2018 | A Blind Quality Measure for Industrial 2D Matrix Symbols Using Shallow Convolutional Neural NetworkabstractIndustrial two-dimensional (2D) matrix symbols are ubiquitous throughout the automatic assembly lines. Most industrial 2D symbols are corrupted by various inevitable artifacts. State-of-the-art decoding algorithms are not able to directly handle low-quality symbols irrespective of problematic artifacts. Degraded symbols require appropriate preprocessing methods, such as morphology filtering, median filtering, or sharpening filtering, according to specific distortion type. In this paper, we first establish a database including 3000 industrial 2D symbols which are degraded by 6 types of distortions. Second, we utilize a shallow convolutional neural network (CNN) to identify the distortion type and estimate the quality grade for 2D symbols. Finally, we recommend an appropriate preprocessing method for low-quality symbol according to its distortion type and quality grade. Experimental results indicate that the proposed method outperforms state-of-the-art methods in terms of PLCC, SRCC and RMSE. It also promotes decoding efficiency at the cost of low extra time spent. Zhaohui Che, Guangtao Zhai, Jing Liu 0002, Ke Gu 0001, Patrick Le Callet, Jiantao Zhou 0001, Xianming Liu 0005 |
ICIP | 3 |
| 2018 | Low-rank regularized multi-view inverse-covariance estimation for visual sentiment distribution prediction
Anan Liu, Yingdi Shi, Peiguang Jing, Jing Liu 0002, Yuting Su 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2018 | Graph regularized low-rank tensor representation for feature selection
Yuting Su 0001, Peiguang Jing, Jing Zhang 0038, Jing Liu 0002 |
J. Vis. Commun. Image Represent. | 6 |
| 2018 | Structured low-rank inverse-covariance estimation for visual sentiment distribution prediction
Anan Liu, Yingdi Shi, Peiguang Jing, Jing Liu 0002, Yuting Su 0001 |
Signal Process. | 4 |
| 2018 | Arrow's Impossibility Theorem inspired subjective image quality assessment approach
Wenhan Zhu, Guangtao Zhai, Menghan Hu, Jing Liu 0002, Xiaokang Yang 0001 |
Signal Process. | 4 |
| 2018 | Low-Rank Regularized Heterogeneous Tensor Decomposition for Subspace ClusteringabstractThis letter proposes a low-rank regularized heterogeneous tensor decomposition (LRRHTD) algorithm for subspace clustering, in which various constrains in different modes are incorporated to enhance the robustness of the proposed model. Specifically, due to the presence of noise and redundancy in the original tensor, LRRHTD seeks a set of orthogonal factor matrices for all but the last mode to map the high-dimensional tensor into a low-dimensional latent subspace. Furthermore, by imposing a low-rank constraint on the last mode, which is relaxed by using a nuclear norm, the lowest rank representation that reveals the global structure of samples is obtained for the purpose of clustering. We develop an effective algorithm based on the augmented Lagrange multiplier to optimize our model. Experiments on two public datasets demonstrate that our method reaches convergence within a small number of iterations and achieves promising results in comparison with the state of the arts. Jing Zhang 0038, Peiguang Jing, Jing Liu 0002, Yuting Su 0001 |
IEEE Signal Process. Lett. | 4 |
| 2018 | IPAD: Intensity Potential for Adaptive De-QuantizationabstractDisplay devices at bit depth of 10 or higher have been mature but the mainstream media source is still at bit depth of eight. To accommodate the gap, the most economic solution is to render source at low bit depth for high bit-depth display, which is essentially the procedure of de-quantization. Traditional methods, such as zero-padding or bit replication, introduce annoying false contour artifacts. To better estimate the least-significant bits, later works use filtering or interpolation approaches, which exploit only limited neighbor information, cannot thoroughly remove the false contours. In this paper, we propose a novel intensity potential (IP) field to model the complicated relationships among pixels. The potential value decreases as the spatial distance to the field source increases and the potentials from different field sources are additive. Based on the proposed IP field, an adaptive de-quantization procedure is then proposed to convert low-bit-depth images to high-bit-depth ones. To the best of our knowledge, this is the first attempt to apply potential field for natural images. The proposed potential field preserves local consistency and models the complicated contexts well. Extensive experiments on natural, synthetic, and high-dynamic range image data sets validate the efficiency of the proposed IP field. Significant improvements have been achieved over the state-of-the-art methods on both the peak signal-to-noise ratio and the structural similarity. Jing Liu 0002, Guangtao Zhai, Anan Liu, Xiaokang Yang 0001, Xibin Zhao, Chang Wen Chen |
IEEE Trans. Image Process. | 1 |
| 2018 | Low-Rank Multi-View Embedding Learning for Micro-Video Popularity PredictionabstractRecently, a prevailing trend of user generated content (UGC) on social media sites is the emerging micro-videos. Microvideos afford many potential opportunities ranging from network content caching to online advertising, yet there are still little efforts dedicated to research on micro-video understanding. In this paper, we focus on popularity prediction of micro-videos by presenting a novel low-rank multi-view embedding learning framework. We name it as transductive low-rank multi-view regression (TLRMVR), and it is capable of boosting the performance of micro-video popularity prediction by jointly considering the intrinsic representations of the source and target samples. In particular, TLRMVR integrates low-rank multi-view embedding and regression analysis into a unified framework such that the lowest-rank representation shared by all views not only captures the global structure of all views, but also indicates the regression requirements. The framework is formulated as a regression model and it seeks a set of view-specific projection matrices with low-rank constraints to map multi-view features into a common subspace. In addition, a multi-graph regularization term is constructed to improve the generalization capability and further prevents the overfitting problem. Extensive experiments conducted on a publicly available dataset demonstrate that our proposed method achieve promising results as compared with state-of-the-art baselines. Peiguang Jing, Yuting Su 0001, Liqiang Nie, Jing Liu 0002, Meng Wang 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2018 | Blind Quality Assessment Based on Pseudo-Reference ImageabstractTraditional full-reference image quality assessment (IQA) metrics generally predict the quality of the distorted image by measuring its deviation from a perfect quality image called reference image. When the reference image is not fully available, the reduced-reference and no-reference IQA metrics may still be able to derive some characteristics of the perfect quality images, and then measure the distorted image's deviation from these characteristics. In this paper, contrary to the conventional IQA metrics, we utilize a new “reference” called pseudo-reference image (PRI) and a PRI-based blind IQA (BIQA) framework. Different from a traditional reference image, which is assumed to have a perfect quality, PRI is generated from the distorted image and is assumed to suffer from the severest distortion for a given application. Based on the PRI-based BIQA framework, we develop distortion-specific metrics to estimate blockiness, sharpness, and noisiness. The PRI-based metrics calculate the similarity between the distorted image's and the PRI's structures. An image suffering from severer distortion has a higher degree of similarity with the corresponding PRI. Through a two-stage quality regression after a distortion identification framework, we then integrate the PRI-based distortion-specific metrics into a general-purpose BIQA method named blind PRI-based (BPRI) metric. The BPRI metric is opinion-unaware (OU) and almost training-free except for the distortion identification process. Comparative studies on five large IQA databases show that the proposed BPRI model is comparable to the state-of-the-art opinion-aware- and OU-BIQA models. Furthermore, BPRI not only performs well on natural scene images, but also is applicable to screen content images. The MATLAB source code of BPRI and other PRI-based distortion-specific metrics will be publicly available. Xiongkuo Min, Ke Gu 0001, Guangtao Zhai, Jing Liu 0002, Xiaokang Yang 0001, Chang Wen Chen |
IEEE Trans. Multim. | 4 |
| 2017 | A-Lamp: Adaptive Layout-Aware Multi-patch Deep Convolutional Neural Network for Photo Aesthetic AssessmentabstractDeep convolutional neural networks (CNN) have recently been shown to generate promising results for aesthetics assessment. However, the performance of these deep CNN methods is often compromised by the constraint that the neural network only takes the fixed-size input. To accommodate this requirement, input images need to be transformed via cropping, warping, or padding, which often alter image composition, reduce image resolution, or cause image distortion. Thus the aesthetics of the original images is impaired because of potential loss of fine grained details and holistic image layout. However, such fine grained details and holistic image layout is critical for evaluating an images aesthetics. In this paper, we present an Adaptive Layout-Aware Multi-Patch Convolutional Neural Network (A-Lamp CNN) architecture for photo aesthetic assessment. This novel scheme is able to accept arbitrary sized images, and learn from both fined grained details and holistic image layout simultaneously. To enable training on these hybrid inputs, we extend the method by developing a dedicated double-subnet neural network structure, i.e. a Multi-Patch subnet and a Layout-Aware subnet. We further construct an aggregation layer to effectively combine the hybrid features from these two subnets. Extensive experiments on the large-scale aesthetics assessment benchmark (AVA) demonstrate significant performance improvement over the state-of-the-art in photo aesthetic assessment. Jing Liu 0002, Chang Wen Chen |
CVPR | 2 |
| 2017 | IPAD: Intensity potential for adaptive de-quantizationabstractDisplay devices at bit-depth of 10 or higher have been mature but the mainstream media source is still at bit-depth as low as 8. To accommodate the gap, the most economic solution is to render source at low bit-depth for high bit-depth display, which is essentially the procedure of de-quantization. Traditional methods, like zero-padding or bit replication, introduce annoying false contour artifacts. To better estimate the least-significant bits, later works use filtering or interpolation approaches, which exploit only limited neighbor information, can not thoroughly remove the false contours. In this paper, we propose a novel intensity potential field to model the complicated relationships among pixels. Then, an adaptive de-quantization algorithm is proposed to convert low bit-depth images to high bit-depth ones. To the best of our knowledge, this is the first attempt to apply potential field for natural images. The proposed potential field preserves local consistency and models the complicated contexts very well. Extensive experiments on natural image datasets validate the efficiency of the proposed intensity potential field. Significant improvements have been achieved over the state-of-the-art methods on both PSNR and SSIM. Jing Liu 0002, Guangtao Zhai, Xiaokang Yang 0001, Menghan Hu, Chang Wen Chen |
ICME | 1 |
| 2017 | Dynamic backlight scaling considering ambient luminance for mobile energy savingabstractThe mobile video playback involves many subsystems of the devices such as computing, rendering and displaying subsystems. Among all subsystems, the displaying subsystem accounts for at least 38% of all consumed power, and it can be up to 68% with the maximum backlight brightness. What is more, lots of people watch videos via mobile devices in various situations, where the ambient luminance condition is different. Therefore, how to save mobile energy and improve the Quality of Experience (QoE) in different situations become significant problems. In this paper, we try to maximally enhance the battery power performance under various ambient luminance conditions through backlight magnitude adjusting, while without negatively impacting users' QoE. In particular, we conduct a series of subject quality assessment experiments to uncover the quantitative relationship among QoE, ambient luminance, video content luminance and backlight level. We first study whether the continuous playback of backlight-scaled shots using the proposed scaling magnitude would cause flicker effect or not. Then motivated by the findings of these subject studies, we implement a Dynamic Backlight Scaling (DBS) strategy. The experiment results demonstrate that the DBS strategy can save more than 40% power at most and can also save 10% power even at a very high ambient luminance. Wei Sun 0029, Guangtao Zhai, Xiongkuo Min, Yutao Liu 0002, Siwei Ma 0001, Jing Liu 0002, Jiantao Zhou 0001, Xianming Liu 0005 |
ICME | 6 |
| 2017 | Perceptual information hiding based on multi-channel visual masking
Guangtao Zhai, Xiaokang Yang 0001, Menghan Hu, Jing Liu 0002 |
Neurocomputing | 5 |
| 2017 | Visual attention analysis and prediction on human faces
Xiongkuo Min, Guangtao Zhai, Ke Gu 0001, Jing Liu 0002, Shiqi Wang 0001, Xinfeng Zhang 0001, Xiaokang Yang 0001 |
Inf. Sci. | 4 |
| 2016 | Principal components analysis-based visual saliency detectionabstractIn this paper, a novel patch-wise saliency detection algorithm is proposed based on Principal Component Analysis (PCA). As a powerful statistical procedure in data analysis, PCA are fully exploited to convert color space and produce compact patch representation. Specifically, images are first converted to linearly uncorrelated channels and divided into non-overlapped patches. Then the patches are represented by the coefficients of principal components using PCA analysis. Based on the compact representation of patches, two types of distinctiveness are introduced: center-surround contrast and global rarity. Experimental results demonstrate that the PCA-based color space conversion and patch representation can improve the accuracy of human fixations prediction, and the proposed algorithm outperforms the mainstream algorithms on predicting human fixations. Bing Yang 0003, Xiaoyun Zhang 0001, Jing Liu 0002, Li Chen 0021 |
ICASSP | 3 |
| 2016 | Visual saliency model based on minimum description lengthabstractIn this paper, a novel patch-wise visual saliency model based on Minimum Length Description (MDL) principle is presented. Visual saliency is measured as the unpredicted information of image patch through an order-adaptive predictor under MDL principle. Specifically, each image patch is estimated with a linear combination of several neighboring patches. The number and location of candidate patches are automatically tuned to local contexts based on MDL. Then the entropy of prediction residuals of center patch, which represents the surprise to the visual system, is used to measure the saliency. Furthermore, a structural redundancy operator is also involved to improve the saliency detection performance. Experimental results demonstrate that the predictor under MDL principle along with the structural redundancy operator can improve the accuracy of human fixations prediction. We show that the proposed model outperforms the mainstream algorithms in predicting human fixations. Jing Liu 0002, Xiaokang Yang 0001, Guangtao Zhai, Chang Wen Chen |
ISCAS | 1 |
| 2015 | Image inpainting with adaptive linear predictorabstractIn this paper, a novel examplar-based inpainting algorithm with adaptive linear predictor is proposed. The patches in the damaged region are sequentially estimated with a linear combination of several nearest neighboring patches. The number of candidate patch is automatically tuned to local contexts based on Bayesian Information Criterion (BIC). The flexibility of the order-adaptive predictor makes the proposed algorithm suitable for both structural regions and detailed textures. The multi-scale framework and a novel propagation order are also involved to further improve the inpainting performance. Compared to the state-of-the-art image inpainting algorithms, experimental results show that the proposed method gives comparative or better performance. Jing Liu 0002, Guangtao Zhai, Xiaokang Yang 0001, Chang Wen Chen |
ICME | 1 |
| 2015 | Spatial Error Concealment With an Adaptive Linear PredictorabstractIn this paper, a novel spatial error concealment (EC) algorithm is proposed. Under the sequential recovery framework, pixels in missing blocks are successively reconstructed based on adaptive linear predictor. The predictor automatically tunes its order and support shape according to local contexts. The predictor order and support shape are determined using Bayesian information criterion, which is able to strike a balance between the bias and variance of the prediction errors. The flexibility of the order-adaptive predictor is able to recover more important features or structures. A novel scan order based on the uncertainty of each pixel is also proposed to alleviate error propagation problem. Compared with the state-of-the-art EC algorithms, experimental results show that the proposed method gives better reconstruction performance in terms of objective and subjective evaluations. Jing Liu 0002, Guangtao Zhai, Xiaokang Yang 0001, Bing Yang 0003, Li Chen 0021 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2014 | Spatial error concealment with adaptive linear predictorabstractIn this paper, we propose a novel spatial error concealment algorithm based on adaptive linear predictor. The predictor automatically tunes its order and support shape according to local contexts. The optimal order is determined using Bayesian Index Criterion (BIC). The order-adaptive predictor is able to recover more important features or structures. Simulation results show that compared to the state-of-the-art error concealment algorithms, the proposed method gives better reconstruction performance in terms of objective and subjective evaluations. Jing Liu 0002, Guangtao Zhai, Xiaokang Yang 0001, Bing Yang 0003 |
ISCAS | 1 |
| 2014 | Lossless Predictive Coding for Images With Bayesian TreatmentabstractAdaptive predictor has long been used for lossless predictive coding of images. Most of existing lossless predictive coding techniques mainly focus on suitability of prediction model for training set with the underlying assumption of local consistency, which may not hold well on object boundaries and cause large predictive error. In this paper, we propose a novel approach based on the assumption that local consistency and patch redundancy exist simultaneously in natural images. We derive a family of linear models and design a new algorithm to automatically select one suitable model for prediction. From the Bayesian perspective, the model with maximum posterior probability is considered as the best. Two types of model evidence are included in our algorithm. One is traditional training evidence, which represents the models’ suitability for current pixel under the assumption of local consistency. The other is target evidence, which is proposed to express the preference for different models from the perspective of patch redundancy. It is shown that the fusion of training evidence and target evidence jointly exploits the benefits of local consistency and patch redundancy. As a result, our proposed predictor is more suitable for natural images with textures and object boundaries. Comprehensive experiments demonstrate that the proposed predictor achieves higher efficiency compared with the state-of-the-art lossless predictors. Jing Liu 0002, Guangtao Zhai, Xiaokang Yang 0001, Li Chen 0021 |
IEEE Trans. Image Process. | 1 |
| 2013 | Hybrid image interpolation with soft-decision kernel regressionabstractParametric linear autoregressive (AR) model has been widely used in image processing but is known to induce unstable results. The recently emerged nonparametric kernel regression is an effective structural method for forestalling outliers but often brings over-smoothed output. This paper introduces a hybrid algorithm for image interpolation through combining the strength of parametric and nonparametric modeling techniques. More specifically, it is a soft-decision kernel regression (SKR) method in which the soft-decision AR model is embedded into the adaptive kernel regression framework. Compared with the state-of-the-art interpolation methods, simulation results show that the proposed SKR algorithm achieves comparative or better results in terms of objective and subjective quality. Jing Liu 0002, Xiaokang Yang 0001, Guangtao Zhai, Li Chen 0021 |
ISCAS | 1 |
| 2013 | Lossless predictive coding with Bayesian treatmentabstractNatural image statistics have been widely exploited for lossless predictive coding and other applications. However, traditional adaptive techniques always focus on the local consistency of training set regardless of what the predicted target looks like. We investigate the problem of introducing the model evidence of predicted target since self-similarity inherent in natural images gives some kind of prior information for the distribution of predicted result. The proposed Bayesian model integrated with both training evidence and target evidence takes full advantages of local structure as well as self-similarity. Experimental results demonstrate that the proposed context model achieves best results compared with the state-of-the-art lossless predictors. Jing Liu 0002, Xiaokang Yang 0001, Guangtao Zhai, Li Chen 0021, Xianghui Sun, Wanhong Chen, Ying Zuo |
VCIP | 1 |