EDBT 2026 Demo / reviewers in the wild / expert
Jun Jia
dblp:75/7384
· DBLP profile ↗
49ranked-venue papers
7as first author
43since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 3 first-author · 30 since 2021Artificial intelligence and machine learning · 11 · 2 first-author · 10 since 2021Computer networks · 7 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VQAThinker: Exploring Generalizable and Explainable Video Quality Assessment via Reinforcement LearningabstractVideo quality assessment (VQA) aims to objectively quantify perceptual quality degradation in alignment with human visual perception. Despite recent advances, existing VQA models still suffer from two critical limitations: poor generalization to out-of-distribution (OOD) videos and limited explainability, which restrict their applicability in real-world scenarios. To address these challenges, we propose VQAThinker, a reasoning-based VQA framework that leverages large multimodal models (LMMs) with reinforcement learning to jointly model video quality understanding and scoring, emulating human perceptual decision-making. Specifically, we adopt group relative policy optimization (GRPO), a rule-guided reinforcement learning algorithm that enables reasoning over video quality under score-level supervision, and introduce three VQA-specific rewards: (1) a bell-shaped regression reward that increases rapidly as the prediction error decreases and becomes progressively less sensitive near the ground truth; (2) a pairwise ranking reward that guides the model to correctly determine the relative quality between video pairs; and (3) a temporal consistency reward that encourages the model to prefer temporally coherent videos over their perturbed counterparts. Extensive experiments demonstrate that VQAThinker achieves state-of-the-art performance on both in-domain and OOD VQA benchmarks, showing strong generalization for video quality scoring. Furthermore, evaluations on video quality understanding tasks validate its superiority in distortion attribution and quality description compared to existing explainable VQA models and LMMs. These findings demonstrate that reinforcement learning offers an effective pathway toward building generalizable and explainable VQA models solely with score-level supervision. Linhan Cao, Wei Sun 0029, Weixia Zhang, Jun Jia, Kaiwei Zhang, Dandan Zhu 0001, Guangtao Zhai, Xiongkuo Min |
AAAI | 5 |
| 2026 | M3DGCQA: A Quality Assessment Dataset for Multi-Object 3D Generated Contents
Farong Wen, Yuanhao Xue, Xiahui Ren, Ziying Wang, Yingjie Zhou 0003, Jun Jia, Jiezhang Cao, Xiaohong Liu 0001, Guangtao Zhai |
QoMEX | 7 |
| 2026 | Enhancing blind video quality assessment with rich quality-aware features
Wei Sun 0029, Linhan Cao, Jun Jia, Xiongkuo Min, Guangtao Zhai |
Expert Syst. Appl. | 3 |
| 2026 | MI3S: A multimodal large language model assisted quality assessment framework for AI-generated talking heads
Yingjie Zhou 0003, Sijing Wu, Jun Jia, Yanwei Jiang, Wei Sun 0029, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai |
Inf. Process. Manag. | 4 |
| 2026 | Diffusion-based adversarial attacks and defenses on template for visual object tracking
Haibo Pang, Chengming Liu, Jun Jia, Qun Jin |
Mach. Vis. Appl. | 4 |
| 2026 | Surveillance Facial Image Quality Assessment: A Multi-Dimensional Dataset and Lightweight ModelabstractSurveillance facial images are often captured under unconstrained conditions, resulting in severe quality degradation due to factors such as low resolution, motion blur, occlusion, and poor lighting. Although recent face restoration techniques applied to surveillance cameras can significantly enhance visual quality, they often compromise fidelity (i.e., identity-preserving features), which directly conflicts with the primary objective of surveillance images -- reliable identity verification. Existing facial image quality assessment (FIQA) predominantly focus on either visual quality or recognition-oriented evaluation, thereby failing to jointly address visual quality and fidelity, which are critical for surveillance applications. To bridge this gap, we propose the first comprehensive study on surveillance facial image quality assessment (SFIQA), targeting the unique challenges inherent to surveillance scenarios. Specifically, we first construct SFIQA-Bench, a multi-dimensional quality assessment benchmark for surveillance facial images, which consists of 5,004 surveillance facial images captured by three widely deployed surveillance cameras in real-world scenarios. A subjective experiment is conducted to collect six dimensional quality ratings, including noise, sharpness, colorfulness, contrast, fidelity and overall quality, covering the key aspects of SFIQA. Furthermore, we propose SFIQA-Assessor, a lightweight multi-task FIQA model that jointly exploits complementary facial views through cross-view feature interaction, and employs learnable task tokens to guide the unified regression of multiple quality dimensions. The experiment results on the proposed dataset show that our method achieves the best performance compared with the state-of-the-art general image quality assessment (IQA) and FIQA methods, validating its effectiveness for real-world surveillance applications. Yanwei Jiang, Wei Sun 0029, Yingjie Zhou 0003, Yuqin Cao, Jun Jia, Sijing Wu, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | SingingHead: A Large-Scale 4D Dataset for Singing Head AnimationabstractSinging, as a common facial movement second only to talking, can be regarded as a universal language across ethnicities and cultures, plays an important role in emotional communication, art, and entertainment. However, it is often overlooked in the field of audio-driven 3D facial animation due to the lack of singing head datasets and the domain gap between singing and talking in rhythm and amplitude. To this end, we collect a large-scale high-quality multi-modal singing head dataset,SingingHead, which consists of more than 27 hours of synchronized singing video, 3D facial motion, singing audio, and background music from 76 individuals and 8 types of music. Along with the SingingHead dataset, we benchmark existing audio-driven 3D facial animation methods and 2D talking head methods on the singing task. Existing 3D facial animation methods and 2D talking head methods fail to produce satisfactory singing results. Focusing on the 3D singing head animation, we first utilize the proposed singing-specific dataset to retrain the 3D facial animation methods, resulting in substantial performance improvements. Besides, considering the absence of background music and the slow generation speed of existing methods, we propose a simple but efficient non-autoregressive VAE-based framework with background music as an input signal to generate diverse and accurate 3D singing facial motions in real time. Extensive experiments demonstrate the significance of the SingingHead dataset in promoting the development of singing head animation. The dataset is released for research purposes at:https://wsj-sjtu.github.io/SingingHead/. Sijing Wu, Weitian Zhang, Jun Jia, Yucheng Zhu, Yichao Yan, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Multim. | 4 |
| 2026 | FreeQR: Free Lunch for Aesthetic QR Codes Emerging From the Latent Space in Diffusion ModelsabstractIn the modern digital age, Quick Response (QR) codes serve as a critical interface for bridging the physical and virtual worlds, widely utilized in multimedia applications. However, traditional binary QR codes often lack the visual appeal desired in contexts. Aesthetic QR codes address this limitation by enabling the customization of QR code patterns to enhance visual attractiveness while retaining compatibility with standard QR decoders. Previous works have explored the use of diffusion models for generating such codes but often require extensive training of ControlNets and face challenges in maintaining scannability. To address these issues, we present FreeQR, a streamlined and effective approach that enables the stable generation of QR code images with diffusion models. Our methodology involves the strategic fusion between the specific channel in the latent space of the denoising process with the noised latent representations of the QR blueprint image at corresponding timesteps. This ensures that the generated images adhere to the brightness distribution required for effective scanning while achieving a balance between aesthetics and functionality. Additionally, we introduce gradient guidance based on scanning errors directly in the latent space, enabling the generation of scannable QR codes in seconds without additional model parameters. Experimental results demonstrate that FreeQR significantly enhances the aesthetics and scannability of QR codes compared to existing methods, making it a lightweight and efficient solution for multimedia applications. Yiwei Yang 0007, Jun Jia, Zheyuan Liu 0011, Zhongpai Gao, Wei Sun 0029, Guangtao Zhai |
IEEE Trans. Multim. | 2 |
| 2025 | 3DGCQA: A Quality Assessment Database for 3D AI-Generated ContentsabstractAlthough 3D generated content (3DGC) offers advantages in reducing production costs and accelerating design timelines, its quality often falls short when compared to 3D professionally generated content. Common quality issues frequently affect 3DGC, highlighting the importance of timely and effective quality assessment. Such evaluations not only ensure a higher standard of 3DGCs for end-users but also provide critical insights for advancing generative technologies. To address existing gaps in this domain, this paper introduces a novel 3DGC quality assessment dataset, 3DGCQA, built using 7 representative Text-to-3D generation methods. During the dataset’s construction, 50 fixed prompts are utilized to generate contents across all methods, resulting in the creation of 313 textured meshes that constitute the 3DGCQA dataset. The visualization intuitively reveals the presence of 6 common distortion categories in the generated 3DGCs. To further explore the quality of the 3DGCs, subjective quality assessment is conducted by evaluators, whose ratings reveal significant variation in quality across different generation methods. Additionally, several objective quality assessment algorithms are tested on the 3DGCQA dataset. The results expose limitations in the performance of existing algorithms and underscore the need for developing more specialized quality assessment methods. To provide a valuable resource for future research and development in 3D content generation and quality assessment, the dataset has been open-sourced in https://github.com/zyj-2000/3DGCQA. Yingjie Zhou 0003, Farong Wen, Jun Jia, Yanwei Jiang, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai |
ICASSP | 4 |
| 2025 | Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking HeadsabstractSpeech-driven methods for portraits are figuratively known as "Talkers" because of their capability to synthesize speaking mouth shapes and facial movements. Especially with the rapid development of the Text-to-Image (T2I) models, AI-Generated Talking Heads (AGTHs) have gradually become an emerging digital human media. However, challenges persist regarding the quality of these talkers and AGTHs they generate, and comprehensive studies addressing these issues remain limited. To address this gap, this paper presents the largest AGTH quality assessment dataset THQA-10K to date, which selects 12 prominent T2I models and 14 advanced talkers to generate AGTHs for 14 prompts. After excluding instances where AGTH generation is unsuccessful, the THQA-10K dataset contains 10,457 AGTHs. Then, volunteers are recruited to subjectively rate the AGTHs and give the corresponding distortion categories. In our analysis for subjective experimental results, we evaluate the performance of talkers in terms of generalizability and quality, and also expose the distortions of existing AGTHs. Finally, an objective quality assessment method based on the first frame, Y-T slice and tone-lip consistency is proposed. Experimental results show that this method can achieve state-of-the-art (SOTA) performance in AGTH quality assessment. The work is released at https://github.com/zyj-2000/Talker. Yingjie Zhou 0003, Jiezhang Cao, Farong Wen, Yanwei Jiang, Jun Jia, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai |
ICCV | 6 |
| 2025 | DiffDeid: High-Quality Face De-identification and Recovery via Diffusion InversionabstractNowadays, personal privacy protection is extremely emphasised. Face de-identification is considered as an effective way to protect the visual privacy through disguising or replacing identity attributes. Existing methods compromise either high fidelity or reversibility. To address these issues, this paper proposes DiffDeid, the first diffusion-based face de-identification and recovery method. Leveraging recent diffusion inversion and control technicques, DiffDeid achieves both high quality imperceptible de-identification and exact recovery with passwords. DiffDeid has three attractions: (1) It can generate de-identified faces with the superior fidelity while maintaining other non-identity attributes for visual tasks. (2) The correct password is powerful enough to restore facial images with extreme details. Meanwhile, incorrect passwords can lead to vastly different decryption results. (3) DiffDeid demands minimal computing resources and instant training time compared to others. We conducted experiments on various face datasets to showcase the superiority of our proposed method. Additional experiments show that DiffDeid is powerful with diverse text prompts and control instructions even beyond human faces. Codes are available at project page. Zheyuan Liu 0011, Jun Jia, Hongyi Miao, Yiwei Yang 0007, Yanwei Jiang, Yingjie Zhou 0003, Zhi Liu 0004, Guangtao Zhai |
ICME | 2 |
| 2025 | CAP: An Advanced No-Reference Quality Assessment Method for AI-Generated 3D MeshesabstractThe advent of generative AI has revolutionized 3D content design, significantly enhancing modelers’ efficiency. However, the quality of generated 3D content, particularly Generated Meshes (GMs), remains a critical concern. GMs pose unique challenges for quality assessment due to their complex geometry, detailed texture mapping, and distortions that differ from traditional meshes. Existing methods fail to address these GM-specific issues. To tackle this gap, we introduce a novel no-reference quality assessment method, CAP, which integrates CT-Slice, prompt Alignment, and Projections. CAP employs a six-face projection to capture external features and a CT-like slicing approach to extract internal quality features. Additionally, it leverages Contrastive Language-Image Pre-Training (CLIP) to measure the alignment between projection embeddings and prompts as a key quality indicator. Experimental results demonstrate that CAP effectively evaluates GM quality by combining internal, external, and alignment features. The code for this work has been open-sourced in https://github.com/zyj-2000/CAP. Yingjie Zhou 0003, Farong Wen, Yanwei Jiang, Jun Jia, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai |
ICME | 5 |
| 2025 | LPerceptual Quality Assessment of AI Generated Content Videos: a Dataset and BenchmarkabstractIn recent years, artificial intelligence (AI) driven video generation has garnered significant attention due to advancements in large language model techniques. Thus, there is a great demand to explore the effectiveness of video quality assessment (VQA) models in evaluating the perceptual quality of AI-generated content (AIGC) videos and in optimizing video generation techniques. Therefore, in this paper, we try to systemically investigate the AIGC-VQA problem from both subjective and objective quality assessment perspectives. For the subjective perspective, we construct a Large-scale Generated Video Quality assessment (LGVQ) dataset, consisting of 2,808 AIGC videos generated by 6 video generation models using 468 carefully selected text prompts. We evaluate the perceptual quality of AIGC videos from three dimensions: spatial quality, temporal quality, and text-to-video alignment, which hold the utmost importance for current video generation techniques. For the objective perspective, we establish a benchmark for evaluating existing quality assessment metrics on the LGVQ dataset, which fully demonstrates the performance of current mainstream VQA methods in evaluating AIGV quality. We hope that this work can contribute to the advancement of AIGC video generation technology as well as the evaluation techniques for AIGC videos. The LGVQ dataset will release publicly. Wei Sun 0029, Xinyue Li 0001, Jun Jia, Xiongkuo Min, Chunyi Li 0001, Zhongpeng Ji, Fengyu Sun, Shangling Jui, Guangtao Zhai |
ISCAS | 5 |
| 2025 | Human-Activity AGV Quality Assessment: A Benchmark Dataset and an Objective Evaluation MetricabstractAI-driven video generation techniques have made significant progress in recent years. However, AI-generated videos (AGVs) involving human activities often exhibit substantial visual and semantic distortions, hindering the practical application of video generation technologies in real-world scenarios. To address this challenge, we conduct a pioneering study on human activity AGV quality assessment, focusing on visual quality evaluation and the identification of semantic distortions. First, we construct the AI-Generated Human activity Video Quality Assessment (Human-AGVQA) dataset, consisting of 6,000 AGVs derived from 15 popular text-to-video (T2V) models using 400 text prompts that describe diverse human activities. We conduct a subjective study to evaluate the human appearance quality, action continuity quality, and overall video quality of AGVs, and identify semantic issues of human body parts. Based on Human-AGVQA, we benchmark the performance of T2V models and analyze their strengths and weaknesses in generating different categories of human activities. Second, we develop an objective evaluation metric, named AI-Generated Human activity Video Quality metric (GHVQ), to automatically analyze the quality of human activity AGVs. GHVQ systematically extracts human-focused quality features, AI-generated content-aware quality features, and temporal continuity features, making it a comprehensive and explainable quality metric for human activity AGVs. The extensive experimental results show that GHVQ outperforms existing quality metrics on the Human-AGVQA dataset by a large margin, demonstrating its efficacy in assessing the quality of human activity AGVs. The Human-AGVQA dataset and GHVQ metric will be released at https://github.com/zczhang-sjtu/GHVQ.git. Wei Sun 0029, Xinyue Li 0001, Qihang Ge, Jun Jia, Zhongpeng Ji, Fengyu Sun, Shangling Jui, Xiongkuo Min, Guangtao Zhai |
ACM Multimedia | 6 |
| 2025 | Subjective and Objective Quality-of-Experience Evaluation Study for Live Video StreamingabstractIn recent years, live video streaming has gained widespread popularity across various social media platforms. Quality of experience (QoE), which reflects end-users’ satisfaction and overall experience, plays a critical role for media service providers to optimize large-scale live compression and transmission strategies to achieve perceptually optimal rate-distortion trade-off. Although many QoE metrics for video-on-demand (VoD) have been proposed, there remain significant challenges in developing QoE metrics for live video streaming. To bridge this gap, we conduct a comprehensive study of subjective and objective QoE evaluations for live video streaming. For the subjective QoE study, we introduce the first live video streaming QoE dataset, TaoLive QoE, which consists of 42 source videos collected from real live broadcasts and 1, 155 corresponding distorted ones degraded due to a variety of streaming distortions, including conventional streaming distortions such as compression, stalling, as well as live streaming-specific distortions like frame skipping, variable frame rate, etc. Subsequently, a human study was conducted to derive subjective QoE scores of videos in the TaoLive QoE dataset. For the objective QoE study, we benchmark existing QoE models on the TaoLive QoE dataset as well as publicly available QoE datasets for VoD scenarios, highlighting that current models struggle to accurately assess video QoE, particularly for live content. Hence, we propose an end-to-end QoE evaluation model, Tao-QoE, which integrates multi-scale semantic features and optical flow-based motion features to predicting a retrospective QoE score, eliminating reliance on statistical quality of service (QoS) features. Extensive experiments demonstrate that Tao-QoE outperforms other models on the TaoLive QoE dataset and five publicly available QoE datasets, showcasing the effectiveness and feasibility of Tao-QoE. Zehao Zhu, Wei Sun 0029, Jun Jia, Jia Wang 0004, Guangtao Zhai |
VCIP | 3 |
| 2025 | Joint Luminance-Chrominance Learning for Image DebandingabstractBanding is a visually annoying artifact that frequently occurs along the chain of video acquisition, production, distribution, and display, showing a significant need for improvement in many fields. Thus far, efforts on banding removal are mainly knowledge-driven or merely learning on RGB space, which is either limited by domain knowledge or lacks the consideration for banding in chrominance channels. In this work, we propose a unified deep neural network that explicitly disentangles the luminance and chrominance channels, and simultaneously recovers intensity gradients and color discontinuity from detection-free measurement in an end-to-end manner. Our debanding model is comprised of a luminance restoration network (LR-Net) and a chrominance restoration network (CR-Net). Each of them follows an encoder-decoder architecture, where a cascade of residual blocks is employed to exploit hierarchical non-local features in spatial dimensions for more powerful feature representation. Moreover, we investigate the characteristics of banding artifacts and apply specific loss functions to guide the debanding in different channels, thus boosting the restoration performance. Both qualitative and quantitative experiments show that our model significantly surpasses the existing method in terms of all 7 metrics. Ultimately, our network trained on simulated data exhibits good adaptiveness under various compression scenarios, which further demonstrates the effectiveness of the proposed model. Zijian Chen 0001, Wei Sun 0029, Jun Jia, Ru Huang 0002, Fangfang Lu, Ying Chen 0011, Xiongkuo Min, Guangtao Zhai, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Study of Subjective and Objective Naturalness Assessment of AI-Generated ImagesabstractThe proliferation of Artificial Intelligence-Generated Images (AIGIs) has greatly expanded the Image Naturalness Assessment (INA) problem. Different from early definitions that mainly focus on tone-mapped images with limited distortions (e.g., exposure, contrast, and color reproduction), INA on AI-generated images is especially challenging as it owns more diverse contents and could be affected by factors from multiple perspectives, including low-level technical distortions and high-level rationality distortions. In this paper, we take the first step to benchmark and assess the visual naturalness of AI-generated images. First, we construct the AI-Generated Image Naturalness (AGIN) dataset by conducting a large-scale subjective study to collect human opinions on the overall naturalness as well as perceptions from the technical quality and rationality perspectives. AGIN verifies several insights for the first time that naturalness is universally and disparately affected by both technical and rational distortions, while its manifestations vary with different generation tasks. Second, to automatically assess the naturalness of AIGIs that align with human opinions, we propose the Joint Objective Image Naturalness evaluaTor (JOINT). Specifically, JOINT imitates human reasoning in naturalness evaluation by jointly learning technical and rationality features with several specific designs to guide model behavior from respective perspectives. Experiments demonstrate that JOINT significantly outperforms existing methods for providing more subjectively consistent results on naturalness assessment. The dataset can be accessed athttps://github.com/zijianchen98/AGIN. Zijian Chen 0001, Wei Sun 0029, Haoning Wu 0001, Jun Jia, Ru Huang 0002, Xiongkuo Min, Guangtao Zhai, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Full-Reference and No-Reference Quality Assessment for Video Frame InterpolationabstractVideo frame interpolation (VFI) synthesizes new frames from original video frames to produce high frame-rate videos and enhance their visual appeal. The quality of these interpolated frames significantly affects the perceptual experience of the synthesized video. Recent research in VFI has increasingly focused on perceptual quality of the interpolated frames and the overall video. However, most existing quality metrics do not align well with human perceptual experiences and often suffer from unnatural artifacts in the interpolated frames. Consequently, there is an urgent need for VFI video quality assessment (VFIVQA) methods to assess the quality of the synthesized videos. In this paper, we propose both a full-reference (FR) method and a no-reference (NR) method for VFIVQA. The FR method employs two feature extraction blocks to measure continuous frame changes, extracting flow features with short temporal spans and motion features with long temporal spans. By calculating multilevel similarities in the temporal dimension of 3D convolutional neural networks and fusing these similarity features, the quality score of the VFI video is obtained from the quality regression network. Since the flow feature extraction block does not utilize the reference VFI video, the proposed NR method consists solely of this feature block. Extensive validation on several VFIVQA datasets demonstrates that the proposed methods outperform state-of-the-art FR and NR methods. Jinliang Han, Xiongkuo Min, Jun Jia, Xiaohong Liu 0001, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Who Is a Better Imitator: Subjective and Objective Quality Assessment of Animated HumansabstractAnimated human (AH) have gained popularity due to their vivid appearance and smooth, natural movements. Various animation methods based on artificial intelligence (AI) have been introduced, which are viewed as “Imitators,” offering new solutions for designing AHs. However, the effectiveness of these AI-generated AHs varies significantly across different categories and within the same category, leading to visual distortions that adversely affect the viewer’s experience. Consequently, it is essential to evaluate the quality of AHs to provide reliable and objective indicators for their further development and to ensure the delivery of higher-quality AH videos to users. In this paper, the first Animated Human Quality Assessment (AHQA) dataset is constructed by selecting 6 advanced and popular imitators and 10 common actions to animate 20 AI-generated characters. The constructed dataset integrates different genders and age groups of character images, and two types of poses, standing and sitting, are selected, highlighting the comprehensiveness and diversity of the AHQA dataset. Subjective experiments reveal significant differences in the quality of AHs produced by different imitators. Finally, we propose a quality assessment method, VIP-QA, incorporating Video quality, Identity consistency, and Posture similarity for the AHQA dataset. Experimental results show that VIP-QA significantly outperforms existing assessment methods on multiple datasets by about 5%, more closely approximates human visual perception, and provides a valid objective metric for assessing imitators. All the work in this paper has been released at https://github.com/zyj-2000/Imitator. Yingjie Zhou 0003, Jun Jia, Yanwei Jiang, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Benchmarking Multi-dimensional AIGC Video Quality Assessment: A Dataset and Unified ModelabstractIn recent years, AI-driven video generation has gained significant attention due to great advancements in visual and language generative techniques. Consequently, there is a growing need for accurate Video Quality Assessment (VQA) metrics to evaluate the perceptual quality of AI-generated content (AIGC) videos and optimize video generation models. However, assessing the quality of AIGC videos remains a significant challenge because these videos often exhibit highly complex distortions, such as unnatural actions and irrational objects. To address this challenge, we systematically investigate the AIGC-VQA problem in this article, considering both subjective and objective quality assessment perspectives. For the subjective perspective, we construct the L arge-scale G enerated V ideo Q uality Assessment (LGVQ) dataset, consisting of \(2,\!808\) AIGC videos generated by six video generation models using 468 carefully curated text prompts. Unlike previous subjective VQA experiments, we evaluate the perceptual quality of AIGC videos from three critical dimensions: spatial quality, temporal quality, and text-video alignment, which hold utmost importance for current video generation techniques. For the objective perspective, we establish a benchmark for evaluating existing quality assessment metrics on the LGVQ dataset. Our findings show that current metrics perform poorly on this dataset, highlighting a gap in effective evaluation tools. To bridge this gap, we propose the U nify G enerated V ideo Q uality Assessment (UGVQ) model, designed to accurately evaluate the multi-dimensional quality of AIGC videos. The UGVQ model integrates the visual and motion features of videos with the textual features of their corresponding prompts, forming a unified quality-aware feature representation tailored to AIGC videos. Experimental results demonstrate that UGVQ achieves state-of-the-art performance on the LGVQ dataset across all three quality dimensions, validating its effectiveness as an accurate quality metric for AIGC videos. We hope that our benchmark can promote the development of AIGC-VQA studies. Both the LGVQ dataset and the UGVQ model are publicly available on https://github.com/zczhang-sjtu/UGVQ.git . Wei Sun 0029, Xinyue Li 0001, Jun Jia, Xiongkuo Min, Chunyi Li 0001, Zijian Chen 0001, Puyi Wang, Fengyu Sun, Shangling Jui, Guangtao Zhai |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | Text2QR: Harmonizing Aesthetic Customization and Scanning Robustness for Text-Guided QR Code GenerationabstractIn the digital era, QR codes serve as a linchpin connecting virtual and physical realms. Their pervasive integration across various applications highlights the demand for aesthetically pleasing codes without compromised scannability. However, prevailing methods grapple with the intrinsic challenge of balancing customization and scannability. Notably, stable-diffusion models have ushered in an epoch of high-quality, customizable content generation. This paper introduces Text2QR, a pioneering approach leveraging these advancements to address a fundamental challenge: concurrently achieving user-defined aesthetics and scanning robustness. To ensure stable generation of aesthetic QR codes, we introduce the QR Aesthetic Blueprint (QAB) module, generating a blueprint image exerting control over the entire generation process. Subsequently, the Scannability Enhancing Latent Refinement (SELR) process refines the output iteratively in the latent space, enhancing scanning robustness. This approach harnesses the potent generation capabilities of stable-diffusion models, navigating the trade-off between image aesthetics and QR code scannability. Our experiments demonstrate the seamless fusion of visual appeal with the practical utility of aesthetic QR codes, markedly outperforming prior methods. Codes are available at https://github.com/mulns/Text2QR Guangyang Wu, Xiaohong Liu 0001, Jun Jia, Xuehao Cui, Guangtao Zhai |
CVPR | 3 |
| 2024 | SG-JND: Semantic-Guided Just Noticeable Distortion Predictor for Image CompressionabstractJust noticeable distortion (JND), representing the threshold of distortion in an image that is minimally perceptible to the human visual system (HVS), is crucial for image compression algorithms to achieve a trade-off between transmission bit rate and image quality. However, traditional JND prediction methods only rely on pixel-level or sub-band level features, lacking the ability to capture the impact of image content on JND. To bridge this gap, we propose a Semantic-Guided JND (SG-JND) network to leverage semantic information for JND prediction. In particular, SG-JND consists of three essential modules: the image preprocessing module extracts semantic-level patches from images, the feature extraction module extracts multi-layer features by utilizing the cross-scale attention layers, and the JND prediction module regresses the extracted features into the final JND value. Experimental results show that SG-JND achieves the state-of-the-art performance on two publicly available JND datasets, which demonstrates the effectiveness of SG-JND and highlight the significance of incorporating semantic information in JND assessment. Linhan Cao, Wei Sun 0029, Xiongkuo Min, Jun Jia, Zijian Chen 0001, Yucheng Zhu, Lizhou Liu, Qiubo Chen, Guangtao Zhai |
ICIP | 4 |
| 2024 | DiffStega: Towards Universal Training-Free Coverless Image Steganography with Diffusion Models
Yiwei Yang 0007, Zheyuan Liu 0011, Jun Jia, Zhongpai Gao, Wei Sun 0029, Xiaohong Liu 0001, Guangtao Zhai |
IJCAI | 3 |
| 2024 | TF-ChineseECE: A Chinese Event Causality Extraction Method Combining Roberta and Bi-FLASH-SRUabstractIn response to the excessive number of filled forms in the existing table-based event relation extraction and the inadequate acquisition of table features. In the paper, TF-ChineseECE, a Chinese event causality extraction method combining Roberta and Bi-FLASH-SRU, is proposed. A table filling strategy is developed to transform the tagged relations in the text into a table with labels; the subject and object event features are obtained with the help of Roberta and Bidirectional Built-in Flash Attention Simple Recurrent Unit (Bi-FLASH-SRU), and then the The table feature recurrent learning module is then used to mine the global features, and finally the table decoding is performed to obtain the event causality triad. The experiments are validated using two public datasets in the financial domain, and the results show that the F1 scores of the method proposed in this paper reach 59.2% and 62.5%, respectively, and the Bi-FLASH-SRU model is faster to train and the number of filled forms is less, which proves the effectiveness of the method. Quanlin Chen, Jun Jia, Shuo Fan |
IJCNN | 2 |
| 2024 | Large Multi-modality Model Assisted AI-Generated Image Quality AssessmentabstractTraditional deep neural network (DNN)-based image quality assessment (IQA) models leverage convolutional neural networks (CNN) or Transformer to learn the quality-aware feature representation, achieving commendable performance on natural scene images. However, when applied to AI-Generated images (AGIs), these DNN-based IQA models exhibit subpar performance. This situation is largely due to the semantic inaccuracies inherent in certain AGIs caused by uncontrollable nature of the generation process. Thus, the capability to discern semantic content becomes crucial for assessing the quality of AGIs. Traditional DNN-based IQA models, constrained by limited parameter complexity and training data, struggle to capture complex fine-grained semantic features, making it challenging to grasp the existence and coherence of semantic content of the entire image. To address the shortfall in semantic content perception of current IQA models, we introduce a large Multi-modality model Assisted AI-Generated Image Quality Assessment (MA-AGIQA) model, which utilizes semantically informed guidance to sense semantic information and extract semantic vectors through carefully designed text prompts. Moreover, it employs a mixture of experts (MoE) structure to dynamically integrate the semantic information with the quality-aware features extracted by traditional DNN-based IQA models. Comprehensive experiments conducted on two AI-generated content datasets and two traditional IQA datasets show that MA-AGIQA achieves state-of-the-art performance, and demonstrate its superior generalization capabilities on assessing the quality of AGIs. The code is available at https://github.com/wangpuyi/MA-AGIQA. Puyi Wang, Wei Sun 0029, Jun Jia, Yanwei Jiang, Xiongkuo Min, Guangtao Zhai |
ACM Multimedia | 4 |
| 2024 | GAIA: Rethinking Action Quality Assessment for AI-Generated VideosabstractAssessing action quality is both imperative and challenging due to its significant impact on the quality of AI-generated videos, further complicated by the inherently ambiguous nature of actions within AI-generated video (AIGV). Current action quality assessment (AQA) algorithms predominantly focus on actions from real specific scenarios and are pre-trained with normative action features, thus rendering them inapplicable in AIGVs. To address these problems, we construct GAIA, a Generic AI-generated Action dataset, by conducting a large-scale subjective evaluation from a novel causal reasoning-based perspective, resulting in 971,244 ratings among 9,180 video-action pairs. Based on GAIA, we evaluate a suite of popular text-to-video (T2V) models on their ability to generate visually rational actions, revealing their pros and cons on different categories of actions. We also extend GAIA as a testbed to benchmark the AQA capacity of existing automatic evaluation methods. Results show that traditional AQA methods, action-related metrics in recent T2V benchmarks, and mainstream video quality methods perform poorly with an average SRCC of 0.454, 0.191, and 0.519, respectively, indicating a sizable gap between current models and human action perception patterns in AIGVs. Our findings underscore the significance of action quality as a unique perspective for studying AIGVs and can catalyze progress towards methods with enhanced capacities for AQA in AIGVs. Zijian Chen 0001, Wei Sun 0029, Yuan Tian 0017, Jun Jia, Ru Huang 0002, Xiongkuo Min, Guangtao Zhai, Wenjun Zhang 0005 |
NeurIPS | 4 |
| 2024 | ReLI-QA: A Multidimensional Quality Assessment Dataset for Relighted Human HeadsabstractLighting conditions significantly affect the quality of both real and AI-generated images. Facial images are particularly sensitive to lighting due to their detailed nature and the importance of facial features in conveying identity. Poor lighting can easily obscure these critical details. To address this issue, various portrait relighting methods have been developed to adjust the lighting in improperly exposed images. However, these methods often encounter challenges such as overexposure, underexposure, and detail loss in the relighted portraits. Consequently, there is a need for effective quality assessment and control of relighted human heads (RHHs). In this study, one proposed simple baseline and three typical relighting methods are applied to six selected human head (HH) images, resulting in the creation of a quality assessment dataset named ReLI-QA, which comprises 840 RHHs. A multidimensional subjective quality assessment method based on visual guidance is proposed to accurately evaluate the visual quality of each RHH in the dataset. By analyzing the results of subjective experiments, the quality of RHHs is shown to be affected by multiple factors. Finally, based on ReLI-QA, some typical image quality assessment (IQA) methods are selected for benchmark experiments. The experimental results show the limitations of the existing methods in RHH quality assessment. The dataset and code for this research has been released at https://github.com/zyj-2000/ReLI-QA. Yingjie Zhou 0003, Farong Wen, Jun Jia, Xiongkuo Min, Jia Wang 0004, Guangtao Zhai |
VCIP | 4 |
| 2024 | BAND-2k: Banding Artifact Noticeable Database for Banding Detection and Quality AssessmentabstractBanding, also known as staircase-like contours, frequently occurs in flat areas of images/videos processed by compression or quantization algorithms. As undesirable artifacts, banding destroys the original image structure, thus inevitably degrading users’ quality of experience (QoE). In this paper, we systematically investigate the banding image quality assessment (IQA) problem, aiming to detect the image banding artifacts and evaluate their perceptual visual quality. Considering that the existing image banding databases only contain limited content sources and banding generation methods, and lack perceptual quality labels (i.e. mean opinion scores), we first build the largest banding IQA database so far, namedBanding Artifact Noticeable Database (BAND-2k), which consists of 2,000 banding images generated by 15 compression and quantization schemes. A total of 23 workers participated in the subjective IQA experiment, yielding over 214,000 patch-level banding class labels and 44,371 reliable image-level quality rating scores. Subsequently, we develop an effective no-reference (NR) banding evaluator for banding detection and quality assessment by leveraging frequency characteristics of banding artifacts. To be more specific, a dual convolutional neural network (CNN) is employed to concurrently learn the feature representation from the high-frequency and low-frequency maps, thereby enhancing the ability to discern banding artifacts. The quality score of a banding image is generated by pooling the banding detection maps masked by the spatial frequency filters. The experimental results demonstrate that our banding evaluator achieves remarkably high accuracy in banding detection and also exhibits high SRCC and PLCC results with the perceptual quality labels, even without directly learning a regression model for banding quality evaluation. These findings unveil the strong correlations between the intensity of banding artifacts and the perceptual visual quality, thus validating the necessity of banding quality assessment. The BAND-2k database and the proposed banding evaluator are available at https://github.com/zijianchen98/BAND-2k. Zijian Chen 0001, Wei Sun 0029, Jun Jia, Fangfang Lu, Jing Liu 0002, Ru Huang 0002, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | LCGNet: Local Sequential Feature Coupling Global Representation Learning for Functional Connectivity Network Analysis With fMRIabstractAnalysis of functional connectivity networks (FCNs) derived from resting-state functional magnetic resonance imaging (rs-fMRI) has greatly advanced our understanding of brain diseases, including Alzheimer's disease (AD) and attention deficit hyperactivity disorder (ADHD). Advanced machine learning techniques, such as convolutional neural networks (CNNs), have been used to learn high-level feature representations of FCNs for automated brain disease classification. Even though convolution operations in CNNs are good at extracting local properties of FCNs, they generally cannot well capture global temporal representations of FCNs. Recently, the transformer technique has demonstrated remarkable performance in various tasks, which is attributed to its effective self-attention mechanism in capturing the global temporal feature representations. However, it cannot effectively model the local network characteristics of FCNs. To this end, in this paper, we propose a novel network structure for Local sequential feature Coupling Global representation learning (LCGNet) to take advantage of convolutional operations and self-attention mechanisms for enhanced FCN representation learning. Specifically, we first build a dynamic FCN for each subject using an overlapped sliding window approach. We then construct three sequential components (i.e., edge-to-vertex layer, vertex-to-network layer, and network-to-temporality layer) with a dual backbone branch of CNN and transformer to extract and couple from local to global topological information of brain networks. Experimental results on two real datasets (i.e., ADNI and ADHD-200) with rs-fMRI data show the superiority of our LCGNet. Biao Jie, Zhengdong Wang, Tongchun Du, Weixin Bian, Yang Yang 0140, Jun Jia |
IEEE Trans. Medical Imaging | 8 |
| 2024 | Hidden Barcode in Sub-Images with Invisible Locating MarkerabstractThe prevalence of the Internet of Things (IoT) has led to the widespread adoption of 2D barcodes as a means of offline-to-online communication. Whereas, 2D barcodes are not ideal for publicity materials, due to their space-consuming nature. Recent works have proposed 2D image barcodes that contain invisible codes or hyperlinks to transmit hidden information from offline to online. However, these methods undermine the purpose of the codes being invisible, due to the the requirement of markers to locate them. The conference version of this work has presents a novel imperceptible information embedding framework for display or print-camera scenarios, which includes not only hiding and recvoery but also locating and correcting. With the assistance of learned invisible markers, hidden codes can be rendered truly imperceptible. A highly effective multi-stage training scheme is proposed to achieve high visual fidelity and retrieval resiliency, wherein information is concealed in a sub-region rather than the entire image. However, our conference version does not address the optimal sub-region for hiding, which is crucial when dealing with local region concealment problems. In this paper extension, we consider human perceptual characteristics and introduce an optimal hiding region recommendation algorithm that comprehensively incorporates Just Noticeable Difference (JND) and visual saliency factors into consideration. Extensive experiments demonstrate superior visual quality and robustness compared to state-of-the-art methods. With the assistance of our proposed hiding region recommendation algorithm, concealed information becomes even less visible than the results of our conference version without compromising robustness. Jun Jia, Zhongpai Gao, Yiwei Yang 0007, Wei Sun 0029, Dandan Zhu 0001, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Subjective and Objective Quality Assessment for in-the-Wild Computer Graphics ImagesabstractComputer graphics images (CGIs) are artificially generated by means of computer programs and are widely perceived under various scenarios, such as games, streaming media, etc. In practice, the quality of CGIs consistently suffers from poor rendering during production, inevitable compression artifacts during the transmission of multimedia applications, and low aesthetic quality resulting from poor composition and design. However, few works have been dedicated to dealing with the challenge of computer graphics image quality assessment (CGIQA). Most image quality assessment (IQA) metrics are developed for natural scene images (NSIs) and validated on databases consisting of NSIs with synthetic distortions, which are not suitable for in-the-wild CGIs. To bridge the gap between evaluating the quality of NSIs and CGIs, we construct a large-scale in-the-wild CGIQA database consisting of 6,000 CGIs (CGIQA-6k) and carry out the subjective experiment in a well-controlled laboratory environment to obtain the accurate perceptual ratings of the CGIs. Then, we propose an effective deep learning–based no-reference (NR) IQA model by utilizing both distortion and aesthetic quality representation. Experimental results show that the proposed method outperforms all other state-of-the-art NR IQA methods on the constructed CGIQA-6k database and other CGIQA-related databases. The database is released at https://github.com/zzc-1998/CGIQA6K . Wei Sun 0029, Yingjie Zhou 0003, Jun Jia, Jing Liu 0002, Xiongkuo Min, Guangtao Zhai |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Chinese Event Extraction with Small-Scale Language Model
Quanlin Chen, Jun Jia, Shuo Fan |
ICIC (4) | 2 |
| 2023 | StableVQA: A Deep No-Reference Quality Assessment Model for Video StabilityabstractVideo shakiness is an unpleasant distortion of User Generated Content (UGC) videos, which is usually caused by the unstable hold of cameras. In recent years, many video stabilization algorithms have been proposed, yet no specific and accurate metric enables comprehensively evaluating the stability of videos. Indeed, most existing quality assessment models evaluate video quality as a whole without specifically taking the subjective experience of video stability into consideration. Therefore, these models cannot measure the video stability explicitly and precisely when severe shakes are present. In addition, there is no large-scale video database in public that includes various degrees of shaky videos with the corresponding subjective scores available, which hinders the development of Video Quality Assessment for Stability (VQA-S). To this end, we build a new database named StableDB that contains 1,952 diversely-shaky UGC videos, where each video has a Mean Opinion Score (MOS) on the degree of video stability rated by 34 subjects. Moreover, we elaborately design a novel VQA-S model named StableVQA, which consists of three feature extractors to acquire the optical flow, semantic, and blur features respectively, and a regression layer to predict the final stability score. Extensive experiments demonstrate that the StableVQA achieves a higher correlation with subjective opinions than the existing VQA-S models and generic VQA models. The database and codes are available at https://github.com/QMME/StableVQA. Tengchuan Kou, Xiaohong Liu 0001, Wei Sun 0029, Jun Jia, Xiongkuo Min, Guangtao Zhai |
ACM Multimedia | 4 |
| 2023 | Perceptual Quality Assessment for Video Frame InterpolationabstractThe quality of frames is significant for both research and application of video frame interpolation (VFI). In recent VFI studies, the methods of full-reference image quality assessment have generally been used to evaluate the quality of VFI frames. However, high frame rate reference videos, necessities for the full-reference methods, are difficult to obtain in most applications of VFI. To evaluate the quality of VFI frames without reference videos, a no-reference perceptual quality assessment method is proposed in this paper. This method is more compatible with VFI application and the evaluation scores from it are consistent with human subjective opinions. A new quality assessment dataset for VFI was constructed through subjective experiments firstly, to assess the opinion scores of interpolated frames. The dataset was created from triplets of frames extracted from high-quality videos using 9 state-of-the-art VFI algorithms. The proposed method evaluates the perceptual coherence of frames incorporating the original pair of VFI inputs. Specifically, the method applies a triplet network architecture, including three parallel feature pipelines, to extract the deep perceptual features of the interpolated frame as well as the original pair of frames. Coherence similarities of the two-way parallel features are jointly calculated and optimized as a perceptual metric. In the experiments, both full-reference and no-reference quality assessment methods were tested on the new quality dataset. The results show that the proposed method achieves the best performance among all compared quality assessment methods on the dataset. Jinliang Han, Xiongkuo Min, Jun Jia, Lei Sun 0009, Zuowei Cao, Yonglin Luo, Guangtao Zhai |
VCIP | 4 |
| 2023 | Application of QR Code Watermarking and Encryption in the Protection of Data Privacy of Intelligent Mouth-Opening TrainerabstractQuick response (QR) codes are widely used in offline to online channels to transfer information from promotional materials to mobile devices. Self-service medical equipment can record the data of each test, so the use of QR codes can realize the data exchange between patients and doctors, medical institutions, and self-service medical equipment, and create a medical information platform for health files. However, since anyone can easily read the information in the QR code, it is not conducive to the protection of patient privacy. Therefore, we propose a QR code encryption and decryption model based on robust digital watermarking. We implement digital watermarking through the generative adversarial networks and increase the robustness of the watermark by adding noise to the model. At the same time, we encrypt and decrypt the QR code information through advanced encryption standards. Experimental results show that the proposed method can well protect the privacy of patients without affecting the data acquisition by patients and doctors. Jiannan Liu, Jun Jia, Dandan Zhu 0001, Guangtao Zhai |
IEEE Internet Things J. | 4 |
| 2023 | RIVIE: Robust Inherent Video Information EmbeddingabstractImagine an interesting situation when watching a movie, we can scan the screen using our smartphones to get some extra information about this movie such as the cast, the release date, the movie's homepage, etc. Our prospect is a world where each video contains invisible information that can be delivered to us through mobile devices with cameras. This paper proposes the first deep learning-based information hiding method for videos to achieve information transmission from screens to cameras. Compared with hiding information in single images, the methods for videos need to maintain visual quality in both spatial and temporal domains. Furthermore, the training of video models builds on a large video dataset, which needs much more computational resources than training models for images. To reduce the computational complexity, we propose to simulate data on-the-fly to generate simulated sequences from single images. Then, we use the simulated data to train a spatio-temporal generator that hides information in videos while maintaining visual quality. During training, a temporal loss function based on the simulated data is exploited to ensure the temporal consistency of generated videos. After embedding, we use a decoder to recover the hidden information. To simulate the imaging pipeline from screens to cameras in the real world, we insert a distortion network between the generator and decoder. The distortion network is based on differentiable 3D rendering to cover possible distortions introduced in the procedure of camera imaging. Experimental results show that the hidden information in videos can be extracted by cameras without impacting the visual quality. Our work can be applied to many fields, such as advertisement, entertainment, and education. Jun Jia, Zhongpai Gao, Dandan Zhu 0001, Xiongkuo Min, Menghan Hu, Guangtao Zhai |
IEEE Trans. Multim. | 1 |
| 2022 | Learning Invisible Markers for Hidden Codes in Offline-to-online PhotographyabstractQR (quick response) codes are widely used as an offline-to-online channel to convey information (e.g., links) from publicity materials (e.g., display and print) to mobile devices. However, QR codes are not favorable for taking up valuable space of publicity materials. Recent works propose invisible codes/hyperlinks that can convey hidden information from offline to online. However, they require markers to locate invisible codes, which fails the purpose of invisible codes to be visible because of the markers. This paper proposes a novel invisible information hiding architecture for display/print-camera scenarios, consisting of hiding, locating, correcting, and recovery, where invisible markers are learned to make hidden codes truly invisible. We hide information in a sub-image rather than the entire image and include a localization module in the end-to-end framework. To achieve both high visual quality and high recovering robustness, an effective multi-stage training strategy is proposed. The experimental results show that the proposed method outperforms the state-of-the-art information hiding methods in both visual quality and robustness. In addition, the automatic localization of hidden codes significantly reduces the time of manually correcting geometric distortions for photos, which is a revolutionary innovation for information hiding in mobile applications. Jun Jia, Zhongpai Gao, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001 |
CVPR | 1 |
| 2022 | RIHOOP: Robust Invisible Hyperlinks in Offline and Online PhotographsabstractIn the era of multimedia and Internet, the quick response (QR) code helps people obtain information from offline to online quickly. However, the QR code is often limited in many scenarios because of its random and dull appearance. Therefore, this article proposes a novel approach to embed hyperlinks into common images, making the hyperlinks invisible for human eyes but detectable for mobile devices equipped with a camera. Our approach is an end-to-end neural network with an encoder to hide messages and a decoder to extract messages. To maintain the hidden message resilient to cameras, we build a distortion network between the encoder and the decoder to augment the encoded images. The distortion network uses differentiable 3-D rendering operations, which can simulate the distortion introduced by camera imaging in both printing and display scenarios. To maintain the visual attraction of the image with hyperlinks, a loss function conforming to the human visual system (HVS) is used to supervise the training of the encoder. Experimental results show that the proposed approach outperforms the previous work on both robustness and quality. Based on the proposed approach, many applications become possible, for example, "image hyperlinks" for advertisement on TV, website, or poster, and "invisible watermark" for copyright protection on digital resources or product packagings. Jun Jia, Zhongpai Gao, Menghan Hu, Xiongkuo Min, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Cybern. | 1 |
| 2021 | Dual-Layer Barcodes
Jun Jia, Guangtao Zhai |
PRCV (2) | 2 |
| 2021 | An Accurate and Efficient 1-D Barcode Detector for Medium of Deployment in IoT SystemsabstractCamera-based 1-D barcode detectors have a lot of applications in Internet-of-Things (IoT) systems (e.g., retail, air travel, post and parcel services, and manufacturing). Based on the observation that 1-D barcodes always come with a 12-digit product code, this article proposes an end-to-end trainable and fully convoluted model that can detect and output accurate localization results of 1-D barcode and product code simultaneously. Our method uses dilated convolutions-based feature extractors which are then combined with systematic feature merging layers to create a U-shaped network. It predicts multichannel feature maps which later yield localization results after thresholding with a confidence map generated by the model and nonmaximum suppression. Furthermore, we use a Taylor series expansion-based criterion to rank and eliminate a subset of least important convolutional filters of the model, which further increases the inference speed to a great extent. This model can act as a preprocessing module for camera-based barcode decoders. Experimental results on the data set combined from public data sets and self-collected retail products data set obtained in more challenging environment conditions demonstrate that our strategy is effective in increasing the decoding rate of existing commercial barcode decoders efficiently. Adnan Sharif, Guangtao Zhai, Jun Jia, Xiongkuo Min |
IEEE Internet Things J. | 3 |
| 2021 | Enhancing Decoding Rate of Barcode Decoders in Complex Scenes for IoT SystemsabstractCamera-based multiclass (1-D and 2-D) barcode detectors that can help in decoding barcodes in different complex scenes have huge potential applications in situations where Internet of Things (IoT) is combined with artificial intelligence (AI) and augmented reality (AR). The decoding rate in such applications under real-life complex scenes is greatly affected by two major factors: first, we cannot accurately localize the barcodes, and second, we cannot decode the blur samples. In this article, we first propose a barcode localization algorithm that is capable of regressing four vertices of barcodes accurately. Our localization method comprises an anchor-free approach that outputs multiscale output prediction maps. These segmentation-like maps of each scale are then further divided into three types of maps (classification, centerness, and 8-D regression). Eight-dimensional localization result of barcodes along with classification result is then obtained after postprocessing. Second, we propose a conditional generative adversarial network-based model for deblurring blur QR codes. Extensive decoding experiments on a challenging complex scene data set show that our localization and deblurring methods can contribute to improving the decoding rate of existing barcode decoders. Adnan Sharif, Guangtao Zhai, Xiongkuo Min, Jun Jia, Kashif Munir |
IEEE Internet Things J. | 4 |
| 2021 | Fine localization and distortion resistant detection of multi-class barcode in complex environments
Xiongkuo Min, Jun Jia, Zehao Zhu, Jia Wang 0004, Guangtao Zhai |
Multim. Tools Appl. | 3 |
| 2021 | Localization of Partial Discharge in Electrical Transformer Considering Multimedia Refraction and DiffractionabstractPartial discharge (PD) is a widely adopted method for the internal insulation detection of electrical transformers. The impact of the refraction and diffraction of the ultrasonic signal in the oil, winding, and core in the transformer makes the localization of the PD source inside the transformer a nontrivial task. The refraction and diffraction factors can introduce additional complexity to the localization equations that can hardly be solved. This article proposes an algorithmic solution based on the time difference of arrival algorithm and semidefinite relaxation convex optimization to improve the PD source localization accuracy. The refraction and diffraction error, measurement error, and their relationship in the process of PD signal propagation are fully considered. The proposed solution is extensively assessed through simulation, testbed, and field experiments, and the results confirm that the proposed solution outperforms the Chan and particle swarm optimization algorithm with the localization error of about 0.1 m. Jun Jia, Chengbo Hu, Qiang Yang 0004, Yuncai Lu, Bo Wang 0047, Hengyang Zhao |
IEEE Trans. Ind. Informatics | 1 |
| 2020 | Joint Barcode and Text Orientation Detection Model for Unmanned Retail SystemabstractThe 1D barcode and specification text on the package of retail products contain rich information, such as date of manufacture and source of product. Acquiring this information quickly and accurately can improve the efficiency of unmanned retail system. However, traditional OCR methods are sensitive to text orientations. Based on the fact that 1D barcodes are usually aligned with text, this paper proposes a joint barcode and text orientation detection method. Our approach first determines the four vertices of arbitrarily aligned 1D barcode by a CNN-based barcode localization network. Then, a post processing module calculates the angle of alignment of text with the help of it. Experiments on combined public and self-collected dataset verify that our approach can localize barcode regions accurately and enhance the robustness for OCR in unmanned retail systems. Adnan Sharif, Jun Jia, Guangtao Zhai |
ISCAS | 2 |
| 2020 | Deep Learning for Search and Recommender Systems in PracticeabstractIn this talk, we will go over the components of personalized search and recommender systems and demonstrate the applications of various deep learning techniques along the way. Zhoutong Fu, Huiji Gao, Weiwei Guo, Sandeep Kumar Jha, Jun Jia, Bo Long, Sida Wang 0002, Mingzhou Zhou |
KDD | 5 |
| 2019 | Protection and Hiding Algorithm of QR Code Based on Multi-channel Visual MaskingabstractQuick Response (QR) Code is extensively used due to its advantages in fast readability as well as large capacity. However, when using QR codes, users may suffer from loss of private information because of peeping and scanning by unauthorized attackers. In this paper, we propose a novel algorithm to hide and protect the QR code based on multichannel visual masking. With our algorithm, the appearance of the QR code is dramatically changed while it maintains the original secret information. Unauthorized users can not extract any information from the protected QR code with the standard QR code reader. For authorized users, we design a truth table based decoder that works with the standard QR code reader. Extensive experiments are performed to evaluate the robustness and effectiveness of our method. The codes of this paper are published at https://github.com/JiaheZhang/Protection- QR-Code-MVM. Jun Jia, Guangtao Zhai |
VCIP | 3 |
| 2019 | EMBDN: An Efficient Multiclass Barcode Detection Network for Complicated EnvironmentsabstractThis article presents a novel method for efficient barcodes detection in real and complicated environments using a convolutional neural network (CNN)-based model. The method is developed as a preprocess-module of existing decoders to enhance decoding rates. Our method is trained as an end-to-end model to determine accurate locations of four barcode vertexes. Our method consists of four modules: 1) base net module; 2) region proposals generator; 3) classification and regression module; and 4) distortion removal module. The feature of barcodes extracted from the base net is fed to the next module. Region proposals are generated and selected as region of interest (ROI). Then the ROI are forward propagated to the classification and regression module to determine the positions and shapes of the barcodes. Finally, the distortion removal module is used to remove the geometric distortion according to regression parameters acquired from the previous step. The accurate position and distorted barcodes shape can be determined and corrected by our method. We validate our method on a challenging large-scale dataset in experiments. Compared with the previous methods, our method provides an end-to-end solution to determine accurate locations of barcode vertexes, which shows an excellent performance on detection accuracy. In addition, our method can enhance decoding rate through distortion removal. Jun Jia, Guangtao Zhai, Zhongpai Gao, Zehao Zhu, Xiongkuo Min, Xiaokang Yang 0001, Guodong Guo |
IEEE Internet Things J. | 1 |
| 2012 | A blink restoration system with contralateral EMG triggered stimulation and real-time software based artifact blankingabstractPatients suffering from facial paralysis are on the hazard of disfigurement and loss of vision due to loss of blink function. Functional-electrical stimulation (FES) is one possible way of restoring blink and other functions in these patients. A blink restoration system for uni-lateral facial paralyzed patients is described in this paper. The system achieves restoration of synchronized blink through processing on EMG signal from eyelid of healthy side in real-time and stimulating the paralyzed eyelid. Design issues are discussed, including real-time artifact blanking, stimulating strategies and EMG processing. An artifact removal algorithm based on software sample and hold technique is proposed. Finally, the whole system has been verified on rabbits. Jun Jia, Mengde Wang, Guoxing Wang, Simin Deng, Guofang Shen |
ISCAS | 1 |
| 2011 | A grey measurement of business synergyabstractMeasurement of business synergy is one of the major directions in research on business strategy. This paper first reviews the development of researches on measurement of business synergy. It then uses grey system methodology to develop an econometrics model for measuring the degree of synergy between different businesses in a company, based on synergy theory. It divided a business system into subsystems such as resources system and output system, and then analyzed synergies for each subsystem. Based on that, a comprehensive model for measuring synergy for the whole system is developed. Finally, the paper used the model to analyze a case of a listed company and verify its application possibility. Jun Jia |
SMC | 2 |