EDBT 2026 Demo / reviewers in the wild / expert
Junle Wang
dblp:89/8371
· DBLP profile ↗
37ranked-venue papers
2as first author
25since 2021 · last 2025
0000-0002-9096-2670ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 32 · 2 first-author · 22 since 2021Artificial intelligence and machine learning · 13 · 11 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GRACE: A Strategic LLM-Enhanced Graph Reinforcement Learning Framework for Adaptive Fault Recovery in Microservice Systems
Ruibo Chen 0001, Yanjun Pu, Ji Xin, Junle Wang, Xingchuang Liao, Wenjun Wu 0001 |
ICSOC (1) | 4 |
| 2025 | CalibCLIP: Contextual Calibration of Dominant Semantics for Text-Driven Image RetrievalabstractExisting Visual Language Models (VLMs) suffer structural limitations where a few low contribution tokens may excessively capture global semantics, dominating the information aggregation process and suppressing the discriminative features in text-driven image retrieval tasks. To address this, we introduce \textbf{CalibCLIP}, a training-free method designed to calibrate the suppressive effect of dominant tokens. Specifically, in the visual space, we propose the Contrastive Visual Enhancer (CVE), which decouples visual features into target and low information regions. Subsequently, it identifies dominant tokens and dynamically suppresses their representations.In the textual space, we introduce the Discriminative Concept Calibrator (DCC), which aims to differentiate between general and discriminative concepts within the text query. By mitigating the challenges posed by generic concepts and improving the representations of discriminative concepts, DCC strengthens the differentiation among similar samples. Finally, extensive experiments demonstrate consistent improvements across seven benchmarks spanning three image retrieval tasks, underscoring the effectiveness of CalibCLIP. Code is available at: https://github.com/kangbin98/CalibCLIP Bin Kang, Bin Chen 0022, Yulin Li 0003, Junzhi Zhao, Junle Wang, Zhuotao Tian |
ACM Multimedia | 6 |
| 2025 | Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for ReasoningabstractProcess reward model (PRM) has been proven effective in test-time scaling of LLM on challenging reasoning tasks. However, the reward hacking induced by PRM hinders its successful applications in reinforcement fine-tuning. We find the primary cause of reward hacking induced by PRM is that: the canonical summation-form credit assignment in reinforcement learning (RL), i.e. cumulative gamma-decayed future rewards, causes the LLM to hack steps with high rewards. Therefore, to unleashing the power of PRM in training-time, we propose PURE: Process sUpervised Reinforcement lEarning. The core of PURE is the min-form credit assignment that defines the value function as the minimum future rewards. This method unifies the optimization objective with respect to process rewards during test-time and training-time, and significantly alleviates reward hacking due to the limits on the range of values of value function and more rational assignment of advantages. Through extensively experiments on 3 base models, we achieve similar reasoning performance using PRM-based approach compared with verifiable reward-based approach if enabling min-form credit assignment. In contrast, the canonical sum-form credit assignment even collapses training at the beginning. Moreover, when we incorporate 1/10th verifiable rewards to auxiliary the PRM-based fine-tuning, it further alleviate reward hacking and results in the best fine-tuned model based on Qwen2.5-Math-7B with 82.5% accuracy on AMC23 and 53.3% average accuracy across 5 benchmarks. Furthermore, we summary the reward hacking cases we encountered during training and analysis the cause of training collapse. Jie Cheng 0009, Gang Xiong 0001, Ruixi Qiao, Chao Guo 0006, Junle Wang, Fei-Yue Wang 0001 |
NeurIPS | 6 |
| 2024 | Attacking Transformers with Feature Diversity Adversarial PerturbationabstractUnderstanding the mechanisms behind Vision Transformer (ViT), particularly its vulnerability to adversarial perturbations, is crucial for addressing challenges in its real-world applications. Existing ViT adversarial attackers rely on labels to calculate the gradient for perturbation, and exhibit low transferability to other structures and tasks. In this paper, we present a label-free white-box attack approach for ViT-based models that exhibits strong transferability to various black-box models, including most ViT variants, CNNs, and MLPs, even for models developed for other modalities. Our inspiration comes from the feature collapse phenomenon in ViTs, where the critical attention mechanism overly depends on the low-frequency component of features, causing the features in middle-to-end layers to become increasingly similar and eventually collapse. We propose the feature diversity attacker to naturally accelerate this process and achieve remarkable performance and transferability. Chenxing Gao, Hang Zhou 0010, Junqing Yu, Yuteng Ye, Jiale Cai, Junle Wang, Wei Yang 0034 |
AAAI | 6 |
| 2024 | Dynamic Feature Pruning and Consolidation for Occluded Person Re-identificationabstractOccluded person re-identification (ReID) is a challenging problem due to contamination from occluders. Existing approaches address the issue with prior knowledge cues, such as human body key points and semantic segmentations, which easily fail in the presence of heavy occlusion and other humans as occluders. In this paper, we propose a feature pruning and consolidation (FPC) framework to circumvent explicit human structure parsing. The framework mainly consists of a sparse encoder, a multi-view feature mathcing module, and a feature consolidation decoder. Specifically, the sparse encoder drops less important image tokens, mostly related to background noise and occluders, solely based on correlation within the class token attention. Subsequently, the matching stage relies on the preserved tokens produced by the sparse encoder to identify k-nearest neighbors in the gallery by measuring the image and patch-level combined similarity. Finally, we use the feature consolidation module to compensate pruned features using identified neighbors for recovering essential information while disregarding disturbance from noise and occlusion. Experimental results demonstrate the effectiveness of our proposed framework on occluded, partial, and holistic Re-ID datasets. In particular, our method outperforms state-of-the-art results by at least 8.6% mAP and 6.0% Rank-1 accuracy on the challenging Occluded-Duke dataset. Yuteng Ye, Hang Zhou 0010, Jiale Cai, Chenxing Gao, Youjia Zhang, Junle Wang, Qiang Hu 0003, Junqing Yu, Wei Yang 0034 |
AAAI | 6 |
| 2024 | FreeMan: Towards Benchmarking 3D Human Pose Estimation Under Real-World ConditionsabstractEstimating the 3D structure of the human body from nat-ural scenes is afundamental aspect of visual perception. 3D human pose estimation is a vital step in advancing fields like AIGC and human-robot interaction, serving as a crucial tech-nique for understanding and interacting with human actions in real-world settings. However, the current datasets, often collected under single laboratory conditions using complex motion capture equipment and unvarying backgrounds, are insufficient. The absence of datasets on variable conditions is stalling the progress of this crucial task. To facilitate the development of 3D pose estimation, we present FreeMan, the first large-scale, multi-view dataset collected under the real-world conditions. FreeMan was captured by synchronizing 8 smartphones across diverse scenarios. It comprises 11M frames from 8000 sequences, viewed from different perspec-tives. These sequences cover 40 subjects across 10 different scenarios, each with varying lighting conditions. We have also established an semi-automated pipeline containing er-ror detection to reduce the workload of manual check and ensure precise annotation. We provide comprehensive eval-uation baselines for a range of tasks, underlining the sig-nificant challenges posed by FreeMan. Further evaluations of standard indoor/outdoor human sensing datasets reveal that FreeMan offers robust representation transferability in real and complex scenes. FreeMan is publicly available at https://wangjiongw.github.io/freeman. Fengyu Yang 0005, Bingliang Li, Wenbo Gou, Danqi Yan 0001, Ailing Zeng, Yijun Gao, Junle Wang, Yanqing Jing, Ruimao Zhang |
CVPR | 8 |
| 2024 | Narrative Action Evaluation with Prompt-Guided Multimodal InteractionabstractIn this paper, we investigate a new problem called narrative action evaluation (NAE). NAE aims to generate professional commentary that evaluates the execution of an action. Unlike traditional tasks such as score-based action qual-ity assessment and video captioning involving superficial sentences, NAE focuses on creating detailed narratives in natural language. These narratives provide intricate descriptions of actions along with objective evaluations. NAE is a more challenging task because it requires both narrative flex-ibility and evaluation rigor. One existing possible solution is to use multi-task learning, where narrative language and evaluative information are predicted separately. However, this approach results in reduced performance for individual tasks because of variations between tasks and differences in modality between language information and evaluation information. To address this, we propose a prompt-guided multimodal interaction framework. This framework utilizes a pair of transformers to facilitate the interaction between different modalities of information. It also uses prompts to transform the score regression task into a video-text matching task, thus enabling task interactivity. To support further research in this field, we re-annotate the MTL-AQA and FineGym datasets with high-quality and comprehensive action narration. Additionally, we establish benchmarks for NAE. Extensive experiment results prove that our method outperforms separate learning methods and naive multi-task learning methods. Data and code are released at here. Sule Bai, Guangyi Chen 0002, Lei Chen 0069, Jiwen Lu, Junle Wang, Yansong Tang |
CVPR | 6 |
| 2023 | REC-MV: REconstructing 3D Dynamic Cloth from Monocular VideosabstractReconstructing dynamic 3D garment surfaces with open boundaries from monocular videos is an important problem as it provides a practical and low-cost solution for clothes digitization. Recent neural rendering methods achieve high-quality dynamic clothed human reconstruction results from monocular video, but these methods cannot separate the garment surface from the body. Moreover, despite existing garment reconstruction methods based on feature curve representation demonstrating impressive results for garment reconstruction from a single image, they struggle to generate temporally consistent surfaces for the video input. To address the above limitations, in this paper, we formulate this task as an optimization problem of 3D garment feature curves and surface reconstruction from monocular video. We introduce a novel approach, called REC-MV, to jointly optimize the explicit feature curves and the implicit signed distance field (SDF) of the garments. Then the open garment meshes can be extracted via garment template registration in the canonical space. Experiments on multiple casually captured datasets show that our approach outperforms existing methods and can produce high-quality dynamic garment surfaces. The source code is available at https://github.com/GAP-LAB-CUHK-SZ/REC-MV. Lingteng Qiu, Guanying Chen, Jiapeng Zhou, Mutian Xu, Junle Wang, Xiaoguang Han 0001 |
CVPR | 5 |
| 2023 | Semantic Human Parsing via Scalable Semantic Transfer Over Multiple Label DomainsabstractThis paper presents Scalable Semantic Transfer (SST), a novel training paradigm, to explore how to leverage the mutual benefits of the data from different label domains (i.e. various levels of label granularity) to train a powerful human parsing network. In practice, two common application scenarios are addressed, termed universal parsing and dedicated parsing, where the former aims to learn homogeneous human representations from multiple label domains and switch predictions by only using different segmentation heads, and the latter aims to learn a specific domain prediction while distilling the semantic knowledge from other domains. The proposed SST has the following appealing benefits: (1) it can capably serve as an effective training scheme to embed semantic associations of human body parts from multiple label domains into the human representation learning process; (2) it is an extensible semantic transfer framework without predetermining the overall relations of multiple label domains, which allows continuously adding human parsing datasets to promote the training. (3) the relevant modules are only used for auxiliary training and can be removed during inference, eliminating the extra reasoning cost. Experimental results demonstrate SST can effectively achieve promising universal human parsing performance as well as impressive improvements compared to its counterparts on three human parsing benchmarks (i.e., PASCAL-Person-Part, ATR, and CIHP). Code is available at https://github.com/yangjie-cv/SST. Chaoqun Wang 0012, Zhen Li 0026, Junle Wang, Ruimao Zhang |
CVPR | 4 |
| 2023 | C2F2NeUS: Cascade Cost Frustum Fusion for High Fidelity and Generalizable Neural Surface ReconstructionabstractThere is an emerging effort to combine the two popular 3D frameworks using Multi-View Stereo (MVS) and Neural Implicit Surfaces (NIS) with a specific focus on the few-shot / sparse view setting. In this paper, we introduce a novel integration scheme that combines the multi-view stereo with neural signed distance function representations, which potentially overcomes the limitations of both methods. MVS uses per-view depth estimation and cross-view fusion to generate accurate surfaces, while NIS relies on a common coordinate volume. Based on this strategy, we propose to construct per-view cost frustum for finer geometry estimation, and then fuse cross-view frustums and estimate the implicit signed distance functions to tackle artifacts that are due to noise and holes in the produced surface reconstruction. We further apply a cascade frustum fusion strategy to effectively captures global-local information and structural consistency. Finally, we apply cascade sampling and a pseudo-geometric loss to foster stronger integration between the two architectures. Extensive experiments demonstrate that our method reconstructs robust surfaces and outperforms existing state-of-the-art methods. Luoyuan Xu, Yuesong Wang 0001, Zhaojie Zeng, Junle Wang, Wei Yang 0011 |
ICCV | 6 |
| 2023 | NeMF: Inverse Volume Rendering with Neural Microflake FieldabstractRecovering the physical attributes of an object’s appearance from its images captured under an unknown illumination is challenging yet essential for photo-realistic rendering. Recent approaches adopt the emerging implicit scene representations and have shown impressive results. However, they unanimously adopt a surface-based representation, and hence can not well handle scenes with very complex geometry, translucent object and etc. In this paper, we propose to conduct inverse volume rendering, in contrast to surface-based, by representing a scene using microflake volume, which assumes the space is filled with infinite small flakes and light reflects or scatters at each spatial location according to microflake distributions. We further adopt the coordinate networks to implicitly encode the microflake volume, and develop a differentiable microflake volume renderer to train the network in an end-to-end way in principle. Our NeMF enables effective recovery of appearance attributes for highly complex geometry and scattering object, enables high-quality relighting, material editing, and especially simulates volume rendering effects, such as scattering, which is infeasible for surface-based approaches. Our data and code are available at: https://github.com/YoujiaZhang/NeMF. Youjia Zhang, Teng Xu 0008, Junqing Yu, Yuteng Ye, Yanqing Jing, Junle Wang, Jingyi Yu 0001, Wei Yang 0034 |
ICCV | 6 |
| 2023 | NPF-200: A Multi-Modal Eye Fixation Dataset and Method for Non-Photorealistic VideosabstractNon-photorealistic videos are in demand with the wave of the metaverse, but lack of sufficient research studies. This work aims to take a step forward to understand how humans perceive non-photorealistic videos with eye fixation (i.e., saliency detection), which is critical for enhancing media production, artistic design, and game user experience. To fill in the gap of missing a suitable dataset for this research line, we present NPF-200, the first large-scale multi-modal dataset of purely non-photorealistic videos with eye fixations. Our dataset has three characteristics: 1) it contains soundtracks that are essential according to vision and psychological studies; 2) it includes diverse semantic content and videos are of high-quality; 3) it has rich motions across and within videos. We conduct a series of analyses to gain deeper insights into this task and compare several state-of-the-art methods to explore the gap between natural images and non-photorealistic data. Additionally, as the human attention system tends to extract visual and audio features with different frequencies, we propose a universal frequency-aware multi-modal non-photorealistic saliency detection model called NPSNet, demonstrating the state-of-the-art performance of our task. The results uncover strengths and weaknesses of multi-modal network design and multi-domain training, opening up promising directions for future works. Our dataset and code can be found at https://github.com/Yangziyu/NPF200 Sucheng Ren, Zongwei Wu, Nanxuan Zhao, Junle Wang, Harry Qin, Shengfeng He |
ACM Multimedia | 5 |
| 2023 | Dance with You: The Diversity Controllable Dancer Generation via Diffusion ModelsabstractRecently, digital humans for interpersonal interaction in virtual environments have gained significant attention. In this paper, we introduce a novel multi-dancer synthesis task called partner dancer generation, which involves synthesizing virtual human dancers capable of performing dance with users. The task aims to control the pose diversity between the lead dancer and the partner dancer. The core of this task is to ensure the controllable diversity of the generated partner dancer while maintaining temporal coordination with the lead dancer. This scenario varies from earlier research in generating dance motions driven by music, as our emphasis is on automatically designing partner dancer postures according to pre-defined diversity, the pose of lead dancer, as well as the accompanying tunes. To achieve this objective, we propose a three-stage framework called Dance-with-You (DanY). Initially, we employ a 3D Pose Collection stage to collect a wide range of basic dance poses as references for motion generation. Then, we introduce a hyper-parameter that coordinates the similarity between dancers by masking poses to prevent the generation of sequences that are over-diverse or consistent. To avoid the rigidity of movements, we design a Dance Pre-generated stage to pre-generate these masked poses instead of filling them with zeros. After that, a Dance Motion Transfer stage is adopted with leader sequences and music, in which a multi-conditional sampling formula is rewritten to transfer the pre-generated poses into a sequence with a partner style. In practice, to address the lack of multi-person datasets, we introduce AIST-M, a new dataset for partner dancer generation, which is publicly availiable at https://github.com/JJessicaYao/AIST-M-Dataset. Comprehensive evaluations on our AIST-M dataset demonstrate that the proposed DanY can synthesize satisfactory partner dancer results with controllable diversity. Siyue Yao, Mingjie Sun, Bingliang Li, Fengyu Yang 0005, Junle Wang, Ruimao Zhang |
ACM Multimedia | 5 |
| 2023 | Reference-based Screentone Transfer via Pattern Correspondence and RegularizationabstractAbstract Adding screentone to initial line drawings is a crucial step for manga generation, but is a tedious and human‐laborious task. In this work, we propose a novel data‐driven method aiming to transfer the screentone pattern from a reference manga image. This not only ensures the quality, but also adds controllability to the generated manga results. The reference‐based screentone translation task imposes several unique challenges. Since manga image often contains multiple screentone patterns interweaved with line drawing, as an abstract art, this makes it even more difficult to extract disentangled style code from the reference. Also, finding correspondence for mapping between the reference and the input line drawing without any screentone is hard. As screentone contains many subtle details, how to guarantee the style consistency to the reference remains challenging. To suit our purpose and resolve the above difficulties, we propose a novel Reference‐based Screentone Transfer Network (RSTN). We encode the screentone style through a 1D stylegram. A patch correspondence loss is designed to build a similarity mapping function for guiding the translation. To mitigate the generated artefacts, a pattern regularization loss is introduced in the patch‐level. Through extensive experiments and a user study, we have demonstrated the effectiveness of our proposed model. Zhansheng Li, Nanxuan Zhao, Zongwei Wu, Yihua Dai, Junle Wang, Yanqing Jing, Shengfeng He |
Comput. Graph. Forum | 5 |
| 2022 | High-resolution Face Swapping via Latent Semantics DisentanglementabstractWe present a novel high-resolution face swapping method using the inherent prior knowledge of a pre-trained GAN model. Although previous research can leverage generative priors to produce high-resolution results, their quality can suffer from the entangled semantics of the latent space. We explicitly disentangle the latent semantics by utilizing the progressive nature of the generator, deriving structure at-tributes from the shallow layers and appearance attributes from the deeper ones. Identity and pose information within the structure attributes are further separated by introducing a landmark-driven structure transfer latent direction. The disentangled latent code produces rich generative features that incorporate feature blending to produce a plausible swapping result. We further extend our method to video face swapping by enforcing two spatio-temporal constraints on the latent space and the image space. Extensive experiments demonstrate that the proposed method outperforms state-of-the-art image/video face swapping methods in terms of hallucination quality and consistency. Code can be found at: https://github.com/cnnlstm/FSLSD_HiRes. Yangyang Xu 0003, Bailin Deng, Junle Wang, Yanqing Jing, Jia Pan 0001, Shengfeng He |
CVPR | 3 |
| 2022 | Considering User Agreement in Learning to Predict the Aesthetic QualityabstractHow to robustly rank the aesthetic quality of given images has been a long-standing ill-posed topic. Such challenge stems mainly from the diverse subjective opinions of different observers about the varied types of content. There is a growing interest in estimating the user agreement by considering the standard deviation (σ) of the scores, instead of only predicting the mean aesthetic opinion score (µ). Nevertheless, when comparing a pair of contents, few studies consider how confident are we regarding the difference in the aesthetic scores. In this paper, we thus propose (1) a re-adapted multi-task attention network to predict both the mean opinion score and the standard deviation in an end-to-end manner; (2) a brand-new confidence interval ranking loss that encourages the model to focus on image-pairs that are less certain about the difference of their aesthetic scores. With such loss, the model is encouraged to learn the uncertainty of the content that is relevant to the diversity of observers’ opinions, i.e., user disagreement. Extensive experiments have demonstrated that the proposed multi-task aesthetic model achieves state-of-the-art performance on two different types of aesthetic datasets, i.e., AVA and TMGA. Suiyi Ling, Andreas Pastor, Junle Wang, Patrick Le Callet |
ICASSP | 3 |
| 2022 | Subjective And Objective Quality Assessment Of Mobile Gaming VideoabstractNowadays, with the vigorous expansion and development of gaming video streaming techniques and services, the expectation of users, especially the mobile phone users, for higher quality of experience is also growing swiftly. As most of the existing research focuses on traditional video streaming, there is a clear lack of both subjective study and objective quality models that are tailored for quality assessment of mobile gaming content. To this end, in this study, we first present a brand new Tencent Gaming Video dataset containing 1293 mobile gaming sequences encoded with three different codecs. Second, we propose an objective quality framework, namely Efficient hard-RAnk Quality Estimator (ERAQUE), that is equipped with (1) a novel hard pairwise ranking loss, which forces the model to put more emphasis on differentiating similar pairs; (2) an adapted model distillation strategy, which could be utilized to compress the proposed model efficiently without causing significant performance drop. Extensive experiments demonstrate the efficiency and robustness of our model. Shaoguo Wen, Suiyi Ling, Junle Wang, Yanqing Jing, Patrick Le Callet |
ICASSP | 3 |
| 2022 | QoEVMA'22: 2nd Workshop on Quality of Experience (QoE) in Visual Multimedia ApplicationsabstractNowadays, people spend dramatically more time on watching videos through different devices. The advanced hardware technology and network allow for the increasing demands of users viewing experience. Thus, enhancing the Quality of Experience of end-users in advanced multimedia is the ultimate goal of service providers, as good services would attract more consumers. Quality assessment is thus important. The second workshop on "Quality of Experience (QoE) in visual multimedia applications" (QoEVMA'22) focuses on the QoE assessment of any visual multimedia applications both subjectively and objectively. The topics include 1) QoE assessment on different visual multimedia applications, including VoD for movies, dramas, variety shows, UGC on social networks, live streaming videos for gaming/shopping/social, etc. 2) QoE assessment for different video formats in multimedia services, including 2D, stereoscopic 3D, High Dynamic Range (HDR), Augmented Reality (AR), Virtual Reality (VR), 360, Free-Viewpoint Video(FVV), etc. 3) Key performance indicators (KPI) analysis for QoE. This summary gives a brief overview of the workshop, which took place on October 14, 2022 in Lisbon, Portugal, as a half-day workshop. The complete QOEVMA'22 workshop proceedings are available at: https://dl.acm.org/doi/proceedings/10.1145/3552469 Jing Li 0026, Patrick Le Callet, Xinbo Gao 0001, Zhi Li 0001, Wen Lu 0004, Junle Wang |
ACM Multimedia | 7 |
| 2022 | Unsupervised knowledge transfer for nonblind image deconvolution
Zhuojie Chen, Yong Xu 0007, Junle Wang, Yuhui Quan |
Pattern Recognit. Lett. | 4 |
| 2022 | Cross-Collaborative Fusion-Encoder Network for Robust RGB-Thermal Salient Object DetectionabstractWith the prevalence of thermal cameras, RGB-T multi-modal data have become more available for salient object detection (SOD) in complex scenes. Most RGB-T SOD works first individually extract RGB and thermal features from two separate encoders and directly integrate them, which pay less attention to the issue of defective modalities. However, such an indiscriminate feature extraction strategy may produce contaminated features and thus lead to poor SOD performance. To address this issue, we propose a novel CCFENet for a perspective to perform robust and accurate multi-modal expression encoding. First, we propose an essential cross-collaboration enhancement strategy (CCE), which concentrates on facilitating the interactions across the encoders and encouraging different modalities to complement each other during encoding. Such a cross-collaborative-encoder paradigm induces our network to collaboratively suppress the negative feature responses of defective modality data and effectively exploit modality-informative features. Moreover, as the network goes deeper, we embed several CCEs into the encoder, further enabling more representative and robust feature generation. Second, benefiting from the proposed robust encoding paradigm, a simple yet effective cross-scale cross-modal decoder (CCD) is designed to aggregate multi-level complementary multi-modal features, and thus encourages efficient and accurate RGB-T SOD. Extensive experiments reveal that our CCFENet outperforms the state-of-the-art models on three RGB-T datasets with a fast inference speed of 62 FPS. In addition, the advantages of our approach in complex scenarios (e.g., bad weather, motion blur, etc.) and RGB-D SOD further verify its robustness and generality. The source code will be publicly available via our project page:https://git.openi.org.cn/OpenVision/CCFENet. Guibiao Liao, Wei Gao 0003, Ge Li 0002, Junle Wang, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Make Your Own Sprites: Aliasing-Aware and Cell-Controllable PixelizationabstractPixel art is a unique art style with the appearance of low resolution images. In this paper, we propose a data-driven pixelization method that can produce sharp and crisp cell effects with controllable cell sizes. Our approach overcomes the limitation of existing learning-based methods in cell size control by introducing a reference pixel art to explicitly regularize the cell structure. In particular, the cell structure features of the reference pixel art are used as an auxiliary input for the pixelization process, and for measuring the style similarity between the generated result and the reference pixel art. Furthermore, we disentangle the pixelization process into specific cell-aware and aliasing-aware stages, mitigating the ambiguities in joint learning of cell size, aliasing effect, and color assignment. To train our model, we construct a dedicated pixel art dataset and augment it with different cell sizes and different degrees of anti-aliasing effects. Extensive experiments demonstrate its superior performance over state-of-the-arts in terms of cell sharpness and perceptual expressiveness. We also show promising results of video game pixelization for the first time. Code and dataset are available at https://github.com/WuZongWei6/Pixelization. Zongwei Wu, Liangyu Chai, Nanxuan Zhao, Bailin Deng, Yongtuo Liu, Junle Wang, Shengfeng He |
ACM Trans. Graph. | 7 |
| 2021 | No-Reference Deep Quality Assessment of Compressed Light Field ImagesabstractUnlike traditional 2D image quality assessment, the structural relationship among sub-aperture images (SAIs) is an essential factor affecting the quality evaluation of light field (LF) images, where the labeled datasets are also not sufficient for improving learning performances. To solve these problems, we present a novel deep neural network-based approach to accurately predict the quality of compressed LF images without pristine images. Two modules dubbed SAI-Fusion and Global Context Perception (GCP) are proposed to obtain the relationship among SAIs. For effective training, we compress LF images from EPFL and HCI datasets and propose a ranking-based method to generate pseudo-labels as equivalents of Mean Opinion Score (MOS), i.e., Ranking-MOS. Therefore, we can pre-train our quality assessment network on compressed LF images with Ranking-MOS, and then fine-tune the model at small-scale datasets with real labels. Experiments demonstrate that the proposed method achieves state-of-the-art performance on compressed LF images of Win5-LID dataset. Zixuan Guo 0002, Wei Gao 0003, Haiqiang Wang, Junle Wang, Songlin Fan |
ICME | 4 |
| 2021 | Dynamic Computational Resource Allocation for Fast Inter Frame Coding in Video Conferencing ApplicationsabstractIn response to the serious increase in complexity brought by the growing trend of video coding tools, this paper proposes a dynamic computational resource allocation algorithm for conference videos to improve the inter frame coding speed, which is mainly aimed at the new coding unit (CU) partitioning modes in AVS3. In the proposed algorithm, the numbers of extended quad tree (EQT) or binary tree (BT) partitions at largest coding unit (LCU) level are firstly predicted based on the evaluation of texture variation. Then the allowed partition for the current area will be dynamically allocated ac-cording to the prediction. Finally, the relatively large size CUs, which are determined by the proposed threshold, will be restricted from using EQT to further reduce complexity. Experimental results show that the proposed algorithm can effectively reduce the computational complexity while ensuring rate-distortion (R-D) performance, and can even obtain bit rate savings for certain thresholds. Wei Gao 0003, Junle Wang |
ICME | 3 |
| 2021 | Multi-Modal Aesthetic Assessment for Mobile Gaming ImageabstractWith the proliferation of various gaming technology, services, game styles, and platforms, multi-dimensional aesthetic assessment of the gaming contents is becoming more and more important for the gaming industry. Depending on the diverse needs of diversified game players, game designers, graphical developers, etc. in particular conditions, multi-modal aesthetic assessment is required to consider different aesthetic dimensions/perspectives. Since there are different underlying relationships between different aesthetic dimensions, e.g., between the ‘Colorfulness’ and ‘Color Harmony’, it could be advantageous to leverage effective information attached in multiple relevant dimensions. To this end, we solve this problem via multi-task learning. Our inclination is to seek and learn the correlations between different aesthetic relevant dimensions to further boost the generalization performance in predicting all the aesthetic dimensions. Therefore, the ‘bottleneck’ of obtaining good predictions with limited labeled data for one individual dimension could be unplugged by harnessing complementary sources of other dimensions, i.e., augment the training data indirectly by sharing training information across dimensions. According to experimental results, the proposed model outperforms state-of-the-art aesthetic metrics significantly in predicting four gaming aesthetic dimensions. Yejing Xie, Suiyi Ling, Andreas Pastor, Junle Wang, Junyu Dong, Patrick Le Callet |
MMSP | 5 |
| 2021 | Re-Visiting Discriminator for Blind Free-Viewpoint Image Quality AssessmentabstractAccurate measurement of perceptual quality is important for various immersive multimedia, which demand real-time quality control or quality-based bench-marking for relevant algorithms. For instance, virtual views rendering in Free-Viewpoint (FV) navigation scenarios is a typical case that introduces challenging distortions, particularly the ones around dis-occluded regions. Existing quality metrics, most of which are targeting for impairments caused by compression or network condition, fail to quantify such non-uniform structure-related distortions. Moreover, the lack of quality databases for such distortions makes it even more challenging to develop robust quality metrics. In this work, a Generative Adversarial Networks based No-Reference (NR) quality Metric, namely GANs-NRM, is proposed. We first present an approach to create masks mimicking dis-occlusions/textureless regions, which is applicable on large-scale 2D image databases publicly available in the computer vision domain. Using these synthetic data, we then train a GANs-based context renderer with the capability of rendering those masked regions. Since the naturalness of the rendered dis-occluded regions strongly relates to the perceptual quality, we assume that the discriminator of the trained GANs has an intrinsic ability for quality assessment. We thus use the features extracted from the discriminator to learn a Bag-of-Distortion-Word (BDW) codebook. We show that a quality predictor can be then well trained using only a small amount of subjective quality data for the FV views rendering. Moreover, in the proposed framework, the discriminator is also adapted as a distortion-detector to locate possible distorted regions. According to the experimental results, the proposed model outperforms significantly the state-of-the-art quality metrics. The corresponding context renderer also shows appealing visualized results over other rendering algorithms. Suiyi Ling, Jing Li 0026, Zhaohui Che, Wei Zhou 0021, Junle Wang, Patrick Le Callet |
IEEE Trans. Multim. | 5 |
| 2020 | Few-Shot Pill RecognitionabstractPill image recognition is vital for many personal/public health-care applications and should be robust to diverse unconstrained real-world conditions. Most existing pill recognition models are limited in tackling this challenging few-shot learning problem due to the insufficient instances per category. With limited training data, neural network-based models have limitations in discovering most discriminating features, or going deeper. Especially, existing models fail to handle the hard samples taken under less controlled imaging conditions. In this study, a new pill image database, namely CURE, is first developed with more varied imaging conditions and instances for each pill category. Secondly, a W2-net is proposed for better pill segmentation. Thirdly, a Multi-Stream (MS) deep network that captures task-related features along with a novel two-stage training methodology are proposed. Within the proposed framework, a Batch All strategy that considers all the samples is first employed for the sub-streams, and then a Batch Hard strategy that considers only the hard samples mined in the first stage is utilized for the fusion network. By doing so, complex samples that could not be represented by one type of feature could be focused and the model could be forced to exploit other domain-related information more effectively. Experiment results show that the proposed model outperforms state-of-the-art models on both the National Institute of Health (NIH) and our CURE database. Suiyi Ling, Andreas Pastor, Jing Li 0026, Zhaohui Che, Junle Wang, Patrick Le Callet |
CVPR | 5 |
| 2020 | A Probabilistic Graphical Model for Analyzing the Subjective Visual Quality Assessment Data from CrowdsourcingabstractThe swift development of the multimedia technology has raised dramatically the users' expectation on the quality of experience. To obtain the ground-truth perceptual quality for model training, subjective assessment is necessary. Crowdsourcing platform provides us a convenient and feasible way to run large-scale experiments. However, the obtained perceptual quality labels are generally noisy. In this paper, we propose a probabilistic graphical annotation model to infer the underlying ground truth and discovering the annotator's behavior. In the proposed model, the ground truth quality label is considered following a categorical distribution rather than a unique number, i.e., different reliable opinions on the perceptual quality are allowed. In addition, different annotator's behaviors in crowdsourcing are modeled, which allows us to identify the possibility that the annotator makes noisy labels during the test. The proposed model has been tested on both simulated data and real-world data, where it always shows superior performance than the other state-of-the-art models in terms of accuracy and robustness. Jing Li 0026, Suiyi Ling, Junle Wang, Patrick Le Callet |
ACM Multimedia | 3 |
| 2019 | Perceptual Representations of Structural Information in Images: Application to Quality Assessment of Synthesized View in FTV ScenarioabstractAs the immersive multimedia techniques like Free-viewpoint TV (FTV) develop at an astonishing rate, user's demand for high-quality immersive contents increases dramatically. Unlike traditional uniform artifacts, the distortions within immersive contents could be non-uniform structure-related and thus are challenging for commonly used quality metrics. Recent studies have demonstrated that the representation of visual features can be extracted from multiple levels of the hierarchy. Inspired by the hierarchical representation mechanism in the human visual system (HVS), in this paper, we explore to adopt structural representations to quantitatively measure the impact of such structure-related distortion on perceived quality in FTV scenario. More specifically, a bio-inspired full reference image quality metric is proposed based on 1) low-level contour descriptor; 2) mid-level contour category descriptor; and 3) task-oriented non-natural structure descriptor. The experimental results show that the proposed model outperforms significantly the state-of-the-art metrics. Suiyi Ling, Jing Li 0026, Patrick Le Callet, Junle Wang |
ICIP | 4 |
| 2018 | Hybrid-MST: A Hybrid Active Sampling Strategy for Pairwise Preference AggregationabstractIn this paper we present a hybrid active sampling strategy for pairwise preference aggregation, which aims at recovering the underlying rating of the test candidates from sparse and noisy pairwise labeling. Our method employs Bayesian optimization framework and Bradley-Terry model to construct the utility function, then to obtain the Expected Information Gain (EIG) of each pair. For computational efficiency, Gaussian-Hermite quadrature is used for estimation of EIG. In this work, a hybrid active sampling strategy is proposed, either using Global Maximum (GM) EIG sampling or Minimum Spanning Tree (MST) sampling in each trial, which is determined by the test budget. The proposed method has been validated on both simulated and real-world datasets, where it shows higher preference aggregation ability than the state-of-the-art methods. Jing Li 0026, Rafal Mantiuk, Junle Wang, Suiyi Ling, Patrick Le Callet |
NeurIPS | 3 |
| 2016 | Saliency-based stereoscopic image retargeting
Yuming Fang 0001, Junle Wang, Yuan Yuan 0029, Jianjun Lei 0001, Weisi Lin, Patrick Le Callet |
Inf. Sci. | 2 |
| 2015 | Perceptual Relevance Based Image RetargetingabstractThe impact of perceptual relevance information on content aware image retargeting is investigated. We integrated fixation density maps and region-of-interest maps into a contemporary image retargeting algorithm to test the hypothesis that the latter result in superior performance given their object level representation. We performed an experimental study with human participants to evaluate the performance gains relative to benchmark conditions. The experiment revealed that the kind of perceptual relevance information, image content, and retargeting ratio all have a strong impact on the overall performance. Recorded response times provided further insight into the difficulty that people experienced when performing the assessment task. These findings are instrumental for the image retargeting research community to further improve their algorithms by augmenting content awareness with perceptual relevance. Ulrich Engelke, Junle Wang, Peter Marendy |
IEEE Signal Process. Lett. | 2 |
| 2014 | Stereoscopic image retargeting based on 3D saliency detectionabstractIn this paper, we propose a novel stereoscopic image retargeting algorithm based on 3D visual saliency detection. A new 3D visual attention model is designed based on 2D visual feature detection, depth feature detection and the modeling of various viewing bias in stereo vision. A geometrically consistent seam carving technique is adopted for retargeting stereo image pair. Experimental results demonstrated that both the proposed visual attention model and the proposed retargeting method outperform the state-of-the-art studies. Junle Wang, Yuming Fang 0001, Manish Narwaria, Weisi Lin, Patrick Le Callet |
ICASSP | 1 |
| 2014 | Saliency Detection for Stereoscopic ImagesabstractMany saliency detection models for 2D images have been proposed for various multimedia processing applications during the past decades. Currently, the emerging applications of stereoscopic display require new saliency detection models for salient region extraction. Different from saliency detection for 2D images, the depth feature has to be taken into account in saliency detection for stereoscopic images. In this paper, we propose a novel stereoscopic saliency detection framework based on the feature contrast of color, luminance, texture, and depth. Four types of features, namely color, luminance, texture, and depth, are extracted from discrete cosine transform coefficients for feature contrast calculation. A Gaussian model of the spatial distance between image patches is adopted for consideration of local and global contrast calculation. Then, a new fusion method is designed to combine the feature maps to obtain the final saliency map for stereoscopic images. In addition, we adopt the center bias factor and human visual acuity, the important characteristics of the human visual system, to enhance the final saliency map for stereoscopic images. Experimental results on eye tracking databases show the superior performance of the proposed model over other existing methods. Yuming Fang 0001, Junle Wang, Manish Narwaria, Patrick Le Callet, Weisi Lin |
IEEE Trans. Image Process. | 2 |
| 2013 | Saliency detection for stereoscopic imagesabstractSaliency detection techniques have been widely used in various 2D multimedia processing applications. Currently, the emerging applications of stereoscopic display require new saliency detection models for stereoscopic images. Different from saliency detection for 2D images, depth features have to be taken into account in saliency detection for stereoscopic images. In this paper, we propose a new stereoscopic saliency detection framework based on the feature contrast of color, intensity, texture, and depth. Four types of features including color, luminance, texture, and depth are extracted from DC-T coefficients to represent the energy for image patches. A Gaussian model of the spatial distance between image patches is adopted for the consideration of local and global contrast calculation. A new fusion method is designed to combine the feature maps for computing the final saliency map for stereoscopic images. Experimental results on a recent eye tracking database show the superior performance of the proposed method over other existing ones in saliency estimation for 3D images. Yuming Fang 0001, Junle Wang, Manish Narwaria, Patrick Le Callet, Weisi Lin |
VCIP | 2 |
| 2013 | How Does Image Content Affect the Added Value of Visual Attention in Objective Image Quality Assessment?abstractOur previous research has demonstrated that adding natural scene saliency (NSS) obtained from eye-tracking data may improve an objective metric's performance in predicting perceived image quality. In this letter, we further investigate the image content dependency of this improvement. Results show that the variation in saliency between observers highly depends on image content, and that this variation predicts the extent to which a certain image may profit from adding saliency in the objective image quality assessment. Hantao Liu, Ulrich Engelke, Junle Wang, Patrick Le Callet, Ingrid Heynderickx |
IEEE Signal Process. Lett. | 3 |
| 2013 | Comparative Study of Fixation Density MapsabstractFixation density maps (FDM) created from eye tracking experiments are widely used in image processing applications. The FDM are assumed to be reliable ground truths of human visual attention and as such, one expects a high similarity between FDM created in different laboratories. So far, no studies have analyzed the degree of similarity between FDM from independent laboratories and the related impact on the applications. In this paper, we perform a thorough comparison of FDM from three independently conducted eye tracking experiments. We focus on the effect of presentation time and image content and evaluate the impact of the FDM differences on three applications: visual saliency modeling, image quality assessment, and image retargeting. It is shown that the FDM are very similar and that their impact on the applications is low. The individual experiment comparisons, however, are found to be significantly different, showing that inter-laboratory differences strongly depend on the experimental conditions of the laboratories. The FDM are publicly available to the research community. Ulrich Engelke, Hantao Liu, Junle Wang, Patrick Le Callet, Ingrid Heynderickx, Hans-Jürgen Zepernick, Anthony J. Maeder |
IEEE Trans. Image Process. | 3 |
| 2013 | Computational Model of Stereoscopic 3D Visual SaliencyabstractMany computational models of visual attention performing well in predicting salient areas of 2D images have been proposed in the literature. The emerging applications of stereoscopic 3D display bring an additional depth of information affecting the human viewing behavior, and require extensions of the efforts made in 2D visual modeling. In this paper, we propose a new computational model of visual attention for stereoscopic 3D still images. Apart from detecting salient areas based on 2D visual features, the proposed model takes depth as an additional visual dimension. The measure of depth saliency is derived from the eye movement data obtained from an eye-tracking experiment using synthetic stimuli. Two different ways of integrating depth information in the modeling of 3D visual attention are then proposed and examined. For the performance evaluation of 3D visual attention models, we have created an eye-tracking database, which contains stereoscopic images of natural content and is publicly available, along with this paper. The proposed model gives a good performance, compared to that of state-of-the-art 2D models on 2D images. The results also suggest that a better performance is obtained when depth information is taken into account through the creation of a depth saliency map, rather than when it is integrated by a weighting method. Junle Wang, Matthieu Perreira Da Silva, Patrick Le Callet, Vincent Ricordel |
IEEE Trans. Image Process. | 1 |