Weixin Li 0001

dblp:18/8653-1 · DBLP profile ↗
← Back
43ranked-venue papers
4as first author
27since 2021 · last 2026
0000-0002-5093-5635ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 33 · 2 first-author · 20 since 2021Artificial intelligence and machine learning · 20 · 2 first-author · 12 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RMAdapter: Reconstruction-based Multi-Modal Adapter for Vision-Language Models
abstract
Pre-trained Vision-Language Models (VLMs), e.g. CLIP, have become essential tools in multimodal transfer learning. However, fine-tuning VLMs in few-shot scenarios poses significant challenges in balancing task-specific adaptation and generalization in the obtained model. Meanwhile, current researches have predominantly focused on prompt-based adaptation methods, leaving adapter-based approaches underexplored and revealing notable performance gaps. To address these challenges, we introduce a novel Reconstruction-based Multimodal Adapter (RMAdapter), which leverages a dual-branch architecture. Unlike conventional single-branch adapters, RMAdapter consists of: (1) an adaptation branch that injects task-specific knowledge through parameter-efficient fine-tuning, and (2) a reconstruction branch that preserves general knowledge by reconstructing latent space features back into the original feature space. This design facilitates a dynamic balance between general and task-specific knowledge. Importantly, although RMAdapter introduces an additional reconstruction branch, it is carefully optimized to remain lightweight. By computing reconstruction loss locally at each layer and sharing projection modules, the overall computational overhead is kept minimal. A consistency constraint is also incorporated to better regulate the trade-off between discriminability and generalization. We comprehensively evaluate the effectiveness of RMAdapter on three representative tasks: generalization to new categories, generalization to new target datasets, and domain generalization. Without relying on data augmentation or duplicate prompt designs, our RMAdapter consistently outperforms state-of-the-art approaches across all evaluation metrics.
Weixin Li 0001, Di Huang 0001
AAAI2
2026 Temporal Group Constrained Transformer With Deformable Landmark Attention for Video Dimensional Emotion Recognition
abstract
Video dimensional emotion recognition aims to map human affect into the dimensional emotion space based on visual signals. Recent works notice that it is beneficial to locate key facial regions related to human emotion perception, as well as establish long-term temporal dependencies. While preliminary attempts have been made, there still exists much space for further improvements. In this paper, to better exploit key facial regions, we propose the Temporal cue guided Deformable Landmark Spatial (TDLS) transformer which attends to key facial regions in a data-dependent manner. We also propose the temporal cue guided frame representation learning to learn the spatial representation of each frame by considering features of other frames together. To better model temporal dependencies, we propose the Multi-layer Group Constrained Temporal (MGCT) transformer to summarize features of frames to multi-layer groups, perform group-to-group communications, and let group-level features guide the frame-level emotion recognition. We also introduce cross-clip representation learning to generate consistent results across different clips and videos. Extensive experiments are conducted on two benchmark datasets and superior results are achieved by our method compared to state-of-the-art approaches.
Weixin Li 0001, Xiangjing Meng, Linmei Hu, Xuan Dong 0001
IEEE Trans. Affect. Comput.1
2025 Generating Editable Head Avatars with 3D Gaussian GANs
abstract
Generating animatable and editable 3D head avatars is essential for various applications in computer vision and graphics. Traditional 3D-aware generative adversarial networks (GANs), often using implicit fields like Neural Radiance Fields (NeRF), achieve photo-realistic and view-consistent 3D head synthesis. However, these methods face limitations in deformation flexibility and editability, hindering the creation of lifelike and easily modifiable 3D heads. We propose a novel approach that enhances the editability and animation control of 3D head avatars by incorporating 3D Gaussian Splatting (3DGS) as an explicit 3D representation. This method enables easier illumination control and improved editability. Central to our approach is the Editable Gaussian Head (EG-Head) model, which combines a 3D Morphable Model (3DMM) with texture maps, allowing precise expression control and flexible texture editing for accurate animation while preserving identity. To capture complex non-facial geometries like hair, we use an auxiliary set of 3DGS and tri-plane features. Extensive experiments demonstrate that our approach delivers high-quality 3D-aware synthesis with state-of-the-art controllability. Our code and models are available at https://github.com/liguohao96/EGG3D.
Guohao Li 0010, Hongyu Yang 0001, Yifang Men, Di Huang 0001, Weixin Li 0001, Ruijie Yang, Yunhong Wang 0001
ICASSP5
2025 GIP: Gated Interaction Prompt for Parameter Efficient Vision-Language Fine-Tuning
abstract
Existing Parameter Efficient Fine-Tuning (PEFT) methods in vision-language (VL) domains, primarily adapted from single-modality approaches, face limitations in modeling cross-modal interactions. These methods often unify visual and textual features without explicit modality-specific processing, or rely on unidirectional interaction, leading to suboptimal task adaptation for pre-trained Vision-Language Models (VLMs). To address this issue, we propose a Gated Interaction Prompt (GIP) module as a plug-and-play adaptation to existing PEFT methods, which effectively enhances the two-way interaction between visual and textual features. Our GIP module integrates learnable prompts alongside visual and textual features into the attention layers of VLMs, serving as a bridge for cross-modal interaction. Furthermore, GIP introduces task-specific gating mechanisms to regulate and adapt the influence of prompts across different tasks, thereby further enhancing model performance. Extensive experiments on four VL tasks demonstrate that our approach can seamlessly integrate with existing methods and achieves significant performance improvements with minimal impact on parameter counts and computational costs. With only a 0.02% increase in trainable parameters, our method achieves performance gains of 0.6%, 0.8%, and 1.2% across four tasks—when applied to VL-PET, VL-Adapter, and LoRA, respectively.
Weixin Li 0001, Di Huang 0001
ICIP2
2025 Modality-Aligned Hierarchical Attention Network for Multi-Modal Popularity Prediction on Social Media
abstract
Social media popularity prediction is essential for content optimization and platform management. Existing approaches often struggle to capture the intricate semantic relationships among heterogeneous content modalities. To solve this problem, we propose a hierarchical attention fusion framework with cross-modal semantic alignment, which integrates text, visual, and user behavior features for enhanced popularity prediction. This design enables the model to adaptively emphasize the most informative features across modalities. We systematically evaluate various regression models and their ensemble strategies on the SMPD dataset, which contains 486,000 social media posts. Experimental results demonstrate that our hierarchical attention fusion consistently outperforms existing fusion methods. These findings highlight the effectiveness of cross-modal semantic alignment and provide valuable insights for advancing multi-modal social media popularity prediction.
Wenzheng Hou, Weixin Li 0001
ACM Multimedia2
2025 Video Demoireing Using Focused-Defocused Dual-Camera System
abstract
Moire patterns, unwanted color artifacts in images and videos, arise from the interference between spatially high-frequency scene contents and the spatial discrete sampling of digital cameras. Existing demoireing methods primarily rely on single-camera image/video processing, which faces two critical challenges: 1) distinguishing moire patterns from visually similar real textures, and 2) preserving tonal consistency and temporal coherence while removing moire artifacts. To address these issues, we propose a dual-camera framework that captures synchronized videos of the same scene: one in focus (retaining high-quality textures but may exhibit moire patterns) and one defocused (with significantly reduced moire patterns but blurred textures). We use the defocused video to help distinguish moire patterns from real texture, so as to guide the demoireing of the focused video. We propose a frame-wise demoireing pipeline, which begins with an optical flow based alignment step to address any discrepancies in displacement and occlusion between the focused and defocused frames. Then, we leverage the aligned defocused frame to guide the demoireing of the focused frame using a multi-scale CNN and a multi-dimensional training loss. To maintain tonal and temporal consistency, our final step involves a joint bilateral filter to leverage the demoireing result from the CNN as the guide to filter the input focused frame to obtain the final output. Experimental results demonstrate that our proposed framework largely outperforms state-of-the-art image and video demoireing methods.
Xuan Dong 0001, Xiangyuan Sun, Ya Li 0001, Weixin Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 Leveraging Predicate and Triplet Learning for Scene Graph Generation
abstract
Scene Graph Generation (SGG) aims to identify entities and predict the relationship tripletsin visual scenes. Given the prevalence of large visual variations of subject-object pairs even in the same predicate, it can be quite challenging to model and refine predicate representations directly across such pairs, which is however a common strategy adopted by most existing SGG methods. We observe that visual variations within the identical triplet are relatively small and certain relation cues are shared in the same type of triplet, which can potentially facilitate the relation learning in SGG. Moreover, for the long-tail problem widely studied in SGG task, it is also crucial to deal with the limited types and quantity of triplets in tail predicates. Accordingly, in this paper, we propose a Dual-granularity Relation Modeling (DRM) network to leverage fine-grained triplet cues besides the coarse-grained predicate ones. DRM utilizes contexts and semantics of predicate and triplet with Dual-granularity Constraints, generating compact and balanced representations from two perspectives to facilitate relation recognition. Furthermore, a Dual-granularity Knowledge Transfer (DKT) strategy is introduced to transfer variation from head predicates/triplets to tail ones, aiming to enrich the pattern diversity of tail classes to alleviate the long-tail problem. Extensive experiments demonstrate the effectiveness of our method, which establishes new state-of-the-art performance on Visual Genome, Open Image, and GQA datasets. Our code is available at https://github.com/jkli1998/DRM
Jiankai Li, Yunhong Wang 0001, Xiefan Guo, Ruijie Yang, Weixin Li 0001
CVPR5
2024 Read, Spell and Repeat: Scene Text Recognition with Vision-Language Circular Refinement
abstract
Scene Text Recognition (STR) has long been considered an important yet challenging task in the field of computer vision. Recent works have demonstrated that utilizing language information is effective for the visually difficult images, like ones with occultation or blurring. However, the use of language information sometimes leads to the over-correction problem. For out-of-vocabulary samples (e.g. "hou" and "0x4a"), some methods have tended to be biased to language side and over-corrected (e.g. over-correct "hou" to "hot"). This imbalance of vision and language has limited the usage of models in practical scenarios, yet it is rarely occurs for human. To address this issue, we rethink the human’s recognition process and propose a model behaving in the order of "Read, Spell and Repeat". It refines the recognition process circularly with vision and language information. With this mechanism, our model integrates vision and language information in a more effective manner, achieving higher accuracy with less parameters compared to baseline and competitive performance with SOTA methods in the standard benchmarks.
Taiwei Zhang, Weixin Li 0001, Qingjie Liu 0001, Yunhong Wang 0001
ICASSP3
2024 ReWiTe: Realistic Wide-angle and Telephoto Dual Camera Fusion Dataset via Beam Splitter Camera Rig
abstract
The fusion of images from dual camera systems featuring a wide-angle and a telephoto camera has become a hotspot problem recently. By integrating simultaneously captured wide-angle and telephoto images from these systems, the resulting fused image achieves a wide field of view (FOV) coupled with high-definition quality. Existing approaches are mostly deep learning methods, and predominantly rely on supervised learning, where the training dataset plays a pivotal role. However, current datasets typically adopt a data synthesis approach, where the wide-angle inputs are synthesized rather than captured using real wide-angle cameras, and the ground-truth image is captured by wide-angle cameras whose quality is substantially lower than that of input telephoto images captured by telephoto cameras. To address these limitations, we introduce a novel hardware setup utilizing a beam splitter to simultaneously capture three images, i.e. input pairs and ground-truth images, from two authentic cellphones equipped with wide-angle and telephoto dual cameras. Specifically, the wide-angle and telephoto images captured by cellphone 2 serve as the input pair, while the telephoto image captured by cellphone 1, which is calibrated to match the optical path of the wide-angle image from cellphone 2, serves as the ground-truth image, maintaining quality on par with the input telephoto image. Experiments validate the efficacy of our newly introduced dataset, named ReWiTe, which can significantly enhance the performance of various existing methods for the real-world wide-angle and telephoto dual image fusion task.
Chunli Peng, Xuan Dong 0001, Zhengqing Li, Weixin Li 0001
ACM Multimedia6
2024 MHRN: A Multimodal Hierarchical Reasoning Network for Topic Detection
abstract
Multimodal topic detection is an important social media analysis task with a wide variety of real-world applications. However, modeling data jointly, and inferring their topics, is challenging due to the semantic gaps between different modalities. Our insights are from the psychological findings pretaining to the hierarchical structure in humans? inherent perception of images and texts. In this paper, we propose a Multimodal Hierarchical Reasoning Network (MHRN) to perform multimodal inference for topic detection. The images and texts are represented in a hierarchical model named the Multimodal Part-whole Aware Graph (MPAG). MHRN then performs reasoning for topic inference based on three modules, which include a Bottom-Up Aggregation (BUA) module for encoding the hierarchical connections and sibling relations in MPAG, a Top-Down Guidance (TDG) module for enriching features of the nodes in MPAG guided by their parents, and a Bottom-Up Cross Aggregation (BUCA) module for capturing and aggregating the cross-modality cues to achieve effective multimodal reasoning. Extensive experiments are conducted on two benchmarks, and the results demonstrate the superiority of our approach.
Jiankai Li, Yunhong Wang 0001, Weixin Li 0001
IEEE Trans. Multim.3
2024 Zero-shot Scene Graph Generation via Triplet Calibration and Reduction
abstract
Scene Graph Generation (SGG) plays a pivotal role in downstream vision-language tasks. Existing SGG methods typically suffer from poor compositional generalizations on unseen triplets. They are generally trained on incompletely annotated scene graphs that contain dominant triplets and tend to bias toward these seen triplets during inference. To address this issue, we propose a Triplet Calibration and Reduction (T-CAR) framework in this article. In our framework, a triplet calibration loss is first presented to regularize the representations of diverse triplets and to simultaneously excavate the unseen triplets in incompletely annotated training scene graphs. Moreover, the unseen space of scene graphs is usually several times larger than the seen space, since it contains a huge number of unrealistic compositions. Thus, we propose an unseen space reduction loss to shift the attention of excavation to reasonable unseen compositions to facilitate the model training. Finally, we propose a contextual encoder to improve the compositional generalizations of unseen triplets by explicitly modeling the relative spatial relations between subjects and objects. Extensive experiments show that our approach achieves consistent improvements for zero-shot SGG over state-of-the-art methods. The code is available at https://github.com/jkli1998/T-CAR .
Jiankai Li, Yunhong Wang 0001, Weixin Li 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2023 Causal Intervention and Counterfactual Reasoning for Multi-modal Fake News Detection
abstract
Due to the rapid upgrade of social platforms, most of today's fake news is published and spread in a multi-modal form.Most existing multi-modal fake news detection methods neglect the fact that some label-specific features learned from the training set cannot generalize well to the testing set, thus inevitably suffering from the harm caused by the latent data bias.In this paper, we analyze and identify the psycholinguistic bias in the text and the bias of inferring news label based on only image features.We mitigate these biases from a causality perspective and propose a Causal intervention and Counterfactual reasoning based Debiasing framework (CCD) for multi-modal fake news detection.To achieve our goal, we first utilize causal intervention to remove the psycholinguistic bias which introduces the spurious correlations between text features and news label.And then, we apply counterfactual reasoning by imagining a counterfactual world where each news has only image features for estimating the direct effect of the image.Therefore we can eliminate the image-only bias by deducting the direct effect of the image from the total effect on labels.Extensive experiments on two real-world benchmark datasets demonstrate the effectiveness of our framework for improving multi-modal fake news detection.
Linmei Hu, Weixin Li 0001, Yingxia Shao, Liqiang Nie
ACL (1)3
2023 Multi-Modal and Multi-Scale Temporal Fusion Architecture Search for Audio-Visual Video Parsing
abstract
The weakly supervised audio-visual video parsing (AVVP) task aims to parse a video into a set of modality-wise events (i.e., audible, visible, or both), recognize categories of these events, and localize their temporal boundaries. Given the prevalence of audio-visual synchronous and asynchronous contents in multi-modal videos, it is crucial to capture and integrate the contextual events occurring at different moments and temporal scales. Although some researchers have made preliminary attempts at modeling event semantics with various temporal lengths, they mostly only perform a late fusion of multi-scale features across modalities. A comprehensive cross-modal and multi-scale temporal fusion strategy remains largely unexplored in the literature. To address this gap, we propose a novel framework named Audio-Visual Fusion Architecture Search (AVFAS) that can automatically find the optimal multi-scale temporal fusion strategy within and between modalities. Our framework generates a set of audio and visual features with distinct temporal scales and employs three modality-wise modules to search multi-scale feature selection and fusion strategies, jointly modeling modality-specific discriminative information. Furthermore, to enhance the alignment of audio-visual asynchrony, we introduce a Position- and Length-Adaptive Temporal Attention (PLATA) mechanism for cross-modal feature fusion. Extensive quantitative and qualitative experimental results demonstrate the effectiveness and efficiency of our framework.
Jiayi Zhang 0012, Weixin Li 0001
ACM Multimedia2
2023 Human Emotion Recognition With Relational Region-Level Analysis
abstract
Recognizing the emotional state of a person within the image in real-world scenarios is a key problem in affective computing and has various promising applications. Local regions in the image, including different objects in the background scene and parts within the foreground body, usually have different contributions to emotion perception of the target person. This, however, has not been well exploited in most existing methods. In this article, we propose to make relational region-level analysis to account for the different contributions of different regions to emotion recognition. For the background scene, we propose a Body-Object Attention (BOA) module to estimate the contributions of background objects to emotion recognition given the target foreground body. Within the foreground body, we propose a Body Part Attention (BPA) module to recalibrate the channel-wise body feature responses to attend on body parts that are more important. Moreover, we propose to model the emotion label dependency in real-world images, considering both the semantic meanings of these labels and their co-occurrence patterns. We evaluate the proposed method on the EMOTIC and CAER-S datasets, and experimental results show the superiority of our method compared with the state-of-the-art algorithms.
Weixin Li 0001, Xuan Dong 0001, Yunhong Wang 0001
IEEE Trans. Affect. Comput.1
2023 Facial Expression Animation by Landmark Guided Residual Module
abstract
We study the problem of facial expression animation from a still image according to a driving video. This is a challenging task as expression motions are non-rigid and very subtle to be captured. Existing methods mostly fail to model these subtle expression motions, leading to the lack of details in their animation results. In this paper, we propose a novel facial expression animation method based on generative adversarial learning. To capture the subtle expression motions, Landmark guided Residual Module (LRM) is proposed to model detailed facial expression features. Specifically, residual learning is conducted at both coarse and fine levels conditioned on facial landmark heatmaps and landmark points respectively. Furthermore, we employ a consistency discriminator to ensure the temporal consistency of the generated video sequence. In addition, a novel metric named Emotion Consistency Metric is proposed to evaluate the consistency of facial expressions in the generated sequences with those in the driving videos. Experiments on MUG-Face, Oulu-CASIA and CAER datasets show that the proposed method can generate arbitrary expression motions on the source still image effectively, which are more photo-realistic and consistent with the driving video compared with results of state-of-the-art methods.
Yunhong Wang 0001, Weixin Li 0001, Zhengyin Du, Di Huang 0001
IEEE Trans. Affect. Comput.3
2023 Dual-Lens HDR using Guided 3D Exposure CNN and Guided Denoising Transformer
abstract
We study the high dynamic range (HDR) imaging problem in dual-lens systems. Existing methods usually treat the HDR imaging problem as an image fusion problem and the HDR result is estimated by fusing the aligned short exposure image and long exposure image. However, the image fusion pipeline depends highly on the image alignment, which is difficult to be perfect. We propose to transfer the dual-lens HDR imaging problem into the disentangled enhancement of exposure correction and denoising for the short exposure image, guided by the long exposure image. In the guided exposure correction module, we make use of the guidance image and 3D color transformation to propose a guided 3D exposure CNN (GEC) to get the rough HDR result from the short exposure image. Then, in the guided denoising module, we make use of the cross-attention mechanism to propose a guided denoising transformer (GDT) to directly use the long exposure image as guidance to denoise the rough HDR result in a pyramid way. And in both modules, we bypass the difficult image alignment processing. Experimental results demonstrate the superiority of our method over the state-of-the-art ones.
Weixin Li 0001, Chang Liu 0071, Xue Tian, Ya Li 0001, Xiaojie Wang 0006, Xuan Dong 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2023 CUR Transformer: A Convolutional Unbiased Regional Transformer for Image Denoising
abstract
Image denoising is a fundamental problem in computer vision and multimedia computation. Non-local filters are effective for image denoising. But existing deep learning methods that use non-local computation structures are mostly designed for high-level tasks, and global self-attention is usually adopted. For the task of image denoising, they have high computational complexity and have a lot of redundant computation of uncorrelated pixels. To solve this problem and combine the marvelous advantages of non-local filter and deep learning, we propose a Convolutional Unbiased Regional (CUR) transformer. Based on the prior that, for each pixel, its similar pixels are usually spatially close, our insights are that (1) we partition the image into non-overlapped windows and perform regional self-attention to reduce the search range of each pixel, and (2) we encourage pixels across different windows to communicate with each other. Based on our insights, the CUR transformer is cascaded by a series of convolutional regional self-attention (CRSA) blocks with U-style short connections. In each CRSA block, we use convolutional layers to extract the query, key, and value features, namely Q , K , and V , of the input feature. Then, we partition the Q , K , and V features into local non-overlapped windows and perform regional self-attention within each window to obtain the output feature of this CRSA block. Among different CRSA blocks, we perform the unbiased window partition by changing the partition positions of the windows. Experimental results show that the CUR transformer outperforms the state-of-the-art methods significantly on four low-level vision tasks, including real and synthetic image denoising, JPEG compression artifact reduction, and low-light image enhancement.
Weixin Li 0001, Xiaoyan Hu 0006, Xiaojie Wang 0006, Xuan Dong 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2022 MGMP: Multimodal Graph Message Propagation Network for Event Detection
Jiankai Li, Yunhong Wang 0001, Weixin Li 0001
MMM (1)3
2022 Spatially Consistent Transformer for Colorization in Monochrome-Color Dual-Lens System
abstract
We study the colorization problem in monochrome-color dual-lens camera systems, i.e. colorizing the gray image from the monochrome camera using the color image from the color camera as reference. In related methods, cost volume based CNN methods achieve the state-of-the-art results, but they are costly in GPU memory due to building the 4D cost volume. Recently, some slice-wise cross-attention based methods are proposed for related problems. The slice-wise cross-attention has much less costs in GPU memory but directly using them for this colorization problem cannot generate competing results. We make use of the non-local computation property of cross-attention to propose a transformer based method. To overcome the limitations of straight-forward slice-wise cross-attention, we propose the spatially consistent cross-attention (SCCA) block to encourage pixels of slices across different epipolar lines in the gray image to find spatially consistent correspondence with pixels of the reference color image. And, to further reduce the memory cost while keeping the colorization accuracy, we design a pyramid processing strategy to cascade a series of SCCA blocks with smaller slice size and perform the colorization from coarse to fine. To extract more powerful image features, we use several regional self-attention (RSA) blocks with U-style connections. Experimental results show that we outperform the state-of-the-art methods largely on the synthesized datasets of Cityscapes, Sintel, and SceneFlow, and the real monochrome-color dual-lens dataset.
Xuan Dong 0001, Chang Liu 0071, Xiaoyan Hu 0006, Weixin Li 0001
IEEE Trans. Image Process.5
2022 A Colorization Framework for Monochrome-Color Dual-Lens Systems Using a Deep Convolutional Network
abstract
In monochrome-color dual-lens systems, the monochrome camera can capture images with higher quality than the color camera. To obtain high quality color images, a better approach is to colorize the gray images from the monochrome camera with the color images from the color camera serving as a reference. In addition, the colorization may fail in some cases, which makes the estimation of the colorization quality a necessary step before outputting the colorization result. To solve these problems, we propose a deep convolutional network based framework. 1) In the colorization module, the proposed colorization CNN uses deep feature representations, attention operation, 3-D regulation and color correction to make use of colors of multiple pixels in the reference image for colorizing each pixel in the input gray image. 2) In the colorization quality estimation module, based on the symmetry property of colorization, we propose to utilize the colorization CNN again to colorize the gray map of the original reference color image using the first-time colorization result from the colorization module as reference. Then, the quality loss of the second-time colorization result can be used for estimating the colorization quality. Experimental results show that our method can largely outperform the state-of-the-art colorization methods and estimate the colorization quality accurately as well.
Xuan Dong 0001, Weixin Li 0001, Xiaoyan Hu 0006, Xiaojie Wang 0006, Yunhong Wang 0001
IEEE Trans. Vis. Comput. Graph.2
2021 MIEHDR CNN: Main Image Enhancement based Ghost-Free High Dynamic Range Imaging using Dual-Lens Systems
abstract
We study the High Dynamic Range (HDR) imaging problem using two Low Dynamic Range (LDR) images that are shot from dual-lens systems in a single shot time with different exposures. In most of the related HDR imaging methods, the problem is usually solved by Multiple Images Merging, i.e. the final HDR image is fused from pixels of all the input LDR images. However, ghost artifacts can be hardly avoided using this strategy. Instead of directly merging the multiple LDR inputs, we use an indirect way which enhances the main image, i.e. the short exposure image IS, using the long exposure image IL serving as guidance. In detail, we propose a new model, named MIEHDR CNN model, which consists of three subnets, i.e. Soft Warp CNN, 3D Guided Denoising CNN and Fusion CNN. The Soft Warp CNN aligns IL to get the aligned result ILA using the soft exposed result of IS as reference. The 3D Guided Denoising CNN denoises the soft exposed result of IS using ILA as guidance, whose result are fed into the Fusion CNN with IS to get the HDR result. The MIEHDR CNN model is implemented by MindSpore and experimental results show that we can outperform related methods largely and avoid ghost artifacts.
Xuan Dong 0001, Xiaoyan Hu 0006, Weixin Li 0001, Xiaojie Wang 0006, Yunhong Wang 0001
AAAI3
2021 Expression-Latent-Space-guided GAN for Facial Expression Animation based on Discrete Labels
abstract
Facial expression animation aims to synthesize face images that correspond to the target expression in a continuum. This is a challenging task because the animation not only cares about the smooth transition in the generated sequence but also needs to take the facial expression and identity details into consideration. Most existing expression animation methods resort to continuous expression labels, e.g. Action Units (AUs) or landmark sequences. Compared with discrete expression labels, the annotations of a considerable part of them are ambiguous and prone to errors. However, how to animate facial expression conditioned on discrete expression labels is less investigated and existing methods cannot generate satisfactory facial details. To tackle these problems, we propose an end-to-end Expression-Latent-Space-guided Generative Adversarial Network (ELS-GAN) model, which utilizes discrete expression labels as input to generate images with expected expressions, and employs expression latent space learning to control the expression changing process. An expression ranking loss is also proposed to strengthen expression intensity learning during generation. Moreover, we put forward a Self-Attention Generator to synthesize face images with fine details by considering both local areas and the long-range dependency of different areas. Extensive experiments show that our method can generate continuous intermediate expression between source and target expressions only conditioned on discrete labels and superior results are achieved compared with state-of-the-art methods.
Weixin Li 0001, Di Huang 0001
FG2
2021 Refining Single Low-Quality Facial Depth Map by Lightweight and Efficient Deep Model
abstract
Consumer depth sensors have become increasingly common, however, the data are rather coarse and noisy, which is problematic to delicate tasks, such as 3D face modeling and 3D face recognition. In this paper, we present a novel and lightweight 3D Face Refinement Model (3D-FRM), to effectively and efficiently improve the quality of such single facial depth maps. 3D-FRM has an encoder-decoder structure, where the encoder applies depth-wise, point-wise convolutions and the fusion of features of different receptive fields to capture original discriminative information, and the decoder exploits sub-pixel convolutions and the combination of low- and high-level features to achieve strong shape recovery. We also propose a joint loss function to smooth facial surfaces and preserve their identities. In addition, we contribute a large dataset with low- and high-quality 3D face pairs to facilitate this research. Extensive experiments are conducted on the Bosphorus and Lock3DFace datasets, and results show the competency of the proposed method at ameliorating both visual quality and recognition accuracy. Code and data will be available at https://github.com/muyouhang/3D-FRM.
Guodong Mu, Di Huang 0001, Weixin Li 0001, Guosheng Hu, Yunhong Wang 0001
IJCB3
2021 Entity Relation Fusion for Real-Time One-Stage Referring Expression Comprehension
abstract
Referring Expression Comprehension (REC) is the task of grounding object which is referred by the language expression. Previous one-stage REC methods usually use one single language feature vector to represent the whole query for grounding and no reasoning between different objects is performed despite the rich relation cues of objects contained in the language expression, which depresses their grounding accuracy. Additionally, these methods mostly use the feature pyramid networks for multi-scale visual object feature extraction but ground on different feature layers separately, neglecting the connections between objects with different scales. To address these problems, we propose a novel one-stage REC method, i.e. the Entity Relation Fusion Network (ERFN) to locate referred object by relation guided reasoning on different objects. In ERFN, instead of grounding objects at each layer separately, we propose a Language Guided Multi-Scale Fusion (LGMSF) model to utilize language to guide the fusion of representations of objects with different scales into one feature map.For modeling connections between different objects, we design a Relation Guided Feature Fusion (RGFF) model that extracts entities in the language expression to enhance the referred entity feature in the visual object feature map, and further extracts relations to guide object feature fusion based on the self-attention mechanism. Experimental results show that our method is competitive with the state-of-the-art one-stage and two-stage REC methods, and can also keep inferring in real time.
Weixin Li 0001, Jiankai Li
MMAsia2
2021 Pyramid convolutional network for colorization in monochrome-color multi-lens camera system
Xuan Dong 0001, Weixin Li 0001, Xiaojie Wang 0006
Neurocomputing2
2021 Spatio-Temporal Encoder-Decoder Fully Convolutional Network for Video-Based Dimensional Emotion Recognition
abstract
Video-based dimensional emotion recognition aims to map human affect into the dimensional emotion space based on visual signals, which is a fundamental challenge in affective computing and human-computer interaction. In this paper, we present a novel encoder-decoder framework to tackle this problem. It adopts a fully convolutional design with the cascaded 2D convolution based spatial encoder and 1D convolution based temporal encoder-decoder for joint spatio-temporal modeling. In particular, to address the key issue of capturing discriminative long-term dynamic dependency, our temporal model, referred to as Temporal Hourglass Convolutional Neural Network (TH-CNN), extracts contextual relationship through integrating both low-level encoded and high-level decoded clues. Temporal Intermediate Supervision (TIS) is then introduced to enhance affective representations generated by TH-CNN under a multi-resolution strategy, which guides TH-CNN to learn macroscopic long-term trend and refined short-term fluctuations progressively. Furthermore, thanks to TH-CNN and TIS, knowledge learnt from the intermediate layers also makes it possible to offer customized solutions to different applications by adjusting the decoder depth. Extensive experiments are conducted on three benchmark databases (RECOLA, SEWA and OMG) and superior results are shown compared to state-of-the-art methods, which indicates the effectiveness of the proposed approach.
Zhengyin Du, Suowei Wu, Di Huang 0001, Weixin Li 0001, Yunhong Wang 0001
IEEE Trans. Affect. Comput.4
2021 Self-Supervised Colorization Towards Monochrome-Color Camera Systems Using Cycle CNN
abstract
Colorization in monochrome-color camera systems aims to colorize the gray image IGfrom the monochrome camera using the color image RCfrom the color camera as reference. Since monochrome cameras have better imaging quality than color cameras, the colorization can help obtain higher quality color images. Related learning based methods usually simulate the monochrome-color camera systems to generate the synthesized data for training, due to the lack of ground-truth color information of the gray image in the real data. However, the methods that are trained relying on the synthesized data may get poor results when colorizing real data, because the synthesized data may deviate from the real data. We present a self-supervised CNN model, named Cycle CNN, which can directly use the real data from monochrome-color camera systems for training. In detail, we use the Weighted Average Colorization (WAC) network to do the colorization twice. First, we colorize IGusing RCas reference to obtain the first-time colorization result IC. Second, we colorize the de-colored map of RC, i.e. RG, using the concatenated image of IGand Cb/Cr channels of the first-time colorization result IC, i.e. ICCband ICCr, as reference to obtain the second-time colorization result RC'. In this way, for the second-time colorization result RC', we use the Cb and Cr channels of the original color map RCas ground-truth and introduce the cycle consistency loss to push RC'Cb/Cr≈ RCCb/Cr. Also, for the Y channel of the first-time colorization result ICY, we propose the Global Curve Adjustment (GCA) network and the structure similarity loss to encourage the structure similarity between ICYand IG. In addition, we introduce a spatial smoothness loss within the WAC network to encourage spatial smoothness of the colorization result. Combining all these losses, we could train the Cycle CNN using the real data in the absence of the ground-truth color information of IG. Experimental results show that we can outperform related methods largely for colorizing real data.
Xuan Dong 0001, Chang Liu 0071, Weixin Li 0001, Xiaoyan Hu 0006, Xiaojie Wang 0006, Yunhong Wang 0001
IEEE Trans. Image Process.3
2020 Cycle-CNN for Colorization towards Real Monochrome-Color Camera Systems
abstract
Colorization in monochrome-color camera systems aims to colorize the gray image IG from the monochrome camera using the color image RC from the color camera as reference. Since monochrome cameras have better imaging quality than color cameras, the colorization can help obtain higher quality color images. Related learning based methods usually simulate the monochrome-color camera systems to generate the synthesized data for training, due to the lack of ground-truth color information of the gray image in the real data. However, the methods that are trained relying on the synthesized data may get poor results when colorizing real data, because the synthesized data may deviate from the real data. We present a new CNN model, named cycle CNN, which can directly use the real data from monochrome-color camera systems for training. In detail, we use the colorization CNN model to do the colorization twice. First, we colorize IG using RC as reference to obtain the first-time colorization result IC. Second, we colorize the de-colored map of RC, i.e. RG, using the first-time colorization result IC as reference to obtain the second-time colorization result R′C. In this way, for the second-time colorization result R′C, we use the original color map RC as ground-truth and introduce the cycle consistency loss to push R′C ≈ RC. Also, for the first-time colorization result IC, we propose a structure similarity loss to encourage the luminance maps between IG and IC to have similar structures. In addition, we introduce a spatial smoothness loss within the colorization CNN model to encourage spatial smoothness of the colorization result. Combining all these losses, we could train the colorization CNN model using the real data in the absence of the ground-truth color information of IG. Experimental results show that we can outperform related methods largely for colorizing real data.
Xuan Dong 0001, Weixin Li 0001, Xiaojie Wang 0006, Yunhong Wang 0001
AAAI2
2020 3D Face Mask Anti-spoofing via Deep Fusion of Dynamic Texture and Shape Clues
abstract
Face anti-spoofing has recently become more important to the wide application of Face Recognition (FR) techniques. Compared to Spoofing Attacks (SAs) of printed photos and replayed videos, 3D masks bring more challenges to FR systems. This paper proposes a novel approach to 3D face mask anti-spoofing, namely Multi-Modal Dynamics Fusion Network (MM-DFN), and different from the overwhelming majority of the methods in the literature that only employ RGB data, it highlights the credit of the geometry information delivered by depth sensors or reconstructed from RGB images. Dynamic texture and shape clues are respectively encoded by a two-branch deep CNN model at different rates so that discriminative details are sufficiently captured, and they are combined at intervals for more comprehensive description. Moreover, a 3D model guided data augmentation method is applied to generate a diversity of samples with various poses, which further enhances the anti-spoofing model. The proposed method is extensively evaluated on three public benchmarks, i.e. 3DMAD, HKBU-MARs V1 and SMAD, and the results achieved are state-of-the-art, demonstrating its effectiveness for this issue.
Weixin Li 0001, Hongyu Yang 0001, Di Huang 0001, Yunhong Wang 0001
FG2
2020 Visual Sentiment Analysis by Leveraging Local Regions and Human Faces
Ruolin Zheng, Weixin Li 0001, Yunhong Wang 0001
MMM (1)2
2019 Learning a Deep Convolutional Network for Colorization in Monochrome-Color Dual-Lens System
abstract
In the monochrome-color dual-lens system, the gray image captured by the monochrome camera has better quality than the color image from the color camera, but does not have color information. To get high-quality color images, it is desired to colorize the gray image with the color image as reference. Related works usually use hand-crafted methods to search for the best-matching pixel in the reference image for each pixel in the input gray image, and copy the color of the best-matching pixel as the result. We propose a novel deep convolution network to solve the colorization problem in an end-to-end way. Based on our observation that, for each pixel in the input image, there usually exist multiple pixels in the reference image that have the correct colors, our method performs weighted average of colors of the candidate pixels in the reference image to utilize more candidate pixels with correct colors. The weight values between pixels in the input image and the reference image are obtained by learning a weight volume using deep feature representations, where an attention operation is proposed to focus on more useful candidate pixels and a 3-D regulation is performed to learn with context information. In addition, to correct wrongly colorized pixels in occlusion regions, we propose a color residue joint learning module to correct the colorization result with the input gray image as guidance. We evaluate our method on the Scene Flow, Cityscapes, Middlebury, and Sintel datasets. Experimental results show that our method largely outperforms the state-of-the-art methods.
Xuan Dong 0001, Weixin Li 0001, Xiaojie Wang 0006, Yunhong Wang 0001
AAAI2
2019 Encoding Visual Behaviors with Attentive Temporal Convolution for Depression Prediction
abstract
Depression is a common and serious medical illness which has a wide negative impact on individuals, families, and society. Automatic Depression Detection (ADD) is increasingly demanded for human healthcare thanks to its objectiveness, convenience, and low cost. Considering that the duration of depressive symptoms varies among different identities and treatment phases, it is essential for ADD methods to have the capability to capture information at various temporal scales. However, most existing ADD methods cannot generate rich contextual cues or utilize long-range temporal dependency effectively. In this paper, we propose a novel approach for depression recognition based on visual behaviors, which employs Atrous Residual Temporal Convolutional Network (DepArt-Net) as well as temporal fusion to capture the long-range dynamic depressive cues. First, the proposed atrous temporal convolution generates multi-scale contextual features from low-level visual behaviors, which are further strengthened by residual blocks across different convolution groups. Second, we introduce the attention mechanism in temporal feature fusion stage, and with the learned attentive distribution, more discriminative video-level depression representation can be acquired. Experimental results on the DAIC-WOZ benchmark demonstrate the effectiveness of the proposed approach and its superiority over other state-of-the-art methods.
Zhengyin Du, Weixin Li 0001, Di Huang 0001, Yunhong Wang 0001
FG2
2019 Discriminative Attention-based Convolutional Neural Network for 3D Facial Expression Recognition
abstract
3D Facial Expression Recognition (FER) is an active research area in computer vision. Although previous methods report promising results, two key issues still remain to be solved. On the one hand, different facial areas contribute unequally to performing various expressions, but most existing methods extract features from the entire 3D surface. On the other hand, the difference between expressions varies, while previous methods generally treat different emotions equally, making some of them extremely hard to be distinguished. To solve these problems, we propose a novel approach for 3D FER, namely Discriminative Attention-based Convolution Neural Network (DA-CNN), to generate more comprehensive expression related representations. DA-CNN introduces an attention module to the CNN models, which helps the deep model selectively focus on emotional salient regions in a learnable way. Furthermore, a novel loss named Dimensional Distribution (DD) loss is proposed to model the inter-expression relationship. Supervised by DD loss, DA-CNN can generate more discriminative expression representation. Extensive experiments are conducted on BU-3DFE dataset, and the results show that DA-CNN achieves significant improvement over the state-of-the-art.
Kangkang Zhu, Zhengyin Du, Weixin Li 0001, Di Huang 0001, Yunhong Wang 0001, Liming Chen 0002
FG3
2019 Continuous Emotion Recognition in Videos by Fusing Facial Expression, Head Pose and Eye Gaze
abstract
Continuous emotion recognition is of great significance in affective computing and human-computer interaction. Most of existing methods for video based continuous emotion recognition utilize facial expression. However, besides facial expression, other clues including head pose and eye gaze are also closely related to human emotion, but have not been well explored in continuous emotion recognition task. On the one hand, head pose and eye gaze could result in different degrees of credibility of facial expression features. On the other hand, head pose and eye gaze carry emotional clues themselves, which are complementary to facial expression. Accordingly, in this paper we propose two ways to incorporate these two clues into continuous emotion recognition. They are respectively an attention mechanism based on head pose and eye gaze clues to guide the utilization of facial features in continuous emotion recognition, and an auxiliary line which helps extract more useful emotion information from head pose and eye gaze. Experiments are conducted on the Recola dataset, a database for continuous emotion recognition, and the results show that our framework outperforms other state-of-the-art methods due to the full use of head pose and eye gaze clues in addition to facial expression for continuous emotion recognition.
Suowei Wu, Zhengyin Du, Weixin Li 0001, Di Huang 0001, Yunhong Wang 0001
ICMI3
2019 Shoot high-quality color images using dual-lens system with monochrome and color cameras
Xuan Dong 0001, Weixin Li 0001
Neurocomputing2
2018 Automatic Facial Attractiveness Prediction by Deep Multi-Task Learning
abstract
Facial Attractiveness Prediction (FAP) is a useful yet challenging problem in the domain of computer vision. In this paper, we propose a deep learning based approach. Different from the existing deep methods, the proposed one models both the texture and shape clues within a multi-task learning framework consisting of attractiveness score prediction and fiducial landmark localization, thus highlighting both of their roles in assessing attractiveness of faces. Considering that the training data are not extensive, a lightweight CNN is designed to jointly learn the facial representation, landmark location, and facial attractiveness score. The proposed method is evaluated on the SCUT-FBP database, and a prediction correlation 0.92, is delivered, which shows the effectiveness of our method. Furthermore, two additional experiments in terms of comparison between facial images before and after make-up or beautification are conducted. The results also prove the advantage of the proposed method.
Lian Gao, Weixin Li 0001, Zehua Huang, Di Huang 0001, Yunhong Wang 0001
ICPR2
2018 Facial Expression Synthesis by U-Net Conditional Generative Adversarial Networks
abstract
High-level manipulation of facial expressions in images such as expression synthesis is challenging because facial expression changes are highly non-linear, and vary depending on the facial appearance. Identity of the person should also be well preserved in the synthesized face. In this paper, we propose a novel U-Net Conditioned Generative Adversarial Network (UC-GAN) for facial expression generation. U-Net helps retain the property of the input face, including the identity information and facial details. We also propose an identity preserving loss, which further improves the performance of our model. Both qualitative and quantitative experiments are conducted on the Oulu-CASIA and KDEF datasets, and the results show that our method can generate faces with natural and realistic expressions while preserve the identity information. Comparison with the state-of-the-art approaches also demonstrates the competency of our method.
Weixin Li 0001, Guodong Mu, Di Huang 0001, Yunhong Wang 0001
ICMR2
2018 Ground-Truth Data Set and Baseline Evaluations for Base-Detail Separation Algorithms at the Part Level
abstract
Base-detail separation is a fundamental image processing problem, which models the image by a smooth base layer for the coarse structure and a detail layer for the texturelike structures. Base-detail separation is hierarchical and can be performed from the fine level to the coarse level. The separation at coarse level, in particular at the part level, is important for many applications, but currently lacks ground-truth data sets that are needed for comparing algorithms quantitatively. Thus, we propose a procedure to construct such data sets and provide two examples: Pascal Part UCLA and Fashionista, containing 1000 and 250 images, respectively. Our assumption is that the base is piecewise smooth, and we label the appearance of each piece by a polynomial model. The pieces are objects and parts of objects obtained from human annotations. Finally, we propose a way to evaluate different separation methods with our data sets and compared the performances of seven state-of-the-art algorithms.
Xuan Dong 0001, Boyan Bonev 0001, Weixin Li 0001, Weichao Qiu, Xianjie Chen, Alan L. Yuille
IEEE Trans. Circuits Syst. Video Technol.3
2017 Joint Image-Text News Topic Detection and Tracking by Multimodal Topic And-Or Graph
abstract
This paper presents a novel method for automatically detecting and tracking news topics from multimodal TV news data. We propose a multimodal topic and-or graph (MT-AOG) to jointly represent textual and visual elements of news stories and their latent topic structures. An MT-AOG leverages a context-sensitive grammar that can describe the hierarchical composition of news topics by semantic elements about people involved, related places, and what happened, and model contextual relationships between elements in the hierarchy. We detect news topics through a cluster sampling process which groups stories about closely related events together. Swendsen-Wang cuts, an effective cluster sampling algorithm, is adopted for traversing the solution space and obtaining optimal clustering solutions by maximizing a Bayesian posterior probability. The detected topics are then continuously tracked and updated with incoming news streams. We generate topic trajectories to show how topics emerge, evolve, and disappear over time. The experimental results show that our method can explicitly describe the textual and visual data in news videos and produce meaningful topic trajectories. Our method also outperforms previous methods for the task of document clustering on Reuters-21578 dataset and our novel dataset, UCLA Broadcast News dataset.
Weixin Li 0001, Jungseock Joo, Hang Qi 0001, Song-Chun Zhu
IEEE Trans. Multim.1
2015 Temporally consistent region-based video exposure correction
abstract
We analyze the problem of temporally consistent video exposure correction. Existing methods usually either fail to evaluate optimal exposure for every region or cannot get temporally consistent correction results. In addition, the contrast is often lost when the detail is not preserved properly during correction. In this paper, we use the block-based energy minimization to evaluate the temporally consistent exposure, which considers 1) the maximization of the visibility of all contents, 2) keeping the relative difference between neighboring regions, and 3) temporally consistent exposure of corresponding contents in different frames. Then, based on Weber contrast definition, we propose a contrast preserving exposure correction method. Experimental results show that our method enables better temporally consistent exposure evaluation and produces contrast preserving outputs.
Xuan Dong 0001, Lu Yuan 0001, Weixin Li 0001, Alan L. Yuille
ICME3
2014 Visual Persuasion: Inferring Communicative Intents of Images
abstract
In this paper we introduce the novel problem of understanding visual persuasion. Modern mass media make extensive use of images to persuade people to make commercial and political decisions. These effects and techniques are widely studied in the social sciences, but behavioral studies do not scale to massive datasets. Computer vision has made great strides in building syntactical representations of images, such as detection and identification of objects. However, the pervasive use of images for communicative purposes has been largely ignored. We extend the significant advances in syntactic analysis in computer vision to the higher-level challenge of understanding the underlying communicative intent implied in images. We begin by identifying nine dimensions of persuasive intent latent in images of politicians, such as "socially dominant, " "energetic, " and "trustworthy, " and propose a hierarchical model that builds on the layer of syntactical attributes, such as "smile" and "waving hand, " to predict the intents presented in the images. To facilitate progress, we introduce a new dataset of 1, 124 images of politicians labeled with ground-truth intents in the form of rankings. This study demonstrates that a systematic focus on visual persuasion opens up the field of computer vision to a new class of investigations around mediated images, intersecting with media analysis, psychology, and political communication.
Jungseock Joo, Weixin Li 0001, Francis F. Steen, Song-Chun Zhu
CVPR2
2012 Combining Tensor Space Analysis and Active Appearance Models for Aging Effect Simulation on Face Images
abstract
Applications of the simulation of adult aging effects are widespread nowadays, whereas the difficulties in certain aspects restrict its development. In this paper, a method is proposed for simulating adult facial aging effects by means of super-resolution. Accounting for the nature of multimodalities in the face image set, multilinear algebra is introduced to represent and process the whole image set in tensor space. To ameliorate the aging simulation results generated by merely the super-resolution method, we further adopt active appearance models to reduce the blurring effects of the results through adding normalization of the faces and postprocessing to the algorithm. To evaluate our aging simulation method, the aged faces obtained are compared with the ground-truth face images of the same individuals and also assessed by several volunteers mainly from two perspectives: the aged faces' perceived age and their preservation effects of the original identities of subjects in the test images. Additionally, objective experiments based on an automatic age estimator and a face recognition method using eigenfaces are also conducted as another way of the evaluation.
Yunhong Wang 0001, Zhaoxiang Zhang 0001, Weixin Li 0001, Fangyuan Jiang
IEEE Trans. Syst. Man Cybern. Part B3
2011 Fast efficient algorithm for enhancement of low lighting video
abstract
We describe a novel and effective video enhancement algorithm for low lighting video. The algorithm works by first inverting an input low-lighting video and then applying an optimized image de-haze algorithm on the inverted video. To facilitate faster computation, temporal correlations between subsequent frames are utilized to expedite the calculation of key algorithm parameters. Simulation results show excellent enhancement results and 4× speed up as compared with the frame-wise enhancement algorithms.
Xuan Dong 0001, Yi Pang, Weixin Li 0001, Jiangtao Wen, Wei Meng 0001, Yao Lu 0006
ICME4