EDBT 2026 Demo / reviewers in the wild / expert
Pengxu Wei
dblp:138/4243
· DBLP profile ↗
57ranked-venue papers
4as first author
47since 2021 · last 2026
0000-0002-2190-0767ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 38 · 4 first-author · 31 since 2021Artificial intelligence and machine learning · 31 · 2 first-author · 26 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Similarity-aware Probabilistic Embeddings Modeling for Video-Text RetrievalabstractVideo-text retrieval is a fundamental task in multi-modal learning, aiming to accurately retrieve videos that match given textual descriptions. While recent contrastive methods have made significant progress by embedding videos and texts into a joint space, they often suffer from semantic over-clustering—a phenomenon where semantically distinct videos are mapped to overly similar embeddings due to dominant but uninformative visual patterns (e.g., recurring backgrounds or common objects). This effect becomes particularly problematic under short or ambiguous queries, where it suppresses fine-grained semantics and degrades retrieval precision. To address this, we propose Similarity-aware Probabilistic Embeddings Modeling (SPEM), a novel framework that refines video representations by modeling them as adaptive probability distributions rather than static vectors. SPEM incorporates cross-modal attention to high-light text-relevant visual content and suppress irrelevant patterns, and leverages multi-level similarity features to dynamically adjust the embedding variance, thereby preserving subtle but critical semantic cues. To further improve alignment, we employ a Semantic-Distribution Contrastive Loss to optimize the alignment structure in the probabilistic space, encouraging more discriminative separation across hard negatives. Extensive experiments on four widely-used video-text retrieval benchmarks—MSRVTT, DiDeMo, VATEX, and MSVD—demonstrate that SPEM consistently outperforms strong CLIP-based baselines. Yuliang Huang, Pengxu Wei, Liang Lin 0004 |
WACV | 2 |
| 2026 | Robust Source-Free Domain Adaptation From Non-Robust Source ModelsabstractA few recent works attempt to train an adversarially robust Unsupervised Domain Adaptation (UDA) model, transferring the robustness from a robust source model or other robust pre-trained models to an unlabeled target domain. However, it is usually impractical to assume the availability of robust source models or robust pre-training, and meanwhile, source data are not always accessible or efficient for adaptation training in many real-world scenarios. In this paper, we dive into a more practical and challenging problem of robust source-free domain adaptation: can we train a robust model on an unlabeled target domain given only a non-robust source model (without source data)? Empirically, we find that applying adversarial training (AT) to the self-supervised adaptation process leads to severe model degradation, as it tends to amplify the inevitable errors of UDA models. To tackle this issue, we propose a novel approach called Source-Free Alternating Optimization (SFAO), which employs a non-robust target model to provide better guidance for the AT of the desired robust target model. The two models are trained in an alternating manner to minimize the discrepancy between the clean source domain and the adversarial target domain. Moreover, we propose Softly-Constrained Adversarial Training (SCAT) to further mitigate the adverse effects of incorrect pseudo-labels in AT. Extensive experimental results demonstrate that the proposed method significantly improves the model performance on both clean and adversarial data. Source code is available at: https://github.com/Coxy7/robust-SFDA. Pengxu Wei, Guangrun Wang, Cong Liu 0001, Liang Lin 0004 |
IEEE Trans. Image Process. | 2 |
| 2026 | CIREC: Causal Intervention-Inspired Policy Learning to Mitigate Exposure Bias for Interactive Recommendation
Yongsen Zheng, Guohua Wang 0005, Jinghui Qin, Ziliang Chen 0001, Junfan Lin, Pengxu Wei, Liang Lin 0004, Kwok-Yan Lam |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2025 | DigitalLLaVA: Incorporating Digital Cognition Capability for Physical World Comprehension in Multimodal LLMsabstractMultimodal Large Language Models (MLLMs) have shown remarkable cognitive capabilities in various cross-modal tasks.However, existing MLLMs struggle with tasks that require physical digital cognition, such as accurately reading an electric meter or pressure gauge. This limitation significantly reduces their effectiveness in practical applications like industrial monitoring and home energy management, where digital sensors are not feasible. For humans, physical digits are artificially defined quantities presented on specific carriers, which require training to recognize. As existing MLLMs are only pre-trained in the manner of object recognition, they fail to comprehend the relationship between digital carriers and their reading. To this end, referring to human behavior, we propose a novel DigitalLLaVA method to explicitly inject digital cognitive abilities into MLLMs in a two-step manner. In the first step, to improve the MLLM's understanding of physical digit carriers, we propose a digit carrier mapping method. This step utilizes object-level text-image pairs to enhance the model's comprehension of objects containing physical digits. For the second step, unlike previous methods that rely on sequential digital prediction or digit regression, we propose a 32 bit floating point simulation approach that treats digit prediction as a whole. Using digit-level text-image pairs, we train three float heads to predict 32-bit floating-point numbers using 0/1 binary classification. This step significantly reduces the search space, making the prediction process more robust and straightforward. Being simple but effective, our method can identify very precise metrics (i.e., accurate to ±0.001) and provide floating-point results, showing its applicability in digital carrier domains. Pengxu Wei, Pengchong Qiao, Chang Liu 0030, Jie Chen 0001 |
AAAI | 2 |
| 2025 | Can We Achieve Efficient Diffusion without Self-Attention? Distilling Self-Attention Into ConvolutionsabstractContemporary diffusion models built upon U-Net or Diffusion Transformer (DiT) architectures have revolutionized image generation through transformer-based attention mechanisms. The prevailing paradigm has commonly employed self-attention with quadratic computational complexity to handle global spatial relationships in complex images, thereby synthesizing high-fidelity images with coherent visual semantics.Contrary to conventional wisdom, our systematic layer-wise analysis reveals an interesting discrepancy: self-attention in pre-trained diffusion models predominantly exhibits localized attention patterns, closely resembling convolutional inductive biases. This suggests that global interactions in self-attention may be less critical than commonly assumed.Driven by this, we propose \(Δ\)ConvFusion to replace conventional self-attention modules with Pyramid Convolution Blocks (\(Δ\)ConvBlocks).By distilling attention patterns into localized convolutional operations while keeping other components frozen, \(Δ\)ConvFusion achieves performance comparable to transformer-based counterparts while reducing computational cost by 6929$\times$ and surpassing LinFusion by 5.42$\times$ in efficiency--all without compromising generative fidelity. ZiYi Dong, Chengxing Zhou, Weijian Deng, Pengxu Wei, Xiangyang Ji, Liang Lin 0004 |
ICCV | 4 |
| 2025 | Towards Understanding the Robustness of Diffusion-Based Purification: A Stochastic PerspectiveabstractDiffusion-Based Purification (DBP) has emerged as an effective defense mechanism against adversarial attacks. The success of DBP is often attributed to the forward diffusion process, which reduces the distribution gap between clean and adversarial images by adding Gaussian noise. Although this explanation is theoretically grounded, the precise contribution of this process to robustness remains unclear. In this paper, through a systematic investigation, we propose that the intrinsic stochasticity in the DBP procedure is the primary factor driving robustness. To explore this hypothesis, we introduce a novel Deterministic White-Box (DW-box) evaluation protocol to assess robustness in the absence of stochasticity, and analyze attack trajectories and loss landscapes. Our results suggest that DBP models primarily leverage stochasticity to evade effective attack directions, and that their ability to purify adversarial perturbations can be weak. To further enhance the robustness of DBP models, we propose Adversarial Denoising Diffusion Training (ADDT), which incorporates classifier-guided adversarial perturbations into diffusion training, thereby strengthening the models' ability to purify adversarial perturbations. Additionally, we propose Rank-Based Gaussian Mapping (RBGM) to improve the compatibility of perturbations with diffusion models. Experimental results validate the effectiveness of ADDT. In conclusion, our study suggests that future research on DBP can benefit from the perspective of decoupling stochasticity-based and purification-based robustness. Kezhao Liu, Ziyi Dong, Xiaogang Xu 0002, Pengxu Wei, Liang Lin 0004 |
ICLR | 6 |
| 2025 | Are High-Quality AI-Generated Images More Difficult for Models to Detect?abstractThe remarkable evolution of generative models has enabled the generation of high-quality, visually attractive images, often perceptually indistinguishable from real photographs to human eyes. This has spurred significant attention on AI-generated image (AIGI) detection. Intuitively, higher image quality should increase detection difficulty. However, our systematic study on cutting-edge text-to-image generators reveals a counterintuitive finding: AIGIs with higher quality scores, as assessed by human preference models, tend to be more easily detected by existing models. To investigate this, we examine how the text prompts for generation and image characteristics influence both quality scores and detector accuracy. We observe that images from short prompts tend to achieve higher preference scores while being easier to detect. Furthermore, through clustering and regression analyses, we verify that image characteristics like saturation, contrast, and texture richness collectively impact both image quality and detector accuracy. Finally, we demonstrate that the performance of off-the-shelf detectors can be enhanced across diverse generators and datasets by selecting input patches based on the predicted scores of our regression models, thus substantiating the broader applicability of our findings. Code and data are available at https://github.com/Coxy7/AIGI-Detection-Quality-Paradox. Zijie Cao, ZiYi Dong, Xiangyang Ji, Liang Lin 0004, Wei Ke 0003, Pengxu Wei |
ICML | 10 |
| 2025 | Delving into Cascaded Instability: A Lipschitz Continuity View on Image Restoration and Object Detection SynergyabstractTo improve detection robustness in adverse conditions (e.g., haze and low light), image restoration is commonly applied as a pre-processing step to enhance image quality for the detector. However, the functional mismatch between restoration and detection networks can introduce instability and hinder effective integration---an issue that remains underexplored. We revisit this limitation through the lens of Lipschitz continuity, analyzing the functional differences between restoration and detection networks in both the input space and the parameter space. Our analysis shows that restoration networks perform smooth, continuous transformations, while object detectors operate with discontinuous decision boundaries, making them highly sensitive to minor perturbations. This mismatch introduces instability in traditional cascade frameworks, where even imperceptible noise from restoration is amplified during detection, disrupting gradient flow and hindering optimization. To address this, we propose Lipschitz-regularized object detection (LROD), a simple yet effective framework that integrates image restoration directly into the detector’s feature learning, harmonizing the Lipschitz continuity of both tasks during training. We implement this framework as Lipschitz-regularized YOLO (LR-YOLO), extending seamlessly to existing YOLO detectors. Extensive experiments on haze and low-light benchmarks demonstrate that LR-YOLO consistently improves detection stability, optimization smoothness, and overall accuracy. Weijian Deng, Pengxu Wei, ZiYi Dong, Hannan Lu, Xiangyang Ji, Liang Lin 0004 |
NeurIPS | 3 |
| 2025 | DreamArtist: Controllable One-Shot Text-to-Image Generation via Positive-Negative Adapter
Ziyi Dong, Pengxu Wei, Liang Lin 0004 |
Int. J. Comput. Vis. | 2 |
| 2025 | SAM-COD+: SAM-Guided Unified Framework for Weakly-Supervised Camouflaged Object DetectionabstractMost Camouflaged Object Detection (COD) methods heavily rely on mask annotations, which are time-consuming and labor-intensive to acquire. Existing weakly-supervised COD approaches exhibit significantly inferior performance compared to fully-supervised methods and struggle to simultaneously support all the existing types of camouflaged object labels, including scribbles, bounding boxes, and points. Even for Segment Anything Model (SAM), it is still problematic to handle the weakly-supervised COD and it typically encounters challenges of prompt compatibility of the scribble labels, extreme response, semantically erroneous response, and unstable feature representations, producing unsatisfactory results in camouflaged scenes. To mitigate these issues, we propose a unified COD framework in this paper, termed SAM-COD, which is capable of supporting arbitrary weakly-supervised labels. Our SAM-COD employs a prompt adapter to handle scribbles as prompts based on SAM. Meanwhile, we introduce response filter and semantic matcher modules to improve the quality of the masks obtained by SAM under COD prompts. To alleviate the negative impacts of inaccurate mask predictions, a new strategy of prompt-adaptive knowledge distillation is utilized to ensure a reliable feature representation. To validate the effectiveness of our approach, we have conducted extensive empirical experiments on three mainstream COD benchmarks. The results demonstrate the superiority of our method against state-of-the-art weakly-supervised and even fully-supervised methods. Our source codes and trained models will be publicly released. Pengxu Wei, Guangqian Guo, Shan Gao 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | SAM-COD: SAM-Guided Unified Framework for Weakly-Supervised Camouflaged Object Detection
Pengxu Wei, Guangqian Guo, Shan Gao 0003 |
ECCV (35) | 2 |
| 2024 | Towards Real-world Continuous Super-Resolution: Benchmark and MethodabstractContinuous Super-Resolution (CSR) has garnered considerable popularity for its capability to reconstruct high-resolution (HR) images from low-resolution (LR) inputs at various scales, thereby holding significant practical value in real-world applications. However, the existing studies have relied solely on synthetic datasets due to the scarcity of real-world continuous datasets. In this paper, we establish a real-world continuous super-resolution dataset, i.e., RealCSR, aiming to explore the realistic continuous image representation in real-world scenarios. Additionally, we introduce a Codebook-embedded Attention Network (CANet) for Real-world CSR to investigate improved information gain and distribution in the presence of real-world image degradation. We employ a codebook-embedded attention module to effectively capture and aggregate local and global information combined with codebook information. Specifically, CANet utilizes a multi-scale SR codebook and a dilated neighborhood attention mechanism to effectively capture and aggregate both local and global information. This enriched information is then fed into a scale-aware implicit fusion module modulated by a scale-guided degradation map. Experimental results demonstrate that our CANet outperforms existing state-of-the-art CSR methods and also achieves visually pleasant SR image predictions with realistic, accurate details. Our dataset, codes, and models are publicly available at https://github.com/doooooithey/CANet. Xingbei Guo, Ziping Ma 0002, Qing Wang 0018, Pengxu Wei |
ICME | 4 |
| 2024 | Kepler codebookabstractA codebook designed for learning discrete distributions in latent space has demonstrated state-of-the-art results on generation tasks. This inspires us to explore what distribution of codebook is better. Following the spirit of Kepler's Conjecture, we cast the codebook training as solving the sphere packing problem and derive a Kepler codebook with a compact and structured distribution to obtain a codebook for image representations. Furthermore, we implement the Kepler codebook training by simply employing this derived distribution as regularization and using the codebook partition method. We conduct extensive experiments to evaluate our trained codebook for image reconstruction and generation on natural and human face datasets, respectively, achieving significant performance improvement. Besides, our Kepler codebook has demonstrated superior performance when evaluated across datasets and even for reconstructing images with different resolutions. Our trained models and source codes will be publicly released. Junrong Lian, Ziyue Dong, Pengxu Wei, Wei Ke 0003, Chang Liu 0030, Qixiang Ye, Xiangyang Ji, Liang Lin 0004 |
ICML | 3 |
| 2024 | Image Restoration Through Generalized Ornstein-Uhlenbeck BridgeabstractDiffusion models exhibit powerful generative capabilities enabling noise mapping to data via reverse stochastic differential equations. However, in image restoration, the focus is on the mapping relationship from low-quality to high-quality images. Regarding this issue, we introduce the Generalized Ornstein-Uhlenbeck Bridge (GOUB) model. By leveraging the natural mean-reverting property of the generalized OU process and further eliminating the variance of its steady-state distribution through the Doob's *h*–transform, we achieve diffusion mappings from point to point enabling the recovery of high-quality images from low-quality ones. Moreover, we unravel the fundamental mathematical essence shared by various bridge models, all of which are special instances of GOUB and empirically demonstrate the optimality of our proposed models. Additionally, we present the corresponding Mean-ODE model adept at capturing both pixel-level details and structural perceptions. Experimental outcomes showcase the state-of-the-art performance achieved by both models across diverse tasks, including inpainting, deraining, and super-resolution. Code is available at https://github.com/Hammour-steak/GOUB. Conghan Yue, Zhengwei Peng, Junlong Ma, Shiyan Du, Pengxu Wei, Dongyu Zhang 0002 |
ICML | 5 |
| 2024 | Decoder-Only LLMs are Better Controllers for Diffusion Models
Ziyi Dong, Pengxu Wei, Liang Lin 0004 |
ACM Multimedia | 3 |
| 2024 | Advancing Real-World Burst Denoising: A New Benchmark and Dual-Branch Burst Denoising Network
Huilei Wu, Zhiyuan Song, Pengxu Wei |
PRCV (8) | 4 |
| 2024 | Integrating instance-level knowledge to see the unseen: A two-stream network for video object segmentation
Hannan Lu, Zhi Tian, Pengxu Wei, Haibing Ren, Wangmeng Zuo |
Neurocomputing | 3 |
| 2024 | Diff-Mosaic: Augmenting Realistic Representations in Infrared Small Target Detection via Diffusion PriorabstractRecently, researchers have proposed various deep learning methods to accurately detect infrared targets with the characteristics of indistinct shape and texture. Due to the limited variety of infrared datasets, training deep learning models with good generalization poses a challenge. To augment the infrared dataset, researchers employ data augmentation techniques, which often involve generating new images by combining images from different datasets. However, these methods are lacking in two respects. In terms of realism, the images generated by mixup-based methods lack realism and are difficult to effectively simulate complex real-world scenarios. In terms of diversity, compared with real-world scenes, borrowing knowledge from another dataset inherently has a limited diversity. Currently, the diffusion model stands out as an innovative generative approach. Large-scale trained diffusion models have a strong generative prior that enables real-world modeling of images to generate diverse and realistic images. In this paper, we propose Diff-Mosaic, a data augmentation method based on the diffusion model. This model effectively alleviates the challenge of diversity and realism of data augmentation methods via diffusion prior. Specifically, our method consists of two stages. Firstly, we introduce an enhancement network called Pixel-Prior, which generates highly coordinated and realistic Mosaic images by harmonizing pixels. In the second stage, we propose an image enhancement strategy named Diff-Prior. This strategy utilizes diffusion priors to model images in the real-world scene, further enhancing the diversity and realism of the images. Extensive experiments have demonstrated that our approach significantly improves the performance of the detection network. The code is available at https://github.com/YupeiLin2388/Diff-Mosaic. Yukai Shi, Yupei Lin, Pengxu Wei, Xiaoyu Xian, Tianshui Chen, Liang Lin 0004 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | IDF-CR: Iterative Diffusion Process for Divide-and-Conquer Cloud Removal in Remote-Sensing ImagesabstractDeep learning technologies have demonstrated their effectiveness in removing cloud cover from optical remote-sensing images. Convolutional Neural Networks (CNNs) exert dominance in the cloud removal tasks. However, constrained by the inherent limitations of convolutional operations, CNNs can address only a modest fraction of cloud occlusion. In recent years, diffusion models have achieved state-of-the-art (SOTA) proficiency in image generation and reconstruction due to their formidable generative capabilities. Inspired by the rapid development of diffusion models, we first present an iterative diffusion process for cloud removal (IDF-CR), which exhibits a strong generative capabilities to achieve component divide-and-conquer cloud removal. IDF-CR consists of a pixel space cloud removal module (Pixel-CR) and a latent space iterative noise diffusion network (IND). Specifically, IDF-CR is divided into two-stage models that address pixel space and latent space. The two-stage model facilitates a strategic transition from preliminary cloud reduction to meticulous detail refinement. In the pixel space stage, Pixel-CR initiates the processing of cloudy images, yielding a suboptimal cloud removal prior to providing the diffusion model with prior cloud removal knowledge. In the latent space stage, the diffusion model transforms low-quality cloud removal into high-quality clean output. We refine the Stable Diffusion by implementing ControlNet. In addition, an unsupervised iterative noise refinement (INR) module is introduced for diffusion model to optimize the distribution of the predicted noise, thereby enhancing advanced detail recovery. Our model performs best with other SOTA methods, including image reconstruction and optical remote-sensing cloud removal on the optical remote-sensing datasets. Meilin Wang, Yexing Song, Pengxu Wei, Xiaoyu Xian, Yukai Shi, Liang Lin 0004 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Multi-Person 3D Pose Estimation With Occlusion ReasoningabstractThe performance of existing methods for multi-person 3D pose estimation in crowded scenes is still limited, due to the challenge of heavy overlapping among persons. Attempt to address this issue, we propose a progressive inference scheme, i.e., Articulation-aware Knowledge Exploration (AKE), to improve the multi-person 3D pose models on those samples with complex occlusions at the inference stage. We argue it is beneficial to explore the underlying articulated information/knowledge of the human body, which helps to further correct the predicted poses in those samples. To exploit such information, we propose an iterative scheme to achieve a self-improving loop for keypoint association. Specifically, we introduce a kinematic validation module for locating unreasonable articulations and an occluded-keypoint discovering module for discovering occluded articulations. Extensive experiments on two challenging benchmarks under both weakly-supervised and fully-supervised settings demonstrate the superiority and generalization ability of our proposed method for crowded scenes. Xipeng Chen, Junzheng Zhang, Keze Wang, Pengxu Wei, Liang Lin 0004 |
IEEE Trans. Multim. | 4 |
| 2024 | CIPL: Counterfactual Interactive Policy Learning to Eliminate Popularity Bias for Online RecommendationabstractPopularity bias, as a long-standing problem in recommender systems (RSs), has been fully considered and explored for offline recommendation systems in most existing relevant researches, but very few studies have paid attention to eliminate such bias in online interactive recommendation scenarios. Bias amplification will become increasingly serious over time due to the existence of feedback loop between the user and the interactive system. However, existing methods have only investigated the causal relations among different factors statically without considering temporal dependencies inherent in the online interactive recommendation system, making them difficult to be adapted to online settings. To address these problems, we propose a novel counterfactual interactive policy learning (CIPL) method to eliminate popularity bias for online recommendation. It first scrutinizes the causal relations in the interactive recommender models and formulates a novel temporal causal graph (TCG) to guide the training and counterfactual inference of the causal interactive recommendation system. Concretely, TCG is used to estimate the causal relations of item popularity on prediction score when the user interacts with the system at each time during model training. Besides, it is also used to remove the negative effect of popularity bias in the test stage. To train the causal interactive recommendation system, we formulated our CIPL by the actor-critic framework with an online interactive environment simulator. We conduct extensive experiments on three public benchmarks and the experimental results demonstrate that our proposed method can achieve the new state-of-the-art performance. Yongsen Zheng, Jinghui Qin, Pengxu Wei, Ziliang Chen 0001, Liang Lin 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Routing User-Interest Markov Tree for Scalable Personalized Knowledge-Aware RecommendationabstractTo facilitate more accurate and explainable recommendation, it is crucial to incorporate side information into user-item interactions. Recently, knowledge graph (KG) has attracted much attention in a variety of domains due to its fruitful facts and abundant relations. However, the expanding scale of real-world data graphs poses severe challenges. In general, most existing KG-based algorithms adopt exhaustively hop-by-hop enumeration strategy to search all the possible relational paths, this manner involves extremely high-cost computations and is not scalable with the increase of hop numbers. To overcome these difficulties, in this article, we propose an end-to-end framework Knowledge-tree-routed UseR-Interest Trajectories Network (KURIT-Net). KURIT-Net employs the user-interest Markov trees (UIMTs) to reconfigure a recommendation-based KG, striking a good balance for routing knowledge between short-distance and long-distance relations between entities. Each tree starts from the preferred items for a user and routes the association reasoning paths along the entities in the KG to provide a human-readable explanation for model prediction. KURIT-Net receives entity and relation trajectory embedding (RTE) and fully reflects potential interests of each user by summarizing all reasoning paths in a KG. Besides, we conduct extensive experiments on six public datasets, our KURIT-Net significantly outperforms state-of-the-art approaches and shows its interpretability in recommendation. Yongsen Zheng, Pengxu Wei, Ziliang Chen 0001, Chengpei Tang, Liang Lin 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Scene Graph to Image Synthesis via Knowledge ConsensusabstractIn this paper, we study graph-to-image generation conditioned exclusively on scene graphs, in which we seek to disentangle the veiled semantics between knowledge graphs and images. While most existing research resorts to laborious auxiliary information such as object layouts or segmentation masks, it is also of interest to unveil the generality of the model with limited supervision, moreover, avoiding extra cross-modal alignments. To tackle this challenge, we delve into the causality of the adversarial generation process, and reason out a new principle to realize a simultaneous semantic disentanglement with an alignment on target and model distributions. This principle is named knowledge consensus, which explicitly describes a triangle causal dependency among observed images, graph semantics and hidden visual representations. The consensus also determines a new graph-to-image generation framework, carried on several adversarial optimization objectives. Extensive experimental results demonstrate that, even conditioned only on scene graphs, our model surprisingly achieves superior performance on semantics-aware image generation, without losing the competence on manipulating the generation through knowledge graphs. Pengxu Wei, Liang Lin 0004 |
AAAI | 2 |
| 2023 | Out-of-Candidate Rectification for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation is typically inspired by class activation maps, which serve as pseudo masks with class-discriminative regions highlighted. Although tremendous efforts have been made to recall precise and complete locations for each class, existing methods still commonly suffer from the unsolicited Out-of-Candidate (OC) error predictions that do not belong to the label candidates, which could be avoidable since the contradiction with image-level class tags is easy to be detected. In this paper, we develop a group ranking-based Out-of-f;Candidate Rectification (OCR) mechanism in a plug-and-play fashion. Firstly, we adaptively split the semantic categories into In-Candidate (IC) and OC groups for each OC pixel according to their prior annotation correlation and posterior prediction correlation. Then, we derive a differentiable rectification loss to force OC pixels to shift to the IC group. Incorporating OCR with seminal baselines (e.g., AffinityNet, SEAM, MCTformer), we can achieve remarkable performance gains on both Pascal VOC (+3.2%, +3.3%, +0.8% mIoU) and MS COCO (+1.0%, +1.3%, +0.5% mIoU) datasets with negligible extra training overhead, which jus-tifies the effectiveness and generality of OCR.††Ŋ github.com/sennnnn/Out-of-Candidate-Rectification Zesen Cheng, Pengchong Qiao, Kehan Li 0002, Siheng Li, Pengxu Wei, Xiangyang Ji, Li Yuan 0007, Chang Liu 0030, Jie Chen 0001 |
CVPR | 5 |
| 2023 | Masked Images Are Counterfactual Samples for Robust Fine-TuningabstractDeep learning models are challenged by the distribution shift between the training data and test data. Recently, the large models pre-trained on diverse data have demonstrated unprecedented robustness to various distribution shifts. However, fine-tuning these models can lead to a trade-off between in-distribution (ID) performance and out-of-distribution (OOD) robustness. Existing methods for tackling this trade-off do not explicitly address the OOD robustness problem. In this paper, based on causal analysis of the aforementioned problems, we propose a novel fine-tuning method, which uses masked images as counterfactual samples that help improve the robustness of the fine-tuning model. Specifically, we mask either the semantics-related or semantics-unrelated patches of the images based on class activation map to break the spurious correlation, and refill the masked patches with patches from other images. The resulting counterfactual samples are used in feature-based distillation with the pre-trained model. Extensive experiments verify that regularizing the fine-tuning with the proposed masked images can achieve a better trade-off between ID and OOD performance, surpassing previous methods on the OOD performance. Our code is available at https://github.com/Coxy7/robust-finetuning. Pengxu Wei, Cong Liu 0001, Liang Lin 0004 |
CVPR | 3 |
| 2023 | Identity-Preserving Talking Face Generation with Landmark and Appearance PriorsabstractGenerating talking face videos from audio attracts lots of research interest. A few person-specific methods can generate vivid videos but require the target speaker's videos for training or fine-tuning. Existing person-generic methods have difficulty in generating realistic and lip-synced videos while preserving identity information. To tackle this problem, we propose a two-stage framework consisting of audio-to-landmark generation and landmark-to-video rendering procedures. First, we devise a novel Transformer-based landmark generator to infer lip and jaw landmarks from the audio. Prior landmark characteristics of the speaker's face are employed to make the generated landmarks coincide with the facial outline of the speaker. Then, a video rendering model is built to translate the generated landmarks into face images. During this stage, prior appearance information is extracted from the lower-half occluded target face and static reference images, which helps generate realistic and identity-preserving visual content. For effectively exploring the prior information of static reference images, we align static reference images with the target face's pose and expression based on motion fields. Moreover, auditory features are reused to guarantee that the generated face images are well synchronized with the audio. Extensive experiments demonstrate that our method can produce more realistic, lip-synced, and identity-preserving videos than existing person-generic talking face generation methods. Weizhi Zhong, Chaowei Fang, Yinqi Cai, Pengxu Wei, Gangming Zhao, Liang Lin 0004, Guanbin Li |
CVPR | 4 |
| 2023 | Learning Task-Aligned Mask Query for Instance SegmentationabstractRecently, query-based instance segmentation methods have achieved comparable performance to previous state-of-the-art methods. However, the query lacks the learning of the consistency between classification and segmentation tasks, which may lead to misalignment between classification score and mask quality (i.e., mask IoU) and can not result in a reliable ranking for predictions. In this work, we propose a novel instance segmentation method, termed AlignMask, which effectively learns task-aligned mask queries for instance end-toend. Specifically, we propose Aligned Query Learning (AQL) to learn task-aligned features for pixel embedding and transformer decoder, which helps segmentation quality estimation of the mask query. We also use Aligned Label Assignment to explicitly align the optimization goals for classification score and mask quality of the query. Extensive experiments on MSCOCO show that our proposed AlignMask achieves competitive performance with state-of-the-art models. Pengxu Wei, Jie Chen 0001 |
ICASSP | 3 |
| 2023 | TopoSeg: Topology-Aware Nuclear Instance SegmentationabstractNuclear instance segmentation has been critical for pathology image analysis in medical science, e.g., cancer diagnosis. Current methods typically adopt pixel-wise optimization for nuclei boundary exploration, where rich structural information could be lost for subsequent quantitative morphology assessment. To address this issue, we develop a topology-aware segmentation approach, termed TopoSeg, which exploits topological structure information to keep the predictions rational, especially in common situations with densely touching and overlapping nucleus instances. Concretely, TopoSeg builds on a topology-aware module (TAM), which encodes dynamic changes of different topology structures within the three-class probability maps (inside, boundary, and background) of the nuclei to persistence barcodes and makes the topology-aware loss function. To efficiently focus on regions with high topological errors, we propose an adaptive topology-aware selection (ATS) strategy to enhance the topology-aware optimization procedure further. Experiments on three nuclear instance segmentation datasets justify the superiority of TopoSeg, which achieves state-of-the-art performance. The code is available at https://github.com/hhlisme/toposeg. Pengxu Wei, Xiangyang Ji, Chang Liu 0030, Jie Chen 0001 |
ICCV | 3 |
| 2023 | Towards Real-World Burst Image Super-Resolution: Benchmark and MethodabstractDespite substantial advances, single-image super-resolution (SISR) is always in a dilemma to reconstruct high-quality images with limited information from one input image, especially in realistic scenarios. In this paper, we establish a large-scale real-world burst super-resolution dataset, i.e., RealBSR, to explore the faithful reconstruction of image details from multiple frames. Furthermore, we introduce a Federated Burst Affinity network (FBAnet) to investigate non-trivial pixel-wise displacements among images under real-world image degradation. Specifically, rather than using pixel-wise alignment, our FBAnet employs a simple homography alignment from a structural geometry aspect and a Federated Affinity Fusion (FAF) strategy to aggregate the complementary information among frames. Those fused informative representations are fed to a Transformer-based module of burst representation decoding. Besides, we have conducted extensive experiments on two versions of our datasets, i.e., RealBSR-RAW and RealBSR-RGB. Experimental results demonstrate that our FBAnet outperforms existing state-of-the-art burst SR methods and also achieves visually-pleasant SR image predictions with model details. Our dataset, codes, and models are publicly available at https://github.com/yjsunnn/FBANet. Pengxu Wei, Yujing Sun 0004, Xingbei Guo, Chang Liu 0030, Guanbin Li, Jie Chen 0001, Xiangyang Ji, Liang Lin 0004 |
ICCV | 1 |
| 2023 | Adversarially Robust Source-free Domain Adaptation with Relaxed Adversarial TrainingabstractUnsupervised Domain Adaptation (UDA) learns a model for an unlabeled target domain, utilizing a labeled source domain. Most existing works on UDA assume the availability of source data and neglect the adversarial robustness of the models, hindering security-sensitive real-world applications. In this paper, we study adversarially robust source-free UDA, aiming to train a robust target model by adapting a non-robust source model without using source data. A basic approach is to train a non-robust teacher model via conventional source-free UDA to predict pseudo-labels for target data, and then train a robust student model via adversarial training (AT). However, AT tends to magnify the errors of the teacher model, reducing the accuracy. Hence, we propose Relaxed Adversarial Training (RAT) that relieves the constraints on the confidence of predictions in AT to balance the robustness and accuracy. Extensive experiments validate that RAT can improve the accuracy on clean and adversarial samples, and is superior to related methods. Our code is available at https://github.com/Coxy7/RAT. Pengxu Wei, Cong Liu 0001, Liang Lin 0004 |
ICME | 2 |
| 2023 | Object-Aware Transfer-Based Black-Box Adversarial Attack on Object Detector
Zhuo Leng, Zesen Cheng, Pengxu Wei, Jie Chen 0001 |
PRCV (12) | 3 |
| 2023 | FIRE: Fine Implicit Reconstruction Enhancement with Detailed Body Part Labels and Geometric Features
Junzheng Zhang, Xipeng Chen, Keze Wang, Pengxu Wei, Liang Lin 0004 |
PRCV (2) | 4 |
| 2023 | Taylor Neural Network for Real-World Image Super-ResolutionabstractDue to the difficulty of collecting paired Low-Resolution (LR) and High-Resolution (HR) images, the recent research on single image Super-Resolution (SR) has often been criticized for the data bottleneck of the synthetic image degradation between LRs and HRs. Recently, the emergence of real-world SR datasets, e.g., RealSR and DRealSR, promotes the exploration of Real-World image Super-Resolution (RWSR). RWSR exposes a more practical image degradation, which greatly challenges the learning capacity of deep neural networks to reconstruct high-quality images from low-quality images collected in realistic scenarios. In this paper, we explore Taylor series approximation in prevalent deep neural networks for image reconstruction, and propose a very general Taylor architecture to derive Taylor Neural Networks (TNNs) in a principled manner. Our TNN builds Taylor Modules with Taylor Skip Connections (TSCs) to approximate the feature projection functions, following the spirit of Taylor Series. TSCs introduce the input connected directly with each layer at different layers, to sequentially produces different high-order Taylor maps to attend more image details, and then aggregate the different high-order information from different layers. Only via simple skip connections, TNN is compatible with various existing neural networks to effectively learn high-order components of the input image with little increase of parameters. Furthermore, we have conducted extensive experiments to evaluate our TNNs in different backbones on two RWSR benchmarks, which achieve a superior performance in comparison with existing baseline methods. Pengxu Wei, Ziwei Xie, Guanbin Li, Liang Lin 0004 |
IEEE Trans. Image Process. | 1 |
| 2023 | Graph-Convolved Factorization Machines for Personalized RecommendationabstractFactorization machines (FMs) and their neural network variants (neural FMs) for modeling second-order feature interactions are effective in building modern recommendation systems. However, feature interactions are based upon pairs of features, whereas multi-features correlations commonly arise in real-world financial product recommendation scenarios. We propose an effective neural recommender system, graph-convolved factorization machine (GCFM), with the spirit of the symbolic graph reasoning principle that provides lightweight and interpretable recommendation suggestions. Given a sample for the recommendation, GCFM constructs the corresponding dense feature embeddings and computes the sample-specific feature relationship graph. Then, a multi-filter graph-convolved feature crossing (GCFC) layer for feature embeddings establishes cross features with their neighboring embeddings. GCFM thus extends the feature interactions from pairs to neighbors to capture more comprehensive and explainable information while simultaneously reaping the advantages of representation learning. To exploit these capabilities, we apply a Graph Bayesian Optimization (GBO). During training, our GBO automatically optimizes our GCFM, including training hyperparameters and architecture hyperparameters. Besides, we conduct extensive experiments on two public financial applications benchmarks, USCFC and OTC, and two real-world datasets that we collect offline. Our GCFM significantly outperforms state-of-the-art algorithms and shows its interpretability in recommendation tasks. Yongsen Zheng, Pengxu Wei, Ziliang Chen 0001, Liang Lin 0004 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Real-World Image Super-Resolution by Exclusionary Dual-LearningabstractReal-world image super-resolution is a practical image restoration problem that aims to obtain high-quality images from in-the-wild input, has recently received considerable attention with regard to its tremendous application potentials. Although deep learning-based methods have achieved promising restoration quality on real-world image super-resolution datasets, they ignore the relationship between L1- and perceptual- minimization and roughly adopt auxiliary large-scale datasets for pre-training. In this paper, we discuss the image types within a corrupted image and the property of perceptual- and Euclidean- based evaluation protocols. Then we propose a method, Real-World image Super-Resolution by Exclusionary Dual-Learning (RWSR-EDL) to address the feature diversity in perceptual- and L1- based cooperative learning. Moreover, a noise-guidance data collection strategy is developed to address the training time consumption in multiple datasets optimization. When an auxiliary dataset is incorporated, RWSR-EDL achieves promising results and repulses any training time increment by adopting the noise-guidance data collection strategy. Extensive experiments show that RWSR-EDL achieves competitive performance over state-of-the-art methods on four in-the-wild image super-resolution datasets. Hao Li 0058, Jinghui Qin, Zhijing Yang, Pengxu Wei, Jinshan Pan, Liang Lin 0004, Yukai Shi |
IEEE Trans. Multim. | 4 |
| 2022 | Cross-modal Contrastive Attention Model for Medical Report GenerationabstractMedical report automatic generation has gained increasing interest recently as a way to help radiologists write reports more efficiently. However, this image-to-text task is rather challenging due to the typical data biases: 1) Normal physiological structures dominate the images, with only tiny abnormalities; 2) Normal descriptions accordingly dominate the reports. Existing methods have attempted to solve these problems, but they neglect to exploit useful information from similar historical cases. In this paper, we propose a novel Cross-modal Contrastive Attention (CMCA) model to capture both visual and semantic information from similar cases, with mainly two modules: a Visual Contrastive Attention Module for refining the unique abnormal regions compared to the retrieved case images; a Cross-modal Attention Module for matching the positive semantic information from the case reports. Extensive experiments on two widely-used benchmarks, IU X-Ray and MIMIC-CXR, demonstrate that the proposed model outperforms the state-of-the-art methods on almost all metrics. Further analyses also validate that our proposed model is able to improve the reports with more accurate abnormal findings and richer descriptions. Xiao Song 0003, Xiaodan Zhang 0003, Junzhong Ji, Pengxu Wei |
COLING | 5 |
| 2022 | Dual Adversarial Adaptation for Cross-Device Real-World Image Super-ResolutionabstractDue to the sophisticated imaging process, an identical scene captured by different cameras could exhibit distinct imaging patterns, introducing distinct proficiency among the super-resolution (SR) models trained on images from different devices. In this paper, we investigate a novel and practical task coded cross-device SR, which strives to adapt a real-world SR model trained on the paired images captured by one camera to low-resolution (LR) images captured by arbitrary target devices. The proposed task is highly challenging due to the absence of paired data from various imaging devices. To address this issue, we propose an unsupervised domain adaptation mechanism for real-world SR, named Dual ADversarial Adaptation (DADA), which only requires LR images in the target domain with available real paired data from a source camera. DADA employs the Domain-Invariant Attention (DIA) module to establish the basis of target model training even without HR supervision. Furthermore, the dual framework of DADA facilitates an Inter-domain Adversarial Adaptation (InterAA) in one branch for two LR input images from two domains, and an Intra-domain Adversarial Adaptation (IntraAA) in two branches for an LR input image. InterAA and IntraAA together improve the model transferability from the source domain to the target. We empirically conduct experiments under six$\text{Real} \rightarrow \text{Real}$adaptation settings among three different cameras, and achieve superior performance compared with existing state-of-the-art approaches. We also evaluate the proposed DADA to address the adaptation to the video camera, which presents a promising re-search topic to promote the wide applications of real-world super-resolution. Our source code is publicly available at https://github.com/lonelyhopeIDADA. Xiaoqian Xu, Pengxu Wei, Weikai Chen 0001, Yang Liu 0267, Mingzhi Mao, Liang Lin 0004, Guanbin Li |
CVPR | 2 |
| 2022 | Adversarially-Aware Robust Object Detector
Ziyi Dong, Pengxu Wei, Liang Lin 0004 |
ECCV (9) | 2 |
| 2022 | Cross-Domain Action Recognition via Prototypical Graph AlignmentabstractCompared with the well-explored cross-domain image recognition, cross-domain action recognition is a more challenging task because not only spatial but also temporal domain gaps exist across domains. Previous works attempt to bridge the temporal domain gap by aligning the domain-related key segments of videos from source and target domains. However, such practice overlooks the heterogeneous temporal domain gaps among different categories and presents temporal alignment strategies in a class-irrelevant manner. To address this issue, we propose to achieve class-wise temporal alignment for cross-domain action recognition via prototypical graph alignment (PGA). Concretely, we generate segment-level prototypes for the classes of both domains to capture per-class temporal dynamics. Furthermore, intra-domain and inter-domain prototypical graphs are established to mine the temporal relationships between each input video and its corresponding intra-domain and inter-domain prototypes. In this way, a discriminative and domain adaptive video representation is obtained by holistically reasoning cross-domain temporal dynamics. To class-wisely align the cross-domain video representations, each action category is equipped with a customized class-specific domain discriminator for temporal alignment via adversarial learning. Extensive experiments on three benchmarks show that PGA yeilds state-of-the-art performance on the task of cross-domain action recognition. Junhao Zhong, Pengxu Wei, Dongyu Zhang 0002, Liang Lin 0004 |
ICME | 3 |
| 2022 | Discovering Implicit Classes Achieves Open Set Domain AdaptationabstractIn Open Set Domain Adaptation (OSDA), large amounts of target samples are drawn from the implicit categories that never appear in the source domain. Due to the lack of their specific belonging, existing methods indiscriminately regard them as a single class “unknown”. We challenge this broadly-adopted practice that may arouse unexpected detrimental ef-fects because the decision boundaries between the implicit categories have been fully ignored. Instead, we propose Self-supervised Class-Discovering Adapter (SCDA) that attempts to achieve OSDA by gradually discovering those implicit classes, then incorporating them to restructure the classifier and update the domain-adaptive features iteratively. SCDA performs two alternate steps to achieve implicit class discov-ery and self-supervised OSDA, respectively. By jointly op-timizing for two tasks, SCDA achieves the state-of-the-art in OSDA and shows a competitive performance to unearth the implicit target classes. Jingyu Zhuang, Ziliang Chen 0001, Pengxu Wei, Guanbin Li, Liang Lin 0004 |
ICME | 3 |
| 2021 | Deductive Learning for Weakly-Supervised 3D Human Pose Estimation via Uncalibrated CamerasabstractWithout prohibitive and laborious 3D annotations, weakly-supervised 3D human pose methods mainly employ the model regularization with geometric projection consistency or geometry estimation from multi-view images. Nevertheless, those approaches explicitly need known parameters of calibrated cameras, exhibiting a limited model generalization in various realistic scenarios. To mitigate this issue, in this paper, we propose a Deductive Weakly-Supervised Learning (DWSL) for 3D human pose machine. Our DWSL firstly learns latent representations on depth and camera pose for 3D pose reconstruction. Since weak supervision usually causes ill-conditioned learning or inferior estimation, our DWSL introduces deductive reasoning to make an inference for the human pose from a view to another and develops a reconstruction loss to demonstrate what the model learns and infers is reliable. This learning by deduction strategy employs the view-transform demonstration and structural rules derived from depth, geometry and angle constraints, which improves the reliability of the model training with weak supervision. On three 3D human pose benchmarks, we conduct extensive experiments to evaluate our proposed method, which achieves superior performance in comparison with state-of-the-art weak-supervised methods. Particularly, our model shows an appealing potential for learning from 2D data captured in dynamic outdoor scenes, which demonstrates promising robustness and generalization in realistic scenarios. Our code is publicly available at https://github.com/Xipeng-Chen/DWSL-3D-pose. Xipeng Chen, Pengxu Wei, Liang Lin 0004 |
AAAI | 2 |
| 2021 | CDNet: Centripetal Direction Network for Nuclear Instance SegmentationabstractNuclear instance segmentation is a challenging task due to a large number of touching and overlapping nuclei in pathological images. Existing methods cannot effectively recognize the accurate boundary owing to neglecting the relationship between pixels (e.g., direction information). In this paper, we propose a novel Centripetal Direction Net-work (CDNet) for nuclear instance segmentation. Specifically, we define centripetal direction feature as a class of adjacent directions pointing to the nuclear center to rep-resent the spatial relationship between pixels within the nucleus. These direction features are then used to construct a direction difference map to represent the similarity within instances and the differences between instances. Finally, we propose a direction-guided refinement module, which acts as a plug-and-play module to effectively integrate auxiliary tasks and aggregate the features of different branches. Experiments on MoNuSeg and CPM17 datasets show that CDNet is significantly better than the other methods and achieves the state-of-the-art performance. The code is available at https://github.com/honglianghe/CDNet. Yao Ding 0006, Guoli Song, Lin Wang 0026, Qian Ren, Pengxu Wei, Jie Chen 0001 |
ICCV | 7 |
| 2021 | Trash to Treasure: Harvesting OOD Data with Cross-Modal Matching for Open-Set Semi-Supervised LearningabstractOpen-set semi-supervised learning (open-set SSL) investigates a challenging but practical scenario where out-of-distribution (OOD) samples are contained in the unlabeled data. While the mainstream technique seeks to completely filter out the OOD samples for semi-supervised learning (SSL), we propose a novel training mechanism that could effectively exploit the presence of OOD data for enhanced feature learning while avoiding its adverse impact on the SSL. We achieve this goal by first introducing a warm-up training that leverages all the unlabeled data, including both the in-distribution (ID) and OOD samples. Specifically, we perform a pretext task that enforces our feature extractor to obtain a high-level semantic understanding of the training images, leading to more discriminative features that can benefit the downstream tasks. Since the OOD samples are inevitably detrimental to SSL, we propose a novel cross-modal matching strategy to detect OOD samples. Instead of directly applying binary classification [39], we train the network to predict whether the data sample is matched to an assigned one-hot class label. The appeal of the proposed cross-modal matching over binary classification is the ability to generate a compatible feature space that aligns with the core classification task. Extensive experiments show that our approach substantially lifts the performance on open-set SSL and outperforms the state-of-the-art by a large margin. Chaowei Fang, Weikai Chen 0001, Zhenhua Chai, Xiaolin Wei, Pengxu Wei, Liang Lin 0004, Guanbin Li |
ICCV | 6 |
| 2021 | Hierarchical Transformer: Unsupervised Representation Learning for Skeleton-Based Human Action RecognitionabstractThe unsupervised representation learning for skeleton-based human action can be utilized in a variety of pose analysis applications. However, previous unsupervised methods focus on modeling the temporal dependencies in sequences, but take less effort in modeling the spatial structure in human action. To this end, we propose a novel unsupervised learning frame-work named Hierarchical Transformer for skeleton-based human action recognition. The Hierarchical Transformer consists of hierarchically aggregated self-attention modules for better capturing the spatial and temporal structure in the skeleton sequences. Furthermore, we propose to predict the motion between adjacent frames as a novel pre-training task for better capturing the long-term dependencies in sequences. Experimental results show that our method outperforms prior state-of-the-art unsupervised methods on NTU RGB+D and NW-UCLA datasets. Besides, our method also achieves state-of-the-art performance when the pre-trained model is transferred to SBU dataset, which demonstrates the generalizability of learned representation. Yi-Bin Cheng, Xipeng Chen, Pengxu Wei, Dongyu Zhang 0002, Liang Lin 0004 |
ICME | 4 |
| 2021 | Robust Real-World Image Super-Resolution against Adversarial AttacksabstractRecently deep neural networks (DNNs) have achieved significant success in real-world image super-resolution (SR). However, adversarial image samples with quasi-imperceptible noises could threaten deep learning SR models. In this paper, we propose a robust deep learning framework for real-world SR that randomly erases potential adversarial noises in the frequency domain of input images or features. The rationale is that on the SR task clean images or features have a different pattern from the attacked ones in the frequency domain. Observing that existing adversarial attacks usually add high-frequency noises to input images, we introduce a novel random frequency mask module that blocks out high-frequency components possibly containing the harmful perturbations in a stochastic manner. Since the frequency masking may not only destroys the adversarial perturbations but also affects the sharp details in a clean image, we further develop an adversarial sample classifier based on the frequency domain of images to determine if applying the proposed mask module. Based on the above ideas, we devise a novel real-world image SR framework that combines the proposed frequency mask modules and the proposed adversarial classifier with an existing super-resolution backbone network. Experiments show that our proposed method is more insensitive to adversarial attacks and presents more stable SR results than existing models and defenses. Jiutao Yue, Haofeng Li, Pengxu Wei, Guanbin Li, Liang Lin 0004 |
ACM Multimedia | 3 |
| 2021 | Deep CockTail Networks
Ziliang Chen 0001, Pengxu Wei, Jingyu Zhuang, Guanbin Li, Liang Lin 0004 |
Int. J. Comput. Vis. | 2 |
| 2021 | Deductive Reinforcement Learning for Visual Autonomous Urban Driving NavigationabstractExisting deep reinforcement learning (RL) are devoted to research applications on video games, e.g., The Open Racing Car Simulator (TORCS) and Atari games. However, it remains under-explored for vision-based autonomous urban driving navigation (VB-AUDN). VB-AUDN requires a sophisticated agent working safely in structured, changing, and unpredictable environments; otherwise, inappropriate operations may lead to irreversible or catastrophic damages. In this work, we propose a deductive RL (DeRL) to address this challenge. A deduction reasoner (DR) is introduced to endow the agent with ability to foresee the future and to promote policy learning. Specifically, DR first predicts future transitions through a parameterized environment model. Then, DR conducts self-assessment at the predicted trajectory to perceive the consequences of current policy resulting in a more reliable decision-making process. Additionally, a semantic encoder module (SEM) is designed to extract compact driving representation from the raw images, which is robust to the changes of the environment. Extensive experimental results demonstrate that DeRL outperforms the state-of-the-art model-free RL approaches on the public CAR Learning to Act (CARLA) benchmark and presents a superior performance on success rate and driving safety for goal-directed navigation. Changxin Huang, Meizi Ouyang, Pengxu Wei, Junfan Lin, Jiang Su, Liang Lin 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2020 | Component Divide-and-Conquer for Real-World Image Super-Resolution
Pengxu Wei, Ziwei Xie, Hannan Lu, Zongyuan Zhan, Qixiang Ye, Wangmeng Zuo, Liang Lin 0004 |
ECCV (8) | 1 |
| 2020 | 3D Human Pose Machines with Self-Supervised LearningabstractDriven by recent computer vision and robotic applications, recovering 3D human poses has become increasingly important and attracted growing interests. In fact, completing this task is quite challenging due to the diverse appearances, viewpoints, occlusions and inherently geometric ambiguities inside monocular images. Most of the existing methods focus on designing some elaborate priors /constraints to directly regress 3D human poses based on the corresponding 2D human pose-aware features or 2D pose predictions. However, due to the insufficient 3D pose data for training and the domain gap between 2D space and 3D space, these methods have limited scalabilities for all practical scenarios (e.g., outdoor scene). Attempt to address this issue, this paper proposes a simple yet effective self-supervised correction mechanism to learn all intrinsic structures of human poses from abundant images. Specifically, the proposed mechanism involves two dual learning tasks, i.e., the 2D-to-3D pose transformation and 3D-to-2D pose projection, to serve as a bridge between 3D and 2D human poses in a type of "free" self-supervision for accurate 3D human pose estimation. The 2D-to-3D pose implies to sequentially regress intermediate 3D poses by transforming the pose representation from the 2D domain to the 3D domain under the sequence-dependent temporal context, while the 3D-to-2D pose projection contributes to refining the intermediate 3D poses by maintaining geometric consistency between the 2D projections of 3D poses and the estimated 2D poses. Therefore, these two dual learning tasks enable our model to adaptively learn from 3D human pose data and external large-scale 2D human pose data. We further apply our self-supervised correction mechanism to develop a 3D human pose machine, which jointly integrates the 2D spatial relationship, temporal smoothness of predictions and 3D geometric knowledge. Extensive evaluations on the Human3.6M and HumanEva-I benchmarks demonstrate the superior performance and efficiency of our framework over all the compared competing methods. Keze Wang, Liang Lin 0004, Chenhan Jiang, Chen Qian 0006, Pengxu Wei |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2019 | Tumor Tissue Segmentation for Histopathological ImagesabstractHistopathological image analysis is considered as a gold standard for cancer identification and diagnosis. Tumor segmentation for histopathological images is one of the most important research topics and its performance directly affects the diagnosis judgment of doctors for cancer categories and their periods. With the remarkable development of deep learning methods, extensive methods have been proposed for tumor segmentation. However, there are few researches on analysis of specific pipeline of tumor segmentation. Moreover, few studies have done detailed research on the hard example mining of tumor segmentation. In order to bridge this gap, this study firstly summarize a specific pipeline of tumor segmentation. Then, hard example mining in tumor segmentation is also explored. Finally, experiments are conducted for evaluating segmentation performance of our method, demonstrating the effects of our method and hard example mining. Xiansong Huang, Pengxu Wei, Juncen Zhang, Jie Chen 0001 |
MMAsia | 3 |
| 2019 | Min-Entropy Latent Model for Weakly Supervised Object DetectionabstractWeakly supervised object detection is a challenging task when provided with image category supervision but required to learn, at the same time, object locations and object detectors. The inconsistency between the weak supervision and learning objectives introduces significant randomness to object locations and ambiguity to detectors. In this paper, a min-entropy latent model (MELM) is proposed for weakly supervised object detection. Min-entropy serves as a model to learn object locations and a metric to measure the randomness of object localization during learning. It aims to principally reduce the variance of learned instances and alleviate the ambiguity of detectors. MELM is decomposed into three components including proposal clique partition, object clique discovery, and object localization. MELM is optimized with a recurrent learning algorithm, which leverages continuation optimization to solve the challenging non-convexity problem. Experiments demonstrate that MELM significantly improves the performance of weakly supervised object detection, weakly supervised object localization, and image classification, against the state-of-the-art approaches. Fang Wan 0001, Pengxu Wei, Zhenjun Han, Jianbin Jiao, Qixiang Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Progressively diffused networks for semantic visual parsing
Ruimao Zhang, Wei Yang 0019, Zhanglin Peng, Pengxu Wei, Xiaogang Wang 0001, Liang Lin 0004 |
Pattern Recognit. | 4 |
| 2018 | Min-Entropy Latent Model for Weakly Supervised Object DetectionabstractWeakly supervised object detection is a challenging task when provided with image category supervision but required to learn, at the same time, object locations and object detectors. The inconsistency between the weak supervision and learning objectives introduces randomness to object locations and ambiguity to detectors. In this paper, a min-entropy latent model (MELM) is proposed for weakly supervised object detection. Min-entropy is used as a metric to measure the randomness of object localization during learning, as well as serving as a model to learn object locations. It aims to principally reduce the variance of positive instances and alleviate the ambiguity of detectors. MELM is deployed as two sub-models, which respectively discovers and localizes objects by minimizing the global and local entropy. MELM is unified with feature learning and optimized with a recurrent learning algorithm, which progressively transfers the weak supervision to object locations. Experiments demonstrate that MELM significantly improves the performance of weakly supervised detection, weakly supervised localization, and image classification, against the state-of-the-art approaches. Fang Wan 0001, Pengxu Wei, Jianbin Jiao, Zhenjun Han, Qixiang Ye |
CVPR | 2 |
| 2017 | Keyword-driven image captioning via Context-dependent Bilateral LSTMabstractImage captioning has recently received much attention. Existing approaches, however, are limited to describing images with simple contextual information, which typically generate one sentence to describe each image with only a single contextual emphasis. In this paper, we address this limitation from a user perspective with a novel approach. Given some keywords as additional inputs, the proposed method would generate various descriptions according to the provided guidance. Hence, descriptions with different focuses can be generated for the same image. Our method is based on a new Context-dependent Bilateral Long Short-Term Memory (CDB-LSTM) model to predict a keyword-driven sentence by considering the word dependence. The word dependence is explored externally with a bilateral pipeline, and internally with a unified and joint training process. Experiments on the MS COCO dataset demonstrate that the proposed approach not only significantly outperforms the baseline method but also shows good adaptation and consistency with various keywords. Xiaodan Zhang 0003, Shengfeng He, Xinhang Song, Pengxu Wei, Shuqiang Jiang, Qixiang Ye, Jianbin Jiao, Rynson W. H. Lau |
ICME | 4 |
| 2017 | Correlated Topic Vector for Scene ClassificationabstractScene images usually involve semantic correlations, particularly when considering large-scale image data sets. This paper proposes a novel generative image representation, correlated topic vector, to model such semantic correlations. Oriented from the correlated topic model, correlated topic vector intends to naturally utilize the correlations among topics, which are seldom considered in the conventional feature encoding, e.g., Fisher vector, but do exist in scene images. It is expected that the involvement of correlations can increase the discriminative capability of the learned generative model and consequently improve the recognition accuracy. Incorporated with the Fisher kernel method, correlated topic vector inherits the advantages of Fisher vector. The contributions to the topics of visual words have been further employed by incorporating the Fisher kernel framework to indicate the differences among scenes. Combined with the deep convolutional neural network (CNN) features and Gibbs sampling solution, correlated topic vector shows great potential when processing large-scale and complex scene image data sets. Experiments on two scene image data sets demonstrate that correlated topic vector improves significantly the deep CNN features, and outperforms existing Fisher kernel-based features. Pengxu Wei, Fang Wan 0001, Yi Zhu 0004, Jianbin Jiao, Qixiang Ye |
IEEE Trans. Image Process. | 1 |
| 2016 | Weakly supervised object detection with correlation and part suppressionabstractIn weakly supervised object detection, conventional methods treat object location in each image as a latent variable and use non-convex optimization to solve the latent variable. However, as the optimization objective is image-level instead of sample-level, the learning procedure tends to choose object parts as false positive samples. Furthermore, when multiple classes of objects appear in the same images, the models could invite class-correlations and lose discriminative capability. In this paper, we propose a simple but effective suppression strategy that mines hard negative samples in the learning procedure to ease the above problems. We propose using a spatial-voting strategy to help finding negative samples to suppress the impact of object parts. We also use regions from class-correlated images as negative samples to suppress the impact of class-correlations. Experiments show that our approach significantly improves the baseline by 6% and achieves state-of-the-art performance. Fang Wan 0001, Pengxu Wei, Zhenjun Han, Kun Fu 0001, Qixiang Ye |
ICIP | 2 |
| 2015 | Pedestrian detection via PCA filters based convolutional channel featuresabstractIn this paper, we propose a kind of image representation, named PCA filters based convolutional channel features (PCA-CCF) for pedestrian detection. The motivation is to use the convolutional network architecture with orthogonal PCA filters to enhance the state-of-the-art aggregate channel features (ACF). In PCA-CCF, the convolutional operation improves the feature robustness to pedestrian local deformation. The learned PCA filters reduce the correlations among features of each channel, and therefore, improve feature discrimination capability. With the proposed PCA-CCF features and cascaded AdaBoost classifiers, we develop a coarse-to-fine pedestrian detection approach. Experiments show that such approach achieves 3.04%, 17.87% and 6.28% performance gain on the INRIA, Caltech Reasonable and Caltech Overall pedestrian datasets, respectively. Wei Ke 0003, Pengxu Wei, Qixiang Ye, Jianbin Jiao |
ICASSP | 3 |