Wenjing Li 0005

dblp:08/6548-5 · DBLP profile ↗
← Back
20ranked-venue papers
4as first author
17since 2021 · last 2026
0000-0003-3201-6675ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 WeatherEdit: Controllable Weather Editing with 4D Gaussian Field
abstract
In this work, we present WeatherEdit, a novel weather editing pipeline for generating realistic weather effects with controllable types and severity in 3D scenes. Our approach is structured into two key components: weather background editing and weather particle construction. For weather background editing, we introduce an all-in-one adapter that integrates multiple weather styles into a single diffusion model, enabling the generation of diverse weather effects in 2D image backgrounds. During inference, we design a Temporal-View (TV-) attention mechanism that follows a specific order to aggregate temporal and spatial information, ensuring consistent editing across multi-frame and multi-view images. To construct the weather particles, we first reconstruct a 3D scene using the edited images and then introduce a 4D Gaussian field to generate snowflakes, raindrops and fog in the scene. The attributes and dynamics of these particles are controlled through attribute modelling and dynamic simulation, ensuring realistic weather representation and flexible severity adjustments. Finally, we integrate the 4D Gaussian field with the 3D scene to render consistent and highly realistic weather effects. Experiments on multiple driving datasets demonstrate that WeatherEdit can generate diverse weather effects with controllable condition severity, highlighting its potential for autonomous driving simulation in adverse weather.
Chenghao Qian, Wenjing Li 0005, Yuhu Guo, Gustav Markkula
AAAI2
2026 Open-World Deepfake Attribution via Confidence-Aware Asymmetric Learning
abstract
The proliferation of synthetic facial imagery has intensified the need for robust Open-World DeepFake Attribution (OW-DFA), which aims to attribute both known and unknown forgeries using labeled data for known types and unlabeled data containing a mixture of known and novel types. However, existing OW-DFA methods face two critical limitations: 1) A confidence skew that leads to unreliable pseudo-labels for novel forgeries, resulting in biased training. 2) An unrealistic assumption that the number of unknown forgery types is known a priori. To address these challenges, we propose a Confidence-aware Asymmetric Learning (CAL) framework, which adaptively balances model confidence across known and novel forgery types. CAL mainly consists of two components: Confidence-aware Consistency Regularization (CCR) and Asymmetric Confidence Reinforcement (ACR). CCR mitigates pseudo-label bias by dynamically scaling sample losses based on normalized confidence, gradually shifting the training focus from high- to low-confidence samples. ACR complements this by separately calibrating confidence for known and novel classes through selective learning on high-confidence samples, guided by their confidence gap. Together, CCR and ACR form a mutually reinforcing loop that significantly improves the model's OW-DFA performance. Moreover, we introduce a Dynamic Prototype Pruning (DPP) strategy that automatically estimates the number of novel forgery types in a coarse-to-fine manner, removing the need for unrealistic prior assumptions and enhancing the scalability of our methods to real-world OW-DFA scenarios. Extensive experiments on the standard and OW-DFA benchmark and a newly extended benchmark incorporating advanced manipulations demonstrate that CAL consistently outperforms previous methods, achieving new state-of-the-art performance on both known and novel forgery attribution.
Haiyang Zheng, Nan Pu, Wenjing Li 0005, Nicu Sebe, Zhun Zhong
AAAI3
2025 Generate, Refine, and Encode: Leveraging Synthesized Novel Samples for On-the-Fly Fine-Grained Category Discovery
abstract
In this paper, we investigate a practical yet challenging task: On-the-fly Category Discovery (OCD). This task focuses on the online identification of newly arriving stream data that may belong to both known and unknown categories, utilizing the category knowledge from only labeled data. Existing OCD methods are devoted to fully mining transferable knowledge from only labeled data. However, the transferability learned by these methods is limited because the knowledge contained in known categories is often insufficient, especially when few annotated data/categories are available in fine-grained recognition. To mitigate this limitation, we propose a diffusion-based OCD framework, dubbed DiffGRE, which integrates Generation, Refinement, and Encoding in a multi-stage fashion. Specifically, we first design an attribute-composition generation method based on cross-image interpolation in the diffusion latent space to synthesize novel samples. Then, we propose a diversity-driven refinement approach to select the synthesized images that differ from known categories for subsequent OCD model training. Finally, we leverage a semi-supervised leader encoding to inject additional category knowledge contained in synthesized data into the OCD models, which can benefit the discovery of both known and unknown categories during the on-the-fly inference process. Extensive experiments demonstrate the superiority of our DiffGRE over previous methods on six fine-grained datasets.
Nan Pu, Haiyang Zheng, Wenjing Li 0005, Nicu Sebe, Zhun Zhong
ICCV4
2025 Weathergs: 3D Scene Reconstruction in Adverse Weather Conditions Via Gaussian Splatting
abstract
D Gaussian Splatting (3DGS) has gained significant attention for 3D scene reconstruction, but still suffers from complex outdoor environments, especially under adverse weather. This is because 3DGS treats the artifacts caused by adverse weather as part of the scene and will directly reconstruct them, largely reducing the clarity of the reconstructed scene. To address this challenge, we propose WeatherGS, a 3DGSbased framework for reconstructing clear scenes from multiview images under different weather conditions. Specifically, we explicitly categorize the multi-weather artifacts into the dense particles and lens occlusions that have very different characters, in which the former are caused by snowflakes and raindrops in the air, and the latter are raised by the precipitation on the camera lens. In light of this, we propose a dense-to-sparse preprocess strategy, which sequentially removes the dense particles by an Atmospheric Effect Filter (AEF) and then extracts the relatively sparse occlusion masks with a Lens Effect Detector (LED). Finally, we train a set of 3D Gaussians by the processed images and generated masks for excluding occluded areas, and accurately recover the underlying clear scene by Gaussian splatting. We conduct a diverse and challenging benchmark to facilitate the evaluation of 3D reconstruction under complex weather scenarios. Extensive experiments on this benchmark demonstrate that our WeatherGS consistently produces high-quality, clean scenes across various weather scenarios, outperforming existing state-of-the-art methods. See project: https://jumponthemoon.github.io/weather-gs.
Chenghao Qian, Yuhu Guo, Wenjing Li 0005, Gustav Markkula
ICRA3
2024 Federated Generalized Category Discovery
abstract
Generalized category discovery (GCD) aims at grouping unlabeled samples from known and unknown classes, given labeled data of known classes. To meet the recent decen-tralization trend in the community, we introduce a practical yet challenging task, Federated GCD (Fed-GCD), where the training data are distributed among local clients and cannot be shared among clients. Fed-GCD aims to train a generic GCD model by client collaboration under the privacy-protected constraint. The Fed-GCD leads to two challenges: 1) representation degradation caused by training each client model with fewer data than centralized GCD learning, and 2) highly heterogeneous label spaces across different clients. To this end, we propose a novel Asso-ciated Gaussian Contrastive Learning (AGCL) framework based on learnable GMMs, which consists of a Client Se-mantics Association (CSA) and a global-local GMM Contrastive Learning (GCL). On the server, CSA aggregates the heterogeneous categories of local-client GMMs to generate a global GMM containing more comprehensive category knowledge. On each client, GCL builds class-level contrastive learning with both local and global GMMs. The local GCL learns robust representation with limited local data. The global GCL encourages the model to produce more discriminative representation with the comprehensive category relationships that may not exist in local data. We build a benchmark based on six visual datasets to facilitate the study of Fed-GCD. Extensive experiments show that our AGCL outperforms multiple baselines on all datasets. Code is available at https://github.com/TPCD/FedGCD.
Nan Pu, Wenjing Li 0005, Xingyuan Ji, Yalan Qin, Nicu Sebe, Zhun Zhong
CVPR2
2024 Learning to Distinguish Samples for Generalized Category Discovery
Fengxiang Yang, Nan Pu, Wenjing Li 0005, Zhiming Luo, Shaozi Li, Nicu Sebe, Zhun Zhong
ECCV (65)3
2024 Textual Knowledge Matters: Cross-Modality Co-teaching for Generalized Visual Class Discovery
Haiyang Zheng, Nan Pu, Wenjing Li 0005, Nicu Sebe, Zhun Zhong
ECCV (52)3
2024 Democratizing Fine-grained Visual Recognition with Large Language Models
abstract
Identifying subordinate-level categories from images is a longstanding task in computer vision and is referred to as fine-grained visual recognition (FGVR). It has tremendous significance in real-world applications since an average layperson does not excel at differentiating species of birds or mushrooms due to subtle differences among the species. A major bottleneck in developing FGVR systems is caused by the need of high-quality paired expert annotations. To circumvent the need of expert knowledge we propose Fine-grained Semantic Category Reasoning (FineR) that internally leverages the world knowledge of large language models (LLMs) as a proxy in order to reason about fine-grained category names. In detail, to bridge the modality gap between images and LLM, we extract part-level visual attributes from images as text and feed that information to a LLM. Based on the visual attributes and its internal world knowledge the LLM reasons about the subordinate-level category names. Our training-free FineR outperforms several state-of-the-art FGVR and language and vision assistant models and shows promise in working in the wild and in new domains where gathering expert annotation is arduous.
Subhankar Roy, Wenjing Li 0005, Zhun Zhong, Nicu Sebe, Elisa Ricci 0001
ICLR3
2024 AllWeather-Net: Unified Image Enhancement for Autonomous Driving Under Adverse Weather and Low-Light Conditions
Chenghao Qian, Mahdi Rezaei 0001, Saeed Anwar, Wenjing Li 0005, Tanveer Hussain 0001, Mohsen Azarmi, Wei Wang 0335
ICPR (30)4
2024 Prototypical Hash Encoding for On-the-Fly Fine-Grained Category Discovery
abstract
In this paper, we study a practical yet challenging task, On-the-fly Category Discovery (OCD), aiming to online discover the newly-coming stream data that belong to both known and unknown classes, by leveraging only known category knowledge contained in labeled data. Previous OCD methods employ the hash-based technique to represent old/new categories by hash codes for instance-wise inference. However, directly mapping features into low-dimensional hash space not only inevitably damages the ability to distinguish classes and but also causes ``high sensitivity'' issue, especially for fine-grained classes, leading to inferior performance. To address these drawbacks, we propose a novel Prototypical Hash Encoding (PHE) framework consisting of Category-aware Prototype Generation (CPG) and Discriminative Category Encoding (DCE) to mitigate the sensitivity of hash code while preserving rich discriminative information contained in high-dimension feature space, in a two-stage projection fashion. CPG enables the model to fully capture the intra-category diversity by representing each category with multiple prototypes. DCE boosts the discrimination ability of hash code with the guidance of the generated category prototypes and the constraint of minimum separation distance. By jointly optimizing CPG and DCE, we demonstrate that these two components are mutually beneficial towards an effective OCD. Extensive experiments show the significant superiority of our PHE over previous methods, e.g. obtaining an improvement of +5.3% in ALL ACC averaged on all datasets. Moreover, due to the nature of the interpretable prototypes, we visually analyze the underlying mechanism of how PHE helps group certain samples into either known or unknown categories. Code is available at https://github.com/HaiyangZheng/PHE.
Haiyang Zheng, Nan Pu, Wenjing Li 0005, Nicu Sebe, Zhun Zhong
NeurIPS3
2024 Strong-Help-Weak: An Online Multi-Task Inference Learning Approach for Robust Advanced Driver Assistance Systems
abstract
Multi-task learning in advanced driver assistance systems aims to endow models with the capacity to jointly handle multiple related tasks, such as object detection, depth estimation, and more. However, existing multi-task learning models largely rely on the extensive number of labelled data. In practice, the process of annotating data for multi-task training proves to be exceedingly costly, yet not always accurate. This study introduces an innovative setting named online multi-task inference learning that updates the multi-task model during inference. And we propose a Strong-Help-Weak (SHW) framework which aims to enhance weaker (or more challenging) tasks by leveraging guidance from closely related stronger (or easier) tasks. Specifically, we first build two benchmarks based on KITTI and BDD with four tasks (object detection, object depth estimation, lane line segmentation, and driving area segmentation). Then, we propose two novel modules inspired by two priors: 1) Detection-guided Depth Inference Learning (DetDis) module that leverages the inverse relationship between object size and distance to refine the predicted object distance; and 2) Area-guided Lane Line Inference Learning (AreaLane) module that utilises inclusion relationship between driving area and lane line to infer more accurate lane line. Both modules are efficient and can provide more reliable supervision for the corresponding weaker tasks (object distance estimation and lane line segmentation), respectively. Extensive experiments on the two benchmarks show that our SHW can obtain consistent improvements on the weaker tasks during the inference stage with low computational costs.
Wenjing Li 0005, Jian Kuang 0005, Jun Zhang 0034, ZhongCheng Wu, Mahdi Rezaei 0001
IEEE Trans. Intell. Transp. Syst.2
2024 MIFI: MultI-Camera Feature Integration for Robust 3D Distracted Driver Activity Recognition
abstract
Distracted driver activity recognition plays a critical role in risk aversion-particularly beneficial in intelligent transportation systems. However, most existing methods make use of only the video from a single view and the difficulty-inconsistent issue is neglected. Different from them, in this work, we propose a novel MultI-camera Feature Integration (MIFI) approach for 3D distracted driver activity recognition by jointly modeling the data from different camera views and explicitly re-weighting examples based on their degree of difficulty. Our contributions are two-fold: (1) We propose a simple but effective multi-camera feature integration framework and provide three types of feature fusion techniques. (2) To address the difficulty-inconsistent problem in distracted driver activity recognition, a periodic learning method, named example re-weighting that can jointly learn the easy and hard samples, is presented. The experimental results on the 3MDAD dataset demonstrate that the proposed MIFI can consistently boost performance compared to single-view models.
Jian Kuang 0005, Wenjing Li 0005, Jun Zhang 0034, ZhongCheng Wu
IEEE Trans. Intell. Transp. Syst.2
2023 YOLOv4-dense: A smaller and faster YOLOv4 for real-time edge-device based object detection in traffic scene
abstract
Abstract Edge‐device‐based object detection is crucial in many real‐world applications, such as self‐driving cars, ADAS, driver behavior analysis. Although deep learning (DL) has become the de‐facto approach for object detection, the limited computing resources of embedded devices and the large model size of current DL‐based methods increase the difficulty of real‐time object detection on edge devices. To overcome these difficulties, in this work a novel YOLOv4‐dense model is proposed to detect objects in an accurate, fast manner, which is built on top of the YOLOv4 framework but with substantial improvements. More specifically, lots of CSP layers are pruned since it will decrease inference speed. And to address the losing small objects problem, a dense block is introduced. In addition, a lightweight two‐stream YOLO head is also designed to further reduce the computational complexity of the model. Experimental results on NVIDIA JETSON TX2 embedded platform demonstrate that YOLOv4‐dense can achieve a higher accuracy, faster speed with smaller model size. For instance, on the KITTI dataset, YOLOv4‐dense obtains 84.3% mAP and 22.6 FPS with only 20.3 M parameters, surpassing the state‐of‐the‐art models with comparable parameter budget such as YOLOv3‐tiny, YOLOv4‐tiny, PP‐YOLO‐tiny by a large margin.
Wenjing Li 0005, Jun Zhang 0034, ZhongCheng Wu
IET Image Process.2
2023 100-Driver: A Large-Scale, Diverse Dataset for Distracted Driver Classification
abstract
Distracted driver classification (DDC) plays an important role in ensuring driving safety. Although many datasets are introduced to support the study of DDC, most of them are small in data size and are short of diversity in environmental variations. This largely limits the development of DDC since many practical problems such as the cross-modality setting cannot be fully studied. In this paper, we introduce 100-Driver, a large-scale, diverse posture-based distracted diver dataset, with more than 470K images taken by 4 cameras observing 100 drivers over 79 hours from 5 vehicles. 100-Driver involves different types of variations that closely meet real-world applications, including changes in the vehicle, person, camera view, lighting, and modality. We provide a detailed analysis of 100-Driver and present 4 settings for investigating practical problems of DDC, including the traditional setting without domain shift and 3 challenging settings (i.e., cross-modality, cross-view, and cross-vehicle) with domain shifts. We conduct comprehensive experiments on these 4 settings with state-the-of-art techniques and show several insights to the future study of DDC. Our 100-Driver will be publicly available offering new opportunities to advance the development of DDC. The 100-driver dataset, source code, and evaluation protocols are available athttps://100-driver.github.io.
Jing Wang 0092, Wenjing Li 0005, Jun Zhang 0034, ZhongCheng Wu, Zhun Zhong, Nicu Sebe
IEEE Trans. Intell. Transp. Syst.2
2023 CarDD: A New Dataset for Vision-Based Car Damage Detection
abstract
Automatic car damage detection has attracted significant attention in the car insurance business. However, due to the lack of high-quality and publicly available datasets, we can hardly learn a feasible model for car damage detection. To this end, we contribute with Car Damage Detection (CarDD), the first public large-scale dataset designed for vision-based car damage detection and segmentation. Our CarDD contains 4,000 high-resolution car damage images with over 9,000 well-annotated instances of six damage categories. We detail the image collection, selection, and annotation processes, and present a statistical dataset analysis. Furthermore, we conduct extensive experiments on CarDD with state-of-the-art deep methods for different tasks and provide comprehensive analyses to highlight the specialty of car damage detection. CarDD dataset and the source code are available athttps://cardd-ustc.github.io.
Xinkuang Wang, Wenjing Li 0005, ZhongCheng Wu
IEEE Trans. Intell. Transp. Syst.2
2022 Learning Accurate, Speedy, Lightweight CNNs via Instance-Specific Multi-Teacher Knowledge Distillation for Distracted Driver Posture Identification
abstract
For deployment on an embedded processor for distracted driver classification, the model should satisfy the demand for both high accuracy, real-time inference, and limited storage resources. Conventional deep CNN models such as VGG, ResNet, DenseNet, often aim for high accuracy, making their model heavy for an embedded system with limited memory space and computing resources. In contrast, lightweight models are greatly compressed but at a significant sacrifice of accuracy. To bridge this gap, we propose an instance-specific multi-teacher knowledge distillation model (IsMt-KD) to learn more accurate, speedy, and lightweight CNNs for distracted driver posture classification. Specifically, in multi-teacher knowledge distillation, most of the current approaches either randomly select a teacher model and apply the prediction of such teacher model as the soft-label or allocate an equal weight to every teacher model and average all the predictions of the teachers as the soft label. In this paper, we observe that, when facing the same instance, the outputs of different teachers vary greatly, in which some teachers can predict it right whereas the others may give pretty high probabilities to the irrelevant classes. Thus, it is inappropriate to set fixed weights or the same weights for teachers. To this end, a simple yet effective instance-specific teacher grading module is designed to dynamically assign weights to teacher models based on individual instances. In this way, we can dynamically distill the knowledge from multiple teachers by considering both instance-specific high-level and instance-specific intermediate-level information. Our extensive experimental results on AUC and StateFarm datasets, and our implementation on edge hardware platforms including HUAWEI MediaPad c5 and Nvidia Jetson TX2, verify the effectiveness and feasibility of our approach.
Wenjing Li 0005, Jing Wang 0092, Tingting Ren, Jun Zhang 0034, ZhongCheng Wu
IEEE Trans. Intell. Transp. Syst.1
2021 Contextual similarity-based multi-level second-order attention network for semi-supervised few-shot learning
Wenjing Li 0005, Tingting Ren, Jun Zhang 0034, ZhongCheng Wu
Neurocomputing1
2020 OVL: One-View Learning for Human Retrieval
abstract
This paper considers a novel problem, named One-View Learning (OVL), in human retrieval a.k.a. person re-identification (re-ID). Unlike fully-supervised learning, OVL only requires pretty cheap annotation cost: labeled training images are only provided from one camera view (source view/domain), while the annotations of training images from other camera views (target views/domains) are not available. OVL is a problem of multi-target open set domain adaptation that is difficult for existing domain adaptation methods to handle. This is because 1) unlabeled samples are drawn from multiple target views in different distributions, and 2) the target views may contain samples of “unknown identity” that are not shared by the source view. To address this problem, this work introduces a novel one-view learning framework for person re-ID. This is achieved by adversarial multi-view learning (AMVL) and adversarial unknown rejection learning (AURL). The former learns a multi-view discriminator by adversarial learning to align the feature distributions between all views. The later is designed to reject unknown samples from target views through adversarial learning with two unknown identity classifiers. Extensive experiments on three large-scale datasets demonstrate the advantage of the proposed method over state-of-the-art domain adaptation and semi-supervised methods.
Wenjing Li 0005, ZhongCheng Wu
AAAI1
2020 LGSim: local task-invariant and global task-specific similarity for few-shot classification
Wenjing Li 0005, ZhongCheng Wu, Jun Zhang 0034, Tingting Ren
Neural Comput. Appl.1
2018 Cross-covariance regularized autoencoders for nonredundant sparse feature representation
Jie Chen 0035, ZhongCheng Wu, Jun Zhang 0034, Wenjing Li 0005
Neurocomputing5