Wenchao Du

dblp:98/9705 · DBLP profile ↗
← Back
18ranked-venue papers
8as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 6 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Hybrid-Domain Adaptative Representation Learning for Gaze Estimation
abstract
Appearance-based gaze estimation, aiming to predict accurate 3D gaze direction from a single facial image, has made promising progress in recent years. However, most methods suffer significant performance degradation in cross-domain evaluation due to interference from gaze-irrelevant factors, such as expressions, wearables, and image quality. To alleviate this problem, we present a novel Hybrid-domain Adaptative Representation Learning (shorted by HARL) framework that exploits multi-source hybrid datasets to learn robust gaze representation. More specifically, we propose to disentangle gaze-relevant representation from low-quality facial images by aligning features extracted from high-quality near-eye images in an unsupervised domain-adaptation manner, which hardly requires any computational or inference costs. Additionally, we analyze the effect of head-pose and design a simple yet efficient sparse graph fusion module to explore the geometric constraint between gaze direction and head-pose, leading to a dense and robust gaze representation. Extensive experiments on EyeDiap, MPIIFaceGaze, and Gaze360 datasets demonstrate that our approach achieves state-of-the-art accuracy of 5.02, 3.36, and 9.26 degrees respectively, and present competitive performances through cross-dataset evaluation.
Qida Tan, Wenchao Du
AAAI3
2026 Applying Fused Data to Predict Vessel Traffic Flow: A Hybrid-Model Deep Learning Approach
abstract
In the research on navigation efficiency optimization in complex waters, accurate vessel traffic flow prediction has emerged as a critical challenge in the field of intelligent maritime navigation. Based on Resolution A.1158(32) of the International Maritime Organization (IMO) and expert knowledge of Vessel Traffic Service (VTS), this study proposes a novel framework for vessel port reporting and VTS decision-making, and constructs a feature-fused database integrated with meteorological data. A hybrid deep learning model based on the DeepAR framework is developed, which extracts the spatial dependencies of vessel traffic via a convolutional neural network (CNN), captures temporal features using a bidirectional long short-term memory network (BiLSTM), and acquires long-range dependencies with a self-attention mechanism (SAM) to achieve multi-dimensional feature fusion for vessel traffic flow prediction. A case study is conducted on the narrow section of the Lüsi inbound waterway in Tongzhou Bay, China, to verify the effectiveness of the proposed model. The results demonstrate that the proposed method can realize high-precision probabilistic prediction of vessel traffic flow. Compared with the benchmark methods, the proposed model reduces both the Mean Absolute Scaled Error (MASE) and the Mean Absolute Error (MAE) simultaneously for medium- and long-term forecasting ( > 24 hours) under specific working conditions (VTS six-shift mode). This study can provide probabilistic decision support for intelligent traffic management in narrow navigable waters, and serve as an important reference for VTS scheduling as well as port and shipping operation dispatching.
Wenchao Du, Tianyou Chen, Chong Ni
IEEE Trans. Intell. Transp. Syst.1
2025 Learning Geometry-Aware Representation for Gaze Estimation
abstract
Appearance-based gaze estimation has achieved remarkable progress in recent years. However, the inherent geometry characteristics of eye and facial areas are not fully explored in existing methods, which limits the generalization and robustness of the model. In this paper, we propose a novel end-to-end framework for cross-domain gaze estimation by integrating latent geometric representation into appearance-based gaze framework. More specifically, we first exploit the 3DMM method to fit unconstrained faces and eyes in the wild, which would generate adaptive normal information with explicit 3D geometry prior. Then we joint the normal map and the corresponding RGB appearance information to infer the 3D gaze direction with carefully-designed spatial-frequency attention and local-global feature interaction modules. The key to our method is to integrate explicit 3D geometry representation into a 2D learning architecture, which leads to a better trade-off between performance and efficiency. Experiments on both MPIIGaze and EyeDiap datasets demonstrate that the proposed method achieves the state-of-the-art accuracy of 3.56° and 5.10° separately, and also presents superior generalization ability on cross-domain dataset evaluations.
Qida Tan, Wenchao Du, Hu Chen 0002, Hongyu Yang 0002
ICIP3
2025 Geometry-Aware Appearance Learning for Generalized Gaze Estimation
Qida Tan, Wenchao Du, Hu Chen 0002, Hongyu Yang 0002
IEEE Signal Process. Lett.3
2025 Solving Zero-Shot Sparse-View CT Reconstruction With Variational Score Solver
abstract
Computed tomography (CT) stands as a ubiquitous medical diagnostic tool. Nonetheless, the radiation-related concerns associated with CT scans have raised public apprehensions. Mitigating radiation dosage in CT imaging poses an inherent challenge as it inevitably compromises the fidelity of CT reconstructions, impacting diagnostic accuracy. While previous deep learning techniques have exhibited promise in enhancing CT reconstruction quality, they remain hindered by the reliance on paired data, which is arduous to procure. In this study, we present a novel approach named Variational Score Solver (VSS) for sparse-view reconstruction without paired data. Our approach entails the acquisition of a probability distribution from densely sampled CT reconstructions, employing a latent diffusion model. High-quality reconstruction outcomes are achieved through an iterative process, wherein the diffusion model serves as the prior term, subsequently integrated with the data consistency term. Notably, rather than directly employing the prior diffusion model, we distill prior knowledge by finding the fixed point of the diffusion model. This framework empowers us to exercise precise control over the process. Moreover, we depart from modeling the reconstruction outcomes as deterministic values, opting instead for a distribution-based approach. This enables us to achieve more accurate reconstructions utilizing a trainable model. Our approach introduces a fresh perspective to the realm of zero-shot CT reconstruction, circumventing the constraints of supervised learning. Extensive qualitative and quantitative experiments unequivocally demonstrate that VSS surpasses other contemporary unsupervised and achieves comparable results compared to the most advanced supervised methods in sparse-view reconstruction tasks. Codes are available in https://github.com/fpsandnoob/vss.
Linchao He, Wenchao Du, Peixi Liao, Fenglei Fan, Hu Chen 0002, Hongyu Yang 0002, Yi Zhang 0018
IEEE Trans. Medical Imaging2
2024 A Progressively Prompt-guided Model for Sparse-View CT Reconstruction
abstract
While sparse-view Computed Tomography (CT) has a remarkable impact on reducing ionizing radiation dose while accelerating data acquisition, the reconstructed images have been compromised by streak-like artifacts, affecting clinical diagnostics. By integrating powerful regularization with deep learning technologies into iterative reconstruction algorithms, the deep-unrolling-based methods have achieved promising results in terms of reconstruction quality and theoretical interpretability. However, leading methods always focus on learning powerful content priors with diverse technologies and ignoring the latent noise distribution prior in the image domain, thereby limiting the ability of structure-preserving and detail reconstructing of the model. To alleviate this problem, we propose a Progressively Prompt-guided Model (shorted by PPM) for sparse-view CT reconstruction. Specifically, we inject the idea of prompt learning into an iterative unrolled neural network, in which a learnable prompt module is inserted into each unrolled block to perceive image content and noise distribution in a self-adaptive manner, which leads to the more powerful priors to guide high-quality CT image reconstruction. Furthermore, we construct a progressively guiding strategy to facilitate high-quality prompt generation while speeding model convergence. Extensive experiments demonstrate that our PPM achieves state-of-the-art performance in artifact suppression, structure fidelity, and visual perception similarity. The code is available at https://github.com/Wenchao-Du/PPM/.
Qiao Mu, Hu Chen 0002, Wenchao Du, Hongyu Yang 0002
BIBM5
2024 Dtpose: Learning Disentangled Token Representation For Effective Human Pose Estimation
abstract
Exploring rich visual clues and spatial geometric constraints to locate keypoints is essential for human pose estimation. Existing Transformer-based methods have presented unique advantages via token representation, where each keypoint is explicitly embedded as a token to learn visual appearance clues and geometric relationships simultaneously from images. However, it is difficult to learn powerful pose representation via self-attention mechanism due to latent interference, e.g., blurring and self-occlusion. To alleviate this challenge, we present a novel framework that Disentangles hybrid Token representation to explore more effective visual and keypoint information for Pose estimation (termed by DTPose). In detail, DTPose contains two key modules. First, the Disentangled Token Representation module is used to explore visual clues and geometry constraints sequentially, which alleviates the noise interference and enables the geometry and appearance clues to be exploited more sufficiently. Furthermore, the Hierarchical Spatial Decoding head is exploited to preserve the 2 D geometric structure information of keypoints as much as possible. Extensive experiments on COCO dataset demonstrate significant performance gains of our DTPose, which achieves 76.5 (+0.7) AP and 75.7 (+0.6) AP than the TokenPose-L on the COCO validation and test-dev sets separately.
Shiyang Ye, Hu Chen 0002, Wenchao Du, Hongyu Yang 0002
ICIP5
2024 Memory Coordinated Cross Perception for Few-Shot Object Detection
abstract
Significant advances have been made in the field of few-shot object detection. Few-shot object detection involves using limited examples to detect new categories. However, as the mainstream approach in FSOD, meta-learning methods face two main challenges: the prototypes they create lack sufficient representativeness and insufficient distinction between different prototypes. These issues stem from two factors. First, the limited quantity and diversity of support instances make it hard for the model to accurately capture the prototype’s essence. Second, similar categories complicate matching query features to the correct prototype. To address these issues, we propose an adaptive memory coordinated generation method. This method distinguishes each similar prototype while generating more representative prototypes using memory prototypes. Specifically, it enhances the representational information of support features through their interaction with memory prototypes. This approach also achieves information perception and adaptive fusion among support features, leading to the generation of easily distinguishable and more identifiable prototypes. Furthermore, we developed a prototype cross-perception module. This module consistently refines prototypes and query features, achieving deep integration of features while preserving the essential information of query features. It enhances prototypes’ directive function and improves query feature robustness. Our model demonstrates leading performance across most shot settings and evaluation metrics on multiple few-shot object detection benchmarks.
Yunfeng Kou, Hu Chen 0002, Wenchao Du
IJCNN5
2024 FSAD:Few-Shot Object Detection via Aggregation and Disentanglement
abstract
Few-shot Object Detection (FSOD) aims to leverage knowledge gained from general object detectors to enhance future detection tasks for novel object categories. In response to the poor performance observed in commonly used attention-based feature fusion methods, particularly in 1-shot, this paper proposes an enhanced Cross-Attention-Like Aggregation (CAL) Module and a GAN-Disentangled-Like Feature Enhancement (GDL) Module.CAL module utilizes an asymmetric mechanism and neutralized features resulting from concatenation for cross-attention, which improves the model’s generalization ability, enabling it to address issues where the target is entirely unrecognizable in 1-shot. The GDL module captures latent independent information from the support set. It employs a discriminator to stabilize the information extraction process, enabling the attention module to discern relevant information more effectively and accurately guide the query features.Extensive experiments conducted on the PASCAL VOC and COCO datasets demonstrate the superior performance of our approach over strong baselines, showcasing substantial advancements in one-shot and two-shot performance.our code is available at https://github.com/yun1232/FSAD
Yunfeng Kou, Kunming Wu, Hu Chen 0002, Wenchao Du
IJCNN5
2024 Hierarchical disentangled representation for image denoising and beyond
Wenchao Du, Hu Chen 0002, Yi Zhang 0018, Hongyu Yang 0002
Image Vis. Comput.1
2022 Learning Interval-Aware Embedding for Macro and Micro-expression Spotting
Wenchao Du, Hu Chen 0002, Hongyu Yang 0002
ACCV (4)3
2022 Depth Completion Using Geometry-Aware Embedding
abstract
Exploiting internal spatial geometric constraints of sparse LiDARs is beneficial to depth completion, however, has been not explored well. This paper proposes an efficient method to learn geometry-aware embedding, which encodes the local and global geometric structure information from 3D points, e.g., scene layout, object's sizes and shapes, to guide dense depth estimation. Specifically, we utilize the dynamic graph representation to model generalized geometric relationship from irregular point clouds in a flexible and efficient manner. Further, we joint this embedding and corresponded RGB appearance information to infer missing depths of the scene with well structure-preserved details. The key to our method is to integrate implicit 3D geometric representation into a 2D learning architecture, which leads to a better trade-off between the performance and efficiency. Extensive experiments demonstrate that the proposed method outperforms previous works and could reconstruct fine depths with crisp boundaries in regions that are over-smoothed by them. The ablation study gives more insights into our method that could achieve significant gains with a simple design, while having better generalization capability and stability. The code is available at https://github.com/Wenchao-Du/GAENet.
Wenchao Du, Hu Chen 0002, Hongyu Yang 0002, Yi Zhang 0018
ICRA1
2020 Learning Invariant Representation for Unsupervised Image Restoration
abstract
Recently, cross domain transfer has been applied for unsupervised image restoration tasks. However, directly applying existing frameworks would lead to domain-shift problems in translated images due to lack of effective supervision. Instead, we propose an unsupervised learning method that explicitly learns invariant presentation from noisy data and reconstructs clear observations. To do so, we introduce discrete disentangling representation and adversarial domain adaption into general domain transfer framework, aided by extra self-supervised modules including background and semantic consistency constraints, learning robust representation under dual domain constraints, such as feature and image domains. Experiments on synthetic and real noise removal tasks show the proposed method achieves comparable performance with other stateof-the-art supervised and unsupervised methods, while having faster and stable convergence than other domain adaption methods.
Wenchao Du, Hu Chen 0002, Hongyu Yang 0002
CVPR1
2019 Boosting Dialog Response Generation
abstract
Neural models have become one of the most important approaches to dialog response generation.However, they still tend to generate the most common and generic responses in the corpus all the time.To address this problem, we designed an iterative training process and ensemble method based on boosting.We combined our method with different training and decoding paradigms as the base model, including mutual-information-based decoding and reward-augmented maximum likelihood learning.Empirical results show that our approach can significantly improve the diversity and relevance of the responses generated by all base models, backed by objective measurements and human evaluation.
Wenchao Du, Alan W. Black
ACL (1)1
2019 Bag-of-Acoustic-Words for Mental Health Assessment: A Deep Autoencoding Approach
Wenchao Du, Louis-Philippe Morency, Jeffrey F. Cohn, Alan W. Black
INTERSPEECH1
2019 Visual Attention Network for Low-Dose CT
abstract
Noise and artifacts are intrinsic to low-dose computed tomography (LDCT) data acquisition, and will significantly affect the imaging performance. Perfect noise removal and image restoration is intractable in the context of LDCT due to the statistical and the technical uncertainties. In this letter, we apply the generative adversarial network (GAN) framework with a visual attention mechanism to deal with this problem in a data-driven/machine learning fashion. Our main idea is to inject visual attention knowledge into the learning process of GAN to provide a powerful prior of the noise distribution. By doing this, both the generator and discriminator networks are empowered with visual attention information so that they will not only pay special attention to noisy regions and surrounding structures but also explicitly assess the local consistency of the recovered regions. Our experiments qualitatively and quantitatively demonstrate the effectiveness of the proposed method with clinic CT images.
Wenchao Du, Hu Chen 0002, Peixi Liao, Hongyu Yang 0002, Ge Wang 0001, Yi Zhang 0018
IEEE Signal Process. Lett.1
2017 Discovering Conversational Dependencies between Messages in Dialogs
abstract
We investigate the task of inferring conversational dependencies between messages in one-on-one online chat, which has become one of the most popular forms of customer service. We propose a novel probabilistic classifier that leverages conversational, lexical and semantic information. The approach is evaluated empirically on a set of customer service chat logs from a Chinese e-commerce website. It outperforms heuristic baselines.
Wenchao Du, Pascal Poupart
AAAI1
2010 Radial acceleration estimation within one pulse echo based on Hough-ambiguity transformation
Guohong Wang, Shuyi Jia, Wenchao Du
Sci. China Inf. Sci.3