EDBT 2026 Demo / reviewers in the wild / expert
Zhixin Wang
dblp:05/2680
· DBLP profile ↗
23ranked-venue papers
6as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RefSTAR: Blind Face Image Restoration with Reference Selection, Transfer, and ReconstructionabstractIntroducing high-quality references can largely alleviate the uncertainty in blind face image restoration tasks, yet the equivocal utilization of reference priors makes it still a struggle to well preserve the human identity. We attribute the identity inconsistency to two deficiencies of existing reference-based face restoration methods, namely the inability to effectively determine which features need to be transferred, and the failure to preserve the structure and details of the selected features. This work mainly focuses on these two issues, and we present a novel blind face image restoration method that considers reference selection, transfer, and reconstruction (RefSTAR) to introduce proper features from reference images. Specifically, we construct a reference selection (RefSel) module, which can generate accurate masks to select reference features. For training the RefSel module, we construct a RefSel-HQ dataset through a mask generation pipeline, which contains annotated masks for 10,000 ground truth-reference pairs. To guarantee the exact introduction of selected reference features, a feature fusion paradigm is designed for reference feature transferring, and a Mask-Compatible Cycle-Consistency Loss is redesigned based on reference reconstruction to further ensure the presence of selected reference image features in the output image. Experiments on various backbone models demonstrate superior performance, showing better identity preservation ability and reference feature transfer quality. Zhicun Yin, Ming Liu 0018, Zhixin Wang, Renjing Pei, Xiaoming Li 0002, Rynson W. H. Lau, Wangmeng Zuo |
AAAI | 4 |
| 2025 | MUC: Mixture of Uncalibrated Cameras for Robust 3D Human Body ReconstructionabstractMultiple cameras can provide comprehensive multi-view video coverage of a person. Fusing this multi-view data is crucial for tasks like behavioral analysis, although it traditionally requires camera calibration—a process that is often complex. Moreover, previous studies have overlooked the challenges posed by self-occlusion under multiple views and the continuity of human body shape estimation. In this study, we introduce a method to reconstruct the 3D human body from multiple uncalibrated camera views. Initially, we utilize a pre-trained human body encoder to process each camera view individually, enabling the reconstruction of human body models and parameters for each view along with predicted camera positions. Rather than merely averaging the models across views, we develop a neural network trained to assign weights to individual views for all human body joints, based on the estimated distribution of joint distances from each camera. Additionally, we focus on the mesh surface of the human body for dynamic fusion, allowing for the seamless integration of facial expressions and body shape into a unified human body model. Our method has shown excellent performance in reconstructing the human body on two public datasets, advancing beyond previous work from the SMPL model to the SMPL-X model. This extension incorporates more complex hand poses and facial expressions, enhancing the detail and accuracy of the reconstructions. Crucially, it supports the flexible ad-hoc deployment of any number of cameras, offering significant potential for various applications. Yitao Zhu, Sheng Wang 0014, Mengjie Xu, Zixu Zhuang, Zhixin Wang, Kaidong Wang, Han Zhang 0002, Qian Wang 0001 |
AAAI | 5 |
| 2025 | Dual Prompting Image Restoration with Diffusion TransformersabstractRecent state-of-the-art image restoration methods mostly adopt latent diffusion models with U-Net backbones, yet still facing challenges in achieving high-quality restoration due to their limited capabilities. Diffusion transformers (DiTs), like SD3, are emerging as a promising alternative because of their better quality with scalability. In this paper, we introduce DPIR (Dual Prompting Image Restoration), a novel image restoration method that effectivly extracts conditional information of low-quality images from multiple perspectives. Specifically, DPIR consits of two branches: a low-quality image conditioning branch and a dual prompting control branch. The first branch utilizes a lightweight module to incorporate image priors into the DiT with high efficiency. More importantly, we believe that in image restoration, textual description alone cannot fully capture its rich visual characteristics. Therefore, a dual prompting module is designed to provide DiT with additional visual cues, capturing both global context and local appearance. The extracted global-local visual prompts as extra conditional control, alongside textual prompts to form dual prompts, greatly enhance the quality of the restoration. Extensive experimental results demonstrate that DPIR delivers superior image restoration performance. Dehong Kong, Zhixin Wang, Renjing Pei, Wenqi Ren |
CVPR | 3 |
| 2025 | Fast Image Super-Resolution via Consistency Rectified Flow
Wenbo Li 0002, Haoze Sun, Zhixin Wang, Long Peng 0003, Xiaowei Hu 0001, Renjing Pei, Pheng-Ann Heng |
ICCV | 5 |
| 2025 | TD-BFR: Truncated Diffusion Model for Efficient Blind Face RestorationabstractDiffusion-based methodologies have shown significant potential in blind face restoration (BFR), leveraging their robust generative capabilities. However, they are often criticized for two significant problems: 1) slow training and inference speed, and 2) inadequate recovery of fine-grained facial details. To address these problems, we propose a novel Truncated Diffusion model for efficient Blind Face Restoration (TD-BFR), a three-stage paradigm tailored for the progressive resolution of degraded images. Specifically, TD-BFR utilizes an innovative truncated sampling method, starting from low-quality (LQ) images at low resolution to enhance sampling speed, and then introduces an adaptive degradation removal module to handle unknown degradations and connect the generation processes across different resolutions. Additionally, we further adapt the priors of pre-trained diffusion models to recover rich facial details. Our method efficiently restores high-quality images in a coarse-to-fine manner and experimental results demonstrate that TD-BFR is, on average, 4.75× faster than current state-of-the-art diffusion-based BFR methods while maintaining competitive quality. Ziying Zhang, Zhixin Wang, Qiang Hu 0003, Xiaoyun Zhang 0001 |
ICME | 3 |
| 2025 | CamEdit: Continuous Camera Parameter Control for Photorealistic Image EditingabstractRecent advances in diffusion models have substantially improved text-driven image editing. However, existing frameworks based on discrete textual tokens struggle to support continuous control over camera parameters and smooth transitions in visual effects. These limitations hinder their applications to realistic, camera-aware, and fine-grained editing tasks. In this paper, we present CamEdit, a diffusion-based framework for photorealistic image editing that enables continuous and semantically meaningful manipulation of common camera parameters such as aperture and shutter speed. CamEdit incorporates a continuous parameter prompting mechanism and a parameter-aware modulation module that guides the model in smoothly adjusting focal plane, aperture, and shutter speed, reflecting the effects of varying camera settings within the diffusion process. To support supervised learning in this setting, we introduce CamEdit50K, a dataset specifically designed for photorealistic image editing with continuous camera parameter settings. It contains over 50k image pairs combining real and synthetic data with dense camera parameter variations across diverse scenes. Extensive experiments demonstrate that CamEdit enables flexible, consistent, and high-fidelity image editing, achieving state-of-the-art performance in camera-aware visual manipulation and fine-grained photographic control. Xinran Qin, Zhixin Wang, Haoyu Chen 0003, Renjing Pei, Wenbo Li 0002, Xiaochun Cao |
NeurIPS | 2 |
| 2025 | PocketSR: The Super-Resolution Expert in Your Pocket MobilesabstractReal-world image super-resolution (RealSR) aims to enhance the visual quality of in-the-wild images, such as those captured by mobile phones. While existing methods leveraging large generative models demonstrate impressive results, the high computational cost and latency make them impractical for edge deployment. In this paper, we introduce PocketSR, an ultra-lightweight, single-step model that brings generative modeling capabilities to RealSR while maintaining high fidelity. To achieve this, we design LiteED, a highly efficient alternative to the original computationally intensive VAE in SD, reducing parameters by 97.5\% while preserving high-quality encoding and decoding. Additionally, we propose online annealing pruning for the U-Net, which progressively shifts generative priors from heavy modules to lightweight counterparts, ensuring effective knowledge transfer and further optimizing efficiency. To mitigate the loss of prior knowledge during pruning, we incorporate a multi-layer feature distillation loss. Through an in-depth analysis of each design component, we provide valuable insights for future research. PocketSR, with a model size of 146M parameters, processes 4K images in just 0.8 seconds, achieving a remarkable speedup over previous methods. Notably, it delivers performance on par with state-of-the-art single-step and even multi-step RealSR models, making it a highly practical solution for edge-device applications. Haoze Sun, Linfeng Jiang, Renjing Pei, Zhixin Wang, Haoyu Chen 0003, Fenglong Song, Yujiu Yang 0001, Wenbo Li 0002 |
NeurIPS | 5 |
| 2025 | Data-inherent vulnerabilities: an invisible framework for model-agnostic backdoor attacks
Zhongguo Yang, Yonglu Jiang, Zhixin Wang |
Vis. Comput. | 4 |
| 2024 | Improving the performance of the ORB-SLAM3 with low-light image enhancementabstractTraditional Visual Simultaneous Localization and Mapping (VSLAM) algorithms demonstrate good accuracy and robustness in well-lighting environments by extracting numerous feature points. However, in low-light conditions, insufficient illumination leads to low-contrast images, which hampers the ability of the front-end to extract adequate feature points, resulting in increased tracking errors or complete tracking failure. To overcome these limitations, this paper improves the performance of the ORB-SLAM3 in low-light environments with image enhancement capability. Specifically, we integrate the Histogram Equalization Prior-based (HEP) image enhancement module into ORB-SLAM3. Furthermore, we compare and analyze the results with two other image enhancement algorithms to ensure the effective fusion of image enhancement and ORB-SLAM3. Experiments were conducted and performance comparisons were made to validate the performance of ORB-SLAM3 with image enhancement in low-light environments. Experimental results indicate that the Root Mean Square Errors (RMSE) of the Absolute Pose Error (APE) are 0.99 cm and 0.76 cm on the two public ETH3D low-light datasets, respectively. Compared to the original ORB-SLAM3, the improvement is over 50%. Similarly, the RMSE of the Relative Pose Error (RPE) are 0.54 cm and 0.65 cm, with an improvement of 29% and 34%, respectively. Tuan Li, Zhixin Wang, Chuang Shi |
IPIN | 3 |
| 2024 | GAABind: a geometry-aware attention-based network for accurate protein-ligand binding pose and binding affinity predictionabstractProtein-ligand interactions are increasingly profiled at high-throughput, playing a vital role in lead compound discovery and drug optimization. Accurate prediction of binding pose and binding affinity constitutes a pivotal challenge in advancing our computational understanding of protein-ligand interactions. However, inherent limitations still exist, including high computational cost for conformational search sampling in traditional molecular docking tools, and the unsatisfactory molecular representation learning and intermolecular interaction modeling in deep learning-based methods. Here we propose a geometry-aware attention-based deep learning model, GAABind, which effectively predicts the pocket-ligand binding pose and binding affinity within a multi-task learning framework. Specifically, GAABind comprehensively captures the geometric and topological properties of both binding pockets and ligands, and employs expressive molecular representation learning to model intramolecular interactions. Moreover, GAABind proficiently learns the intermolecular many-body interactions and simulates the dynamic conformational adaptations of the ligand during its interaction with the protein through meticulously designed networks. We trained GAABind on the PDBbindv2020 and evaluated it on the CASF2016 dataset; the results indicate that GAABind achieves state-of-the-art performance in binding pose prediction and shows comparable binding affinity prediction performance. Notably, GAABind achieves a success rate of 82.8% in binding pose prediction, and the Pearson correlation between predicted and experimental binding affinities reaches up to 0.803. Additionally, we assessed GAABind's performance on the severe acute respiratory syndrome coronavirus 2 main protease cross-docking dataset. In this evaluation, GAABind demonstrates a notable success rate of 76.5% in binding pose prediction and achieves the highest Pearson correlation coefficient in binding affinity prediction compared with all baseline methods. Huishuang Tan, Zhixin Wang, Guang Hu 0001 |
Briefings Bioinform. | 2 |
| 2023 | DR2: Diffusion-Based Robust Degradation Remover for Blind Face RestorationabstractBlind face restoration usually synthesizes degraded low-quality data with a pre-defined degradation model for training, while more complex cases could happen in the real world. This gap between the assumed and actual degradation hurts the restoration performance where artifacts are often observed in the output. However, it is expensive and infeasible to include every type of degradation to cover real-world cases in the training data. To tackle this robustness issue, we propose Diffusion-based Robust Degradation Remover (DR2) to first transform the degraded image to a coarse but degradation-invariant prediction, then employ an enhancement module to restore the coarse prediction to a high-quality image. By leveraging a well-performing denoising diffusion probabilistic model, our DR2 diffuses input images to a noisy status where various types of degradation give way to Gaussian noise, and then captures semantic information through iterative denoising steps. As a result, DR2 is robust against common degradation (e.g. blur, resize, noise and compression) and compatible with different designs of enhancement modules. Experiments in various settings show that our framework outperforms state-of-the-art methods on heavily degraded synthetic and real-world datasets. Zhixin Wang, Ziying Zhang, Xiaoyun Zhang 0001, Huangjie Zheng, Mingyuan Zhou, Ya Zhang 0002, Yanfeng Wang 0001 |
CVPR | 1 |
| 2022 | DFIL: an Efficient Web Service Prediction Method based on Deep Feature Interactive LearningabstractAs a representative of nonfunctional features, the Quality-of-Service (QoS) plays an important role in recommending the best service to users. To obtain the QoS value of web services, researchers have proposed many web service QoS prediction methods. However, most existing methods, e.g., the collaborative filtering method, leverage the historical user service invocation information to predict the QoS value, and only consider the QoS information of similar users or services, and ignore the attributes and characteristics of users and services. In this paper, we propose a Deep Feature Interaction Learning method, called DFIL for short, for web service QoS prediction. Specifically, (1) DFIL constructs a feature extraction method to effectively obtain the features of users and services. (2) DFIL trains a novel neural network for QoS prediction, which can accurately find the potential relationship between users and services based on the feature information, and provide good web service recommendation. Ke Zhu 0003, Zhixin Wang, Yingyuan Xiao, Wenguang Zheng, Ching-Hsien Hsu |
CSCWD | 2 |
| 2022 | A narrowband active noise control system with autoregressive model and linear cascaded adaptive notch filter
Jian Liu 0054, Zhixin Wang |
Signal Process. | 2 |
| 2020 | Object as Hotspots: An Anchor-Free 3D Object Detection Approach via Firing of Hotspots
Qi Chen 0014, Lin Sun 0004, Zhixin Wang, Kui Jia, Alan L. Yuille |
ECCV (21) | 3 |
| 2020 | Location-Aware Feature Interaction Learning for Web Service RecommendationabstractWith the increasing prevalence of web services on the World Wide Web, a large number of functionally equivalent web services are provided by different providers. Quality-of-Service (QoS), representing the nonfunctional characteristics, plays an important role in dealing with how to recommend the optimal services to users among these candidates. Many existing methods for predicting QoS values of web services show that QoS values are intensively relevant to location due to the great influence of network distance and the internet connection between users and services. In this paper, we propose a novel location-aware feature interaction learning (LAFIL) method for predicting the QoS values of the user-service matrix and then making the recommendation by learning the underlying relation, which is hidden in the features concerning with location information. LAFIL can effectively solve the problems of data sparsity and cold-start by leveraging the location features of both users and services. To evaluate the performance of our proposed method, comprehensive experiments are conducted using a real-world dataset and the results show that our method achieves better QoS prediction accuracy compared to state-of-the-art approaches. Zhixin Wang, Yingyuan Xiao, Chenchen Sun, Wenguang Zheng, Xu Jiao |
ICWS | 1 |
| 2020 | Part-Aware Fine-Grained Object Categorization Using Weakly Supervised Part Detection NetworkabstractFine-grained object categorization aims for distinguishing objects of subordinate categories that belong to the same entry-level object category. It is a rapidly developing subfield in multimedia content analysis. The task is challenging due to the facts that (1) training images with ground-truth labels are difficult to obtain, and (2) variations among different subordinate categories are subtle. It is well established that characterizing features of different subordinate categories are located on local parts of object instances. However, manually annotating object parts requires expertise, which is also difficult to generalize to new fine-grained categorization tasks. In this work, we propose a Weakly Supervised Part Detection Network (PartNet) that is able to detect discriminative local parts for the use of fine-grained categorization. A vanilla PartNet builds on top of a base subnetwork two parallel streams of upper network layers, which respectively compute scores of classification probabilities (over subordinate categories) and detection probabilities (over a specified number of discriminative part detectors) for local regions of interest (RoIs). The image-level prediction is obtained by aggregating element-wise products of these region-level probabilities, and meanwhile diverse part detectors can be learned in an end-to-end fashion under the image-level supervision. To generate a diverse set of RoIs as inputs of PartNet, we propose a simple Discretized Part Proposals module (DPP) that directly targets for proposing candidates of discriminative local parts, with no bridging via object-level proposals. Experiments on benchmark datasets of CUB-200-2011, Oxford Flower 102 and Oxford-IIIT Pet show the efficacy of our proposed method for both discriminative part detection and fine-grained categorization. In particular, we achieve the new state-of-the-art performance on CUB-200-2011 and Oxford-IIIT Pet datasets when ground-truth part annotations are not available. Yabin Zhang 0001, Kui Jia, Zhixin Wang |
IEEE Trans. Multim. | 3 |
| 2019 | Frustum ConvNet: Sliding Frustums to Aggregate Local Point-Wise Features for AmodalabstractIn this work, we propose a novel method termed Frustum ConvNet (F-ConvNet) for amodal 3D object detection from point clouds. Given 2D region proposals in an RGB image, our method first generates a sequence of frustums for each region proposal, and uses the obtained frustums to group local points. F-ConvNet aggregates point-wise features as frustum-level feature vectors, and arrays these feature vectors as a feature map for use of its subsequent component of fully convolutional network (FCN), which spatially fuses frustum-level features and supports an end-to-end and continuous estimation of oriented boxes in the 3D space. We also propose component variants of F-ConvNet, including an FCN variant that extracts multi-resolution frustum features, and a refined use of F-ConvNet over a reduced 3D space. Careful ablation studies verify the efficacy of these component variants. F-ConvNet assumes no prior knowledge of the working 3D environment and is thus dataset-agnostic. We present experiments on both the indoor SUN-RGBD and outdoor KITTI datasets. F-ConvNet outperforms all existing methods on SUN-RGBD, and at the time of submission it outperforms all published works on the KITTI benchmark. Code has been made available at: https://github.com/zhixinwang/frustum-convnet. Zhixin Wang, Kui Jia |
IROS | 1 |
| 2015 | Continuous Function Modeling of Head-Related Impulse ResponseabstractA continuous function model is developed for efficient head-related impulse response (HRIR) representation. In this model, HRIR is modeled as the response of an IIR filter via common factor decomposition analysis, where HRIRs at the same azimuth but different elevations share the same IIR pole part and HRIRs at the same elevation but different azimuthes share the same IIR zero part. The pole and zero coefficients of the proposed common factor IIR (CF-IIR) filters are further modeled using sinusoidal functions of azimuth and elevation, respectively. A fast algorithm for evaluating sinusoidal function for efficient synthesis of virtual 3D sound is also developed. Zhixin Wang, Cheung-Fat Chan |
IEEE Signal Process. Lett. | 1 |
| 2013 | Two-Dimension Common Factor Decomposition of Head-Related Impulse ResponseabstractIn this letter, a Two-Dimension Common Factor Decomposition (2D-CFD) algorithm is proposed to represent the 3D (time, azimuth and elevation) head-related impulse response (HRIR) dataset with two factorized matrices. One matrix is made up of a set of common factor responses extracted from HRIRs with the same azimuth angle and carries azimuthal information contained in original HRIR dataset. The other matrix carries elevation information similarly. By using this dimension reduction, the storage requirement of HRIRs is reduced remarkably. In addition, the common factor responses in the matrix are further modeled with efficient low order IIR filters to reduce the computation complexity in 3D sound synthesis. Zhixin Wang, Cheung-Fat Chan |
IEEE Signal Process. Lett. | 1 |
| 2011 | Efficient implementation of virtual 3D sound synthesis based on combining grouped PCA and BMTabstractIn this paper, Principal Component Analysis (PCA) is performed on HRIRs grouped according to azimuth and covariance to improve distortion performance while reducing storage requirement. By further combining Balanced Model Truncation (BMT) using a joint optimization scheme, a virtual 3D sound synthesis scheme with low computation complexity, little storage requirement and efficient linear interpolation is proposed. A smaller model distortion is achieved as high PCA model error of contralateral HRIR is decreased by grouping PCA strategy. Zhixin Wang, Cheung-Fat Chan |
ICASSP | 1 |
| 2007 | A Model for Analyzing and Evaluating the Return on Investment in e-LearningabstractAs investment in e-learning accelerates rapidly worldwide, it is important to improve the economic performance in the existing e-learning initiatives. This article presents a model for analyzing and evaluating the return on investment (ROI) in e-learning. The article also explains each component of the model in details. The model highlights the important issues that must be addressed to optimize the investment strategy. The model also further identifies the return on the investment by the mathematical method of fuzzy comprehensive evaluation. Lingyun Yi, Mingzhang Zuo, Zhixin Wang |
ICALT | 3 |
| 2007 | Design and implementation of a smart system for personalization and accurate selection of mobile services
Qusay H. Mahmoud, Eyhab Al-Masri, Zhixin Wang |
Requir. Eng. | 3 |
| 2006 | Customizing and delivering mobile services using software agents and CC/PPabstractA downside brought about by the explosive growth of online information and the shear amount of data traffic moving through our networks is the modern day headaches of information overload. The growth of handheld wireless devices can only mean that the end-users will continually be bombarded by emails and real-time information feeds, but now, even as they are on the go. Software agents provide mechanisms that have good potentials to lower the amount of information that an end- user has to deal with. Agent-enabled applications allow the end- user to personalize the way online resources are presented and can also filter out irrelevant or unwanted information. In this paper, we present our experience in using software agents and the Composite Capabilities/Preference Profiles (CC/PP) for customizing and delivering mobile services for Java 2 Micro Edition (J2ME) enabled and Wireless Application Protocol (WAP) enabled devices. We use software agents because such autonomous software entities have characteristics that can benefit mobile devices and the wireless environment, and the CC/PP is a standard for defining profiles for user preferences and device capabilities. Qusay H. Mahmoud, Zhixin Wang |
CCNC | 2 |