EDBT 2026 Demo / reviewers in the wild / expert
Weize Quan
dblp:185/6242
· DBLP profile ↗
31ranked-venue papers
5as first author
26since 2021 · last 2026
0000-0003-0892-581XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 2 first-author · 22 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 8 since 2021Security and privacy · 2 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | M2HF: Multi-Branch Multi-Modal Hybrid Fusion for Text-Video RetrievalabstractVideos contain multi-modal content, and exploring multi-branch cross-modal interactions with natural language queries can be of benefit to the text-video retrieval task (TVR). However, recent methods applying the large-scale pre-trained CLIP model for TVR only focus on visual cues in videos. Furthermore, traditional methods of simply concatenating multimodal features do not exploit fine-grained cross-modal information in videos. In this paper, we propose a multi-branch multi-modal hybrid fusion (M2HF) network to hierarchically explore interaction between text queries and other modality content in videos. Specifically, M2HF first fuses visual features extracted by CLIP with audio and motion features extracted from videos to obtain fused audio-visual features and motion-visual features respectively. The multi-modal completion problem is also considered and solved in this process. Then, visual features, audio-visual features, motion-visual features, and text extracted from the video are used to establish cross-modal relationships with caption text queries using a multibranch approach. The retrieval outputs from all branches are then fused to obtain the final text-video retrieval results. Our framework provides two kinds of training strategies, using an ensemble approach and an end-to-end approach. Moreover, a novel multi-modal loss function is proposed to balance the contributions of each modality for efficient end-to-end training. M2HF allows us to obtain state-of-the-art results on various benchmarks: Rank@1 of 66.0%, 68.6%, 33.9%, 57.4%, and 57.3% on MSR-VTT, MSVD, LSMDC, DiDeMo, and ActivityNet, respectively. Weize Quan, Zhe Zhao 0006, Kimmo Yan, Chen Chen 0001, Dong-Ming Yan 0001 |
Comput. Vis. Media | 2 |
| 2026 | Facade parsing via joint structural priors and phased deep learning
Yuning Huang, Weize Quan, Dong-Ming Yan 0001, Jie Jiang 0017, Yingmei Wei |
Neurocomputing | 3 |
| 2026 | CasLayout: Cascaded 3D Layout Diffusion for Indoor Scene Synthesis with Implicit Relation ModelingabstractSynthesizing realistic 3D indoor scenes remains challenging due to data scarcity and the difficulty of simultaneously enforcing global architectural constraints and local semantic consistency. Existing approaches often overlook structural boundaries or rely on fully connected relation graphs that introduce redundant generation errors. Inspired by human design cognition, we present CasLayout, a cascaded diffusion framework that decomposes the joint scene generation task into four conditional sub-stages with explicit physical and semantic roles: (1) predicting furniture quantity and categories, (2) refining object sizes and feature embeddings, (3) modeling spatial relationships in a latent space, and (4) generating Oriented Bounding Boxes (OBBs). This decoupled architecture reduces data requirements and enables flexible integration of Large Language Models (LLMs) and Vision Language Models (VLMs) for zero-shot tasks such as image-to-scene generation. To maintain physical validity within complex floor plans, we explicitly model building elements ( e.g. , walls, doors, and windows) as conditional constraints. Furthermore, to address the high entropy of dense relation graphs, we introduce a sparse relation graph formulation aligned with human spatial descriptions. By encoding these sparse graphs into a compact latent space using a bidirectional Variational Autoencoder (VAE), the proposed framework provides enhanced relational controllability, allowing generated layouts to better respect functional organization. Experiments demonstrate that CasLayout achieves state-of-the-art performance in fidelity and diversity while enabling improved controllability in practical applications. Yingrui Wu, Youkang Kong, Mingyang Zhao 0001, Weize Quan, Dong-Ming Yan 0001, Yang Liu 0014 |
ACM Trans. Graph. | 4 |
| 2026 | E$^{3}$3-Net: Efficient E(3)-Equivariant Normal Estimation NetworkabstractPoint cloud normal estimation is a fundamental task in 3D geometry processing, playing a crucial role in applications such as 3D reconstruction, object recognition, and surface analysis. While recent learning-based methods achieve notable advancements in normal prediction, they often overlook the critical aspect of equivariance. This oversight leads to inefficient learning of symmetric patterns inherent in geometric data. To address this issue, we propose E$^{3}$3-Net, an innovative neural network architecture designed to inherently achieve equivariance for normal estimation. We introduce an efficient random frame method, which significantly reduces the training resources required for this task to just 1/8 of previous work, while simultaneously enhancing prediction accuracy. Furthermore, we design a Gaussian-weighted loss function and a receptive-aware inference strategy that effectively leverage the local properties of point clouds, ensuring more precise and reliable normal estimation. Our method demonstrates superior performance across both synthetic and real-world datasets, consistently outperforming current state-of-the-art techniques by a substantial margin. Specifically, we achieve a 4% improvement in RMSE on the PCPNet dataset, 2.67% on the SceneNN dataset, and 2.44% on the FamousShape dataset, highlighting the robustness and scalability of E$^{3}$3-Net in diverse environments. Mingyang Zhao 0001, Weize Quan, Zhen Chen 0013, Dong-Ming Yan 0001, Peter Wonka |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | PointCFormer: A Relation-Based Progressive Feature Extraction Network for Point Cloud CompletionabstractPoint cloud completion aims to reconstruct the complete 3D shape from incomplete point clouds, and it is crucial for tasks such as 3D object detection and segmentation. Despite the continuous advances in point cloud analysis techniques, feature extraction methods are still confronted with apparent limitations. The sparse sampling of point clouds, used as inputs in most methods, often results in a certain loss of global structure information. Meanwhile, traditional local feature extraction methods usually struggle to capture the intricate geometric details. To overcome these drawbacks, we introduce PointCFormer, a transformer framework optimized for robust global retention and precise local detail capture in point cloud completion. This framework embraces several key advantages. First, we propose a relation-based local feature extraction method to perceive local delicate geometry characteristics. This approach establishes a fine-grained relationship metric between the target point and its k-nearest neighbors, quantifying each neighboring point's contribution to the target point's local features. Secondly, we introduce a progressive feature extractor that integrates our local feature perception method with self-attention. Starting with a denser sampling of points as input, it iteratively queries long-distance global dependencies and local neighborhood relationships. This extractor maintains enhanced global structure and refined local details, without generating substantial computational overhead. Additionally, we develop a correction module after generating point proxies in the latent space to reintroduce denser information from the input points, enhancing the representation capability of the point proxies. PointCFormer demonstrates state-of-the-art performance on several widely used benchmarks. Weize Quan, Dong-Ming Yan 0001, Jie Jiang 0017, Yingmei Wei |
AAAI | 2 |
| 2025 | GoHD: Gaze-oriented and Highly Disentangled Portrait Animation with Rhythmic Poses and Realistic ExpressionsabstractAudio-driven talking head generation necessitates seamless integration of audio and visual data amidst the challenges posed by diverse input portraits and intricate correlations between audio and facial motions. In response, we propose a robust framework GoHD designed to produce highly realistic, expressive, and controllable portrait videos from any reference identity with any motion. GoHD innovates with three key modules: Firstly, an animation module utilizing latent navigation is introduced to improve the generalization ability across unseen input styles. This module achieves high disentanglement of motion and identity, and it also incorporates gaze orientation to rectify unnatural eye movements that were previously overlooked. Secondly, a conformer-structured conditional diffusion model is designed to guarantee head poses that are aware of prosody. Thirdly, to estimate lip-synchronized and realistic expressions from the input audio within limited training data, a two-stage training strategy is devised to decouple frequent and frame-wise lip motion distillation from the generation of other more temporally dependent but less audio-related motions, e.g., blinks and frowns. Extensive experiments validate GoHD's advanced generalization capabilities, demonstrating its effectiveness in generating realistic talking face results on arbitrary subjects. Weize Quan, Hailin Shi, Lili Wang 0006, Dong-Ming Yan 0001 |
AAAI | 2 |
| 2025 | Concept-Edge Fusion: Background Generation for Product Presentation Based on Text-to-Image Model
Pengfei Deng, Weize Quan, Hanyu Wang 0002, Qinglin Lu, Zhifeng Li 0001, Dong-Ming Yan 0001 |
CVM (2) | 3 |
| 2025 | Diffused Poses and Distilled Expressions for Controllable Audio-driven Talking Face GenerationabstractAudio-driven portrait animation is an emerging field in multi-modal generation that aims to create lifelike talking face videos from audio input. While significant progress has been made, accurately modeling the relationship between audio signals and various facial motions, such as head poses and expressions, remains a challenge. Existing methods have primarily focused on generating lip-synchronized movements, often neglecting the intricate correlations between audio and other facial dynamics like head movements and eye blinks. More recent approaches have attempted to address these limitations by introducing latent disentanglement of facial motions, though this often comes at the cost of reduced flexibility in motion control. In this work, we propose a novel framework for audio-driven talking portrait animation that allows for precise and controllable generation of head poses and facial expressions. Our approach includes two key components: an audio-conditional diffusion model for generating prosody-aware head poses and a noise-conditional, lip-distilling transformer for predicting synchronized facial expressions. We further introduce an innovative animation model that uses these generated poses and expressions to produce highly realistic and controllable talking head videos. Extensive experiments demonstrate that our method not only achieves superior performance in generating natural and synchronized facial motions but also outperforms state-of-the-art techniques in the field. Weize Quan, Zhaojin Lu, Dong-Ming Yan 0001 |
ICASSP | 2 |
| 2025 | MPIC: Exploring alternative approach to standard convolution in deep neural networks
Jie Jiang 0017, Ruoli Yang, Weize Quan, Dong-Ming Yan 0001 |
Neural Networks | 4 |
| 2025 | BrepGPT: Autoregressive B-rep Generation with Voronoi Half-PatchabstractBoundary representation (B-rep) is the de facto standard for CAD model representation in modern industrial design. The intricate coupling between geometric and topological elements in B-rep structures has forced existing generative methods to rely on cascaded multi-stage networks, resulting in error accumulation and computational inefficiency. We present BrepGPT, a single-stage autoregressive framework for B-rep generation. Our key innovation lies in the Voronoi Half-Patch (VHP) representation, which decomposes B-reps into unified local units by assigning geometry to nearest half-edges and sampling their next pointers. Unlike hierarchical representations that require multiple distinct encodings for different structural levels, our VHP representation facilitates unifying geometric attributes and topological relations in a single, coherent format. We further leverage dual VQ-VAEs to encode both vertex topology and Voronoi Half-Patches into vertex-based tokens, achieving a more compact sequential encoding. A decoder-only Transformer is then trained to autoregressively predict these tokens, which are subsequently mapped to vertex-based features and decoded into complete B-rep models. Experiments demonstrate that BrepGPT achieves state-of-the-art performance in unconditional B-rep generation. The framework also exhibits versatility in various applications, including conditional generation from category labels, point clouds, text descriptions, and images, as well as B-rep autocompletion and interpolation. Weize Quan, Biao Zhang 0005, Peter Wonka, Dong-Ming Yan 0001 |
ACM Trans. Graph. | 3 |
| 2024 | CMG-Net: Robust Normal Estimation for Point Clouds via Chamfer Normal Distance and Multi-Scale GeometryabstractThis work presents an accurate and robust method for estimating normals from point clouds. In contrast to predecessor approaches that minimize the deviations between the annotated and the predicted normals directly, leading to direction inconsistency, we first propose a new metric termed Chamfer Normal Distance to address this issue. This not only mitigates the challenge but also facilitates network training and substantially enhances the network robustness against noise. Subsequently, we devise an innovative architecture that encompasses Multi-scale Local Feature Aggregation and Hierarchical Geometric Information Fusion. This design empowers the network to capture intricate geometric details more effectively and alleviate the ambiguity in scale selection. Extensive experiments demonstrate that our method achieves the state-of-the-art performance on both synthetic and real-world datasets, particularly in scenarios contaminated by noise. Our implementation is available at https://github.com/YingruiWoo/CMG-Net_Pytorch. Yingrui Wu, Mingyang Zhao 0001, Keqiang Li 0005, Weize Quan, Tianqi Yu, Xiaohong Jia 0001, Dong-Ming Yan 0001 |
AAAI | 4 |
| 2024 | Neural Parametric Human Hand Modeling with Point Cloud RepresentationabstractRecently, multi-layer perceptron-based implicit representations have achieved remarkable successes in hand modeling. Compared with previous explicit mesh-based representation methods, implicit methods are more compact shape representations. However, it is expensive to obtain explicit geometry surfaces from implicit functions with Marching Cubes, which limits the real-time performance in surface reconstruction applications. To explore a more effective and efficient hand representation, we present a skeleton-driven method to represent a human hand with a point cloud. To achieve this goal, we propose a Tri-Axis Modeling method to model the motion pattern of the xyz coordinate of a patch of point cloud, and an Order Encoding strategy to construct a parameter-sharing and geometry-disentangled network. These two effective strategies make our method run in real-time and has super-high fidelity close to implicit methods. Qualitative and quantitative experiments on public datasets demonstrate the efficiency, effectiveness, and robustness of our method against state-of-the-art approaches. Jian Yang 0035, Weize Quan, Zhen Shen 0004, Dong-Ming Yan 0001 |
ICMR | 2 |
| 2024 | VQ-CAD: Computer-Aided Design model generation with vector quantized diffusion
Mingyang Zhao 0001, Yiqun Wang 0001, Weize Quan, Dong-Ming Yan 0001 |
Comput. Aided Geom. Des. | 4 |
| 2024 | Deep Learning-Based Image and Video Inpainting: A Survey
Weize Quan, Jiaxi Chen, Dong-Ming Yan 0001, Peter Wonka |
Int. J. Comput. Vis. | 1 |
| 2024 | CGFormer: ViT-Based Network for Identifying Computer-Generated Images With Token LabelingabstractThe advanced graphics rendering techniques and image generation algorithms significantly improve the visual quality of computer-generated (CG) images, and this makes it more challenging to distinguish between CG images and natural images (NIs) for a forensic detector. For the identification of CG images, human beings often need to inspect and evaluate the entire image and its local region as well. In addition, we observe that the distributions of both near and far patch-wise correlation have differences between CG images and NIs. Current mainstream methods adopt the CNN-based architecture with the classical cross entropy loss, however, there are several limitations: 1) the weakness of long-distance relationship modeling of image content due to the local receptive field of CNN; 2) the pixel sensitivity due to the convolutional computation; 3) the insufficient supervision due to the training loss on the whole image. In this paper, we propose a novel vision transformer (ViT)-based network with token labeling for CG image identification. Our network, called CGFormer, consists of patch embedding, feature modeling, and token prediction. We apply patch embedding to sequence the input image and weaken the pixel sensitivity. Stacked multi-head attention-based transformer blocks are utilized to model the patch-wise relationship and introduce a certain level of adaptability. Besides the conventional classification loss on class token of the whole image, we additionally introduce a soft cross entropy loss on patch tokens to comprehensively exploit the supervision information from local patches. Extensive experiments demonstrate that our method achieves the state-of-the-art forensic performance on six publicly available datasets in terms of classification accuracy, generalization, and robustness. Code is available athttps://github.com/feipiefei/CGFormer. Weize Quan, Pengfei Deng, Kai Wang 0002, Dong-Ming Yan 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2024 | Layout-aware Single-image Document FlatteningabstractSingle image rectification of document deformation is a challenging task. Although some recent deep learning-based methods have attempted to solve this problem, they cannot achieve satisfactory results when dealing with document images with complex deformations. In this article, we propose a new efficient framework for document flattening. Our main insight is that most layout primitives in a document have rectangular outline shapes, making unwarping local layout primitives essentially homogeneous with unwarping the entire document. The former task is clearly more straightforward to solve than the latter due to the more consistent texture and relatively smooth deformation. On this basis, we propose a layout-aware deep model working in a divide-and-conquer manner. First, we employ a transformer-based segmentation module to obtain the layout information of the input document. Then a new regression module is applied to predict the global and local UV maps. Finally, we design an effective merging algorithm to correct the global prediction with local details. Both quantitative and qualitative experimental results demonstrate that our framework achieves favorable performance against state-of-the-art methods. In addition, the current publicly available document flattening datasets have limited 3D paper shapes without layout annotation and also lack a general geometric correction metric. Therefore, we build a new large-scale synthetic dataset by utilizing a fully automatic rendering method to generate deformed documents with diverse shapes and exact layout segmentation labels. We also propose a new geometric correction metric based on our paired document UV maps. Code and dataset will be released at https://github.com/BunnySoCrazy/LA-DocFlatten . Weize Quan, Jianwei Guo 0003, Dong-Ming Yan 0001 |
ACM Trans. Graph. | 2 |
| 2023 | DPE: Disentanglement of Pose and Expression for General Video Portrait EditingabstractOne-shot video-driven talking face generation aims at producing a synthetic talking video by transferring the facial motion from a video to an arbitrary portrait image. Head pose and facial expression are always entangled in facial motion and transferred simultaneously. However, the entanglement sets up a barrier for these methods to be used in video portrait editing directly, where it may require to modify the expression only while maintaining the pose unchanged. One challenge of decoupling pose and expression is the lack of paired data, such as the same pose but different expressions. Only a few methods attempt to tackle this challenge with the feat of 3D Morphable Models (3DMMs) for explicit disentanglement. But 3DMMs are not accurate enough to capture facial details due to the limited number of Blend-shapes, which has side effects on motion transfer. In this paper, we introduce a novel self-supervised disentanglement framework to decouple pose and expression without 3DMMs and paired data, which consists of a motion editing module, a pose generator, and an expression generator. The editing module projects faces into a latent space where pose motion and expression motion can be disentangled, and the pose or expression transfer can be performed in the latent space conveniently via addition. The two generators render the modified latent codes to images, respectively. Moreover, to guarantee the disentanglement, we propose a bidirectional cyclic training strategy with well-designed constraints. Evaluations demonstrate our method can control pose or expression independently and be used for general video editing. Code: https://github.com/Carlyx/DPE Youxin Pang, Yong Zhang 0034, Weize Quan, Yanbo Fan, Xiaodong Cun, Ying Shan, Dong-Ming Yan 0001 |
CVPR | 3 |
| 2023 | Dense Modality Interaction Network for Audio-Visual Event LocalizationabstractHuman perception systems can integrate audio and visual information automatically to obtain a profound understanding of real-world events. Accordingly, fusing audio and visual contents is important to solve the audio-visual event (AVE) localization problem. Although most existing works have fused audio and visual modalities to explore their relationship with attention-based networks, we can delve into their relationship more deeply to improve the fusion capability of the two modalities. In this paper, we propose a dense modality interaction network (DMIN) to elegantly leverage audio and visual information by integrating two novel modules, namely, the audio-guided triplet attention (AGTA) module and the dense inter-modality attention (DIMA) module. The AGTA module enables audio information to guide the network to pay more attention to event-relevant visual regions. This guidance is conducted in the channel, temporal, and spatial dimensions, which emphasize informative features, temporal relationships and spatial regions, to boost the capacity of representations. Furthermore, the DIMA module establishes the dense-relationship between audio and visual modalities. Specifically, the DIMA module leverages the information of all channel pairs of audio and visual features to formulate the cross-modality attention weight, which is superior to the multi-head attention module that uses limited information. Moreover, a novel unimodal discrimination loss (UDL) is introduced to exploit the unimodal and fused features together for more exact AVE localization. The experimental results show that our method is remarkably superior to the state-of-the-art methods in fully- and weakly-supervised AVE settings. To further evaluate the model's ability to build audio-visual connections, we design a dense cross modality relation network (DCMR) to solve the cross-modality localization task. DCMR is a simple deformation of a DMIN, and the experimental results further illustrate that DIMA can explore denser relationships between the two modalities. Code is available at https://github.com/weizequan/DMIN.git. Weize Quan, Bin Liu 0041, Dong-Ming Yan 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | W-Net: Structure and Texture Interaction for Image InpaintingabstractRecent literature has developed two advanced tools for image inpainting: appearance propagation and attention matching. However, given the ineffective feature reorganization and vulnerable attention maps, existing works yield suboptimal results with distorted structures and inconsistent contents. Furthermore, we observe that deep sampling layers (DSL) and shallow skip connections (SSC) in U-Net separately promote image structure inference and texture synthesis. To address the above two issues, we devise a W-shaped network (W-Net), which consists of two key components: a texture spatial attention (TSA) module in SSC and a structure channel excitation (SCE) module in DSL. W-Net is a two-stage network, with coarse and refined structures derived at each stage. Meanwhile, the TSA module fills incomplete textures with reliable attention scores under the guidance of coarse structures, which effectively diminishes inconsistency from appearance to semantics. The SCE module rectifies structures according to the difference between coarse structures and refined structures enhanced by texture features. Then the module motivates them to produce more reasonable shapes. Complete textures and refined structures constitute desired inpainted images, as the output of W-Net. Experiments on multiple datasets demonstrate the superior performance of W-Net. Ruisong Zhang, Weize Quan, Yong Zhang 0034, Jue Wang 0001, Dong-Ming Yan 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | Bi-Directional Modality Fusion Network For Audio-Visual Event LocalizationabstractAudio and visual signals stimulate many audio-visual sensory neurons of persons to generate audio-visual contents, helping humans perceive the world. Most of the existing audio-visual event localization approaches focus on generating audio-visual features by fusing the audio and visual modalities for final predictions. However, an audio-visual adjustment mechanism exists in a complicated multi-modal perception system. Inspired by this observation, we propose a novel bi-directional modality fusion network (BMFN), which not only simply fuses audio and visual features, but also adjusts the fused features to increase their representativeness with the help of the original audio and visual contents. The high-level audio-visual features achieved from two directions with two forward-backward fusion modules and a mean operation are summarized for the final event localization. Experimental results demonstrate that our method outperforms state-of-the-art works in both fully- and weakly-supervised learning settings. The code is available at https://github.com/weizequan/BMFN.git. Weize Quan, Dong-Ming Yan 0001 |
ICASSP | 2 |
| 2022 | Scene text removal via cascaded text stroke detection and erasingabstractRecent learning-based approaches show promising performance improvement for the scene text removal task but usually leave several remnants of text and provide visually unpleasant results. In this work, a novel end-to-end framework is proposed based on accurate text stroke detection. Specifically, the text removal problem is decoupled into text stroke detection and stroke removal; we design separate networks to solve these two subproblems, the latter being a generative network. These two networks are combined as a processing unit, which is cascaded to obtain our final model for text removal. Experimental results demonstrate that the proposed method substantially outperforms the state-of-the-art for locating and erasing scene text. A new large-scale real-world dataset with 12,120 images has been constructed and is being made available to facilitate research, as current publicly available datasets are mainly synthetic so cannot properly measure the performance of different methods. Xuewei Bian, Weize Quan, Juntao Ye, Xiaopeng Zhang 0001, Dong-Ming Yan 0001 |
Comput. Vis. Media | 3 |
| 2022 | Neural texture transfer assisted video coding with adaptive up-sampling
Li Yu 0004, Wenshuai Chang, Weize Quan, Jimin Xiao, Dong-Ming Yan 0001, Moncef Gabbouj |
Signal Process. Image Commun. | 3 |
| 2022 | Image Inpainting With Local and Global RefinementabstractImage inpainting has made remarkable progress with recent advances in deep learning. Popular networks mainly follow an encoder-decoder architecture (sometimes with skip connections) and possess sufficiently large receptive field, i.e., larger than the image resolution. The receptive field refers to the set of input pixels that are path-connected to a neuron. For image inpainting task, however, the size of surrounding areas needed to repair different kinds of missing regions are different, and the very large receptive field is not always optimal, especially for the local structures and textures. In addition, a large receptive field tends to involve more undesired completion results, which will disturb the inpainting process. Based on these insights, we rethink the process of image inpainting from a different perspective of receptive field, and propose a novel three-stage inpainting framework with local and global refinement. Specifically, we first utilize an encoder-decoder network with skip connection to achieve coarse initial results. Then, we introduce a shallow deep model with small receptive field to conduct the local refinement, which can also weaken the influence of distant undesired completion results. Finally, we propose an attention-based encoder-decoder network with large receptive field to conduct the global refinement. Experimental results demonstrate that our method outperforms the state of the arts on three popular publicly available datasets for image inpainting. Our local and global refinement network can be directly inserted into the end of any existing networks to further improve their inpainting performance. Code is available at https://github.com/weizequan/LGNet.git. Weize Quan, Ruisong Zhang, Yong Zhang 0034, Zhifeng Li 0001, Jue Wang 0001, Dong-Ming Yan 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | Deep Video Decaptioning
Pengpeng Chu, Weize Quan, Tong Wang 0013, Pan Wang 0008, Peiran Ren, Dong-Ming Yan 0001 |
BMVC | 2 |
| 2021 | Text-Aware Single Image Specular Highlight Removal
Shiyu Hou, Weize Quan, Jingen Jiang 0001, Dong-Ming Yan 0001 |
PRCV (4) | 3 |
| 2021 | Efficient Center Voting for Object Detection and 6D Pose Estimation in 3D Point CloudabstractWe present a novel and efficient approach to estimate 6D object poses of known objects in complex scenes represented by point clouds. Our approach is based on the well-known point pair feature (PPF) matching, which utilizes self-similar point pairs to compute potential matches and thereby cast votes for the object pose by a voting scheme. The main contribution of this paper is to present an improved PPF-based recognition framework, especially a new center voting strategy based on the relative geometric relationship between the object center and point pair features. Using this geometric relationship, we first generate votes to object centers resulting in vote clusters near real object centers. Then we group and aggregate these votes to generate a set of pose hypotheses. Finally, a pose verification operator is performed to filter out false positives and predict appropriate 6D poses of the target object. Our approach is also suitable to solve the multi-instance and multi-object detection tasks. Extensive experiments on a variety of challenging benchmark datasets demonstrate that the proposed algorithm is discriminative and robust towards similar-looking distractors, sensor noise, and geometrically simple shapes. The advantage of our work is further verified by comparing to the state-of-the-art approaches. Jianwei Guo 0003, Xuejun Xing, Weize Quan, Dong-Ming Yan 0001, Qingyi Gu, Yang Liu 0014, Xiaopeng Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | Pixel-wise Dense Detector for Image InpaintingabstractAbstract Recent GAN‐based image inpainting approaches adopt an average strategy to discriminate the generated image and output a scalar, which inevitably lose the position information of visual artifacts. Moreover, the adversarial loss and reconstruction loss (e.g., ℓ1loss) are combined with tradeoff weights, which are also difficult to tune. In this paper, we propose a novel detection‐based generative framework for image inpainting, which adopts the min‐max strategy in an adversarial process. The generator follows an encoder‐decoder architecture to fill the missing regions, and the detector using weakly supervised learning localizes the position of artifacts in a pixel‐wise manner. Such position information makes the generator pay attention to artifacts and further enhance them. More importantly, we explicitly insert the output of the detector into the reconstruction loss with a weighting criterion, which balances the weight of the adversarial loss and reconstruction loss automatically rather than manual operation. Experiments on multiple public datasets show the superior performance of the proposed framework. The source code is available at https://github.com/Evergrow/GDN_Inpainting. Rui-Song Zhang, Weize Quan, Baoyuan Wu, Zhifeng Li 0001, Dong-Ming Yan 0001 |
Comput. Graph. Forum | 2 |
| 2020 | Distinguishing Computer-Generated Images from Natural Images Using Channel and Pixel Correlation
Rui-Song Zhang, Weize Quan, Lu-Bin Fan, Liming Hu, Dong-Ming Yan 0001 |
J. Comput. Sci. Technol. | 2 |
| 2018 | Learning 3D Keypoint Descriptors for Non-rigid Shape Matching
Hanyu Wang 0002, Jianwei Guo 0003, Dong-Ming Yan 0001, Weize Quan, Xiaopeng Zhang 0001 |
ECCV (8) | 4 |
| 2018 | Distinguishing Between Natural and Computer-Generated Images Using Convolutional Neural NetworksabstractDistinguishing between natural images (NIs) and computer-generated (CG) images by naked human eyes is difficult. In this paper, we propose an effective method based on a convolutional neural network (CNN) for this fundamental image forensic problem. Having observed the rather limited performance of training existing CCNs from scratch or fine-tuning pre-trained network, we design and implement a new and appropriate network with two cascaded convolutional layers at the bottom of a CNN. Our network can be easily adjusted to accommodate different sizes of input image patches while maintaining a fixed depth, a stable structure of CNN, and a good forensic performance. Considering the complexity of training CNNs and the specific requirement of image forensics, we introduce the so-called local-to-global strategy in our proposed network. Our CNN derives a forensic decision on local patches, and a global decision on a full-sized image can be easily obtained via simple majority voting. This strategy can also be used to improve the performance of existing methods that are based on hand-crafted features. Experimental results show that our method outperforms existing methods, especially in a challenging forensic scenario with NIs and CG images of heterogeneous origins. Our method also has good robustness against typical post-processing operations, such as resizing and JPEG compression. Unlike previous attempts to use CNNs for image forensics, we try to understand what our CNN has learned about the differences between NIs and CG images with the aid of adequate and advanced visualization tools. Weize Quan, Kai Wang 0002, Dong-Ming Yan 0001, Xiaopeng Zhang 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2016 | Analyzing surface sampling patterns using the localized pair correlation functionabstractPoint distributions with different characteristics have a crucial influence on graphics applications. Various analysis tools have been developed in recent years, mainly for blue noise sampling in Euclidean domains. In this paper, we present a new method to analyze the properties of general sampling patterns that are distributed on mesh surfaces. The core idea is to generalize to surfaces the pair correlation function (PCF) which has successfully been employed in sampling pattern analysis and synthesis in 2D and 3D. Experimental results demonstrate that the proposed approach can reveal correlations of point sets generated by a wide range of sampling algorithms. An acceleration technique is also suggested to improve the performance of the PCF. Weize Quan, Jianwei Guo 0003, Dong-Ming Yan 0001, Weiliang Meng, Xiaopeng Zhang 0001 |
Comput. Vis. Media | 1 |