Yishun Dou

dblp:273/9779 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
10since 2021 · last 2025
0009-0008-8345-8258ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 7 since 2021
YearPublicationVenuePosition
2025 Neural Block Compression: Variable Bitrates Feature Blocks for Texture Representation
abstract
The imperative for compression of material textures emerges from the critical demand for high-quality rendering, which necessitates sophisticated textures that, in turn, require substantial storage and memory resources. Thus, low-bitrate compression is crucial, especially in modern games demanding higher texture resolutions. Concurrent methodologies in texture compression predominantly employ a block-based paradigm based on color space, which inevitably leads to representational redundancies and a limited compression scope, particularly at lower bitrates. In the context of mobile devices, bandwidth during texture loading and runtime memory are major bottlenecks, making existing compression algorithms inadequate for high-resolution textures. To mitigate these limitations, we propose a novel multi-resolution texture compression scheme, Neural Block Compression (NBC), developed within the neural feature domain. Our encoding scheme is constructed on a hierarchy of multi-resolution neural feature blocks, and the key ingredient is the variable bitrates quantization scheme. It allocates higher bitrates to higher feature mip-levels and lower bitrates to lower feature mip-levels, thereby extending the concept of block compression from color domain into neural feature domain. Extensive experiments demonstrate the superior texture compression quality achieved by the proposed scheme, especially at low bitrates.
Yishun Dou, Xiangzhong Fang, Wenjun Zhang 0001, Bingbing Ni
AAAI2
2025 InstantSticker: Realistic Decal Blending via Disentangled Object Reconstruction
abstract
We present InstantSticker, a disentangled reconstruction pipeline based on Image-Based Lighting (IBL), which focuses on highly realistic decal blending, simulates stickers attached to the reconstructed surface, and allows for instant editing and real-time rendering. To achieve stereoscopic impression of the decal, we introduce shadow factor into IBL, which can be adaptively optimized during training. This allows the shadow brightness of surfaces to be accurately decomposed rather than baked into the diffuse color, ensuring that the edited texture exhibits authentic shading. To address the issues of warping and blurriness in previous methods, we apply As-Rigid-As-Possible (ARAP) parameterization to pre-unfold a specified area of the mesh and use the local UV mapping combined with a neural texture map to enhance the ability to express high-frequency details in that area. For instant editing, we utilize the Disney BRDF model, explicitly defining material colors with 3-channel diffuse albedo. This enables instant replacement of albedo RGB values during the editing process, avoiding the prolonged optimization required in previous approaches. In our experiment, we introduce the Ratio Variance Warping (RVW) metric to evaluate the local geometric warping of the decal area. Extensive experimental results demonstrate that our method surpasses previous decal blending methods in terms of editing quality, editing speed and rendering speed, achieving the state-of-the-art.
Yishun Dou, Ye Chen 0006, Bingbing Ni, Wenjun Zhang 0001
AAAI3
2024 FocalDreamer: Text-Driven 3D Editing via Focal-Fusion Assembly
abstract
While text-3D editing has made significant strides in leveraging score distillation sampling, emerging approaches still fall short in delivering separable, precise and consistent outcomes that are vital to content creation. In response, we introduce FocalDreamer, a framework that merges base shape with editable parts according to text prompts for fine-grained editing within desired regions. Specifically, equipped with geometry union and dual-path rendering, FocalDreamer assembles independent 3D parts into a complete object, tailored for convenient instance reuse and part-wise control. We propose geometric focal loss and style consistency regularization, which encourage focal fusion and congruent overall appearance. Furthermore, FocalDreamer generates high-fidelity geometry and PBR textures which are compatible with widely-used graphics engines. Extensive experiments have highlighted the superior editing capabilities of FocalDreamer in both quantitative and qualitative evaluations.
Yuhan Li 0003, Yishun Dou, Xuanhong Chen, Peng Zhou 0010, Bingbing Ni
AAAI2
2024 Real-Time Neural BRDF with Spherically Distributed Primitives
abstract
We propose a neural reflectance model (NeuBRDF) that offers highly versatile material representation, yet with light memory and neural computation consumption towards achieving real-time rendering. The results depicted in Fig. 1, rendered at full HD resolution on a contemporary desktop machine, demonstrate that our system achieves real-time performance with a wide variety of appearances, which is approached by the following two designs. Firstly, recognizing that the bidirectional reflectance is distributed in a sparse high-dimensional space, we propose to project the BRDF into two low-dimensional components, i.e. two hemisphere feature-grids for incoming and outgoing directions, respectively. Secondly, we distribute learnable neural reflectance primitives on our highly-tailored spherical surface grid. These primitives offer informative features for each hemisphere component and reduce the complexity of the feature learning network, leading to fast evaluation. These primitives are centrally stored in a codebook and can be shared across multiple grids and even across materials, based on low-cost indices stored in material-specific spher-ical surface grids. Our NeuBRDF, agnostic to the material, provides a unified framework for representing a variety of materials consistently. Comprehensive experimental results on measured BRDF compression, Monte Carlo simulated BRDF acceleration, and extension to spatially varying effects demonstrate the superior quality and generalizability achieved by the proposed scheme.
Yishun Dou, Qiaoqiao Jin, Bingbing Ni, Yugang Chen, Junxiang Ke
CVPR1
2024 Differentiable Micro-Mesh Construction
abstract
Micro-mesh (μ-mesh.) is a new graphics primitive for compact representation of extreme geometry, consisting of a low-polygon base mesh enriched by per micro-vertex displacement. A new generation of GPUs supports this structure with hardware evolution on μ-mesh ray tracing, achieving real-time rendering in pixel level geometric details. In this article, we present a differentiable framework to convert standard meshes into this efficient format, offering a holistic scheme in contrast to the previous stage-based methods. In our construction context, a μ-mesh is defined where each base triangle is a parametric primitive, which is then reparameterized with Laplacian operators for efficient geometry optimization. Our framework offers numerous advantages for high-quality μ-mesh production: (i) end-to-end geometry optimization and displacement baking; (ii) enabling the differentiation of renderings with respect to μ-mesh for faithful reprojectability; (iii) high scalability for integrating useful features for μ-mesh production and rendering, such as minimizing shell volume, maintaining the isotropy of the base mesh, and visual-guided adaptive level of detail. Extensive experiments on μ-mesh construction for a large set of high-resolution meshes demonstrate the superior quality achieved by the proposed scheme.
Yishun Dou, Qiaoqiao Jin, Yuhan Li 0003, Bingbing Ni
CVPR1
2024 AU-vMAE: Knowledge-Guide Action Units Detection via Video Masked Autoencoder
Qiaoqiao Jin, Yishun Dou, Bingbing Ni
PRCV (15)3
2023 Multiplicative Fourier Level of Detail
abstract
We develop a simple yet surprisingly effective implicit representing scheme called Multiplicative Fourier Level of Detail (MFLOD) motivated by the recent success of multiplicative filter network. Built on multi-resolution feature grid/volume (e.g., the sparse voxel octree), each level's feature is first modulated by a sinusoidal function and then element-wisely multiplied by a linear transformation of previous layer's representation in a layer-to-layer recursive manner, yielding the scale-aggregated encodings for a subsequent simple linear forward to get final output. In contrast to previous hybrid representations relying on interleaved multilevel fusion and nonlinear activation-based decoding, MFLOD could be elegantly characterized as a linear combination of sine basis functions with varying amplitude, frequency, and phase upon the learned multilevel features, thus offering great feasibility in Fourier analysis. Comprehensive experimental results on implicit neural representation learning tasks including image fitting, 3D shape representation, and neural radiance fields well demonstrate the superior quality and generalizability achieved by the proposed MFLOD scheme.
Yishun Dou, Qiaoqiao Jin, Bingbing Ni
CVPR1
2023 Generalized Deep 3D Shape Prior via Part-Discretized Diffusion Process
abstract
We develop a generalized 3D shape generation prior model, tailored for multiple 3D tasks including unconditional shape generation, point cloud completion, and cross-modality shape generation, etc. On one hand, to precisely capture local fine detailed shape information, a vector quantized variational autoencoder (VQ-VAE) is utilized to index local geometry from a compactly learned code-book based on a broad set of task training data. On the other hand, a discrete diffusion generator is introduced to model the inherent structural dependencies among different tokens. In the meantime, a multi-frequency fusion module (MFM) is developed to suppress high-frequency shape feature fluctuations, guided by multi-frequency contextual information. The above designs jointly equip our proposed 3D shape prior model with high-fidelity, diverse features as well as the capability of cross-modality alignment, and extensive experiments have demonstrated superior performances on various 3D shape generation tasks.
Yuhan Li 0003, Yishun Dou, Xuanhong Chen, Bingbing Ni, Yilin Sun, Yutian Liu 0004, Fuzhen Wang
CVPR2
2021 Learning Discriminative Features for Semi-Supervised Anomaly Detection
abstract
Anomaly detection is the task of identifying unusual samples in data. Typically anomaly detection is defined on an unlabeled dataset that is assumed most of the samples are normal and others are anomalies. However, in industrial practice, one may have access to a part of annotated data. This gives us the potential for semi-supervised learning. In addition, existing methods assume all training data is normal and neglect the impact of a small number of anomalous samples. In this paper, we consolidate the model’s discriminative power by introducing a transfer learning scheme to anomaly detection, thereby the model suffers less perturbation caused by pollution. We also propose a novel loss function to further adapt to semi-supervised data scenario. We ensure that the contribution of pollution can be well suppressed and reach a harmonious balance in magnitude of loss/gradient between unlabeled and labeled samples. Experiments on three publicly available datasets show that our method achieves state-of-the-art results.
Jie Tang 0006, Yishun Dou, Gangshan Wu
ICASSP3
2021 Imitating Arbitrary Talking Style for Realistic Audio-Driven Talking Face Synthesis
abstract
People talk with diversified styles. For one piece of speech, different talking styles exhibit significant differences in the facial and head pose movements. For example, the "excited" style usually talks with the mouth wide open, while the "solemn" style is more standardized and seldomly exhibits exaggerated motions. Due to such huge differences between different styles, it is necessary to incorporate the talking style into audio-driven talking face synthesis framework. In this paper, we propose to inject style into the talking face synthesis framework through imitating arbitrary talking style of the particular reference video. Specifically, we systematically investigate talking styles with our collected Ted-HD dataset and construct style codes as several statistics of 3D morphable model (3DMM) parameters. Afterwards, we devise a latent-style-fusion (LSF) model to synthesize stylized talking faces by imitating talking styles from the style codes. We emphasize the following novel characteristics of our framework: (1) It doesn't require any annotation of the style, the talking style is learned in an unsupervised manner from talking videos in the wild. (2) It can imitate arbitrary styles from arbitrary videos, and the style codes can also be interpolated to generate new styles. Extensive experiments demonstrate that the proposed framework has the ability to synthesize more natural and expressive talking styles compared with baseline methods.
Haozhe Wu, Jia Jia 0001, Haoyu Wang 0009, Yishun Dou, Qingshan Deng
ACM Multimedia4
2020 Belief Map Enhancement Network for Accurate Human Pose Estimation
Jie Liu 0040, Yishun Dou, Wenjie Zhang 0006, Jie Tang 0006, Gangshan Wu
ECAI2