VLDB 2026 Research / reviewers in the wild / expert
Bowen Cai 0001
dblp:179/8578-1
· DBLP profile ↗
13ranked-venue papers
5as first author
6since 2021 · last 2024
0000-0002-5945-9911ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 2 first-authorArtificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
3D vision · 66% Segmentation and scene understanding · 21% Transfer learning and domain adaptation · 10% | |
| Computer graphics and multimedia
4 papers |
Rendering · 85% Visual content generation and editing · 15% |
Topics — the 17 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Rendering › neural rendering
neural field rendering |
0.8 | 1 | 2024 | 3D Scene Creation and Rendering via Rough Meshes: A Lighting Transfer Avenue · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Rendering
physically based rendering |
0.8 | 1 | 2024 | 3D Scene Creation and Rendering via Rough Meshes: A Lighting Transfer Avenue · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Rendering
neural radiance fields |
0.7 | 2 | 2022 | Ray Priors through Reprojection: Improving Neural Radiance Fields for Novel View Extrapolation · CVPR 2022 Digging into Radiance Grid for Real-Time View Synthesis with Detail Preservation · ECCV (15) 2022 |
Computer vision › 3D vision
3d reconstruction |
0.7 | 1 | 2023 | NeuDA: Neural Deformable Anchor for High-Fidelity Implicit Surface Reconstruction · CVPR 2023 |
Computer vision › 3D vision
implicit neural representation |
0.7 | 1 | 2023 | NeuDA: Neural Deformable Anchor for High-Fidelity Implicit Surface Reconstruction · CVPR 2023 |
Computer vision › 3D vision › 3d reconstruction › surface reconstruction › neural surface reconstruction
neural implicit surface reconstruction |
0.7 | 1 | 2023 | NeuDA: Neural Deformable Anchor for High-Fidelity Implicit Surface Reconstruction · CVPR 2023 |
Computer vision › 3D vision
neural radiance field |
0.7 | 1 | 2023 | NeuDA: Neural Deformable Anchor for High-Fidelity Implicit Surface Reconstruction · CVPR 2023 |
Rendering
novel view synthesis |
0.6 | 1 | 2022 | Ray Priors through Reprojection: Improving Neural Radiance Fields for Novel View Extrapolation · CVPR 2022 |
Rendering › novel view synthesis
real-time view synthesis |
0.6 | 1 | 2022 | Digging into Radiance Grid for Real-Time View Synthesis with Detail Preservation · ECCV (15) 2022 |
Rendering › novel view synthesis
view extrapolation |
0.6 | 1 | 2022 | Ray Priors through Reprojection: Improving Neural Radiance Fields for Novel View Extrapolation · CVPR 2022 |
Computer vision › 3D vision
3d scene understanding |
0.5 | 1 | 2021 | 3D-FRONT: 3D Furnished Rooms with layOuts and semaNTics · ICCV 2021 |
Computer vision › Segmentation and scene understanding › semantic segmentation › transfer learning for semantic segmentation
domain adaptive semantic segmentation |
0.5 | 1 | 2021 | Exploiting Diverse Characteristics and Adversarial Ambivalence for Domain Adaptive Segmentation · AAAI 2021 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.5 | 1 | 2021 | Exploiting Diverse Characteristics and Adversarial Ambivalence for Domain Adaptive Segmentation · AAAI 2021 |
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation |
0.5 | 1 | 2021 | Exploiting Diverse Characteristics and Adversarial Ambivalence for Domain Adaptive Segmentation · AAAI 2021 |
Visual content generation and editing
layout generation |
0.5 | 1 | 2021 | 3D-FRONT: 3D Furnished Rooms with layOuts and semaNTics · ICCV 2021 |
Visual content generation and editing › scene authoring
3d scene design |
0.2 | 1 | 2024 | 3D Scene Creation and Rendering via Rough Meshes: A Lighting Transfer Avenue · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Machine learning › Trustworthy machine learning
robustness |
0.1 | 1 | 2021 | Exploiting Diverse Characteristics and Adversarial Ambivalence for Domain Adaptive Segmentation · AAAI 2021 |
Methods — techniques the papers use, named apart from their topics
recommender system · 1.0perceptual constraints · 0.8lighting transfer network · 0.8neural deformable anchor · 0.7hierarchical positional encoding · 0.7differentiable ray casting · 0.7ray casting · 0.6ray atlas · 0.6radiance grid · 0.6multi-view consistency · 0.6detail preservation · 0.6self-training · 0.5recommender systems · 0.5attentive feature matching · 0.5adversarial training · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | 3D Scene Creation and Rendering via Rough Meshes: A Lighting Transfer AvenueabstractThis paper studies how to flexibly integrate reconstructed 3D models into practical 3D modeling pipelines such as 3D scene creation and rendering. Due to the technical difficulty, one can only obtain rough 3D models (R3DMs) for most real objects using existing 3D reconstruction techniques. As a result, physically-based rendering (PBR) would render low-quality images or videos for scenes that are constructed by R3DMs. One promising solution would be representing real-world objects as Neural Fields such as NeRFs, which are able to generate photo-realistic renderings of an object under desired viewpoints. However, a drawback is that the synthesized views through Neural Fields Rendering (NFR) cannot reflect the simulated lighting details on R3DMs in PBR pipelines, especially when object interactions in the 3D scene creation cause local shadows. To solve this dilemma, we propose a lighting transfer network (LighTNet) to bridge NFR and PBR, such that they can benefit from each other. LighTNet reasons about a simplified image composition model, remedies the uneven surface issue caused by R3DMs, and is empowered by several perceptual-motivated constraints and a new Lab angle loss which enhances the contrast between lighting strength and colors. Comparisons demonstrate that LighTNet is superior in synthesizing impressive lighting, and is promising in pushing NFR further in practical 3D modeling workflows. Bowen Cai 0001, Yuqin Liang, Rongfei Jia, Binqiang Zhao, Mingming Gong, Huan Fu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | NeuDA: Neural Deformable Anchor for High-Fidelity Implicit Surface ReconstructionabstractThis paper studies implicit surface reconstruction leveraging differentiable ray casting. Previous works such as IDR [34] and NeuS [27] overlook the spatial context in 3D space when predicting and rendering the surface, thereby may fail to capture sharp local topologies such as small holes and structures. To mitigate the limitation, we propose a flexible neural implicit representation leveraging hierarchical voxel grids, namely Neural Deformable Anchor (NeuDA), for high-fidelity surface reconstruction. NeuDA maintains the hierarchical anchor grids where each vertex stores a 3D position (or anchor) instead of the direct embedding (or feature). We optimize the anchor grids such that different local geometry structures can be adaptively encoded. Besides, we dig into the frequency encoding strategies and introduce a simple hierarchical positional encoding method for the hierarchical anchor structure to flexibly exploit the properties of high-frequency and low-frequency geometry and appearance. Experiments on both the DTU [8] and BlendedMVS [32] datasets demonstrate that NeuDA can produce promising mesh surfaces. Bowen Cai 0001, Jinchi Huang, Rongfei Jia, Chengfei Lv, Huan Fu |
CVPR | 1 |
| 2022 | Ray Priors through Reprojection: Improving Neural Radiance Fields for Novel View ExtrapolationabstractNeural Radiance Fields (NeRF) [22] have emerged as a potent paradigm for representing scenes and synthesizing photo-realistic images. A main limitation of conventional NeRFs is that they often fail to produce high-quality renderings under novel viewpoints that are significantly different from the training viewpoints. In this paper, instead of ex-ploiting few-shot image synthesis, we study the novel view extrapolation setting that (1) the training images can well describe an object, and (2) there is a notable discrepancy between the training and test viewpoints' distributions. We present RapNeRF (RAy Priors) as a solution. Our insight is that the inherent appearances of a 3D surface's arbitrary visible projections should be consistent. We thus propose a random ray casting policy that allows training unseen views using seen views. Furthermore, we show that a ray atlas pre-computed from the observed rays' viewing directions could further enhance the rendering quality for ex-trapolated views. A main limitation is that RapNeRF would remove the strong view-dependent effects because it lever-ages the multi-view consistency property. Yuanqing Zhang, Huan Fu, Xiaowei Zhou 0001, Bowen Cai 0001, Jinchi Huang, Rongfei Jia, Binqiang Zhao |
CVPR | 5 |
| 2022 | Digging into Radiance Grid for Real-Time View Synthesis with Detail Preservation
Jinchi Huang, Bowen Cai 0001, Huan Fu, Mingming Gong, Chaohui Wang, Hongchen Luo, Rongfei Jia, Binqiang Zhao |
ECCV (15) | 3 |
| 2021 | Exploiting Diverse Characteristics and Adversarial Ambivalence for Domain Adaptive SegmentationabstractAdapting semantic segmentation models to new domains is an important but challenging problem. Recently enlightening progress has been made, but the performance of existing methods is unsatisfactory on real datasets where the new target domain comprises of heterogeneous sub-domains (e.g. diverse weather characteristics). We point out that carefully reasoning about the multiple modalities in the target domain can improve the robustness of adaptation models. To this end, we propose a condition-guided adaptation framework that is empowered by a special attentive progressive adversarial training (APAT) mechanism and a novel self-training policy. The APAT strategy progressively performs condition-specific alignment and attentive global feature matching. The new self-training scheme exploits the adversarial ambivalences of easy and hard adaptation regions and the correlations among target sub-domains effectively. We evaluate our method (DCAA) on various adaptation scenarios where the target images vary in weather conditions. The comparisons against baselines and the state-of-the-art approaches demonstrate the superiority of DCAA over the competitors. Bowen Cai 0001, Huan Fu, Rongfei Jia, Binqiang Zhao |
AAAI | 1 |
| 2021 | 3D-FRONT: 3D Furnished Rooms with layOuts and semaNTicsabstractWe introduce 3D-FRONT (3D Furnished Rooms with layOuts and semaNTics), a new, large-scale, and comprehensive repository of synthetic indoor scenes highlighted by professionally designed layouts and a large number of rooms populated by high-quality textured 3D models with style compatibility. From layout semantics down to texture details of individual objects, our dataset is freely available to the academic community and beyond. Currently, 3D-FRONT contains 6,813 CAD houses, where 18,968 rooms diversely furnished by 3D objects, far surpassing all publicly available scene datasets. The 13,151 furniture objects all come with high-quality textures. While the floorplans and layout designs (i.e., furniture arrangements) are directly sourced from professional creations, the interior designs in terms of furniture styles, color, and textures have been carefully curated based on a recommender system we develop to attain consistent styles as expert designs. Furthermore, we release Trescope, a light-weight rendering tool, to support benchmark rendering of 2D images and annotations from 3D-FRONT. We demonstrate two applications, interior scene synthesis and texture synthesis, that are especially tailored to the strengths of our new dataset. Huan Fu, Bowen Cai 0001, Lin Gao 0004, Lingxiao Zhang, Cao Li, Qixun Zeng, Chengyue Sun, Rongfei Jia, Binqiang Zhao, Hao (Richard) Zhang |
ICCV | 2 |
| 2018 | Inshore Ship Detection Based on Mask R-CNNabstractInshore ship detection is a popular research domain for optical remote sensing image understanding with many applications in harbor management. However, recent approaches on inshore ship detection depend heavily on hand-crafted features, which need a complicated procedure. In this paper, we propose a new method to achieve inshore ship detection based on Mask R-CNN. We introduce Soft-Non-Maximum Suppression (Soft-NMS) into our framework to improve the robustness to nearby inshore ships. Both battleships and merchantships can be detected in our framework. Furthermore, our framework can also obtain the binary masks of inshore ships. Experimental results on a dataset collected from Google Earth have quantitatively and qualitatively demonstrated the effectiveness of our approach. Shanlan Nie, Zhiguo Jiang 0001, Haopeng Zhang 0001, Bowen Cai 0001 |
IGARSS | 4 |
| 2018 | Star Image Simulation and Subpixel Centroiding for an Earth Observing SensorabstractIn this paper, a novel solution is introduced for accurate subpixel star centroiding of the focused geostationary Earth observing sensor on a three-axis stabilized satellite. A small 2-dimensional array is utilized to better capture the star spot than the linear array of detectors. The popular center of mass method is used to compute star centroid in a single frame. Then the subpixel accuracy of star centroiding can be improved by fitting the linear trajectory of the observed star according to exact imaging time produced from time awarding system of the satellite. Experimental results on simulated star images in various conditions validate the effectiveness and robustness of our star centroiding method. Haopeng Zhang 0001, Bowen Cai 0001, Zhiguo Jiang 0001 |
IGARSS | 3 |
| 2018 | Online Exemplar-Based Fully Convolutional Network for Aircraft Detection in Remote Sensing ImagesabstractConvolutional neural network obtains remarkable achievements on target detection, due to its prominent capability on feature extraction. However, it still needs further study for aircraft detection task, since intraclass variation still restricts the accuracy of aircraft detection in remote sensing images. In this letter, we adopt regularity of aircraft circle response to design our end-to-end fully convolutional network (FCN), and embed online exemplar mining into our network to handle intraclass variation. The mined exemplars are employed to capture different intraclass characteristics, which effectively reduces the burden of network training. Specifically, we first select basic exemplars based on labeled information and initialize the relationships between exemplars and aircraft examples. Then, these relationships will be updated by the similarity of these examples in high-level features space. Finally, aircraft examples will be used to train different exemplar detectors according to updated relationships. Motivated by the geometric shape of aircraft, a circle response map is developed to construct our FCN to achieve more efficient aircraft detection. The comparative experiments indicate that superior performance of our network in accurate and efficient aircraft detection. Bowen Cai 0001, Zhiguo Jiang 0001, Haopeng Zhang 0001, Shanlan Nie |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2017 | Training deep convolution neural network with hard example mining for airport detectionabstractThe geometrical characteristic and low-level manually designed features are usually used to detect airports in optical remote sensing images. But it is insufficient to describe airport in low resolution and illumination environment. This paper presents a hard example mining algorithm to train the end-to-end deep convolutional neural network for airport detection in complex situation. Compared with conventional airport detection methods which design specific low-level manually designed features for high-resolution remote sensing images, an end-to-end network can mine the general characteristic among the training samples and learn high-level features in multi-scale and multi-view remote sensing images. Meanwhile, an automatic hard example mining principle is introduced to make training more efficiently and accurately. The proposed method is validated on a multi-scale and multi-view dataset collected from Google Earth. The experimental results demonstrate that the proposed method is robust and efficient, and superior to the state-of-the-art airport detection models. Bowen Cai 0001, Zhiguo Jiang 0001, Haopeng Zhang 0001 |
IGARSS | 1 |
| 2017 | Region proposal for ship detection based on structured forests edge methodabstractRemote sensing images are with the characteristics of large width and sparse distribution of specific targets, so that the extraction of region proposal is necessary before detection. In this paper, we propose a new ship detection method on sea-background remote sensing images, which are generally influenced by clouds, waves and other inhomogeneities. Instead of exhaustive search, the core of our method is that the region proposals are obtained from edge detection based on structured forests, which makes our method accurate and efficient. This edge detection method only demands a small training set and then produces contours with the background suppressed. After some morphological processing on the contours, we obtained ship proposals by connected domain detection. Adopting support vector machine(SVM) as classifier, we finally acquire ship detection results. The remote sensing images in our datasets are downloaded from Google Earth map. In our experiments, the proposed method is feasible and effective, and it shows better performance than other methods especially in various illumination and interference conditions. Zhiguo Jiang 0001, Haopeng Zhang 0001, Bowen Cai 0001 |
IGARSS | 4 |
| 2017 | Chimney and condensing tower detection based on faster R-CNN in high resolution remote sensing imagesabstractThe persistent haze weather in North China has aroused extensive attention to environmental protection. Among all pollution resources, the anthropogenic emission by fossil fuel power plants plays an important role. To assist the environmental protection administration monitoring fossil fuel power plants, we propose an effective approach in this paper to learn an integrated model for chimney and condensing tower detection based on Faster R-CNN in high resolution remote sensing images. Our method can detect chimneys and condensing towers under different imaging condition efficiently and accurately. Experimental results on a self-collected dataset demonstrate the effectiveness of the proposed method. Zhiguo Jiang 0001, Haopeng Zhang 0001, Bowen Cai 0001, Gang Meng, Deshan Zuo |
IGARSS | 4 |
| 2016 | Semi-supervised Conditional Random Field for hyperspectral remote sensing image classificationabstractConditional Random Field(CRF) has been successfully applied to the hyperspectral image classification. However, it suffers from the availability of large amount of labeled pixels, which is labor- and time-consuming to obtain in practice. In this paper, a semi-supervised CRF(ssCRF) is proposed for hyperspectral image classification with limited labeled pixels. Laplacian Support Vector Machine(LapSVM), after extended into the composite kernel type, is defined as the association potential. And the Potts model is utilized as the interaction potential. The ssCRF is evaluated on the two benchmarks and the results show the effectiveness of ssCRF. Zhiguo Jiang 0001, Haopeng Zhang 0001, Bowen Cai 0001, Quanmao Wei |
IGARSS | 4 |