EDBT 2026 Demo / reviewers in the wild / expert
Zhuoyuan Chen
dblp:75/801
· DBLP profile ↗
18ranked-venue papers
7as first author
1since 2021 · last 2022
0000-0002-5814-925XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 7 first-authorArtificial intelligence and machine learning · 9 · 5 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorSystems, architecture and hardware · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
3D vision · 33% Reinforcement learning · 26% Representation and self-supervised learning · 14% | |
| Computer graphics and multimedia
3 papers |
Visual content generation and editing · 53% Image and video processing · 47% | |
| Human-computer interaction and pervasive computing
1 paper |
Games and playful interaction · 100% |
Topics — the 25 heaviest of 27, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision › 3d generation
3d scene generation |
0.6 | 1 | 2022 | GAUDI: A Neural Architect for Immersive 3D Scene Generation · NeurIPS 2022 |
Machine learning › Reinforcement learning › deep reinforcement learning
alphazero |
0.4 | 1 | 2019 | ELF OpenGo: an analysis and open reimplementation of AlphaZero · ICML 2019 |
Machine learning › Reinforcement learning
deep reinforcement learning |
0.4 | 1 | 2019 | ELF OpenGo: an analysis and open reimplementation of AlphaZero · ICML 2019 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
game playing |
0.4 | 1 | 2019 | ELF OpenGo: an analysis and open reimplementation of AlphaZero · ICML 2019 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
self-play |
0.4 | 1 | 2019 | ELF OpenGo: an analysis and open reimplementation of AlphaZero · ICML 2019 |
Visual content generation and editing › 3d content generation
3d generative modeling |
0.4 | 1 | 2019 | Order-Aware Generative Modeling Using the 3D-Craft Dataset · ICCV 2019 |
Games and playful interaction
board games |
0.4 | 1 | 2019 | ELF OpenGo: an analysis and open reimplementation of AlphaZero · ICML 2019 |
Computer vision › 3D vision › stereo vision › stereo matching
matching cost |
0.2 | 1 | 2015 | A Deep Visual Correspondence Embedding Model for Stereo Matching Costs · ICCV 2015 |
Computer vision › 3D vision › stereo vision
stereo matching |
0.2 | 1 | 2015 | A Deep Visual Correspondence Embedding Model for Stereo Matching Costs · ICCV 2015 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding
dictionary learning |
0.2 | 1 | 2013 | Robust Dictionary Learning by Error Source Decomposition · ICCV 2013 |
Computer vision › 3D vision › motion estimation › optical flow
large displacement optical flow |
0.2 | 1 | 2013 | Large Displacement Optical Flow from Nearest Neighbor Fields · CVPR 2013 |
Computer vision › 3D vision › motion estimation
optical flow |
0.2 | 1 | 2013 | Large Displacement Optical Flow from Nearest Neighbor Fields · CVPR 2013 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding › dictionary learning
robust dictionary learning |
0.2 | 1 | 2013 | Robust Dictionary Learning by Error Source Decomposition · ICCV 2013 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
sparse coding |
0.2 | 1 | 2013 | Robust Dictionary Learning by Error Source Decomposition · ICCV 2013 |
Computer vision › Video understanding and tracking › action recognition
3d action recognition |
0.1 | 1 | 2012 | Robust 3D Action Recognition with Random Occupancy Patterns · ECCV (2) 2012 |
Image and video processing
motion estimation |
0.1 | 1 | 2012 | Decomposing and regularizing sparse/non-sparse components for motion field estimation · CVPR 2012 |
Computer vision › Video understanding and tracking
action recognition |
0.1 | 1 | 2011 | Action recognition with multiscale spatio-temporal contexts · CVPR 2011 |
Machine learning › Representation and self-supervised learning › visual representation › image representation
bag of visual words |
0.1 | 1 | 2011 | Action recognition with multiscale spatio-temporal contexts · CVPR 2011 |
Computer vision › Video understanding and tracking › spatio-temporal modeling
spatio-temporal context modeling |
0.1 | 1 | 2011 | Action recognition with multiscale spatio-temporal contexts · CVPR 2011 |
Image and video processing
image matting |
0.1 | 1 | 2009 | Auto-cut for web images · ACM Multimedia 2009 |
Image and video processing
image segmentation |
0.1 | 1 | 2009 | Auto-cut for web images · ACM Multimedia 2009 |
Computer vision › 3D vision
depth estimation |
0.1 | 1 | 2015 | A Deep Visual Correspondence Embedding Model for Stereo Matching Costs · ICCV 2015 |
Computer vision › Video understanding and tracking
motion segmentation |
0.0 | 1 | 2013 | Large Displacement Optical Flow from Nearest Neighbor Fields · CVPR 2013 |
Computer vision › 3D vision
3d shape representation |
0.0 | 1 | 2012 | Robust 3D Action Recognition with Random Occupancy Patterns · ECCV (2) 2012 |
Information retrieval
image retrieval |
0.0 | 1 | 2009 | Auto-cut for web images · ACM Multimedia 2009 |
Methods — techniques the papers use, named apart from their topics
monte carlo tree search · 0.8imitation learning · 0.8deep neural network · 0.8radiance field · 0.6latent representation learning · 0.6voxel CNN · 0.4VoxelCNN · 0.4embedding · 0.2convolutional neural network · 0.2hierarchical clustering · 0.2global cost function · 0.2fuzzy matting · 0.2nearest neighbor fields · 0.2local deformation · 0.2robust statistics · 0.1optimization · 0.1influence function · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | GAUDI: A Neural Architect for Immersive 3D Scene GenerationabstractWe introduce GAUDI, a generative model capable of capturing the distribution of complex and realistic 3D scenes that can be rendered immersively from a moving camera. We tackle this challenging problem with a scalable yet powerful approach, where we first optimize a latent representation that disentangles radiance fields and camera poses. This latent representation is then used to learn a generative model that enables both unconditional and conditional generation of 3D scenes. Our model generalizes previous works that focus on single objects by removing the assumption that the camera pose distribution can be shared across samples. We show that GAUDI obtains state-of-the-art performance in the unconditional generative setting across multiple datasets and allows for conditional generation of 3D scenes given conditioning variables like sparse image observations or text that describes the scene. Miguel Ángel Bautista 0001, Pengsheng Guo, Samira Abnar, Walter Talbott, Alexander Toshev, Zhuoyuan Chen, Laurent Dinh, Shuangfei Zhai, Hanlin Goh, Daniel Ulbricht, Afshin Dehghan, Joshua M. Susskind |
NeurIPS | 6 |
| 2019 | Order-Aware Generative Modeling Using the 3D-Craft DatasetabstractResearch on 2D and 3D generative models typically focuses on the final artifact being created, e.g., an image or a 3D structure. Unlike 2D image generation, the generation of 3D objects in the real world is commonly constrained by the process and order in which the object is constructed. For instance, gravity needs to be taken into account when building a block tower. In this paper, we explore the prediction of ordered actions to construct 3D objects. Instead of predicting actions based on physical constraints, we propose learning through observing human actions. To enable large-scale data collection, we use the Minecraft1 environment. We introduce 3D-Craft, a new dataset of 2,500 Minecraft houses each built by human players sequentially from scratch. To learn from these human action sequences, we propose an order-aware 3D generative model called VoxelCNN. In contrast to other 3D generative models which either have no explicit order (e.g. holistic generation with 3DGAN [35]), or follow a simple heuristic order (e.g. raster-scan), VoxelCNN is trained to imitate human building order with spatial awareness. We also transferred the order to other dataset such as ShapeNet[10]. The 3D-Craft dataset, models, and benchmark system will be made publicly available, which may inspire new directions for future research exploration. https://github.com/facebookresearch/VoxelCNN. Zhuoyuan Chen, Kavya Srinet, Charles R. Qi, Haoqi Fan 0001, Jerry Ma, C. Lawrence Zitnick, Demi Guo, Tong Xiao 0003, Saining Xie, Xinlei Chen, Arthur Szlam, Shubham Tulsiani, Haonan Yu, Jonathan Gray |
ICCV | 1 |
| 2019 | ELF OpenGo: an analysis and open reimplementation of AlphaZeroabstractThe AlphaGo, AlphaGo Zero, and AlphaZero series of algorithms are remarkable demonstrations of deep reinforcement learning’s capabilities, achieving superhuman performance in the complex game of Go with progressively increasing autonomy. However, many obstacles remain in the understanding of and usability of these promising approaches by the research community. Toward elucidating unresolved mysteries and facilitating future research, we propose ELF OpenGo, an open-source reimplementation of the AlphaZero algorithm. ELF OpenGo is the first open-source Go AI to convincingly demonstrate superhuman performance with a perfect (20:0) record against global top professionals. We apply ELF OpenGo to conduct extensive ablation studies, and to identify and analyze numerous interesting phenomena in both the model training and in the gameplay inference procedures. Our code, models, selfplay datasets, and auxiliary data are publicly available. Yuandong Tian, Jerry Ma, Qucheng Gong, Shubho Sengupta, Zhuoyuan Chen, James Pinkerton, C. Lawrence Zitnick |
ICML | 5 |
| 2015 | A Deep Visual Correspondence Embedding Model for Stereo Matching CostsabstractThis paper presents a data-driven matching cost for stereo matching. A novel deep visual correspondence embedding model is trained via Convolutional Neural Network on a large set of stereo images with ground truth disparities. This deep embedding model leverages appearance data to learn visual similarity relationships between corresponding image patches, and explicitly maps intensity values into an embedding feature space to measure pixel dissimilarities. Experimental results on KITTI and Middlebury data sets demonstrate the effectiveness of our model. First, we prove that the new measure of pixel dissimilarity outperforms traditional matching costs. Furthermore, when integrated with a global stereo framework, our method ranks top 3 among all two-frame algorithms on the KITTI benchmark. Finally, cross-validation results show that our model is able to make correct predictions for unseen data which are outside of its labeled training set. Zhuoyuan Chen, Liang Wang 0001, Yinan Yu, Chang Huang |
ICCV | 1 |
| 2013 | Large Displacement Optical Flow from Nearest Neighbor FieldsabstractWe present an optical flow algorithm for large displacement motions. Most existing optical flow methods use the standard coarse-to-fine framework to deal with large displacement motions which has intrinsic limitations. Instead, we formulate the motion estimation problem as a motion segmentation problem. We use approximate nearest neighbor fields to compute an initial motion field and use a robust algorithm to compute a set of similarity transformations as the motion candidates for segmentation. To account for deviations from similarity transformations, we add local deformations in the segmentation process. We also observe that small objects can be better recovered using translations as the motion candidates. We fuse the motion results obtained under similarity transformations and under translations together before a final refinement. Experimental validation shows that our method can successfully handle large displacement motions. Although we particularly focus on large displacement motions in this work, we make no sacrifice in terms of overall performance. In particular, our method ranks at the top of the Middlebury benchmark. Zhuoyuan Chen, Hailin Jin, Zhe Lin 0001, Scott Cohen, Ying Wu 0001 |
CVPR | 1 |
| 2013 | Robust Dictionary Learning by Error Source DecompositionabstractSparsity models have recently shown great promise in many vision tasks. Using a learned dictionary in sparsity models can in general outperform predefined bases in clean data. In practice, both training and testing data may be corrupted and contain noises and outliers. Although recent studies attempted to cope with corrupted data and achieved encouraging results in testing phase, how to handle corruption in training phase still remains a very difficult problem. In contrast to most existing methods that learn the dictionary from clean data, this paper is targeted at handling corruptions and outliers in training data for dictionary learning. We propose a general method to decompose the reconstructive residual into two components: a non-sparse component for small universal noises and a sparse component for large outliers, respectively. In addition, further analysis reveals the connection between our approach and the ``partial'' dictionary learning approach, updating only part of the prototypes (or informative code words) with remaining (or noisy code words) fixed. Experiments on synthetic data as well as real applications have shown satisfactory performance of this new robust dictionary learning approach. Zhuoyuan Chen, Ying Wu 0001 |
ICCV | 1 |
| 2012 | Decomposing and regularizing sparse/non-sparse components for motion field estimationabstractRegularizing motion field is critical to achieve accurate estimation of the motion field. As the motion field may include discontinuity (e.g., at the motion boundaries), traditional smoothness regularization may not work well. Among many approaches to handling motion discontinuity, recent attempts pursued a sparse representation of the motion field for regularization, and achieved quite encouraging results. However, statistics show that these methods tend to over-sparsify the motion field, and thus confronted by the non-sparse noise in practice. In this paper, we propose to decompose the motion field into sparse and non-sparse components for the motion boundaries and small universal noises, respectively. This separation approach regularizes these two sources differently. We propose a novel and efficient optimization algorithm to solve this problem. In addition, our study reveals the in-depth connection between this noise separation approach and the influence function approach in robust statistics. We validate and evaluate our new approach on the Middlebury benchmark, and have achieved outstanding testing performance. Zhuoyuan Chen, Jiang Wang 0001, Ying Wu 0001 |
CVPR | 1 |
| 2012 | Robust 3D Action Recognition with Random Occupancy Patterns
Jiang Wang 0001, Zicheng Liu 0001, Jan Chorowski, Zhuoyuan Chen, Ying Wu 0001 |
ECCV (2) | 4 |
| 2011 | Action recognition with multiscale spatio-temporal contextsabstractThe popular bag of words approach for action recognition is based on the classifying quantized local features density. This approach focuses excessively on the local features but discards all information about the interactions among them. Local features themselves may not be discriminative enough, but combined with their contexts, they can be very useful for the recognition of some actions. In this paper, we present a novel representation that captures contextual interactions between interest points, based on the density of all features observed in each interest point's mutliscale spatio-temporal contextual domain. We demonstrate that augmenting local features with our contextual feature significantly improves the recognition performance. Jiang Wang 0001, Zhuoyuan Chen, Ying Wu 0001 |
CVPR | 2 |
| 2011 | A Compressive Sensing Reconstruction Algorithm for Trinary and Binary Sparse Signals Using Pre-mappingabstractIn this paper, we first analyze impact of the distribution of sparse signals on reconstruction quality in compressive sensing through experimental results and heuristic analysis. We suggest that trinary/binary sparse signals are one of the most difficult signals to reconstruct in terms of error bounds. We then show that by incorporating linear or non-linear mapping prior to sensing, significant improvement in the recovery performance can be achieved. Zhuoyuan Chen, Jiangtao Wen, Jianwei Ma 0006, Yuxing Han 0001, John D. Villasenor |
DCC | 2 |
| 2010 | Image Compression Using the DCT and Noiselets: A New Algorithm and Its Rate Distortion PerformanceabstractWe describe an image coding algorithm combining the DCT and noiselet information. The algorithm first transmits DCT information sufficient to reproduce a "low-quality" version of the image at the decoder. This image is then used both at the decoder and encoder to create a mutually known list of locations of likely significant noiselet coefficients. The coefficient values themselves are then transmitted to the decoder differentially, by subtracting, at the encoder, the low-quality image from the original image, obtaining the noiselet values and subjecting them to quantization and entropy coding. There remain significant opportunities for further work combining CS-inspired information theoretic techniques with the rate-distortion considerations that are critical in practical image communications. Zhuoyuan Chen, Jiangtao Wen, Shiqiang Yang, Yuxing Han 0001, John D. Villasenor |
DCC | 1 |
| 2010 | Reconstruction of Sparse Binary Signals Using Compressive SensingabstractSummary form only given. This paper has described an improved algorithm for reconstructing sparse binary signals using compressive sensing. The algorithm is based on the reweighted lqnorm optimization algorithm, but with the important additional operation of bounding in each round of the interior-point method iteration, and progressive reduction of q. Experimental results confirm that the algorithm performs well both in terms of the ability to recover an input signal as well as in terms of speed. We also found that both the progressive reduction and the bounding are integral to the improvement in performance. Future work includes extending this approach to Gaussian distributed, as opposed to binary inputs. Jiangtao Wen, Zhuoyuan Chen, Shiqiang Yang, Yuxing Han 0001, John D. Villasenor |
DCC | 2 |
| 2010 | A compressive sensing image compression algorithm using quantized DCT and noiselet informationabstractInspired by recent theoretical advances in compressive sensing (CS), we propose a new framework that combines the classical local discrete cosine transform used in image compression algorithms such as JPEG with a global noiselet measure which is solved using second order cone programming (SOCP). Jiangtao Wen, Zhuoyuan Chen, Yuxing Han 0001, John D. Villasenor, Shiqiang Yang |
ICASSP | 2 |
| 2009 | Robust inter-scale non-blind image motion deblurringabstractKernel estimate errors and image noise are major causes of visual artifacts in image motion deblurring. We propose an inter-scale non-blind image motion deblurring approach that significantly reduces those artifacts. We use Gaussian Scale Mixture Field of Experts (GSM FOE) model as image prior. The inter-scale smoothness constraint is adopted to suppress the ringing artifacts. In each scale, image details are recovered by the residual deconvolution and the cross bilateral filter (CBF). We further propose a std-controlled CBF to denoise the result. The experimental results are much better than those of previous methods. Chao Wang 0063, Lifeng Sun, Zhuoyuan Chen, Shiqiang Yang, Jianwei Zhang 0001 |
ICIP | 3 |
| 2009 | High-quality non-blind motion deblurringabstractTraditional non-blind motion deblurring methods are sensitive to kernel estimate errors and image noise, thus suffering from either ringing artifacts, enlarged image noise, or over-smoothed image details. We introduce a robust non-blind deblurring algorithm that produces high quality results even from many challenging images with noisy kernels. We adopt the Gaussian Scale Mixture Fields of Experts (GSM FOE) model and the smoothness constraint as image prior, and use the iterative re-weight least-square (IRLS) algorithm to produce the temporal result. The residual deconvolution suite is used to restore the lost image details. We denoise the result using our std-controlled cross bilateral filter. The experimental results are much better than those of previous approaches. Chao Wang 0063, Lifeng Sun, Zhuoyuan Chen, Shiqiang Yang, Jianwei Zhang 0001 |
ICIP | 3 |
| 2009 | The Bilateral Wavelet Pyramid (BWP): A Novel Image RepresentationabstractIn this paper, we study the properties of the bilateral pyramid, based on which a novel image representation-the bilateral wavelet pyramid (BWP) is proposed. The BWP provides a multiresolution technique to manipulate image variances simultaneously in radiometric, spatial and frequency domains. Application to texture retrieval is discussed to demonstrate the effectiveness of the BWP. Chao Wang 0063, Lifeng Sun, Zhuoyuan Chen, Jianwei Zhang 0001, Shiqiang Yang |
ISCAS | 3 |
| 2009 | Statistics in Bilateral Domain: Novel Statistics of Natural ImagesabstractRecently, there has been a great deal of interest in statistics of natural images in linear domain. However, a detail study of image statistics in non-linear domain on a very large dataset is still missing. In this paper, we analyze natural images' statistics in ldquobilateral domainrdquo on different sub bands which are generated by the classical bilateral filter. Compared with those in linear domain, the statistics in bilateral domain are significantly superior in generalization, de-correlation, and local feature extraction. We also observe some entirely new but interesting features. Chao Wang 0063, Lifeng Sun, Zhuoyuan Chen, Jianwei Zhang 0001, Shiqiang Yang |
ISCAS | 3 |
| 2009 | Auto-cut for web imagesabstractIn this paper, we propose a novel automatic algorithm for foreground/background labeling. We aim to generate ROI cutout automatically for further processing such as image editing, classification and information retrieval. Different from traditional semi-supervised segmentation method, we use a rather weak prior on boundary label. Accordingly, a global cost function is proposed to combine our prior knowledge with pixel-level feature. We compute fuzzy matting components as building blocks to construct semantically meaningful mattes. Finally, these mattes are hierarchically clustered and ranked by central preference. Experimental results on a large benchmark data set demonstrate the performance of our algorithm. Zhuoyuan Chen, Lifeng Sun, Shiqiang Yang |
ACM Multimedia | 1 |