VLDB 2026 Research / reviewers in the wild / expert
Shaodi You
dblp:72/7950
· DBLP profile ↗
58ranked-venue papers
6as first author
26since 2021 · last 2026
0000-0001-8973-645XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 42 · 6 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 2 first-author · 14 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Atlantis++: Enabling Underwater Depth Estimation with Stable Diffusion and Beyond
Fan Zhang 0123, Shaodi You, Yu Li 0003, Ying Fu 0001 |
Int. J. Comput. Vis. | 2 |
| 2025 | Sequential Joint Dependency Aware Human Pose Estimation with State Space ModelabstractIn this paper, we present a sequential joint dependency aware model for monocular 2D-to-3D human pose estimation. While existing estimators leverage the (bi)directional joint dependency with graph convolutions and attention, we further propose to exploit the sequential dependency between joints with state space model (SSM). Our sequential dependency takes into consideration the information of kinematic chain, joint hierarchy and the body part. We design a sequential dependency aware representation to transform the pose data into sequential data for our pose SSM module. We tailor the SSM layer in the pose SSM module for pose estimation by learning joint-dependent parameters and introducing pose aware hidden state initialization. Extensive experiments are conducted on two datasets to validate the effectiveness of our proposed SSM module, and the results demonstrate that our pose estimator can deliver impressive performance. Hanxi Yin, Shaodi You, Jungong Han, Zhixiang Chen 0003 |
AAAI | 2 |
| 2025 | Automatic Spectral Calibration of Hyperspectral Images: Method, Dataset and BenchmarkabstractHyperspectral images (HSI) densely sample the world in both the space and frequency domains and, therefore, are more distinctive than RGB images. Usually, HSI needs to be calibrated to minimize the impact of various illumination conditions. The traditional way to calibrate HSI utilizes a physical reference, which involves manual operations, occlusions, and/or limits camera mobility. These limitations inspire this paper to automatically calibrate HSIs using a learning-based method. Towards this goal, a large-scale HSI calibration dataset, which has 765 high-quality HSI pairs covering diversified natural scenes and illuminations, is created. The dataset is further expanded to 7650 pairs by combining with 10 different physically measured illuminations. A spectral illumination transformer (SIT) together with an illumination attention module is proposed. Extensive benchmarks demonstrate the SoTA performance of the proposed SIT. The benchmarks also indicate that low-light conditions are more challenging than normal conditions. The dataset and codes are available online: https://github.com/duranze/Automatic-spectral-calibration-of-HSI. Zhuoran Du, Shaodi You, Shikui Wei |
CVPR | 2 |
| 2025 | LucIE: Language-Guided Local Image Editing for Fashion ImagesabstractLanguage-guided fashion image editing is challenging, as fashion image editing is local and requires high precision, while natural language cannot provide precise visual information for guidance. In this paper, we propose LucIE, a novel unsupervised language-guided local image editing method for fashion images. LucIE adopts and modifies recent text-to-image synthesis network, DF-GAN, as its backbone. However, the synthesis backbone often changes the global structure of the input image, making local image editing impractical. To increase structural consistency between input and edited images, we propose Content-Preserving Fusion Module (CPFM). Different from existing fusion modules, CPFM prevents iterative refinement on visual feature maps and accumulates additive modifications on RGB maps. LucIE achieves local image editing explicitly with language-guided image segmentation and mask-guided image blending while only using image and text pairs. Results on the DeepFashion dataset shows that LucIE achieves state-of-the-art results. Compared with previous methods, images generated by LucIE also exhibit fewer artifacts. We provide visualizations and perform ablation studies to validate LucIE and the CPFM. We also demonstrate and analyze limitations of LucIE, to provide a better understanding of LucIE. Huanglu Wen, Shaodi You, Ying Fu 0001 |
Comput. Vis. Media | 2 |
| 2025 | Learning Rain Location Prior for Nighttime Deraining and BeyondabstractMost deraining methods work on day scenes while leaving nighttime deraining underexplored, where darkness and non-uniform illuminations pose additional challenges. Consequently, night rain has a quite different appearance varying by location and cannot be effectively handled. To accommodate this issue, we propose a Rain Location Prior (RLP) by implicitly learning it from rainy images to reflect rain location information and boost the performance of deraining models by prior injection. Then, we introduce a Rain Prior Injection Module (RPIM) with a multi-scale scheme to modulate it by attention and emphasize the features of rain streak areas for better injection efficiency. Finally, to alleviate the data scarcity issue and facilitate the research on nighttime deraining, we propose the GTAV-NightRain dataset by considering the interaction between rain streaks and non-uniform illuminations, and provide detailed instructions on data collection pipeline which is highly replicable and flexible to integrate challenging factors of rainy night in the future. Our method outperforms state-of-the-art backbone by 1.3 dB in PSNR and generalizes better on real data such as heavy rain and the presence of glow and glaring lights. Ablation studies are conducted to validate the effectiveness of each component and we visualize RLP to show good interpretability. Moreover, we apply our method to daytime deraining and desnow to show good generalizability on other location-dependent degradations. Our method is a step forward in nighttime deraining and the GTAV-NightRain dataset may become a good complement to previous datasets. Fan Zhang 0123, Shaodi You, Yu Li 0003, Ying Fu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | 3D human pose estimation and action recognition using fisheye cameras: A survey and benchmarkabstract3D human pose estimation based on visual information aims to predict 3D poses of humans in images or videos. The aim of human action recognition is to classify what kind of actions people do. Both topics are widely studied in the field of computer vision. Existing methods mainly focus on 3D human pose estimation and human action recognition using images/videos recorded by perspective cameras. In contrast to perspective cameras, fisheye cameras use wide-angle lenses capturing wider field-of-views (FOV). Fisheye cameras are used in many applications such as surveillance and autonomous driving. In this paper, a survey is given on monocular 3D human pose estimation and action recognition. A new benchmark dataset is proposed using a fisheye camera to quantitatively compare and analyze existing methods. Shaodi You, Sezer Karaoglu, Theo Gevers |
Pattern Recognit. | 2 |
| 2024 | Learning Generalized Segmentation for Foggy-Scenes by Bi-directional Wavelet GuidanceabstractLearning scene semantics that can be well generalized to foggy conditions is important for safety-crucial applications such as autonomous driving. Existing methods need both annotated clear images and foggy images to train a curriculum domain adaptation model. Unfortunately, these methods can only generalize to the target foggy domain that has seen in the training stage, but the foggy domains vary a lot in both urban-scene styles and fog styles. In this paper, we propose to learn scene segmentation well generalized to foggy-scenes under the domain generalization setting, which does not involve any foggy images in the training stage and can generalize to any arbitrary unseen foggy scenes. We argue that an ideal segmentation model that can be well generalized to foggy-scenes need to simultaneously enhance the content, de-correlate the urban-scene style and de-correlate the fog style. As the content (e.g., scene semantic) rests more in low-frequency features while the style of urban-scene and fog rests more in high-frequency features, we propose a novel bi-directional wavelet guidance (BWG) mechanism to realize the above three objectives in a divide-and-conquer manner. With the aid of Haar wavelet transformation, the low frequency component is concentrated on the content enhancement self-attention, while the high frequency component is shifted to the style and fog self-attention for de-correlation purpose. It is integrated into existing mask-level Transformer segmentation pipelines in a learnable fashion. Large-scale experiments are conducted on four foggy-scene segmentation datasets under a variety of interesting settings. The proposed method significantly outperforms existing directly-supervised, curriculum domain adaptation and domain generalization segmentation methods. Source code is available at https://github.com/BiQiWHU/BWG. Qi Bi, Shaodi You, Theo Gevers |
AAAI | 2 |
| 2024 | Learning Content-Enhanced Mask Transformer for Domain Generalized Urban-Scene SegmentationabstractDomain-generalized urban-scene semantic segmentation (USSS) aims to learn generalized semantic predictions across diverse urban-scene styles. Unlike generic domain gap challenges, USSS is unique in that the semantic categories are often similar in different urban scenes, while the styles can vary significantly due to changes in urban landscapes, weather conditions, lighting, and other factors. Existing approaches typically rely on convolutional neural networks (CNNs) to learn the content of urban scenes. In this paper, we propose a Content-enhanced Mask TransFormer (CMFormer) for domain-generalized USSS. The main idea is to enhance the focus of the fundamental component, the mask attention mechanism, in Transformer segmentation models on content information. We have observed through empirical analysis that a mask representation effectively captures pixel segments, albeit with reduced robustness to style variations. Conversely, its lower-resolution counterpart exhibits greater ability to accommodate style variations, while being less proficient in representing pixel segments. To harness the synergistic attributes of these two approaches, we introduce a novel content-enhanced mask attention mechanism. It learns mask queries from both the image feature and its down-sampled counterpart, aiming to simultaneously encapsulate the content and address stylistic variations. These features are fused into a Transformer decoder and integrated into a multi-resolution content-enhanced mask attention learning scheme. Extensive experiments conducted on various domain-generalized urban-scene segmentation datasets demonstrate that the proposed CMFormer significantly outperforms existing CNN-based methods by up to 14.0% mIoU and the contemporary HGFormer by up to 1.7% mIoU. The source code is publicly available at https://github.com/BiQiWHU/CMFormer. Qi Bi, Shaodi You, Theo Gevers |
AAAI | 2 |
| 2024 | Atlantis: Enabling Underwater Depth Estimation with Stable DiffusionabstractMonocular depth estimation has experienced significant progress on terrestrial images in recent years thanks to deep learning advancements. But it remains inadequate for underwater scenes primarily due to data scarcity. Given the inherent challenges of light attenuation and backscat-ter in water, acquiring clear underwater images or precise depth is notably difficult and costly. To mitigate this issue, learning-based approaches often rely on synthetic data or turn to self- or unsupervised manners. Nonetheless, their performance is often hindered by domain gap and looser constraints. In this paper, we propose a novel pipeline for generating photorealistic underwater images using accurate terrestrial depth. This approach facilitates the supervised training of models for underwater depth estimation, effectively reducing the performance disparity between ter-restrial and underwater environments. Contrary to previous synthetic datasets that merely apply style transfer to terres-trial images without scene content change, our approach uniquely creates vivid non-existent underwater scenes by leveraging terrestrial depth data through the innovative Stable Diffusion model. Specifically, we introduce a specialized Depth2Underwater ControlNet, trained on prepared {Underwater, Depth, Text} data triplets, for this generation task. Our newly developed dataset, Atlantis, enables terres-trial depth estimation models to achieve considerable improvements on unseen underwater scenes, surpassing their terrestrial pretrained counterparts both quantitatively and qualitatively. Moreover, we further show its practical utility by applying the improved depth in underwater image enhancement, and its smaller domain gap from the LLVM perspective. Code and dataset are publicly available at https://github.com/zkawfanx/Atlantis. Fan Zhang 0123, Shaodi You, Yu Li 0003, Ying Fu 0001 |
CVPR | 2 |
| 2024 | Kinship similarity for open sets
Wei Wang 0469, Shaodi You, Sezer Karaoglu, Theo Gevers |
Pattern Recognit. | 2 |
| 2023 | Learning Rain Location Prior for Nighttime DerainingabstractRain can significantly degrade image quality and visibility, making deraining a critical area of research in computer vision. Despite recent progress in learning-based deraining methods, there is a lack of focus on nighttime deraining due to the unique challenges posed by non-uniform local illuminations from artificial light sources. Rain streaks in these scenes have diverse appearances that are tightly related to their relative positions to light sources, making it difficult for existing deraining methods to effectively handle them. In this paper, we highlight the importance of rain streak location information in nighttime deraining. Specifically, we propose a Rain Location Prior (RLP) that is learned implicitly from rainy images using a recurrent residual model. This learned prior contains location information of rain streaks and, when injected into deraining models, can significantly improve their performance. To further improve the effectiveness of the learned prior, we also propose a Rain Prior Injection Module (RPIM) to modulate the prior before injection, increasing the importance of features within rain streak areas. Experimental results demonstrate that our approach outperforms existing state-of-the-art methods by about 1dB and effectively improves the performance of deraining models. We also evaluate our method on real night rainy images to show the capability to handle real scenes with fully synthetic data for training. Our method represents a significant step forward in the area of nighttime deraining and highlights the importance of location information in this challenging problem. The code is publicly available at https://github.com/zkawfanx/RLP. Fan Zhang 0123, Shaodi You, Yu Li 0003, Ying Fu 0001 |
ICCV | 2 |
| 2023 | Learning rotation equivalent scene representation from instance-level semantics: A novel top-down perspectiveabstractThis paper focuses on rotation variant scene recognition. Different from existing rotation invariant recognition approaches which learn from either rotated images or rotated convolutional filters in a bottom-up manner, a new top-down perspective by learning is explored from instance-level semantic representation. The goal is to eliminate the convolutional feature differences in bottom-up feature propagation caused by the rotation sensitive nature of convolution operation. Our rotation equivalent convolutional neural network (RE-CNN) scheme consists of three components. Firstly, our key instance selection module highlights the instances strongly related to the scene scheme regardless of their orientation. Secondly, our key instance aggregation module builds a scene representation invariant to the position change of each instance caused by rotation. Finally, our semantic fusion module allows the framework to be organized as a whole and implements rotation regularization. Notably, our RE-CNN scheme can be adapted to existing CNNs in a plug-in-and-play manner. Extensive experiments on rotation variant scene recognition benchmarks from four domains demonstrate the state-of-the-art performance and generalization capability of the proposed RE-CNN. Qi Bi, Shaodi You, Wei Ji 0011, Theo Gevers |
Comput. Vis. Image Underst. | 2 |
| 2023 | A survey on kinship verificationabstractIn this survey, kinship verification is defined as the automatic process of verifying whether two or more persons are blood relatives (kin) by analyzing images of their faces. Kinship verification is an important research field in computer vision with many applications such as finding missing persons, family album organization, and online image search. Although substantial progress has been made in kinship verification in the past decade, there are still challenges such as intrinsic (face i.e., differences in facial appearance) and extrinsic (acquisition i.e., varying imaging conditions) problems. And there is still a demand for more diverse datasets. Therefore, this paper provides a survey on kinship verification methods and datasets. The survey starts with the definition of kinship verification and its corresponding intrinsic and extrinsic challenges. Then, an overview of kinship verification methods and datasets is given. Finally, a new multi-modal dataset (Nemo-Kinship Dataset) is proposed as a benchmark dataset addressing large inter-subject age variations consisting of 4216 videos of 248 persons from 85 families. The newly collected dataset is used to systematically test and analyze state-of-the-art methods. Wei Wang 0469, Shaodi You, Sezer Karaoglu, Theo Gevers |
Neurocomputing | 2 |
| 2023 | Interactive Learning of Intrinsic and Extrinsic Properties for All-Day Semantic SegmentationabstractScene appearance changes drastically throughout the day. Existing semantic segmentation methods mainly focus on well-lit daytime scenarios and are not well designed to cope with such great appearance changes. Naively using domain adaption does not solve this problem because it usually learns a fixed mapping between the source and target domain and thus have limited generalization capability on all-day scenarios (i. e., from dawn to night). In this paper, in contrast to existing methods, we tackle this challenge from the perspective of image formulation itself, where the image appearance is determined by both intrinsic (e. g., semantic category, structure) and extrinsic (e. g., lighting) properties. To this end, we propose a novel intrinsic-extrinsic interactive learning strategy. The key idea is to interact between intrinsic and extrinsic representations during the learning process under spatial-wise guidance. In this way, the intrinsic representation becomes more stable and, at the same time, the extrinsic representation gets better at depicting the changes. Consequently, the refined image representation is more robust to generate pixel-wise predictions for all-day scenarios. To achieve this, we propose an All-in-One Segmentation Network (AO-SegNet) in an end-to-end manner. Large scale experiments are conducted on three real datasets (Mapillary, BDD100K and ACDC) and our proposed synthetic All-day CityScapes dataset. The proposed AO-SegNet shows a significant performance gain against the state-of-the-art under a variety of CNN and ViT backbones on all the datasets. Qi Bi, Shaodi You, Theo Gevers |
IEEE Trans. Image Process. | 2 |
| 2022 | Pose Guided Human Motion Transfer by Exploiting 2D and 3D InformationabstractHuman motion transfer aims to animate the pose of a human in a source image driven by the poses of a human in a target video. To warp (transfer) human poses, most of the existing methods are based on optical flow or affine transformations as an intermediate representation followed by a generator module to perform the motion transfer. Existing methods perform well in terms of reconstruction quality. However, the quality of the human pose transfer has received less attention although it is an important part of the motion transfer process. Therefore, in this paper, we propose a method focusing on both the reconstruction quality as well as pose consistency. In contrast to existing methods, performing warping procedures in 2D- or 3D-space, we introduce a strategy to combine the warped features in both 2D- and 3D-space to alleviate the self-occlusion problem. In this way, our method benefits from 2D (robustness) and 3D (steering) information to guide the generation process. To reduce the pose error caused by inaccurate 3D estimation, a method is proposed to maintain semantic consistency between predictions and target images at arm and leg regions. Experiments conducted on large scale datasets show that the proposed method outperforms existing methods. Ablation studies clarify the benefits of using feature fusion and semantic consistency. Shaodi You, Sezer Karaoglu, Theo Gevers |
3DV | 2 |
| 2022 | Multi-person 3D pose estimation from a single image captured by a fisheye cameraabstractMulti-person 3D pose estimation with absolute depths for a fisheye camera is a challenging task but with valuable applications in daily life, especially for video surveillance. However, to the best of our knowledge, such problem has not been explored so far, leaving a gap in practical applications. In this work, we first propose a method for multi-person 3D pose estimation from a single image taken by a fisheye camera. Our method consists of two branches to estimate absolute 3D human poses: (1) a 2D-to-3D lifting module to predict root-relative 3D human poses (HPoseNet); (2) a root regression module to estimate absolute root locations in the camera coordinate (HRootNet). Finally, we propose a fisheye re-projection module without using ground-truth camera parameters to connect two branches, alleviating the impact of image distortions on 3D pose estimation and further regularizing prediction absolute 3D poses. Experimental results demonstrate that our method achieves the state-of-the-art performance on two public multi-person 3D pose datasets with synthetic fisheye images and our newly collected dataset with real fisheye images. The code and new dataset will be made publicly available. Shaodi You, Sezer Karaoglu, Theo Gevers |
Comput. Vis. Image Underst. | 2 |
| 2022 | Artificial Intelligence for Dunhuang Cultural Heritage Protection: The Project and the DatasetabstractAbstract In this work, we introduce our project on Dunhuang cultural heritage protection using artificial intelligence. The Dunhuang Mogao Grottoes in China, also known as the Grottoes of the Thousand Buddhas, is a religious and cultural heritage located on the Silk Road. The grottoes were built from the 4th century to the 14th century. After thousands of years, the in grottoes decaying is serious. In addition, numerous historical records were destroyed throughout the years, making it difficult for archaeologists to reconstruct history. We aim to use modern computer vision and machine learning technologies to solve such challenges. First, we propose to use deep networks to automatically perform the restoration. Through out experiments, we find the automated restoration can provide comparable quality as those manually restored from an archaeologist. This can significantly speed up the restoration given the enormous size of the historical paintings. Second, we propose to use detection and retrieval for further analyzing the tremendously large amount of objects because it is unreasonable to manually label and analyze them. Several state-of-the-art methods are rigorously tested and quantitatively compared in different criteria and categorically. In this work, we created a new dataset, namely, AI for Dunhuang, to facilitate the research. Version v1.0 of the dataset comprises of data and label for the restoration, style transfer, detection, and retrieval. Specifically, the dataset has 10,000 images for restoration, 3455 for style transfer, and 6147 for property retrieval. Lastly, we propose to use style transfer to link and analyze the styles over time, given that the grottoes were build over 1000 years by numerous artists. This enables the possibly to analyze and study the art styles over 1000 years and further enable future researches on cross-era style analysis. We benchmark representative methods and conduct a comparative study on the results for our solution. The dataset will be publicly available along with this paper. Tianxiu Yu, Chunxue Wang, Xiaohong Ding, Huili An, Xiaoxiang Liu, Ting Qu 0002, Shaodi You, Jiawan Zhang |
Int. J. Comput. Vis. | 10 |
| 2022 | Hybrid supervised instance segmentation by learning label noise suppression
Ying Fu 0001, Shaodi You, Hongzhe Liu 0001 |
Neurocomputing | 3 |
| 2022 | LE-GAN: Unsupervised low-light image enhancement network using attention module and identity invariant loss
Ying Fu 0001, Shaodi You |
Knowl. Based Syst. | 4 |
| 2022 | Semantic Guided Single Image Reflection RemovalabstractReflection is common when we see through a glass window, which not only is a visual disturbance but also influences the performance of computer vision algorithms. Removing the reflection from a single image, however, is highly ill-posed since the color at each pixel needs to be separated into two values belonging to the clear background and the reflection, respectively. To solve this, existing methods use additional priors such as reflection layer smoothness, double reflection effect, and color consistency to distinguish the two layers. However, these low-level priors may not be consistently valid in real cases. In this paper, inspired by the fact that human beings can separate the two layers easily by recognizing the objects and understanding the scene, we propose to use the object semantic cue, which is high-level information, as the guidance to help reflection removal. Based on the data analysis, we develop a multi-task end-to-end deep learning method with a semantic guidance component, to solve reflection removal and semantic segmentation jointly. Extensive experiments on different datasets show significant performance gain when using high-level object-oriented information. We also demonstrate the application of our method to other computer vision tasks. Yunfei Liu 0001, Yu Li 0003, Shaodi You, Feng Lu 0005 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2021 | Learning Temporal Consistency for Low Light Video Enhancement From Single ImagesabstractSingle image low light enhancement is an important task and it has many practical applications. Most existing methods adopt a single image approach. Although their performance is satisfying on a static single image, we found, however, they suffer serious temporal instability when handling low light videos. We notice the problem is because existing data-driven methods are trained from single image pairs where no temporal information is available. Unfortunately, training from real temporally consistent data is also problematic because it is impossible to collect pixel-wisely paired low and normal light videos under controlled environments in large scale and diversities with noise of identical statistics. In this paper, we propose a novel method to enforce the temporal stability in low light video enhancement with only static images. The key idea is to learn and infer motion field (optical flow) from a single image and synthesize short range video sequences. Our strategy is general and can extend to large scale datasets directly. Based on this idea, we propose our method which can infer motion prior for single image low light video enhancement and enforce temporal consistency. Rigorous experiments and user study demonstrate the state-of-the-art performance of our proposed method. Our code and model will be publicly available at https://github.com/zkawfanx/StableLLVE. Fan Zhang 0123, Yu Li 0003, Shaodi You, Ying Fu 0001 |
CVPR | 3 |
| 2021 | Multitask AET with Orthogonal Tangent Regularity for Dark Object DetectionabstractDark environment becomes a challenge for computer vision algorithms owing to insufficient photons and undesirable noise. To enhance object detection in a dark environment, we propose a novel multitask auto encoding transformation (MAET) model which is able to explore the intrinsic pattern behind illumination translation. In a self-supervision manner, the MAET learns the intrinsic visual structure by encoding and decoding the realistic illumination-degrading transformation considering the physical noise model and image signal processing (ISP). Based on this representation, we achieve the object detection task by decoding the bounding box coordinates and classes. To avoid the over-entanglement of two tasks, our MAET disentangles the object and degrading features by imposing an orthogonal tangent regularity. This forms a parametric manifold along which multitask predictions can be geometrically formulated by maximizing the orthogonality between the tangents along the outputs of respective tasks. Our framework can be implemented based on the mainstream object detection architecture and directly trained end-to-end using normal target detection datasets, such as VOC and COCO. We have achieved the state-of-the-art performance using synthetic and real-world datasets. Codes will be released at https://github.com/cuiziteng/MAET. Ziteng Cui, Guo-Jun Qi, Lin Gu 0003, Shaodi You, Zenghui Zhang, Tatsuya Harada |
ICCV | 4 |
| 2021 | Automatic Calibration of the Fisheye Camera for Egocentric 3D Human Pose Estimation from a Single ImageabstractWe propose a method for egocentric 3D human pose estimation from a single image captured by a fisheye camera. The problem of estimating the egocentric 3D pose for a fisheye camera is that images may be subject to strong image distortions (e.g. 2D poses on the image plane that pass through the line of sight of the fisheye lens).Therefore, in this paper, we approach this problem by an automatic calibration module. Given a single image, our network first estimates 3D joint locations of a human in camera coordinates. To alleviate the impact of image distortions on 3D human pose estimation, we then use the automatic calibration to further regularize the 3D predictions. Experimental results demonstrate that the proposed method achieves state-of-the-art performance. Shaodi You, Theo Gevers |
WACV | 2 |
| 2021 | Real-time foreground object segmentation networks using long and short skip connectionsabstractForeground object segmentation is an important task with various applications in outdoor surveillance and navigation. Most existing methods focus on accuracy and therefore, are computationally expensive and low in speed, making them difficult to use in actual applications. In this study, we aim to address the issue of accuracy and efficiency trade-off. In particular, in contrast with existing methods that use fine-tuning routine on heavyweight pretrained models and/or optimization techniques to enhance results, we propose a lightweight end-to-end network that can be trained from scratch effectively and efficiently. First, long and short skip connections are used among convolutional blocks and within the bottleneck block. By doing so, information flow within the networks is enhanced during the training stage, and thus, the use rate of parameters in the model is increased, allowing a more compact and efficient network design. Second, we use feature fusions based on element-wise summing before each up-sampling layer to reduce the size of the decoder, accelerate the up-sampling process, and stabilize training convergence. Our proposed method is tested rigorously. In particular, we achieved 1000 times higher speed compared with state-of-the-art methods on CD2014 and SBI2015 datasets with comparable accuracy. Shaodi You, Xiaoxiang Liu |
Inf. Sci. | 3 |
| 2021 | Cross-modal dynamic convolution for multi-modal emotion recognition
Huanglu Wen, Shaodi You, Ying Fu 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2021 | Cross-modal context-gated convolution for multi-modal sentiment analysisabstractWhen inferring sentiments, using verbal clues only is problematic because of the ambiguity. Adding related vocal and visual contexts as complements for verbal clues can be helpful. To infer sentiments from multi-modal temporal sequences, we need to identify both sentiment-related clues and their cross-modal interactions. However, sentiment-related behaviors of different modalities may not occur at the same time. These behaviors and their interactions are also sparse in time, making it hard to infer the correct sentiments. Besides, unaligned sequences from sensors also have varying sampling rates , which amplify the misalignment and sparsity mentioned above. While most previous multi-modal sentiment analysis works only focus on word-aligned sequences, we propose cross-modal context-gated convolution for unaligned sequences. Cross-modal context-gated convolution captures the local cross-modal interactions, dealing with the misalignment while reducing the effect of unrelated information. Cross-modal context-gated convolution introduces the concept of cross-modal context gate, enabling itself to catch useful cross-modal interactions more effectively. Cross-modal context-gated convolution also brings more possibilities to the layer design for multi-modal sequential modeling. Experiments on multi-modal sentiment analysis datasets under both word-aligned and unaligned conditions show the validity of our approach. Huanglu Wen, Shaodi You, Ying Fu 0001 |
Pattern Recognit. Lett. | 2 |
| 2020 | Unsupervised Learning for Intrinsic Image Decomposition From a Single ImageabstractIntrinsic image decomposition, which is an essential task in computer vision, aims to infer the reflectance and shading of the scene. It is challenging since it needs to separate one image into two components. To tackle this, conventional methods introduce various priors to constrain the solution, yet with limited performance. Meanwhile, the problem is typically solved by supervised learning methods, which is actually not an ideal solution since obtaining ground truth reflectance and shading for massive general natural scenes is challenging and even impossible. In this paper, we propose a novel unsupervised intrinsic image decomposition framework, which relies on neither labeled training data nor hand-crafted priors. Instead, it directly learns the latent feature of reflectance and shading from unsupervised and uncorrelated data. To enable this, we explore the independence between reflectance and shading, the domain invariant content constraint and the physical constraint. Extensive experiments on both synthetic and real image datasets demonstrate consistently superior performance of the proposed method. Yunfei Liu 0001, Yu Li 0003, Shaodi You, Feng Lu 0005 |
CVPR | 3 |
| 2020 | Kinship Identification Through Joint Learning Using Kinship Verification Ensembles
Wei Wang 0469, Shaodi You, Theo Gevers |
ECCV (22) | 2 |
| 2020 | Orthographic Projection Linear Regression for Single Image 3D Human Pose Estimationabstract3D human pose estimation from a single 2D image in the wild is an important computer vision task but yet extremely challenging. Unlike images taken from indoor and well constrained environments, 2D outdoor images in the wild are extremely complex because of varying imaging conditions. Furthermore, 2D images usually do not have corresponding 3D pose ground truth making a supervised approach ill-constrained. Therefore, in this paper, we propose to associate the 3D human pose, the 2D human pose projection and the 2D image appearance through a new orthographic projection based linear regression module. Unlike existing reprojection based approaches, our orthographic projection and regression do not suffer from small angle problems, which usually lead to overfitting in the depth dimension. Hence, we propose a deep neural network which adopts the 2D pose, 3D pose regression and orthographic projection linear regression module. The proposed method shows state-of-the-art performance on the Human3.6M dataset and generalizes well to in-the-wild images. Shaodi You, Theo Gevers |
ICPR | 2 |
| 2020 | iFlask: Isolate flask security system from dangerous execution environment by using ARM TrustZone
Diming Zhang, Shaodi You |
Future Gener. Comput. Syst. | 2 |
| 2020 | Deep clustering for weakly-supervised semantic segmentation in autonomous driving scenes
Xiang Wang 0003, Huimin Ma 0001, Shaodi You |
Neurocomputing | 3 |
| 2020 | Unidirectional Representation-Based Efficient Dictionary LearningabstractDictionary learning (DL) has been widely studied for pattern classification. Most existing methods introduce multiple discriminative terms into objective functions for accuracy improvement, leading to complex learning frameworks and high computational burdens. This paper proposes a simple yet effective DL algorithm for classification, namely unidirectional representation dictionary learning (URDL). Unidirectional constraint is proposed to guide coefficient directions in the representation to be discriminative. Besides, direction-thresholding is proposed to exploit the direction property in the classification scheme. It suppresses the disturbance from undesired non-zero coefficients, and improves the representation discriminability. We adopt squared ℓ2-norm-based regularization for efficient coding, and systematically analyze the mechanism of the proposed method. Extensive experiments on five data sets are conducted, including object categorization, scene classification, face recognition, and fine-grained flower classification. The experimental results demonstrate that the proposed approach not only outperforms the state-of-the-art DL algorithms in terms of recognition accuracy significantly, but also exhibits a much higher computational efficiency. Xiudong Wang, Yali Li 0001, Shaodi You, Hongdong Li, Shengjin Wang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Learning to Minify Photometric StereoabstractPhotometric stereo estimates the surface normal given a set of images acquired under different illumination conditions. To deal with diverse factors involved in the image formation process, recent photometric stereo methods demand a large number of images as input. We propose a method that can dramatically decrease the demands on the number of images by learning the most informative ones under different illumination conditions. To this end, we use a deep learning framework to automatically learn the critical illumination conditions required at input. Furthermore, we present an occlusion layer that can synthesize cast shadows, which effectively improves the estimation accuracy. We assess our method on challenging real-world conditions, where we outperform techniques elsewhere in the literature with a significantly reduced number of light conditions. Antonio Robles-Kelly, Shaodi You, Yasuyuki Matsushita |
CVPR | 3 |
| 2019 | Classification-Reconstruction Learning for Open-Set RecognitionabstractOpen-set classification is a problem of handling `unknown' classes that are not contained in the training dataset, whereas traditional classifiers assume that only known classes appear in the test environment. Existing open-set classifiers rely on deep networks trained in a supervised manner on known classes in the training set; this causes specialization of learned representations to known classes and makes it hard to distinguish unknowns from knowns. In contrast, we train networks for joint classification and reconstruction of input data. This enhances the learned representation so as to preserve information useful for separating unknowns from knowns, as well as to discriminate classes of knowns. Our novel Classification-Reconstruction learning for Open-Set Recognition (CROSR) utilizes latent representations for reconstruction and enables robust unknown detection without harming the known-class classification accuracy. Extensive experiments reveal that the proposed method outperforms existing deep open-set classifiers in multiple standard datasets and is robust to diverse outliers. Ryota Yoshihashi, Wen Shao, Rei Kawakami, Shaodi You, Makoto Iida, Takeshi Naemura |
CVPR | 4 |
| 2019 | Non-Local Intrinsic Decomposition With Near-Infrared PriorsabstractIntrinsic image decomposition is a highly under-constrained problem that has been extensively studied by computer vision researchers. Previous methods impose additional constraints by exploiting either empirical or data-driven priors. In this paper, we revisit intrinsic image decomposition with the aid of near-infrared (NIR) imagery. We show that NIR band is considerably less sensitive to textures and can be exploited to reduce ambiguity caused by reflectance variation, promoting a simple yet powerful prior for shading smoothness. With this observation, we formulate intrinsic decomposition as an energy minimisation problem. Unlike existing methods, our energy formulation decouples reflectance and shading estimation, into a convex local shading component based on NIR-RGB image pair, and a reflectance component that encourages reflectance homogeneity both locally and globally. We further show the minimisation process can be approached by a series of multi-dimensional kernel convolutions, each within linear time complexity. To validate the proposed algorithm, a NIR-RGB dataset is captured over real-world objects, where our NIR-assisted approach demonstrates clear superiority over RGB methods. Ziang Cheng, Yinqiang Zheng, Shaodi You, Imari Sato |
ICCV | 3 |
| 2019 | Cross-Connected Networks for Multi-Task Learning of Detection and SegmentationabstractMulti-task learning improves generalization performance in neural networks by sharing knowledge among related tasks. Existing models are for task combinations annotated on the same dataset; research on how to utilize the knowledge of successful single-task convolutional neural networks (CNNs) that are trained on individual datasets is limited. We propose a cross-connected CNN, an architecture that connects single-task CNNs through convolutional layers that transfer useful information to their counterparts. We evaluated the architecture with a combination of detection and segmentation using datasets of two targets: pedestrians and wild birds. Experiments demonstrate how well our CNN learns general representations from multi-task learning. Rei Kawakami, Ryota Yoshihashi, Seiichiro Fukuda, Shaodi You, Makoto Iida, Takeshi Naemura |
ICIP | 4 |
| 2018 | Detail Preserving Depth Estimation from a Single Image Using Attention Guided NetworksabstractConvolutional Neural Networks have demonstrated superior performance on single image depth estimation in recent years. These works usually use stacked spatial pooling or strided convolution to get high-level information which are common practices in classification task. However, depth estimation is a dense prediction problem and low-resolution feature maps usually generate blurred depth map which is undesirable in application. In order to produce high quality depth map, say clean and accurate, we propose a network consists of a Dense Feature Extractor (DFE) and a Depth Map Generator (DMG). The DFE combines ResNet and dilated convolutions. It extracts multi-scale information from input image while keeping the feature maps dense. As for DMG, we use attention mechanism to fuse multi-scale features produced in DFE. Our Network is trained end-to-end and does not need any post-processing. Hence, it runs fast and can predict depth map in about 15 fps. Experiment results show that our method is competitive with the state-of-the-art in quantitative evaluation, but can preserve better structural details of the scene depth. Zhixiang Hao, Yu Li 0003, Shaodi You, Feng Lu 0005 |
3DV | 3 |
| 2018 | Loss Guided Activation for Action Recognition in Still Images
Robby T. Tan, Shaodi You |
ACCV (5) | 3 |
| 2018 | JTAV: Jointly Learning Social Media Content Representation by Fusing Textual, Acoustic, and Visual FeaturesabstractLearning social media content is the basis of many real-world applications, including information retrieval and recommendation systems, among others. In contrast with previous works that focus mainly on single modal or bi-modal learning, we propose to learn social media content by fusing jointly textual, acoustic, and visual information (JTAV). Effective strategies are proposed to extract fine-grained features of each modality, that is, attBiGRU and DCRNN. We also introduce cross-modal fusion and attentive pooling techniques to integrate multi-modal information comprehensively. Extensive experimental evaluation conducted on real-world datasets demonstrate our proposed model outperforms the state-of-the-art approaches by a large margin. Hongru Liang, Haozheng Wang, Jun Wang 0023, Shaodi You, Zhe Sun 0009, Jinmao Wei 0001, Zhenglu Yang |
COLING | 4 |
| 2018 | Weakly-Supervised Semantic Segmentation by Iteratively Mining Common Object FeaturesabstractWeakly-supervised semantic segmentation under image tags supervision is a challenging task as it directly associates high-level semantic to low-level appearance. To bridge this gap, in this paper, we propose an iterative bottom-up and top-down framework which alternatively expands object regions and optimizes segmentation network. We start from initial localization produced by classification networks. While classification networks are only responsive to small and coarse discriminative object regions, we argue that, these regions contain significant common features about objects. So in the bottom-up step, we mine common object features from the initial localization and expand object regions with the mined features. To supplement non-discriminative regions, saliency maps are then considered under Bayesian framework to refine the object regions. Then in the top-down step, the refined object regions are used as supervision to train the segmentation network and to predict object masks. These object masks provide more accurate localization and contain more regions of object. Further, we take these object masks as initial localization and mine common object features from them. These processes are conducted iteratively to progressively produce fine object masks and optimize segmentation networks. Experimental results on Pascal VOC 2012 dataset demonstrate that the proposed method outperforms previous state-of-the-art methods by a large margin. Xiang Wang 0003, Shaodi You, Xi Li 0010, Huimin Ma 0001 |
CVPR | 2 |
| 2018 | Single Image Water Hazard Detection Using FCN with Reflection Attention Units
Chuong V. Nguyen, Shaodi You, Jianfeng Lu 0003 |
ECCV (6) | 3 |
| 2018 | Deep Texture and Structure Aware Filtering Network for Image Smoothing
Kaiyue Lu, Shaodi You, Nick Barnes |
ECCV (4) | 2 |
| 2018 | A Frequency Domain Neural Network for Fast Image Super-resolutionabstractIn this paper, we present a frequency domain neural network for image super-resolution. The network employs the convolution theorem so as to cast convolutions in the spatial domain as products in the frequency domain. Moreover, the non-linearity in deep nets, of ten achieved by a rectifier unit, is here cast as a convolution in the frequency domain. This not only yields a network which is very computationally efficient at testing, but also one whose parameters can all be learnt accordingly. The network can be trained using back propagation and is devoid of complex numbers due to the use of the Hartley transform as an alternative to the Fourier transform. Moreover, the network is potentially applicable to other problems elsewhere in computer vision and image processing which are of ten cast in the frequency domain. We show results on super-resolution and compare against alternatives elsewhere in the literature. In our experiments, our network is one to two orders of magnitude faster than the alternatives with a marginal loss of performance. Shaodi You, Antonio Robles-Kelly |
IJCNN | 2 |
| 2018 | VBMq: pursuit baremetal performance by embracing block I/O parallelism in virtualization
Diming Zhang, Shaodi You |
Frontiers Comput. Sci. | 4 |
| 2018 | Multiview Rectification of Folded DocumentsabstractDigitally unwrapping images of paper sheets is crucial for accurate document scanning and text recognition. This paper presents a method for automatically rectifying curved or folded paper sheets from a few images captured from multiple viewpoints. Prior methods either need expensive 3D scanners or model deformable surfaces using over-simplified parametric representations. In contrast, our method uses regular images and is based on general developable surface models that can represent a wide variety of paper deformations. Our main contribution is a new robust rectification method based on ridge-aware 3D reconstruction of a paper sheet and unwrapping the reconstructed surface using properties of developable surfaces via conformal mapping. We present results on several examples including book pages, folded letters and shopping receipts. Shaodi You, Yasuyuki Matsushita, Sudipta N. Sinha, Yusuke Bou, Katsushi Ikeuchi |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Semisupervised and Weakly Supervised Road Detection Based on Generative Adversarial NetworksabstractRoad detection is a key component of autonomous driving; however, most fully supervised learning road detection methods suffer from either insufficient training data or high costs of manual annotation. To overcome these problems, we propose a semisupervised learning (SSL) road detection method based on generative adversarial networks (GANs) and a weakly supervised learning (WSL) method based on conditional GANs. Specifically, in our SSL method, the generator generates the road detection results of labeled and unlabeled images, and then they are fed into the discriminator, which assigns a label on each input to judge whether it is labeled. Additionally, in WSL method we add another network to predict road shapes of input images and use them in both generator and discriminator to constrain the learning progress. By training under these frameworks, the discriminators can guide a latent annotation process on the unlabeled data; therefore, the networks can learn better representations of road areas and leverage the feature distributions on both labeled and unlabeled data. The experiments are carried out on KITTI ROAD benchmark, and the results show our methods achieve the state-of-the-art performances. Jianfeng Lu 0003, Chunxia Zhao, Shaodi You, Hongdong Li |
IEEE Signal Process. Lett. | 4 |
| 2018 | Edge Preserving and Multi-Scale Contextual Neural Network for Salient Object DetectionabstractIn this paper, we propose a novel edge preserving and multi-scale contextual neural network for salient object detection. The proposed framework is aiming to address two limits of the existing CNN based methods. First, region-based CNN methods lack sufficient context to accurately locate salient object since they deal with each region independently. Second, pixel-based CNN methods suffer from blurry boundaries due to the presence of convolutional and pooling layers. Motivated by these, we first propose an end-to-end edge-preserved neural network based on Fast R-CNN framework (named RegionNet) to efficiently generate saliency map with sharp object boundaries. Later, to further improve it, multi-scale spatial context is attached to RegionNet to consider the relationship between regions and the global scenes. Furthermore, our method can be generally applied to RGB-D saliency detection by depth refinement. The proposed framework achieves both clear detection boundary and multi-scale contextual robustness simultaneously for the first time, and thus achieves an optimized performance. Experiments on six RGB and two RGB-D benchmark datasets demonstrate that the proposed method achieves state-of-the-art performance. Xiang Wang 0003, Huimin Ma 0001, Xiaozhi Chen, Shaodi You |
IEEE Trans. Image Process. | 4 |
| 2017 | Single Image Action Recognition Using Semantic Body Part ActionsabstractIn this paper, we propose a novel single image action recognition algorithm based on the idea of semantic part actions. Unlike existing part-based methods, we argue that there exists a mid-level semantic, the semantic part action; and human action is a combination of semantic part actions and context cues. In detail, we divide human body into seven parts: head, torso, arms, hands and lower body. For each of them, we define a few semantic part actions (e.g. head: laughing). Finally, we exploit these part actions to infer the entire body action (e.g. applauding). To make the proposed idea practical, we propose a deep network-based framework which consists of two subnetworks, one for part localization and the other for action prediction. The action prediction network jointly learns part-level and body-level action semantics and combines them for the final decision. Extensive experiments demonstrate our proposal on semantic part actions as elements for entire body action. Our method reaches mAP of 93.9% and 91.2% on PASCAL VOC 2012 and Stanford-40, which outperforms the state-of-the-art by 2.3% and 8.6%. Zhichen Zhao, Huimin Ma 0001, Shaodi You |
ICCV | 3 |
| 2017 | Survival-Oriented Reinforcement Learning Model: An Effcient and Robust Deep Reinforcement Learning Algorithm for Autonomous Driving Problem
Changkun Ye, Huimin Ma 0001, Shaodi You |
ICIG (2) | 5 |
| 2017 | Automatic Generation of Grounded Visual QuestionsabstractIn this paper, we propose the first model to be able to generate visually grounded questions with diverse types for a single image. Visual question generation is an emerging topic which aims to ask questions in natural language based on visual input. To the best of our knowledge, it lacks automatic methods to generate meaningful questions with various types for the same visual input. To circumvent the problem, we propose a model that automatically generates visually grounded questions with varying types. Our model takes as input both images and the captions generated by a dense caption model, samples the most probable question types, and generates the questions in sequel. The experimental results on two real world datasets show that our model outperforms the strongest baseline in terms of both correctness and diversity with a wide margin. Lizhen Qu, Shaodi You, Zhenglu Yang, Jiawan Zhang |
IJCAI | 3 |
| 2017 | Haze visibility enhancement: A Survey and quantitative benchmarking
Yu Li 0003, Shaodi You, Michael S. Brown, Robby T. Tan |
Comput. Vis. Image Underst. | 2 |
| 2017 | Think locally, fit globally: Robust and fast 3D shape matching via adaptive algebraic fitting
Shaodi You, Diming Zhang |
Neurocomputing | 1 |
| 2016 | Local Background Enclosure for RGB-D Salient Object DetectionabstractRecent work in salient object detection has considered the incorporation of depth cues from RGB-D images. In most cases, depth contrast is used as the main feature. However, areas of high contrast in background regions cause false positives for such methods, as the background frequently contains regions that are highly variable in depth. Here, we propose a novel RGB-D saliency feature. Local Background Enclosure (LBE) captures the spread of angular directions which are background with respect to the candidate region and the object that it is part of. We show that our feature improves over state-of-the-art RGB-D saliency approaches as well as RGB methods on the RGBD1000 and NJUDS2000 datasets. David Feng 0002, Nick Barnes, Shaodi You, Chris McCarthy |
CVPR | 3 |
| 2016 | Adherent Raindrop Modeling, Detectionand Removal in VideoabstractRaindrops adhered to a windscreen or window glass can significantly degrade the visibility of a scene. Modeling, detecting and removing raindrops will, therefore, benefit many computer vision applications, particularly outdoor surveillance systems and intelligent vehicle systems. In this paper, a method that automatically detects and removes adherent raindrops is introduced. The core idea is to exploit the local spatio-temporal derivatives of raindrops. To accomplish the idea, we first model adherent raindrops using law of physics, and detect raindrops based on these models in combination with motion and intensity temporal derivatives of the input video. Having detected the raindrops, we remove them and restore the images based on an analysis that some areas of raindrops completely occludes the scene, and some other areas occlude only partially. For partially occluding areas, we restore them by retrieving as much as possible information of the scene, namely, by solving a blending function on the detected partially occluding areas using the temporal intensity derivative. For completely occluding areas, we recover them by using a video completion technique. Experimental results using various real videos show the effectiveness of our method. Shaodi You, Robby T. Tan, Rei Kawakami, Yasuhiro Mukaigawa, Katsushi Ikeuchi |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2014 | Raindrop Detection and Removal from Long Range Trajectories
Shaodi You, Robby T. Tan, Rei Kawakami, Yasuhiro Mukaigawa, Katsushi Ikeuchi |
ACCV (2) | 1 |
| 2013 | Adherent Raindrop Detection and Removal in VideoabstractRaindrops adhered to a windscreen or window glass can significantly degrade the visibility of a scene. Detecting and removing raindrops will, therefore, benefit many computer vision applications, particularly outdoor surveillance systems and intelligent vehicle systems. In this paper, a method that automatically detects and removes adherent raindrops is introduced. The core idea is to exploit the local spatio-temporal derivatives of raindrops. First, it detects raindrops based on the motion and the intensity temporal derivatives of the input video. Second, relying on an analysis that some areas of a raindrop completely occludes the scene, yet the remaining areas occludes only partially, the method removes the two types of areas separately. For partially occluding areas, it restores them by retrieving as much as possible information of the scene, namely, by solving a blending function on the detected partially occluding areas using the temporal intensity change. For completely occluding areas, it recovers them by using a video completion technique. Experimental results using various real videos show the effectiveness of the proposed method. Shaodi You, Robby T. Tan, Rei Kawakami, Katsushi Ikeuchi |
CVPR | 1 |
| 2011 | Manifold topological multi-resolution analysis method
Shaodi You, Huimin Ma 0001 |
Pattern Recognit. | 1 |
| 2009 | A Solution to Efficient Viewpoint Space Partition in 3D Object RecognitionabstractViewpoint Space Partition based on Aspect Graph is one of the core techniques of 3D object recognition. Projection images obtained from critical viewpoint following this approach can efficiently provide topological information of an object. Computational complexity has been a huge challenge for obtaining the representation viewpoints used in 3D recognition. In this paper, we discuss inefficiency of calculation due to redundant nonexistent visual events; propose a systematic criterion for edge selection involved in EEE events. Pruning algorithm based on concave-convex property is demonstrated. We further introduce intersect relation into our pruning algorithm. These two methods not only enable the calculation of EEE events, but also can be implemented before viewpoint calculation, hence realizes view-independent pruning algorithm. Finally, analysis on simple representative models supports the effectiveness of our methods. Further investigations on Princeton Models, including airplane, automobile, etc, show a two orders of magnitude reduction in the number of EEE events on average. Huimin Ma 0001, Shaodi You, Ze Yuan |
ICIG | 3 |