VLDB 2026 Research / reviewers in the wild / expert
Shichao Zhao
dblp:172/1180
· DBLP profile ↗
10ranked-venue papers
3as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cultural Heritage and Resilience in Immigrant Communities: Exploring Interactive Technology Opportunities for Mutual AcculturationabstractThis study explores how British-Chinese immigrants sustain cultural resilience through heritage practices and the role of interactive technologies in this process. Using a two-phase mixed-methods approach, it begins with a scoping study (n=150) to map patterns of heritage engagement, and then employed interviews with 11 immigrants and scenario-based workshops with 12 stakeholders to investigate technology-mediated dissemination. Findings show that intangible cultural heritage, particularly foodways, language and festivals, plays a central role in identity formation and intergenerational cohesion, with community networks mediating resilience. However, heritage practices and technologies remain largely internally-focused, while fragmented platforms and linguistic barriers limit intercultural exchange. To address these gaps, the paper proposes the Community-Centred Mutual Acculturation (CCMA) framework, emphasising community authorship, plural representation, and cross-cultural engagement through modular platforms, youth-led co-creation, and immersive multi-sensory interaction. This research contributes empirical insights and design-oriented recommendations for interactive systems that support inclusive, heritage-based engagement and mutual acculturation. Qinqing Fu, Shichao Zhao, Jo Barnes, Lise Jaillant |
DIS | 2 |
| 2026 | Monocular Mesh Recovery and Body Measurement of Female Saanen GoatsabstractThe lactation performance of Saanen dairy goats, renowned for their high milk yield, is intrinsically linked to their body size, making accurate 3D body measurement essential for assessing milk production potential, yet existing reconstruction methods lack goat-specific authentic 3D data. To address this limitation, we establish the FemaleSaanenGoat dataset containing synchronized eight-view RGBD videos of 55 female Saanen goats (6-18 months). Using multi-view DynamicFusion, we fuse noisy, non-rigid point cloud sequences into high-fidelity 3D scans, overcoming challenges from irregular surfaces and rapid movement. Based on these scans, we develop SaanenGoat, a parametric 3D shape model specifically designed for female Saanen goats. This model features a refined template with 41 skeletal joints and enhanced udder representation, registered with our scan data. A comprehensive shape space constructed from 48 goats enables precise representation of diverse individual variations. With the help of SaanenGoat model, we get high-precision 3D reconstruction from single-view RGBD input, and achieve automated measurement of six critical body dimensions: body length, height, chest width, chest girth, hip width, and hip height. Experimental results demonstrate the superior accuracy of our method in both 3D reconstruction and body measurement, presenting a novel paradigm for large-scale 3D vision applications in precision livestock farming. Shichao Zhao, Jin Lyu, Tao Yu 0007, Liang An 0001, Yebin Liu, Meili Wang 0002 |
AAAI | 2 |
| 2026 | Enhanced Multi-Scale PoseNet for Self-Supervised Monocular Depth EstimationabstractMonocular depth estimation is essential for 3D perception in applications such as autonomous driving and robotics. Self-supervised methods avoid depth labels but often rely on shallow pose networks with weak temporal modeling, leading to unstable predictions. We propose EMSP-Net, an Enhanced Multi-Scale PoseNet for self-supervised monocular depth estimation. It introduces a hierarchical feature fusion encoder, a temporal attention-context decoder, and a pose consistency loss to jointly improve feature extraction, temporal stability, and geometric constraints. On the KITTI dataset, EMSP-Net achieved an absolute relative error of 0.105 and a squared relative error of 0.708. In the Make3D cross-domain test, its strong robustness was further demonstrated. Chao Zhang 0105, Cheng Han 0002, Tiancheng Shao, Shichao Zhao |
IEEE Signal Process. Lett. | 6 |
| 2025 | From Focal to Scattered: Designing Culturally Adaptive VR for Chinese Architectural Painting
Yuting Cheng 0006, Shichao Zhao, Jiashu Yang, Kamarin Merritt |
INTERACT (1) | 2 |
| 2023 | Involving British-Chinese Immigrants in Participatory Action Research: Lessons Learnt from the FieldabstractBritish-Chinese communities in the United Kingdom have experienced an increase in discriminatory behaviour with other communities due to the COVID-19 pandemic and the stigmatisation it has brought about as a result of the speculated COVID-19 origins. Therefore, as a pilot study, this paper investigates how Participatory Action Research (PAR), principally the integration of interactive technology with co-design activities, can be applied to support the producing and sharing of community-based immigrant heritage for British-Chinese citizens. In addition, the reasoning behind why British-Chinese communities have faced cross-cultural barriers when sharing their values and significance of their heritage more widely within British society during the COVID-19 pandemic has also been explored. This study potentially makes a significant contribution to the literature because design-led inquiry was used to explore design strategies and considerations of interactive technology that improved the participation of British-Chinese immigrants in sharing the significance of their intangible heritage socially, equally, and coherently during the COVID-19 pandemic. Shichao Zhao |
Conference on Designing Interactive Systems | 1 |
| 2023 | Overlap Loss: Rethinking Weakly Supervised Instance Segmentation in Crowded ScenesabstractWeakly supervised instance segmentation (WSIS) has gained increasing popularity in recent years due to low labelling cost. However, its performance deteriorates dramatically in more challenging crowded scenario, which is caused by overlapping among similar objects. To ameliorate the negative effects of instance overlapping, we propose a new loss, i.e., OverlapLoss, which achieves instance disentanglement between masks according to the degree of overlapping among instances. Besides, a new dataset of CrowdHuman Instance Segmentation (CIS) is presented to bridge the gap in crowded scenes. Experiments on the CIS and COCO datasets validate that the proposed loss can improve the baseline in typical crowed scenes by at least 2% and in uncrowded scenes by more than 0.3% w.r.t. absolute AP. The code and dataset are available at: https://github.com/shanghangjiang/CIS. Shanghang Jiang, Shichao Zhao, Le Zhang 0001 |
ICIP | 2 |
| 2021 | MagFace: A Universal Representation for Face Recognition and Quality AssessmentabstractThe performance of face recognition system degrades when the variability of the acquired faces increases. Prior work alleviates this issue by either monitoring the face quality in pre-processing or predicting the data uncertainty along with the face feature. This paper proposes MagFace, a category of losses that learn a universal feature embedding whose magnitude can measure the quality of the given face. Under the new loss, it can be proven that the magnitude of the feature embedding monotonically increases if the subject is more likely to be recognized. In addition, Mag-Face introduces an adaptive mechanism to learn a well-structured within-class feature distributions by pulling easy samples to class centers while pushing hard samples away. This prevents models from overfitting on noisy low-quality samples and improves face recognition in the wild. Extensive experiments conducted on face recognition, quality assessments as well as clustering demonstrate its superiority over state-of-the-arts. The code is available at https://github.com/IrvingMeng/MagFace. Shichao Zhao, Zhida Huang, Feng Zhou 0002 |
CVPR | 2 |
| 2018 | Pooling the Convolutional Layers in Deep ConvNets for Video Action RecognitionabstractDeep ConvNets have shown their good performance in image classification tasks. However, there still remains problems in deep video representations for action recognition. On one hand, current video ConvNets are relatively shallow compared with image ConvNets, which limits their capability of capturing the complex video action information; on the other hand, temporal information of videos is not properly utilized to pool and encode the video sequences. Toward these issues, in this paper we utilize two state-of-the-art ConvNets, i.e., the very deep spatial net (VGGNet [1]) and the temporal net from Two-Stream ConvNets [2], for action representation. The convolutional layers and the proposed new layer, called frame-diff layer, are extracted and pooled with two temporal pooling strategies: Trajectory pooling and Line pooling. The pooled local descriptors are then encoded with vector of locally aggregated descriptors (VLAD) [3] to form the video representations. In order to verify the effectiveness of the proposed framework, we conduct experiments on UCF101 and HMDB51 data sets. It achieves accuracy of 92.08% on UCF101, which is the state-of-the-art, and the accuracy of 65.62% on HMDB51, which is comparable to the state-of-the-art. In addition, we propose the new Line pooling strategy, which can speed up the extraction of feature and achieve the comparable performance of the Trajectory pooling. Shichao Zhao, Yanbin Liu 0003, Yahong Han, Richang Hong, Qinghua Hu, Qi Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | Top attention in line with time: A light-weight strategyabstractFor video representation, dense sampling along trajectories or optical flow stacking are both heavy-cost computations. This paper aims to develop a light-weight strategy which could skip the computations of optical flow and trajectories. Particularly, taking frames as inputs to a pre-trained ConvNet, we extract top layers as video feature maps. Instead of trajectory pooling, we directly pooled these feature maps in line with time, which is named Line Pooling. We utilize the proposed Line-pooled Deep-convolutional Descriptors (LDDs) to weight regions with high motion saliency, which turns out to pay attention to actions in line with time. Experiments on UCF101 and HMDB51 demonstrate the efficiency, effectiveness, and promising performance of our method. Youjiang Xu, Shichao Zhao, Yahong Han, Qinghua Hu, Fei Wu 0001 |
ICME | 2 |
| 2016 | Large-Scale E-Commerce Image Retrieval with Top-Weighted Convolutional Neural NetworksabstractSeveral recent researches have shown that image features produced by Convolutional Neural Networks (CNNs) provide the state-of-the-art performance for image classification and retrieval. Moreover, some researchers have found that the features extracted from the deep convolutional layers of CNNs perform better than that from the fully-connected layers. Features extracted from the convolutional layers have a natural interpretation: descriptors of local image regions correspond well to the receptive fields of the particular features. In order to obtain both representative and discriminative descriptors for large-scale e-commerce image retrieval, we come up with a new feature extraction framework. At first, we propose the Top-Weight method to detect the interesting area of e-commerce images automatically. With the estimated weight, we then aggregate local deep features and produce high-quality global representation for e-commerce image retrieval. We have conducted experiments on an e-commerce dataset ALISC [1] released by Alibaba Group. Experimental results show that our method outperforms other deep learning based methods. Shichao Zhao, Youjiang Xu, Yahong Han |
ICMR | 1 |