VLDB 2026 Research / reviewers in the wild / expert
Zhenzhou Shao
dblp:63/9966
· DBLP profile ↗
19ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0002-9166-9468ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 4 first-author · 7 since 2021Systems, architecture and hardware · 9 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Trustworthy machine learning · 50% Vision and language · 22% Face, body and person analysis · 22% | |
| Computer graphics and multimedia
2 papers |
Visual content generation and editing · 67% Visualization and visual analytics · 22% Rendering · 11% | |
| Human-computer interaction and pervasive computing
1 paper |
Ubiquitous computing and smart environments · 100% |
Topics — the 13 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
fairness |
1.0 | 1 | 2026 | Debiased Multimodal Personality Understanding through Dual Causal Intervention · SIGIR 2026 |
Computer vision › Vision and language › multimodal reasoning
multimodal mathematical reasoning |
1.0 | 1 | 2026 | MathSight: A Benchmark Exploring Have Vision-Language Models Really Seen in University-Level Mathematical Reasoning? · ACL (1) 2026 |
Computer vision › Face, body and person analysis › personality assessment
multimodal personality recognition |
1.0 | 1 | 2026 | Debiased Multimodal Personality Understanding through Dual Causal Intervention · SIGIR 2026 |
Visual content generation and editing › stylization
image stylization |
0.6 | 1 | 2022 | Non-Local Representation Based Mutual Affine-Transfer Network for Photorealistic Stylization · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Visual content generation and editing
style transfer |
0.6 | 1 | 2022 | Non-Local Representation Based Mutual Affine-Transfer Network for Photorealistic Stylization · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Machine learning › Trustworthy machine learning › fairness › bias mitigation
demographic bias mitigation |
0.3 | 1 | 2026 | Debiased Multimodal Personality Understanding through Dual Causal Intervention · SIGIR 2026 |
Rendering › volume rendering
direct volume rendering |
0.2 | 1 | 2014 | Geometry constrained sparse embedding for multi-dimensional transfer function design in direct volume rendering · ICRA 2014 |
Visualization and visual analytics › volume visualization
transfer function design |
0.2 | 1 | 2014 | Geometry constrained sparse embedding for multi-dimensional transfer function design in direct volume rendering · ICRA 2014 |
Visualization and visual analytics
volume visualization |
0.2 | 1 | 2014 | Geometry constrained sparse embedding for multi-dimensional transfer function design in direct volume rendering · ICRA 2014 |
Ubiquitous computing and smart environments › context recognition
activity recognition |
0.1 | 1 | 2011 | Pedestrian positioning with physical activity classification for indoors · ICRA 2011 |
Ubiquitous computing and smart environments
indoor localization |
0.1 | 1 | 2011 | Pedestrian positioning with physical activity classification for indoors · ICRA 2011 |
Algorithms and data structures › numerical linear algebra
dimensionality reduction |
0.1 | 1 | 2014 | Geometry constrained sparse embedding for multi-dimensional transfer function design in direct volume rendering · ICRA 2014 |
Algorithms and data structures › metric embedding
sparse embedding |
0.1 | 1 | 2014 | Geometry constrained sparse embedding for multi-dimensional transfer function design in direct volume rendering · ICRA 2014 |
Methods — techniques the papers use, named apart from their topics
structural causal model · 1.0prototype-based confounder dictionary · 1.0controlled visual ablation · 1.0causal intervention · 1.0benchmark construction · 1.0mutual affine-transfer network · 0.6dictionary-based decomposition · 0.6stochastic gradient descent · 0.4geometry constrained sparse embedding · 0.4dictionary learning · 0.4zero velocity update · 0.1hidden markov model · 0.1heading estimation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MathSight: A Benchmark Exploring Have Vision-Language Models Really Seen in University-Level Mathematical Reasoning?abstractRecent advances in Vision-Language Models (VLMs) have achieved impressive progress in multimodal mathematical reasoning.Yet, how much visual information truly contributes to reasoning remains unclear.Existing benchmarks report strong overall performance but seldom isolate the role of the image modality, leaving open whether VLMs genuinely leverage visual understanding or merely depend on linguistic priors.To address this, we present MathSight, a university-level multimodal mathematical reasoning benchmark designed to disentangle and quantify the effect of visual input.Each problem includes multiple visual variants-original, hand-drawn, photocaptured-and a text-only condition for controlled comparison.Experiments on state-ofthe-art VLMs reveal a consistent trend: the contribution of visual information diminishes with increasing problem difficulty.Remarkably, Qwen3-VL without any image input surpasses both its multimodal variants and GPT-5, underscoring the need for benchmarks like MathSight to advance genuine vision-grounded reasoning in future models.The project page is available at https Yuandong Wang 0002, Yao Cui, Zhen Yang 0034, Yangfu Zhu, Zhenzhou Shao |
ACL (1) | 6 |
| 2026 | Debiased Multimodal Personality Understanding through Dual Causal InterventionabstractMultimodal personality understanding plays a critical role in human-centered artificial intelligence. Previous work mainly focus on learning rich multimodal representations for video personality understanding. However, they often suffer from potential harm caused by subject bias (e.g., observable age and unobservable mental states), as subjects originate from diverse demographic backgrounds. Learning such spurious associations between multimodal features and traits may lead to unfair personality understanding. In this work, we construct a Structural Causal Model (SCM) to analyze the impact of these biases from a causal perspective, and propose a novel Dual Causal Adjustment Network (DCAN) to mitigate the interference of subject attributes on personality understanding. Specifically, we design a Back-door Adjustment Causal Learning (BACL) module to block spurious correlations from observable demographic factors via a prototype-based confounder dictionary, and subsequently apply a Front-door Adjustment Causal Learning (FACL) module to address latent and unobservable biases through a learned mediator dictionary intervention, thereby achieving causal disentanglement of representations for deconfounded reasoning. Importantly, we construct a Demographic-annotated Multimodal Student Personality (DMSP) dataset to support the analysis and discussion of fairness-related factors. Extensive experiments on the benchmark dataset CFI-V2 and our DMSP dataset demonstrate that DCAN consistently improves prediction accuracy, reaching 92.11% and 92.90%, respectively. Meanwhile, the improvements in the fairness metrics of equal opportunity and demographic parity are 6.57% and 7.97% on CFI-V2, and 15.38% and 20.06% on the DMSP dataset. Our code and DMSP dataset are available at https://github.com/Sabrina-han/DCAN Yangfu Zhu, Zitong Han, Nianwen Ning, Yuandong Wang 0002, Hang Feng, Zhenzhou Shao |
SIGIR | 7 |
| 2024 | Weakly Aligned Multi-spectral Pedestrian Detection via Cross-Modality Differential Enhancement and Multi-scale Spatial Alignment
Zhenzhou Shao, Yibo Zou |
ICPR (30) | 1 |
| 2024 | Unsupervised Approach for Multimodality Telerobotic Trajectory SegmentationabstractThe significance of telerobotic trajectory segmentation has been demonstrated on a range of skill training and robotic automation tasks. However, the trajectories of telerobot are characterized by complexity and high dimensionality, making it difficult to segment them accurately. Existing methods are often plagued by feature inefficiency, and the clustering methods are also difficult to properly model the trajectories. In addition, over-segmentation also affects the accuracy of the clustering-based segmentation methods. To address above problems, this article presents a new unsupervised approach that automatically segments multimodal trajectory data of a telerobot and solves the problem of over-segmentation through a postpromoting procedure. First, we present an unsupervised visual feature-extraction network, that is, a dense connection spatial convolution network (DCSC) to generate more discriminative features for clustering. The dense convolution and spatial convolution facilitates information flow, enhances feature propagation, and avoids manual annotation. Next, we develop an unsupervised trajectory segmentation method that is called multimodality clustering with Chinese restaurant process (MC-CRP). The model utilizes data from different modalities and segments the trajectory through hierarchical clustering. MC-CRP obtains more accurate results in a short period of time. To further improve the precision of trajectory segmentation, we merge over-segments based on predefined similarity measurements. Extensive experiments on the publicly available data set JIGSAWS show that the presented approach achieves 70.1% silhouette coefficient, 25.1% normalized mutual information, and 71.4% segmentation accuracy. These metrics demonstrate that the presented segmentation approach provides deeper insight into the trajectory features and improve the accuracy of segmentation more efficiently than others. Jiexin Xie, Haitao Yan, Jiaxin Wang 0003, Zhenzhou Shao, Shijie Guo, Jinhua She |
IEEE Internet Things J. | 4 |
| 2023 | Parallelizable Simple Recurrent Units with Hierarchical Memory
Yu Qiao 0001, Hengyi Zhang, Yuan Tian 0017, Zhenzhou Shao, Zhi-Ping Shi 0002 |
ICONIP (15) | 6 |
| 2023 | 6D Object Pose Estimation with Attention Aware Bi-gated Fusion
Laichao Wang, Weiding Lu, Yuan Tian 0017, Zhenzhou Shao, Zhi-Ping Shi 0002 |
ICONIP (2) | 5 |
| 2022 | Cross-Guided Feature Fusion with Intra-Modality Reweighting for Multi-Spectral Pedestrian DetectionabstractMulti-spectral pedestrian detection has gained extensive attention over the past decade. To alleviate the problem of modality imbalance in the multi-spectral tasks, a novel cross-guided feature fusion network based on the auto-encoder framework is proposed using RGB-thermal image pairs as inputs. To obtain the complementary features, a cross-guided loss is designed, so that the output images are balanced with both modalities in an unsupervised manner. An intra-modality reweighting module is implemented to filter the redundant features before the fusion. Finally, YOLOv3 is chosen as the detector fed by the fused features. The proposed method is verified using the public KAIST and VOT-RGBT datasets. Experimental results demonstrate that the proposed method can outperform the state-of-the-art methods, the miss rate of pedestrian detection reaches 48.57% and 4.52% using KAIST and VOT-RGBT datasets, respectively. Zhenzhou Shao, Ying Qu 0001, Jun Zhang 0031, Zhi-Ping Shi 0002 |
ICPR | 4 |
| 2022 | Non-Local Representation Based Mutual Affine-Transfer Network for Photorealistic StylizationabstractPhotorealistic stylization aims to transfer the style of a reference photo onto a content photo in a natural fashion, such that the stylized image looks like a real photo taken by a camera. State-of-the-art methods stylize the image locally within each matched semantic region and are prone to global color inconsistency across semantic objects/parts, making the stylized image less photorealistic. To tackle the challenging issues, we propose a non-local representation scheme, constrained with a mutual affine-transfer network (NL-MAT). Through a dictionary-based decomposition, NL-MAT is able to successfully decouple matched non-local representations and color information of the image pair, such that the context correspondence between the image pair is incorporated naturally, which largely facilitates local style transfer in a global-consistent fashion. To the best of our knowledge, this is the first attempt to address the photorealistic stylization problem with a non-local representation scheme, such that no additional models or steps for semantic matching are required during stylization. Experimental results demonstrate that, the proposed method is able to generate photorealistic results with local style transfer while preserving both the spatial structure and global color consistency of the content image. Ying Qu 0001, Zhenzhou Shao, Hairong Qi 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Fast and Unsupervised Non-Local Feature Learning for Direct Volume Rendering of 3D Medical ImagesabstractTo improve the efficiency of medical visualization for computer aided surgery, we propose a fast and unsupervised 3D-CNN based non-local feature learning network. The proposed network consists of an encoder structure and a decoder structure. The encoder of the network projects the cube into a high-dimensional feature space, and the decoder of the network reconstructs the cube from the feature space. The decoder of the network serves as a dictionary shared by the cube to enforce the features for similar parts to be similar although they may distribute at disjointed locations. With such structures, the network is able to extract non-local features of the entire data. Moreover, a sparse constraint is incorporated into the network to increase the discriminative of the non-local features. Then the extracted non-local features of each voxel are fused with the corresponding position matrix and Hessian matrix for the voxel classification using Random Forest. Finally, a multidimensional transfer function is designed to enable the volume rendering. Experimental results demonstrate that the proposed method outperforms the state-of-the-art methods with much less training time. Xinmei Fu, Zhenzhou Shao, Ying Qu 0001, Yibo Zou, Zhi-Ping Shi 0002, Jindong Tan |
IROS | 2 |
| 2020 | Batch Normalization Masked Sparse Autoencoder for Robotic Grasping DetectionabstractTo improve the accuracy of the grasping detection, this paper proposes a novel detector with batch normalization masked evaluation model. It is designed with a two-layer sparse autoencoder, and a Batch Normalization based mask is incorporated into the second layer of the model to effectively reduce the features with weak correlation. The extracted features from such model are more distinctive, which guarantees the higher accuracy of the grasping detection. Extensive experiments show that the proposed evaluation model outperforms the state-of- the-art, and the recognition accuracy can reach 95.51% for robotic grasping detection. Zhenzhou Shao, Ying Qu 0001, Guangli Ren, Zhi-Ping Shi 0002, Jindong Tan |
IROS | 1 |
| 2020 | Deep Visible and Thermal Image Fusion with Cross-Modality Feature Selection for Pedestrian Detection
Zhenzhou Shao, Zhi-Ping Shi 0002 |
NPC | 2 |
| 2019 | Inverse Dynamics Modeling of Robotic Manipulator with Hierarchical Recurrent NetworkabstractInverse dynamics modeling is a critical problem for the computed-torque control of robotic manipulator. This paper presents a novel recurrent network based on the modified Simple Recurrent Unit (SRU) with hierarchical memory (SRU-HM), which is achieved by the nested SRU structure. In this way, it enables the capability to retain the long-term information in the distant past, compared with the conventional stacked structure. The hidden state of SRU is able to provide more complete information relevant to current prediction. Experimental results demonstrate that the proposed method can improve the accuracy of dynamics model greatly, and outperforms the state-of-the-art methods. Zhenzhou Shao, Ying Qu 0001, Jindong Tan |
IROS | 2 |
| 2018 | Unsupervised Trajectory Segmentation and Promoting of Multi-Modal Surgical DemonstrationsabstractTo improve the efficiency of surgical trajectory segmentation for robot learning in robot-assisted minimally invasive surgery, this paper presents a fast unsupervised method using video and kinematic data, followed by a promoting procedure to address the over-segmentation issue. Unsupervised deep learning network, stacking convolutional auto-encoder, is employed to extract more discriminative features from videos in an effective way. To further improve the accuracy of segmentation, on one hand, wavelet transform is used to filter out the noises existed in the features from video and kinematic data. On the other hand, the segmentation result is promoted by identifying the adjacent segments with no state transition based on the predefined similarity measurements. Extensive experiments on a public dataset JIGSAWS show that our method achieves much higher accuracy of segmentation than state-of-the-art methods in the shorter time. Zhenzhou Shao, Hongfa Zhao, Jiexin Xie, Ying Qu 0001, Jindong Tan |
IROS | 1 |
| 2017 | A fast search algorithm based on image pyramid for robotic graspingabstractTo improve the search efficiency of robotic grasping detection, this paper presents a novel search algorithm based on the image pyramid. It significantly reduces the search space for grasping position detection using the coarse-to-fine strategy. The proposed method searches the positions from the top layer of the pyramid, and initializes the search area at the next layer. The sparse automatic encoder is employed to construct the model which is used to evaluate the grasp quality. The experimental results demonstrate that the proposed search algorithm can improve efficiency of the robotic grasping detection with the comparative performance on the grasp quality. Guangli Ren, Zhenzhou Shao, Ying Qu 0001, Jindong Tan, Hongxing Wei, Guofeng Tong |
IROS | 2 |
| 2016 | RT-ROS: A real-time ROS architecture on multi-core processors
Hongxing Wei, Zhenzhou Shao, Renhai Chen, Jindong Tan, Zili Shao |
Future Gener. Comput. Syst. | 2 |
| 2015 | Recognition of Car Makes and Models From a Single Traffic-Camera ImageabstractThis paper proposes the recognition framework of car makes and models from a single image captured by a traffic camera. Due to various configurations of traffic cameras, a traffic image may be captured in different viewpoints and lighting conditions, and the image quality varies in resolution and color depth. In the framework, cars are first detected using a part-based detector, and license plates and headlamps are detected as cardinal anchor points to rectify projective distortion. Car features are extracted, normalized, and classified using an ensemble of neural-network classifiers. In the experiment, the performance of the proposed method is evaluated on a data set of practical traffic images. The results prove the effectiveness of the proposed method in vehicle detection and model recognition. Hongsheng He, Zhenzhou Shao, Jindong Tan |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2014 | Geometry constrained sparse embedding for multi-dimensional transfer function design in direct volume renderingabstractDirect volume rendering (DVR) is commonly employed for the medical visualization. Multi-dimensional transfer functions are used in DVR to emphasize the region of interest in details. However, it is impractical to interact directly with the functions in more than three dimension. This paper proposes a novel framework called geometry constrained sparse embedding (GCSE) for dimensionality reduction (DR). GCSE allows the conventional DR methods to be applied to a dictionary with much smaller atoms instead. The mapping derived from the dictionary feeds to the original features to obtain the ones in the reduced dimension. To obtain a good dictionary, the intrinsic structure of features is encoded in the sparse embedding based on a geometry distance. In addition, stochastic gradient descent algorithm is employed to speed up the dictionary learning. Various experiments have been conducted using both synthetic and real CT data sets. Compared with conventional methods, GCSE not only produces the comparable results, but also performs well with the capability to handle the large data set more powerfully. The rendering results using the real CT data has demonstrated the effectiveness of GCSE. Zhenzhou Shao, Hongsheng He, Jindong Tan |
ICRA | 1 |
| 2011 | A Real-Time Cardiac Arrhythmia Classification System with Wearable ElectrocardiogramabstractLong term continuous monitoring of electrocardiogram (ECG) in a free living environment provides valuable information for prevention on the heart attack and other high risk diseases. A design of a real-time wearable ECG monitoring system with cardiac arrhythmia classification is proposed in this paper. One of the striking advantages is that ECG analog front-end and on-node digital processing are designed to remove most of the noise and bias. In addition, a novel layered hidden Markov model is seamlessly integrated to classify multiple cardiac arrhythmias in real time. Last, human activities by an accelerometer can be identified to reduce the chance of false alarm in classification due to the motion artifacts. Zhenzhou Shao, Jindong Tan |
BSN | 2 |
| 2011 | Pedestrian positioning with physical activity classification for indoorsabstractThis paper presents a wearable Inertial Measurement Unit pedestrian positioning system for indoors. Hidden Markov Model (HMM) is introduced to pre-process the sensor data and classify common activities. HMM also complements local minimum angular rate value for capturing the onset/end of each step. ZUPT algorithm are implemented to correct the walking velocity at step stance phase when errors existed. A novel acceleration-based approach combined with gyroscope data is developed to achieve a better heading estimation. Proposed method is able to reduce drift errors from gyroscopes and avoid electromagnetic perturbance to magnetometers when estimate subject's position. Experiment results show the positioning system achieves approximately 99% accuracy. Zhenzhou Shao, Jindong Tan |
ICRA | 3 |