EDBT 2026 Demo / reviewers in the wild / expert
Sanyi Zhang
dblp:214/3977
· DBLP profile ↗
19ranked-venue papers
6as first author
16since 2021 · last 2025
0000-0003-3786-2299ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 14 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MixLGN: Mixed Local-Global Network for 3D Human Pose GenerationabstractAccurate 3D human pose generation plays a vital role in human-centric AI tasks. Although condition-guided (e.g. 2D pose) solutions have obtained suitable 3D human poses, the performance is insufficient while facing complex scenarios, e.g. occlusion or complicated actions. In this paper, we propose a novel Mixed Local-Global Network (MixLGN) that takes both intactness and plausibility into account for 3D human pose generation. Considering the local connectivity of adjacent keypoints, a novel dual convolution module is introduced to enhance the local ability for better spatial feature extraction. To guarantee the intactness of the 3D pose, we propose a part-to-body integrating mechanism that fuses local and global features together by taking guidance from human body hierarchy properties. Specifically, a part-aware pose feature encoding module is employed to learn structure-aware representation for each body part, and then put them together for global localization. We have verified the proposed MixLGN on two popular benchmark datasets, i.e., Human3.6M and MPI-INF-3DHP. The experimental results show that our model achieves the best performance both on single-hypothesis and multi-hypothesis. Sanyi Zhang, Chixuan Wei, Yinghao Yang 0002, Long Ye |
ICME | 2 |
| 2025 | MonoBite: Scale-Aware 3D Reconstruction and Volume Estimation from Monocular Multi-food Images
Songen Gu, Lina Liu 0010, Binjie Liu, Sanyi Zhang, Yanwei Fu 0001 |
PRCV (10) | 4 |
| 2025 | Full-Body Pose Motion Tracking From Sparse Data via Morphology-Aware Constraints
Yinghao Yang 0002, Sanyi Zhang, Chixuan Wei, Chenxi Feng, Long Ye |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Point Evolution Hierarchy Network for Weak Single-Point Human ParsingabstractSparsely single-point human parsing aims at segmenting the human body into fine-grained categories via weak point-level labels (e.g., point-level, scribble-level, or image-level, etc). The point-level label, especially single-point supervision, can simultaneously preserve spatial positions as well as take light annotation time, which is particularly advantageous in alleviating the human labeling burden. However, how to obtain satisfactory parsing performance under limited sparse point annotations is challenging, which requires further investigation. In this paper, we propose a novel end-to-end Point Evolution Hierarchy human parsing Network (PEHNet) for fine-grained human parsing task that just leverages single-point supervision. Motivated by the concept of a divide-and-conquer strategy, we partition all pixels into three distinct groups, i.e., single-point labels, pseudo-region labels, and unlabeled pixels, then optimize each group with suitable mechanisms. To expand the coverage of single-point labels, we introduce a point dissemination module that generates high-quality pseudo-region labels. Furthermore, the point-level spatial position information inherently preserves the structural characteristics of the human body. Inspired by this hierarchical property, we devise a point-level human hierarchy-wise constraint that guides the prediction probabilities to align with the inherent hierarchy of the human body. Experimental results demonstrate that the proposed PEHNet outperforms state-of-the-art parsing methods on two popular human parsing benchmark datasets (LIP and ATR) and one semantic segmentation dataset (Pascal VOC 2012). Sanyi Zhang, Xiaochun Cao, Long Ye, Zhanjie Song, Guo-Jun Qi, Jie Zhou 0001 |
IEEE Trans. Image Process. | 1 |
| 2024 | HPattack: An Effective Adversarial Attack for Human Parsing
Xin Dong 0015, Rui Wang 0032, Sanyi Zhang, Lihua Jing |
MMM (2) | 3 |
| 2024 | PoseVR: Structure-Aware Hybrid Full-Body Pose Estimation in Virtual Reality
Yinghao Yang 0002, Sanyi Zhang, Long Ye, Neng Rao |
PRCV (11) | 2 |
| 2024 | S2CL-Leaf Net: Recognizing Leaf Images Like Human BotanistsabstractAutomatically classifying plant leaves is a challenging fine-grained classification task because of the diversity in leaf morphology, including size, texture, shape, and venation. Although powerful deep learning-based methods have achieved great improvement in leaf classification, these methods still require a large number of well-labeled samples for supervised training, which is difficult to get. In contrast, relying on the specific coarse-to-fine classification strategy, human botanists only require a small number of samples for accurate leaf recognition. Inspired by the classification strategy of human botanists, we propose a novel S 2 CL-Leaf Net , which exploits multi-granularity clues with a hierarchical attention mechanism and boosts the learning ability with the supervised sampling contrastive learning with limited training samples to classify plant leaves as human botanists do. Specifically, to fully explore and exploit the subtle details of the leaves, a novel sampling transformation mechanism is combined with the supervised contrastive learning to enhance the network’s perception of details by amplifying the discriminative regions with a weighted sampling of different regions. Furthermore, we construct the hierarchical attention mechanism to produce attention maps of different granularity, which helps to discover details in leaves that are important for classification. Experiments are conducted on the open-access leaf datasets, including Flavia, Swedish, and LeafSnap, which prove the effectiveness of the proposed S 2 CL-Leaf Net . Cong Zou, Rui Wang 0032, Cheng Jin 0001, Sanyi Zhang, Xin Wang 0019 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | A Database for Multi-Modal Short Video Quality AssessmentabstractThe short video has gained increasing attention in information sharing and commercial promotions due to the fast development of social platforms. Accompanying, it introduces great requirements for assessing the quality of short videos for efficient information acquirement and propagation. However, existing video quality assessment researches focus on assessing video content with five rating scores, limiting the assessment to a one-dimension and simplified criterion. In this paper, we establish a novel database dubbed MMSVD-Douyin for assessing multi-modal short video quality under consideration of three evaluation criteria. It includes 4,684 short videos, three kinds of modalities, six kinds of data formats, and three assessment criteria. To conduct the short video quality assessment, we set up an all-around multi-modal short video quality assessment benchmark (MulSVQA) that dynamically fuses representations from three modalities and produces numbers of "likes", "shares" and "comments" of short videos. Chuan Wang 0002, Sanyi Zhang, Xiaochun Cao |
ICASSP | 3 |
| 2023 | Frequency Perception Network for Camouflaged Object DetectionabstractCamouflaged object detection (COD) aims to accurately detect objects hidden in the surrounding environment. However,the existing COD methods mainly locate camouflaged objects in the RGB domain, their performance has not been fully exploited in many challenging scenarios. Considering that the features of the camouflaged object and the background are more discriminative in the frequency domain, we propose a novel learnable and separable frequency perception mechanism driven by the semantic hierarchy in the frequency domain. Our entire network adopts a two-stage model, including a frequency-guided coarse localization stage and a detail-preserving fine localization stage.With the multi-level features extracted by the backbone, we design a flexible frequency perception module based on octave convolution for coarse positioning. Then, we design the correction fusion module to step-by-step integrate the high-level features through the prior-guided correction and cross-layer feature channel association, and finally combine them with the shallow features to achieve the detailed correction of the camouflaged objects. Compared with the currently existing models, our proposed method achieves competitive performance in three popular benchmark datasets both qualitatively and quantitatively. The code will be released at https://github.com/rmcong/FPNet_ACMMM23. Runmin Cong, Mengyao Sun 0003, Sanyi Zhang, Xiaofei Zhou 0003, Wei Zhang 0021, Yao Zhao 0001 |
ACM Multimedia | 3 |
| 2023 | Exploring the Robustness of Human Parsers Toward Common CorruptionsabstractHuman parsing aims to segment each pixel of the human image with fine-grained semantic categories. However, current human parsers trained with clean data are easily confused by numerous image corruptions such as blur and noise. To improve the robustness of human parsers, in this paper, we construct three corruption robustness benchmarks, termed LIP-C, ATR-C, and Pascal-Person-Part-C, to assist us in evaluating the risk tolerance of human parsing models. Inspired by the data augmentation strategy, we propose a novel heterogeneous augmentation-enhanced mechanism to bolster robustness under commonly corrupted conditions. Specifically, two types of data augmentations from different views, i.e., image-aware augmentation and model-aware image-to-image transformation, are integrated in a sequential manner for adapting to unforeseen image corruptions. The image-aware augmentation can enrich the high diversity of training images with the help of common image operations. The model-aware augmentation strategy that improves the diversity of input data by considering the model's randomness. The proposed method is model-agnostic, and it can plug and play into arbitrary state-of-the-art human parsing frameworks. The experimental results show that the proposed method demonstrates good universality which can improve the robustness of the human parsing models and even the semantic segmentation models when facing various image common corruptions. Meanwhile, it can still obtain approximate performance on clean data. Sanyi Zhang, Xiaochun Cao, Rui Wang 0032, Guo-Jun Qi, Jie Zhou 0001 |
IEEE Trans. Image Process. | 1 |
| 2022 | PMP-NET: Rethinking Visual Context for Scene Graph GenerationabstractScene graph generation aims to describe the contents in scenes by identifying the objects and their relationships. In previous works, visual context is widely utilized in message passing networks to generate the representations for classification. However, the noisy estimation of visual context limits model performance. In this paper, we revisit the concept of incorporating visual context via a randomly ordered bidirectional Long Short Temporal Memory (biLSTM) based baseline, and show that noisy estimation is worse than random. To alleviate the problem, we propose a new method, dubbed Progressive Message Passing Network (PMP-Net) that better estimates the visual context in a coarse to fine manner. Specifically, we first estimate the visual context with a random initiated scene graph, then refine it with multi-head attention. The experimental results on the benchmark dataset Visual Genome show that PMP-Net achieves better or comparable performance on all three tasks: scene graph generation (SGGen), scene graph classification (SGCls), and predicate classification (PredCls). Xuezhi Tong, Rui Wang 0032, Chuan Wang 0002, Sanyi Zhang, Xiaochun Cao |
ICASSP | 4 |
| 2022 | ACE: Anchor-Free Corner Evolution for Real-Time Arbitrarily-Oriented Object DetectionabstractObjects with different orientations are ubiquitous in the real world (e.g., texts/hands in the scene image, objects in the aerial image, etc.), and the widely-used axis-aligned bounding box does not compactly enclose the oriented objects. Thus arbitrarily-oriented object detection has attracted rising attention in recent years. In this paper, we propose a novel and effective model to detect arbitrarily-oriented objects. Instead of directly predicting the angles of oriented bounding boxes like most existing methods, we evolve the axis-aligned bounding box to the oriented quadrilateral box with the assistance of dynamically gathering contour information. More specifically, we first obtain the axis-aligned bounding box in an anchor-free manner. After that, we set the key points based on the sampled contour points of the axis-aligned bounding box. To improve the localization performance, we enrich the feature representations of these key points by exploiting a dynamic information gathering mechanism. This technique propagates the geometrical and semantic information along the sampled contour points, and fuses the information from the semantic neighbors of each sampled point, which varies for different locations. Finally, we estimate the offsets between the axis-aligned bounding box key points and the oriented quadrilateral box corner points. Extensive experiments on two frequently-used aerial image benchmarks HRSC2016 and DOTA, as well as scene text/hand datasets ICDAR2015, TD500, and Oxford-Hand, demonstrate the effectiveness and advantage of our proposed model. Pengwen Dai, Siyuan Yao, Zekun Li 0007, Sanyi Zhang, Xiaochun Cao |
IEEE Trans. Image Process. | 4 |
| 2022 | AIParsing: Anchor-Free Instance-Level Human ParsingabstractMost state-of-the-art instance-level human parsing models adopt two-stage anchor-based detectors and, therefore, cannot avoid the heuristic anchor box design and the lack of analysis on a pixel level. To address these two issues, we have designed an instance-level human parsing network which is anchor-free and solvable on a pixel level. It consists of two simple sub-networks: an anchor-free detection head for bounding box predictions and an edge-guided parsing head for human segmentation. The anchor-free detector head inherits the pixel-like merits and effectively avoids the sensitivity of hyper-parameters as proved in object detection applications. By introducing the part-aware boundary clue, the edge-guided parsing head is capable to distinguish adjacent human parts from among each other up to 58 parts in a single human instance, even overlapping instances. Meanwhile, a refinement head integrating box-level score and part-level parsing quality is exploited to improve the quality of the parsing results. Experiments on two multiple human parsing datasets (i.e., CIHP and LV-MHP-v2.0) and one video instance-level human parsing dataset (i.e., VIP) show that our method achieves the best global-level and instance-level performance over state-of-the-art one-stage top-down alternatives. Sanyi Zhang, Xiaochun Cao, Guo-Jun Qi, Zhanjie Song, Jie Zhou 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | Progressive Contour Regression for Arbitrary-Shape Scene Text DetectionabstractState-of-the-art scene text detection methods usually model the text instance with local pixels or components from the bottom-up perspective and, therefore, are sensitive to noises and dependent on the complicated heuristic post-processing especially for arbitrary-shape texts. To relieve these two issues, instead, we propose to progressively evolve the initial text proposal to arbitrarily shaped text contours in a top-down manner. The initial horizontal text proposals are generated by estimating the center and size of texts. To reduce the range of regression, the first stage of the evolution predicts the corner points of oriented text proposals from the initial horizontal ones. In the second stage, the contours of the oriented text proposals are iteratively regressed to arbitrarily shaped ones. In the last iteration of this stage, we rescore the confidence of the final localized text by utilizing the cues from multiple contour points, rather than the single cue from the initial horizontal proposal center that may be out of arbitrary-shape text regions. Moreover, to facilitate the progressive contour evolution, we design a contour information aggregation mechanism to enrich the feature representation on text contours by considering both the circular topology and semantic context. Experiments conducted on CTW1500, Total-Text, ArT, and TD500 have demonstrated that the proposed method especially excels in line-level arbitrary-shape texts. Code is available at https://github.com/dpengwen/PCR. Pengwen Dai, Sanyi Zhang, Hua Zhang 0008, Xiaochun Cao |
CVPR | 2 |
| 2021 | Few-Shot Classification with Multi-task Self-supervised Learning
Rui Wang 0032, Sanyi Zhang, Xiaochun Cao |
ICONIP (4) | 3 |
| 2021 | Human Parsing With Pyramidical Gather-Excite ContextabstractHuman parsing, especially in the wild, has attracted a lot of attention due to its great potential in many real-world applications. The Pyramid Spatial Parsing (PSP) module has shown superior performances in scene and human parsing tasks. However, the basic AvgPool operation in PSP equally aggregates spatial clues of a local region, and thus mixes up influences of different human parts presented in this region. It results in failures in capturing useful contexts relevant to parsing different parts. To address this problem, a suitable mechanism to collect spatial clues aligning with different human parts is proposed in this paper. We employ a Gather-Excite (GE) operation, a replacement of the AvgPool-Upsample operation in a pyramidical structure, to accurately reflect relevant human parts of various scales. The GE operation contains two steps: the gather operation that adaptively aggregates spatial clues to relevant human parts, and the excite operation that generates new feature maps with the gathered contextual information. This results in a novel Pyramidical Gather-Excite Context (PGEC) module to solve the multi-scale problem and parse person at various scales. The PGEC module is composed of multiple GE operations with different spatial extents and aggregates local and global spatial clues for better modeling multi-scale contextual information in parallel. Moreover, we integrate the PGEC module with fine-grained details, edge preserving module and deep supervision to formulate a novel PGEC Network (PGECNet) for human parsing. The proposed PGECNet has achieved state-of-the-art performance on four single-person human parsing datasets (i.e., LIP, PPSS, ATR and Fashion Clothing) and two multi-person human parsing datasets (i.e., PASCAL-Person-Part and CIHP). The experimental results show that the proposed PGEC is superior to the PSP and ASPP modules especially in single-human parsing task. The source code is publicly available at https://github.com/31sy/PGECNet. Sanyi Zhang, Guo-Jun Qi, Xiaochun Cao, Zhanjie Song, Jie Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Task-Aware Attention Model for Clothing Attribute PredictionabstractClothing attribute recognition, especially in unconstrained street images, is a challenging task for multimedia. Existing methods for multi-task clothing attribute prediction often ignore the relation between specific attributes and positions. However, the attribute response is always location-sensitive, i.e., different spatial locations have various contributions to attributes. Inspired by the locality of clothing attributes, in this paper, we introduce the attention mechanism to incorporate the impact of positions for clothing attribute prediction with only image-level annotations. However, the performance improvement is limited if we directly use the traditional spatial attention model for each task since it does not take the influence from other tasks into account. Instead, we propose a novel task-aware attention mechanism, which estimates the importance of each position across different tasks. We first evaluate a task attention network with an end-to-end multi-task clothing attribute learning architecture on the shop domain. And then, we employ curriculum learning strategy, which transfers the well-trained shop domain attribute knowledge to the street domain attribute prediction. Experiments are conducted on three clothing benchmarks, i.e., cross-domain clothing attribute dataset, woman clothing dataset, and man clothing dataset. The performance of attribute prediction demonstrates the superiority of the proposed task-aware attention mechanism over several state-of-the-art methods both in shop and street domains. Sanyi Zhang, Zhanjie Song, Xiaochun Cao, Hua Zhang 0008, Jie Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Nested Network With Two-Stream Pyramid for Salient Object Detection in Optical Remote Sensing ImagesabstractArising from the various object types and scales, diverse imaging orientations, and cluttered backgrounds in optical remote sensing image (RSI), it is difficult to directly extend the success of salient object detection for nature scene image to the optical RSI. In this paper, we propose an end-to-end deep network called LV-Net based on the shape of network architecture, which detects salient objects from optical RSIs in a purely data-driven fashion. The proposed LV-Net consists of two key modules, i.e., a two-stream pyramid module (L-shaped module) and an encoder-decoder module with nested connections (V-shaped module). Specifically, the L-shaped module extracts a set of complementary information hierarchically by using a two-stream pyramid structure, which is beneficial to perceiving the diverse scales and local details of salient objects. The V-shaped module gradually integrates encoder detail features with decoder semantic features through nested connections, which aims at suppressing the cluttered backgrounds and highlighting the salient objects. In addition, we construct the first publicly available optical RSI data set for salient object detection, including 800 images with varying spatial resolutions, diverse saliency types, and pixel-wise ground truth. Experiments on this benchmark data set demonstrate that the proposed method outperforms the state-of-the-art salient object detection methods both qualitatively and quantitatively. Chongyi Li, Runmin Cong, Junhui Hou, Sanyi Zhang, Sam Kwong |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2018 | Watch fashion shows to tell clothing attributes
Sanyi Zhang, Si Liu 0001, Xiaochun Cao, Zhanjie Song, Jie Zhou 0001 |
Neurocomputing | 1 |