EDBT 2026 Demo / reviewers in the wild / expert
Xiangzeng Zhou
dblp:136/5317
· DBLP profile ↗
9ranked-venue papers
5as first author
2since 2021 · last 2021
0000-0003-3123-6561ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Deep learning architectures and training · 13% Graph learning · 13% Knowledge representation and reasoning · 13% | |
| Computer graphics and multimedia
2 papers |
Multimedia analysis and retrieval · 100% |
Topics — the 10 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Graph learning › graph neural network
graph convolutional network |
0.5 | 1 | 2021 | Extremely Compact Non-local Representation Learning · KDD 2021 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › automated reasoning
high-order reasoning |
0.5 | 1 | 2021 | Extremely Compact Non-local Representation Learning · KDD 2021 |
Machine learning › Efficient and distributed learning › distributed training › distributed training systems
large-scale distributed training |
0.5 | 1 | 2021 | Once and for All: Self-supervised Multi-modal Co-training on One-billion Videos at Alibaba · ACM Multimedia 2021 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
multimodal self-supervised learning |
0.5 | 1 | 2021 | Once and for All: Self-supervised Multi-modal Co-training on One-billion Videos at Alibaba · ACM Multimedia 2021 |
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels |
0.4 | 1 | 2020 | Weakly Supervised Learning with Side Information for Noisy Labeled Images · ECCV (30) 2020 |
Machine learning › Learning paradigms
weakly supervised learning |
0.4 | 1 | 2020 | Weakly Supervised Learning with Side Information for Noisy Labeled Images · ECCV (30) 2020 |
Computer vision › Video understanding and tracking
object tracking |
0.2 | 1 | 2015 | Online Object Tracking Based on CNN with Metropolis-Hasting Re-Sampling · ACM Multimedia 2015 |
Computer vision › Video understanding and tracking › object tracking
online tracking |
0.2 | 1 | 2015 | Online Object Tracking Based on CNN with Metropolis-Hasting Re-Sampling · ACM Multimedia 2015 |
Multimedia analysis and retrieval
object tracking |
0.2 | 1 | 2015 | Tennis Ball Tracking Using a Two-Layered Data Association Approach · IEEE Trans. Multim. 2015 |
Multimedia analysis and retrieval › multimedia feature representation
video representation learning |
0.1 | 1 | 2021 | Once and for All: Self-supervised Multi-modal Co-training on One-billion Videos at Alibaba · ACM Multimedia 2021 |
Methods — techniques the papers use, named apart from their topics
sliding-window subset sampling · 1.0cross-modal pseudo-label consistency · 1.0contrastive co-training · 1.0coarse-to-fine clustering · 1.0graph convolution network · 0.5global hadamard pooling · 0.5side information · 0.4shift token transfer · 0.2particle filtering · 0.2metropolis-hastings re-sampling · 0.2dynamic programming · 0.2CNN · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Extremely Compact Non-local Representation LearningabstractIn contrast to regular convolutions with local receptive fields, non-local operations have widely proven an effective method for modeling long-range dependencies. Although lots of prior works have been proposed, prohibitive computation and GPU memory occupation are still the major concerns. Different from that carrying out non-local operations pixel-wise or channel-wise in a computation intensive way, we argue that we can achieve effective non-local operation using a more compact high-order statistic, which can be computed more efficiently and may convey some high-level information. In this paper, we propose an extremely compact non-local learning module (CoNL) with high-order reasoning based on a graph convolution as the core. In our CoNL, a global Hadamard pooling (GHP) as a non-local operation is used to extract a compact second-order feature vector from the input tensor. With the help of a light-weight graph convolution network (GCN), this high-order compact vector is further refined with high-level reasoning. After the GCN refinement, the compact high-order vector intuitively indicates some global semantic characteristics, and is eventually applied to enhance the input tensor through a channel scaling operation. The CoNL module is designed easily pluggable to upgrade existing networks. Extensive experiments on a wide range of tasks demonstrate the effectiveness and efficiency of our work. The proposed CoNL can achieve comparable or superior performance over previous state-of-the-art baselines on video recognition, semantic segmentation, object detection and instance segmentation tasks. For a 96 x 96 x 2048 input, our block consumes 13.6 x less in computational cost than non-local block while 7.6 x smaller in GPU memory occupation. Ansheng You, Xiangzeng Zhou, Yingya Zhang |
KDD | 2 |
| 2021 | Once and for All: Self-supervised Multi-modal Co-training on One-billion Videos at AlibabaabstractVideos grow to be one of the largest mediums on the Internet. E-commerce platforms like Alibaba need to process millions of video data across multimedia (e.g., visual, audio, image, and text) and on a variety of tasks (e.g., retrieval, tagging, and summary) every day. In this work, we aim to develop a once and for all pretraining technique for diverse modalities and downstream tasks. To achieve this, we make the following contributions: (1) We propose a self-supervised multi-modal co-training framework. It takes cross-modal pseudo-label consistency as the supervision and can jointly learn representations of multiple modalities. (2) We introduce several novel techniques (e.g., sliding-window subset sampling, coarse-to-fine clustering, fast spatial-temporal convolution and parallel data transmission and processing) to optimize the training process, making billion-scale stable training feasible. (3) We construct a large-scale multi-modal dataset consisting of 1.4 billion videos (~0.5 PB) and train our framework on it. The training takes only 4.6 days on an in-house 256 GPUs cluster, and it simultaneously produces pretrained video, audio, image, motion, and text networks. (4) Finetuning from our pretrained models, we obtain significant performance gains and faster convergence on diverse multimedia tasks at Alibaba. Furthermore, we also validate the learned representation on public datasets. Despite the domain gap between our commodity-centric pretraining and the action-centric evaluation data, we show superior results against state-of-the-arts. Lianghua Huang, Yu Liu 0063, Xiangzeng Zhou, Ansheng You, Yingya Zhang |
ACM Multimedia | 3 |
| 2020 | Large Scale Long-tailed Product Recognition System at AlibabaabstractA practical large scale product recognition system suffers from the phenomenon of long-tailed imbalanced training data under the E-commercial circumstance at Alibaba. Besides product images at Alibaba, plenty of image related side information (e.g. title, tags) reveal rich semantic information about images. Prior works mainly focus on addressing the long tail problem in visual perspective only, but lack of consideration of leveraging the side information. In this paper, we present a novel side information based large scale visual recognition co-training~(SICoT) system to deal with the long tail problem by leveraging the image related side information. In the proposed co-training system, we firstly introduce a bilinear word attention module aiming to construct a semantic embedding over the noisy side information. A visual feature and semantic embedding co-training scheme is then designed to transfer knowledge from classes with abundant training data (head classes) to classes with few training data (tail classes) in an end-to-end fashion. Extensive experiments on four challenging large scale datasets, whose numbers of classes range from one thousand to one million, demonstrate the scalable effectiveness of the proposed SICoT system in alleviating the long tail problem. In the visual search platform Pailitao\footnote{http://www.pailitao.com} at Alibaba, we settle a practical large scale product recognition application driven by the proposed SICoT system, and achieve a significant gain of unique visitor~(UV) conversion rate. Xiangzeng Zhou, Rong Jin 0001 |
CIKM | 1 |
| 2020 | Weakly Supervised Learning with Side Information for Noisy Labeled Images
Lele Cheng, Xiangzeng Zhou, Dangwei Li, Hong Shang |
ECCV (30) | 2 |
| 2015 | Online Object Tracking Based on CNN with Metropolis-Hasting Re-SamplingabstractTracking-by-learning strategies have been effective in solving many challenging problems in visual tracking, in which the learning sample generation and labeling play important roles for final performance. Since the concern of deep learning based approaches has shown an impressive performance in different vision tasks, how to properly apply the learning model, such as CNN, to an online tracking framework is still challenging. In this paper, to overcome the overfitting problem caused by straight-forward incorporation, we propose an online tracking framework by constructing a CNN based adaptive appearance model to generate more reliable training data over time. With a reformative Metropolis-Hastings re-sampling scheme to reshape particles for a better state posterior representation during online learning, the proposed tracking outperforms most of the state-of-art trackers on challenging benchmark video sequences. Xiangzeng Zhou, Lei Xie 0001, Peng Zhang 0005, Yanning Zhang 0001 |
ACM Multimedia | 1 |
| 2015 | Tennis Ball Tracking Using a Two-Layered Data Association ApproachabstractBall tracking is a key technology in processing and analyzing a ball game. Because of the complexity of visual scenes, a large number of objects are often selected as candidates for the ball, leading to incorrect identification, and conversely, the true position of the ball may sometimes be missed because of occlusion and blur, which can both be frequent and severe. Several tennis ball tracking algorithms have been reported in literature. In this paper, we propose a two-layered data association method to improve the robustness of tennis ball tracking. At the local layer, a shift token transfer method is proposed, based on shift window processing, to generate a set of short trajectories or “ trajectorylets .” At the global layer, a unique ball trajectory is obtained by applying a dynamic programming based splice method to a directed acyclic graph consisting of trajectorylets. We evaluated our approach on tennis matches from the Australian Open and the U.S. Open, and the results obtained show that our approach outperforms current state-of-the-art approach. Xiangzeng Zhou, Lei Xie 0001, Qiang Huang 0006, Stephen J. Cox, Yanning Zhang 0001 |
IEEE Trans. Multim. | 1 |
| 2014 | Unsupervised broadcast news story segmentation using distance dependent Chinese restaurant processesabstractTraditional unsupervised broadcast news story segmentation approaches have to set the segmentation number manually, while this number is often unknown in real-world applications. In this paper, we solve this problem by modeling the generative process of stories as distance dependent Chinese restaurant process (dd-CRP) mixtures. We cut a news program into fixed-size text blocks and consider these blocks in the same story are generated from a story-specific topic. Specifically, we add a dd-CRP prior which has an essential bias that the blocks' topic is more likely to be the same with the nearby blocks. Subsequently, story boundaries can be found by detecting the changes of topics. Experiments show that our approach outperforms both supervised and unsupervised approaches and the segmentation number can be automatically learned from data. Chao Yang 0031, Lei Xie 0001, Xiangzeng Zhou |
ICASSP | 3 |
| 2014 | An ensemble of deep neural networks for object trackingabstractObject tracking in complex backgrounds with dramatic appearance variations is a challenging problem in computer vision. We tackle this problem by a novel approach that incorporates a deep learning architecture with an on-line AdaBoost framework. Inspired by its multi-level feature learning ability, a stacked denoising autoencoder (SDAE) is used to learn multi-level feature descriptors from a set of auxiliary images. Each layer of the SDAE, representing a different feature space, is subsequently transformed to a discriminative object/background deep neural network (DNN) classifier by adding a classification layer. By an on-line AdaBoost feature selection framework, the ensemble of the DNN classifiers is then updated on-line to robustly distinguish the target from the background. Experiments on an open tracking benchmark show promising results of the proposed tracker as compared with several state-of-the-art approaches. Xiangzeng Zhou, Lei Xie 0001, Peng Zhang 0005, Yanning Zhang 0001 |
ICIP | 1 |
| 2013 | A two layered data association approach for ball trackingabstractBall-tracking is a key technology in processing and analyzing a ball game. Because of the complexity of visual scenes, a large number of objects are usually selected as candidates for the ball, leading to incorrect identification, and conversely, the true position of the ball may sometimes be missed. In this paper, we propose a two layered data association method to improve the robustness of ball-tracking. At a local layer, we use a sliding window based Token Transfer method to generate a set of sub-trajectory candidates. At a global layer, a single ball trajectory is obtained by applying a dynamic programming based splice method to a graph consisting of the sub-trajectory candidates. We evaluated our approach on tennis matches from the Australian Open and the U.S. Open, and the results obtained show that our approach outperforms the state-of-art approach by around 30%. Xiangzeng Zhou, Qiang Huang 0006, Lei Xie 0001, Stephen J. Cox |
ICASSP | 1 |