EDBT 2026 Demo / reviewers in the wild / expert
Fangyi Zhang
dblp:10/8496
· DBLP profile ↗
12ranked-venue papers
3as first author
7since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 4 since 2021Systems, architecture and hardware · 5 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Robot manipulation · 46% Graph learning · 16% Video understanding and tracking · 10% | |
| Computer graphics and multimedia
1 paper |
Image and video coding · 100% | |
| Computer networks
1 paper |
Wireless sensing and localization · 50% Physical-layer communications · 50% |
Topics — the 20 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Robotics › Robot manipulation
deformable object manipulation |
0.8 | 1 | 2024 | Learning Fabric Manipulation in the Real World with Human Videos · ICRA 2024 |
Robotics › Robot manipulation › deformable object manipulation
fabric manipulation |
0.8 | 1 | 2024 | Learning Fabric Manipulation in the Real World with Human Videos · ICRA 2024 |
Robotics › Robot manipulation
learning from demonstration |
0.8 | 1 | 2024 | Learning Fabric Manipulation in the Real World with Human Videos · ICRA 2024 |
Robotics › Robot manipulation › learning from demonstration
learning from human video |
0.8 | 1 | 2024 | Learning Fabric Manipulation in the Real World with Human Videos · ICRA 2024 |
Computer vision › Face, body and person analysis
face clustering |
0.6 | 1 | 2022 | Ada-NETS: Face Clustering via Adaptive Neighbour Discovery in the Structure Space · ICLR 2022 |
Machine learning › Graph learning
graph structure learning |
0.6 | 1 | 2022 | Robust Graph Structure Learning via Multiple Statistical Tests · NeurIPS 2022 |
Machine learning › Representation and self-supervised learning › representation learning
metric learning |
0.6 | 1 | 2022 | Robust Graph Structure Learning via Multiple Statistical Tests · NeurIPS 2022 |
Image and video coding
image compression |
0.5 | 1 | 2021 | Interpolation Variable Rate Image Compression · ACM Multimedia 2021 |
Image and video coding › image compression
learned image compression |
0.5 | 1 | 2021 | Interpolation Variable Rate Image Compression · ACM Multimedia 2021 |
Image and video coding
variable-rate coding |
0.5 | 1 | 2021 | Interpolation Variable Rate Image Compression · ACM Multimedia 2021 |
Computer vision › Video understanding and tracking
object tracking |
0.4 | 1 | 2019 | SiamRPN++: Evolution of Siamese Visual Tracking With Very Deep Networks · CVPR 2019 |
Computer vision › Video understanding and tracking › object tracking › deep tracking
siamese tracking |
0.4 | 1 | 2019 | SiamRPN++: Evolution of Siamese Visual Tracking With Very Deep Networks · CVPR 2019 |
Robotics › Robot manipulation
grasping |
0.3 | 1 | 2017 | The ACRV picking benchmark: A robotic shelf picking benchmark to foster reproducible research · ICRA 2017 |
Machine learning › Deep learning architectures and training
reproducible benchmarking |
0.3 | 1 | 2017 | The ACRV picking benchmark: A robotic shelf picking benchmark to foster reproducible research · ICRA 2017 |
Wireless sensing and localization
indoor localization |
0.2 | 1 | 2015 | Asynchronous blind signal decomposition using tiny-length code for Visible Light Communication-based indoor localization · ICRA 2015 |
Physical-layer communications › signal processing for communications › signal representation
signal decomposition |
0.2 | 1 | 2015 | Asynchronous blind signal decomposition using tiny-length code for Visible Light Communication-based indoor localization · ICRA 2015 |
Physical-layer communications
signal processing for communications |
0.2 | 1 | 2015 | Asynchronous blind signal decomposition using tiny-length code for Visible Light Communication-based indoor localization · ICRA 2015 |
Wireless sensing and localization › optical sensing › optical localization
visible light communication localization |
0.2 | 1 | 2015 | Asynchronous blind signal decomposition using tiny-length code for Visible Light Communication-based indoor localization · ICRA 2015 |
Computer vision › Face, body and person analysis
person re-identification |
0.2 | 1 | 2022 | Robust Graph Structure Learning via Multiple Statistical Tests · NeurIPS 2022 |
Image and video coding
quality assessment |
0.1 | 1 | 2021 | Interpolation Variable Rate Image Compression · ACM Multimedia 2021 |
Methods — techniques the papers use, named apart from their topics
pick-and-place policy learning · 0.8imitation learning · 0.8statistical hypothesis testing · 0.6attention mechanism · 0.6adaptive neighbour discovery · 0.6linear interpolation · 0.5interpolation channel attention · 0.5spatial aware sampling · 0.4layer-wise aggregation · 0.4depth-wise aggregation · 0.4evaluation protocol · 0.3benchmark design · 0.3gold sequences · 0.2correlation-based decomposition · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Learning Fabric Manipulation in the Real World with Human VideosabstractFabric manipulation is a long-standing challenge in robotics due to the enormous state space and complex dynamics. Learning approaches stand out as promising for this domain as they allow us to learn behaviours directly from data. Most prior methods however rely heavily on simulation, which is still limited by the large sim-to-real gap of deformable objects or rely on large datasets. A promising alternative is to learn fabric manipulation directly from watching humans perform the task. In this work, we explore how demonstrations for fabric manipulation tasks can be collected directly by humans, providing an extremely natural and fast data collection pipeline. Then, using only a handful of such demonstrations, we show how a pick-and-place policy can be learned and deployed on a real robot, without any robot data collection at all. We demonstrate our approach on a fabric smoothing and folding task, showing that our policy can reliably reach folded states from crumpled initial configurations. Code, video and data are available on the project website: https://sites.google.com/view/foldingbyhand Robert Lee, Jad Abou-Chakra, Fangyi Zhang, Peter I. Corke |
ICRA | 3 |
| 2023 | Re-Evaluating Parallel Finger-Tip Tactile Sensing for Inferring Object Adjectives: An Empirical StudyabstractFinger-tip tactile sensors are increasingly used for robotic sensing to establish stable grasps and to infer object properties. Promising performance has been shown in a number of works for inferring adjectives that describe the object, but there remains a question about how each taxel contributes to the performance. This paper explores this question with empirical experiments, leading insights for future finger-tip tactile sensor usage and design: one tactile sensor instead of a pair of sensors is sufficient for symmetric objects and interaction motions; dense taxels are beneficial for texture-related adjectives, but can be distracting to non-texture-related ones; and a frame-rate much lower than the BioTac sensor can satisfy the demand of inferring object adjectives in the PHAC-2 dataset. Fangyi Zhang, Peter I. Corke |
IROS | 1 |
| 2023 | A Linkage-based Doubly Imbalanced Graph Learning Framework for Face ClusteringabstractIn recent years, benefiting from the expressive power of Graph Convolutional Networks (GCNs), significant breakthroughs have been made in face clustering area. However, rare attention has been paid to GCN-based clustering on imbalanced data. Although imbalance problem has been extensively studied, the impact of imbalanced data on GCN- based linkage prediction task is quite different, which would cause problems in two aspects: imbalanced linkage labels and biased graph representations. The former is similar to that in classic image classification task, but the latter is a particular problem in GCN-based clustering via linkage prediction. Significantly biased graph representations in training can cause catastrophic over-fitting of a GCN model. To tackle these challenges, we propose a linkage-based doubly imbalanced graph learning framework for face clustering. In this framework, we evaluate the feasibility of those existing methods for imbalanced image classification problem on GCNs, and present a new method to alleviate the imbal- anced labels and also augment graph representations using a Reverse-Imbalance Weighted Sampling (RIWS) strategy. With the RIWS strategy, probability-based class balancing weights could ensure the overall distribution of positive and negative samples; In addition, weighted random sampling provides diverse subgraph structures, which effectively alleviates the over-fitting problem and improves the representation ability of GCNs. Extensive experiments on series of imbalanced benchmark datasets synthesized from MS-Celeb-1M and DeepFashion demonstrate the effectiveness and generality of our proposed method. Our implementation and the synthesized datasets will be openly available on https://github.com/espectre/GCNs_on_imbalanced_datasets. Huafeng Yang, Qijie Shen, Xingjian Chen, Fangyi Zhang |
SDM | 4 |
| 2022 | Jmpnet: Joint Motion Prediction for Learning-Based Video CompressionabstractIn recent years, more attention is attracted by learning-based approaches in the field of video compression. Recent methods of this kind normally consist of three major components: intra-frame network, motion prediction network, and residual network, among which the motion prediction part is particularly critical for video compression. Benefiting from the optical flow which enables dense motion prediction, recent methods have shown competitive performance compared with traditional codecs. However, problems such as tail shadow and background distortion in the predicted frame remain unsolved. To tackle these problems, JMPNet is introduced in this paper to provide more accurate motion information by using both optical flow and dynamic local filter as well as an attention map to further fuse these motion information in a smarter way. Experimental results show that the proposed method surpasses state-of-the-art (SOTA) rate-distortion (RD) performance in the most data-sets. Zhenhong Sun, Zhiyu Tan, Xiuyu Sun, Fangyi Zhang, Yichen Qian, Hao Li 0030 |
ICASSP | 5 |
| 2022 | Ada-NETS: Face Clustering via Adaptive Neighbour Discovery in the Structure Space
Yaobin Zhang, Fangyi Zhang, Senzhang Wang, Ming Lin 0002, YuQi Zhang, Xiuyu Sun |
ICLR | 3 |
| 2022 | Robust Graph Structure Learning via Multiple Statistical TestsabstractGraph structure learning aims to learn connectivity in a graph from data. It is particularly important for many computer vision related tasks since no explicit graph structure is available for images for most cases. A natural way to construct a graph among images is to treat each image as a node and assign pairwise image similarities as weights to corresponding edges. It is well known that pairwise similarities between images are sensitive to the noise in feature representations, leading to unreliable graph structures. We address this problem from the viewpoint of statistical tests. By viewing the feature vector of each node as an independent sample, the decision of whether creating an edge between two nodes based on their similarity in feature representation can be thought as a ${\it single}$ statistical test. To improve the robustness in the decision of creating an edge, multiple samples are drawn and integrated by ${\it multiple}$ statistical tests to generate a more reliable similarity measure, consequentially more reliable graph structure. The corresponding elegant matrix form named $\mathcal{B}$$\textbf{-Attention}$ is designed for efficiency. The effectiveness of multiple tests for graph structure learning is verified both theoretically and empirically on multiple clustering and ReID benchmark datasets. Source codes are available at https://github.com/Thomas-wyh/B-Attention. Fangyi Zhang, Ming Lin 0002, Senzhang Wang, Xiuyu Sun, Rong Jin 0001 |
NeurIPS | 2 |
| 2021 | Interpolation Variable Rate Image CompressionabstractCompression standards have been used to reduce the cost of image storage and transmission for decades. In recent years, learned image compression methods have been proposed and achieved compelling performance to the traditional standards. However, in these methods, a set of different networks are used for various compression rates, resulting in a high cost in model storage and training. Although some variable-rate approaches have been proposed to reduce the cost by using a single network, most of them brought some performance degradation when applying fine rate control. To enable variable-rate control without sacrificing the performance, we propose an efficient Interpolation Variable-Rate (IVR) network, by introducing a handy Interpolation Channel Attention (InterpCA) module in the compression network. With the use of two hyperparameters for rate control and linear interpolation, the InterpCA achieves a fine PSNR interval of 0.001 dB and a fine rate interval of 0.0001 Bits-Per-Pixel (BPP) with 9000 rates in the IVR network. Experimental results demonstrate that the IVR network is the first variable-rate learned method that outperforms VTM 9.0 (intra) in PSNR and Multiscale Structural Similarity (MS-SSIM). Zhenhong Sun, Zhiyu Tan, Xiuyu Sun, Fangyi Zhang, Yichen Qian, Hao Li 0030 |
ACM Multimedia | 4 |
| 2019 | Relation-aware Multiple Attention Siamese Networks for Robust Visual Tracking
Fangyi Zhang, Bingpeng Ma, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001 |
BMVC | 1 |
| 2019 | SiamRPN++: Evolution of Siamese Visual Tracking With Very Deep NetworksabstractSiamese network based trackers formulate tracking as convolutional feature cross-correlation between target template and searching region. However, Siamese trackers still have accuracy gap compared with state-of-the-art algorithms and they cannot take advantage of feature from deep networks, such as ResNet-50 or deeper. In this work we prove the core reason comes from the lack of strict translation invariance. By comprehensive theoretical analysis and experimental validations, we break this restriction through a simple yet effective spatial aware sampling strategy and successfully train a ResNet-driven Siamese tracker with significant performance gain. Moreover, we propose a new model architecture to perform depth-wise and layer-wise aggregations, which not only further improves the accuracy but also reduces the model size. We conduct extensive ablation studies to demonstrate the effectiveness of the proposed tracker, which obtains currently the best results on four large tracking benchmarks, including OTB2015, VOT2018, UAV123, and LaSOT. Our model will be released to facilitate further studies based on this problem. Bo Li 0114, Wei Wu 0021, Qiang Wang 0051, Fangyi Zhang, Junliang Xing |
CVPR | 4 |
| 2017 | The ACRV picking benchmark: A robotic shelf picking benchmark to foster reproducible researchabstractRobotic challenges like the Amazon Picking Challenge (APC) or the DARPA Challenges are an established and important way to drive scientific progress. They make research comparable on a well-defined benchmark with equal test conditions for all participants. However, such challenge events occur only occasionally, are limited to a small number of contestants, and the test conditions are very difficult to replicate after the main event. We present a new physical benchmark challenge for robotic picking: the ACRV Picking Benchmark. Designed to be reproducible, it consists of a set of 42 common objects, a widely available shelf, and exact guidelines for object arrangement using stencils. A well-defined evaluation protocol enables the comparison of complete robotic systems - including perception and manipulation - instead of sub-systems only. Our paper also describes and reports results achieved by an open baseline system based on a Baxter robot. Jürgen Leitner, Adam W. Tow, Niko Sünderhauf, Jake E. Dean, Joseph W. Durham, Matthew Cooper 0005, Markus Eich, Chris Lehnert, Ruben Mangels, Chris McCool, Peter Kujala, Lachlan Nicholson, Trung Pham, James Sergeant, Liao Wu, Fangyi Zhang, Ben Upcroft, Peter I. Corke |
ICRA | 16 |
| 2015 | Asynchronous blind signal decomposition using tiny-length code for Visible Light Communication-based indoor localizationabstractIndoor localization is a fundamental capability for service robots and indoor applications on mobile devices. To realize that, the cost and performance are of great concern. In this paper, we introduce a lightweight signal encoding and decomposition method for a low-cost and low-power Visible Light Communication (VLC)-based indoor localization system. Firstly, a Gold-sequence-based tiny-length code selection method is introduced for light encoding. Then a correlation-based asynchronous blind light-signal decomposition method is developed for the decomposition of the lights mixed with modulated light sources. It is able to decompose the mixed light-signal package in real-time. The average decomposition time-cost for each frame is 20 ms. By using the decomposition results, the localization system achieves accuracy at 0.56 m. These features outperform other existing low-cost indoor localization approaches, such as WiFiSLAM. Fangyi Zhang, Kejie Qiu, Ming Liu 0001 |
ICRA | 1 |
| 2015 | Visible Light Communication-based indoor localization using Gaussian ProcessabstractFor mobile robots and position-based services, such as healthcare service, precise localization is the most fundamental capability while low-cost localization solutions are with increasing need and potentially have a wide market. A low-cost localization solution based on a novel Visible Light Communication (VLC) system for indoor environments is proposed in this paper. A number of modulated LED lights are used as beacons to aid indoor localization additional to illumination. A Gaussian Process(GP) is used to model the intensity distributions of the light sources. A Bayesian localization framework is constructed using the results of the GP, leading to precise localization. Path-planning is hereby feasible by only using the GP variance field, rather than using a metric map. Dijkstra's algorithm-based path-planner is adopted to cope with the practical situations. We demonstrate our localization system by real-time experiments performed on a tablet PC in an indoor environment. Kejie Qiu, Fangyi Zhang, Ming Liu 0001 |
IROS | 2 |