EDBT 2026 Demo / reviewers in the wild / expert
Yiding Yang
dblp:151/9483
· DBLP profile ↗
19ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0001-8290-9805ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSystems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deployment Prior Injection for Run-Time Re-Biasable Object Detection
Yiding Yang, Vishal M. Patel, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | UGG: Unified Generative Grasping
Jiaxin Lu 0001, Hao Kang, Bo Liu 0043, Yiding Yang, Qixing Huang, Gang Hua 0001 |
ECCV (67) | 5 |
| 2023 | Deep Graph ReprogrammingabstractIn this paper, we explore a novel model reusing task tailored for graph neural networks (GNNs), termed as “deep graph reprogramming”. We strive to reprogram a pretrained GNN, without amending raw node features nor model parameters, to handle a bunch of cross-level downstream tasks in various domains. To this end, we propose an innovative Data Reprogramming paradigm alongside a Model Reprogramming paradigm. The former one aims to address the challenge of diversified graph feature dimensions for various tasks on the input side, while the latter alleviates the dilemma of fixed per-task-per-model behavior on the model side. For data reprogramming, we specifically devise an elaborated Meta-FeatPadding method to deal with heterogeneous input dimensions, and also develop a transductive Edge-Slimming as well as an inductive Meta-GraPadding approach for diverse homogenous samples. Meanwhile, for model reprogramming, we propose a novel task-adaptive Reprogrammable-Aggregator, to endow the frozen model with larger expressive capacities in handling cross-domain tasks. Experiments on fourteen datasets across node/graph classification/regression, 3D object recognition, and distributed action recognition, demonstrate that the proposed methods yield gratifying results, on par with those by re-training from scratch. Yongcheng Jing, Chongbin Yuan, Yiding Yang, Xinchao Wang, Dacheng Tao |
CVPR | 4 |
| 2022 | Learning Graph Neural Networks for Image Style Transfer
Yongcheng Jing, Yining Mao, Yiding Yang, Yibing Zhan, Mingli Song, Xinchao Wang, Dacheng Tao |
ECCV (7) | 3 |
| 2021 | Overcoming Catastrophic Forgetting in Graph Neural NetworksabstractCatastrophic forgetting refers to the tendency that a neural network ``forgets'' the previous learned knowledge upon learning new tasks. Prior methods have been focused on overcoming this problem on convolutional neural networks (CNNs), where the input samples like images lie in a grid domain, but have largely overlooked graph neural networks (GNNs) that handle non-grid data. In this paper, we propose a novel scheme dedicated to overcoming catastrophic forgetting problem and hence strengthen continual learning in GNNs. At the heart of our approach is a generic module, termed as topology-aware weight preserving (TWP), applicable to arbitrary form of GNNs in a plug-and-play fashion. Unlike the main stream of CNN-based continual learning methods that rely on solely slowing down the updates of parameters important to the downstream task, TWP explicitly explores the local structures of the input graph, and attempts to stabilize the parameters playing pivotal roles in the topological aggregation. We evaluate TWP on different GNN backbones over several datasets, and demonstrate that it yields performances superior to the state of the art. Code is publicly available at https://github.com/hhliu79/TWP. Yiding Yang, Xinchao Wang |
AAAI | 2 |
| 2021 | Turning Frequency to Resolution: Video Super-Resolution via Event CamerasabstractState-of-the-art video super-resolution (VSR) methods focus on exploiting inter- and intra-frame correlations to estimate high-resolution (HR) video frames from low-resolution (LR) ones. In this paper, we study VSR from an exotic perspective, by explicitly looking into the role of temporal frequency of video frames. Through experiments, we observe that a higher frequency, and hence a smaller pixel displacement between consecutive frames, tends to de-liver favorable super-resolved results. This discovery motivates us to introduce Event Cameras, a novel sensing de-vice that responds instantly to pixel intensity changes and produces up to millions of asynchronous events per second, to facilitate VSR. To this end, we propose an Event-based VSR framework (E-VSR), of which the key component is an asynchronous interpolation (EAI) module that reconstructs a high-frequency (HF) video stream with uniform and tiny pixel displacements between neighboring frames from an event stream. The derived HF video stream is then encoded into a VSR module to recover the desired HR videos. Furthermore, an LR bi-directional interpolation loss and an HR self-supervision loss are also introduced to respectively regulate the EAI and VSR modules. Experiments on both real-world and synthetic datasets demonstrate that the proposed approach yields results superior to the state of the art. Yongcheng Jing, Yiding Yang, Xinchao Wang, Mingli Song, Dacheng Tao |
CVPR | 2 |
| 2021 | Amalgamating Knowledge From Heterogeneous Graph Neural NetworksabstractIn this paper, we study a novel knowledge transfer task in the domain of graph neural networks (GNNs). We strive to train a multi-talented student GNN, without accessing human annotations, that “amalgamates” knowledge from a couple of teacher GNNs with heterogeneous architectures and handling distinct tasks. The student derived in this way is expected to integrate the expertise from both teachers while maintaining a compact architecture. To this end, we propose an innovative approach to train a slimmable GNN that enables learning from teachers with varying feature dimensions. Meanwhile, to explicitly align topological semantics between the student and teachers, we introduce a topological attribution map (TAM) to highlight the structural saliency in a graph, based on which the student imitates the teachers’ ways of aggregating information from neighbors. Experiments on seven datasets across various tasks, including multi-label classification and joint segmentation-classification, demonstrate that the learned student, with a lightweight architecture, achieves gratifying results on par with and sometimes even superior to those of the teachers in their specializations. Our code is publicly available at https://github.com/ycjing/AmalgamateGNN.PyTorch. Yongcheng Jing, Yiding Yang, Xinchao Wang, Mingli Song, Dacheng Tao |
CVPR | 2 |
| 2021 | Scene EssenceabstractWhat scene elements, if any, are indispensable for recognizing a scene? We strive to answer this question through the lens of an exotic learning scheme. Our goal is to identify a collection of such pivotal elements, which we term as Scene Essence, to be those that would alter scene recognition if taken out from the scene. To this end, we devise a novel approach that learns to partition the scene objects into two groups, essential ones and minor ones, under the supervision that if only the essential ones are kept while the minor ones are erased in the input image, a scene recognizer would preserve its original prediction. Specifically, we introduce a learnable graph neural network (GNN) for labelling scene objects, based on which the minor ones are wiped off by an off-the-shelf image inpainter. The features of the inpainted image derived in this way, together with those learned from the GNN with the minor-object nodes pruned, are expected to fool the scene discriminator. Both subjective and objective evaluations on Places365, SUN397, and MIT67 datasets demonstrate that, the learned Scene Essence yields a visually plausible image that convincingly retains the original scene category. Jiayan Qiu, Yiding Yang, Xinchao Wang, Dacheng Tao |
CVPR | 2 |
| 2021 | Learning Dynamics via Graph Neural Networks for Human Pose Estimation and TrackingabstractMulti-person pose estimation and tracking serve as crucial steps for video understanding. Most state-of-the-art approaches rely on first estimating poses in each frame and only then implementing data association and refinement. Despite the promising results achieved, such a strategy is inevitably prone to missed detections especially in heavily-cluttered scenes, since this tracking-by-detection paradigm is, by nature, largely dependent on visual evidences that are absent in the case of occlusion. In this paper, we propose a novel online approach to learning the pose dynamics, which are independent of pose detections in current fame, and hence may serve as a robust estimation even in challenging scenarios including occlusion. Specifically, we derive this prediction of dynamics through a graph neural network (GNN) that explicitly accounts for both spatial-temporal and visual information. It takes as input the historical pose tracklets and directly predicts the corresponding poses in the following frame for each tracklet. The predicted poses will then be aggregated with the detected poses, if any, at the same frame so as to produce the final pose, potentially recovering the occluded joints missed by the estimator. Experiments on PoseTrack 2017 and Pose-Track 2018 datasets demonstrate that the proposed method achieves results superior to the state of the art on both human pose estimation and tracking tasks. Yiding Yang, Zhou Ren, Chunluan Zhou, Xinchao Wang, Gang Hua 0001 |
CVPR | 1 |
| 2021 | Meta-Aggregator: Learning to Aggregate for 1-bit Graph Neural NetworksabstractIn this paper, we study a novel meta aggregation scheme towards binarizing graph neural networks (GNNs). We begin by developing a vanilla 1-bit GNN framework that binarizes both the GNN parameters and the graph features. Despite the lightweight architecture, we observed that this vanilla framework suffered from insufficient discriminative power in distinguishing graph topologies, leading to a dramatic drop in performance. This discovery motivates us to devise meta aggregators to improve the expressive power of vanilla binarized GNNs, of which the aggregation schemes can be adaptively changed in a learnable manner based on the binarized features. Towards this end, we propose two dedicated forms of meta neighborhood aggregators, an exclusive meta aggregator termed as Greedy Gumbel Neighborhood Aggregator (GNA), and a diffused meta aggregator termed as Adaptable Hybrid Neighborhood Aggregator (ANA). GNA learns to exclusively pick one single optimal aggregator from a pool of candidates, while ANA learns a hybrid aggregation behavior to simultaneously retain the benefits of several individual aggregators. Furthermore, the proposed meta aggregators may readily serve as a generic plugin module into existing full-precision GNNs. Experiments across various domains demonstrate that the proposed method yields results superior to the state of the art. Yongcheng Jing, Yiding Yang, Xinchao Wang, Mingli Song, Dacheng Tao |
ICCV | 2 |
| 2020 | VOLDOR: Visual Odometry From Log-Logistic Dense Optical Flow ResidualsabstractWe propose a dense indirect visual odometry method taking as input externally estimated optical flow fields instead of hand-crafted feature correspondences. We define our problem as a probabilistic model and develop a generalized-EM formulation for the joint inference of camera motion, pixel depth, and motion-track confidence. Contrary to traditional methods assuming Gaussian-distributed observation errors, we supervise our inference framework under an (empirically validated) adaptive log-logistic distribution model. Moreover, the log-logistic residual model generalizes well to different state-of-the-art optical flow methods, making our approach modular and agnostic to the choice of optical flow estimators. Our method achieved top-ranking results on both TUM RGB-D and KITTI odometry benchmarks. Our open-sourced implementation is inherently GPU-friendly with only linear computational and storage growth. Zhixiang Min, Yiding Yang, Enrique Dunn |
CVPR | 2 |
| 2020 | Distilling Knowledge From Graph Convolutional NetworksabstractExisting knowledge distillation methods focus on convolutional neural networks (CNNs), where the input samples like images lie in a grid domain, and have largely overlooked graph convolutional networks (GCN) that handle non-grid data. In this paper, we propose to our best knowledge the first dedicated approach to distilling knowledge from a pre-trained GCN model. To enable the knowledge transfer from the teacher GCN to the student, we propose a local structure preserving module that explicitly accounts for the topological semantics of the teacher. In this module, the local structure information from both the teacher and the student are extracted as distributions, and hence minimizing the distance between these distributions enables topology-aware knowledge transfer from the teacher, yielding a compact yet high-performance student model. Moreover, the proposed approach is readily extendable to dynamic graph models, where the input graphs for the teacher and the student may differ. We evaluate the proposed method on two different datasets using GCN models of different architectures, and demonstrate that our method achieves the state-of-the-art knowledge distillation performance for GCN models. Yiding Yang, Jiayan Qiu, Mingli Song, Dacheng Tao, Xinchao Wang |
CVPR | 1 |
| 2020 | Hallucinating Visual Instances in Total Absentia
Jiayan Qiu, Yiding Yang, Xinchao Wang, Dacheng Tao |
ECCV (5) | 2 |
| 2020 | Learning Propagation Rules for Attribution Map Generation
Yiding Yang, Jiayan Qiu, Mingli Song, Dacheng Tao, Xinchao Wang |
ECCV (20) | 1 |
| 2020 | Factorizable Graph Convolutional NetworksabstractGraphs have been widely adopted to denote structural connections between entities. The relations are in many cases heterogeneous, but entangled together and denoted merely as a single edge between a pair of nodes. For example, in a social network graph, users in different latent relationships like friends and colleagues, are usually connected via a bare edge that conceals such intrinsic connections. In this paper, we introduce a novel graph convolutional network (GCN), termed as factorizable graph convolutional network (FactorGCN), that explicitly disentangles such intertwined relations encoded in a graph. FactorGCN takes a simple graph as input, and disentangles it into several factorized graphs, each of which represents a latent and disentangled relation among nodes. The features of the nodes are then aggregated separately in each factorized latent space to produce disentangled features, which further leads to better performances for downstream tasks. We evaluate the proposed FactorGCN both qualitatively and quantitatively on the synthetic and real-world datasets, and demonstrate that it yields truly encouraging results in terms of both disentangling and feature aggregation. Code is publicly available at https://github.com/ihollywhy/FactorGCN.PyTorch. Yiding Yang, Zunlei Feng, Mingli Song, Xinchao Wang |
NeurIPS | 1 |
| 2019 | SPAGAN: Shortest Path Graph Attention NetworkabstractGraph convolutional networks (GCN) have recently demonstrated their potential in analyzing non-grid structure data that can be represented as graphs. The core idea is to encode the local topology of a graph, via convolutions, into the feature of a center node. In this paper, we propose a novel GCN model, which we term as Shortest Path Graph Attention Network (SPAGAN). Unlike conventional GCN models that carry out node-based attentions, on either first-order neighbors or random higher-order ones, the proposed SPAGAN conducts path-based attention that explicitly accounts for the influence of a sequence of nodes yielding the minimum cost, or shortest path, between the center node and its higher-order neighbors. SPAGAN therefore allows for a more informative and intact exploration of the graph structure and further the more effective aggregation of information from distant neighbors, as compared to node-based GCN methods. We test SPAGAN for the downstream classification task on several standard datasets, and achieve performances superior to the state of the art. Yiding Yang, Xinchao Wang, Mingli Song, Junsong Yuan 0001, Dacheng Tao |
IJCAI | 1 |
| 2017 | M-FCN: Effective Fully Convolutional Network-Based Airplane Detection FrameworkabstractAirplane detection is a challenging problem in complex remote sensing imaging. In this letter, an effective airplane detection framework called Markov random field-fully convolutional network (M-FCN) is proposed. The M-FCN uses a cascade strategy that consists of an FCN-based coarse candidate extraction stage, a multi-Markov random field (multi-MRF)-based region proposal (RP) generation stage, and a final classification stage. In the first stage, the FCN model is trained to be sensitive to airplanes, and a coarse candidate map is generated. This model is scale-, direction-, and color-invariant and does not require many training examples. After the first stage, the coarse candidate map is used as the initial labeling field for a multi-MRF algorithm, and RPs are generated according to the multi-MRF output. This RP-generating strategy can yield more accurate locations with fewer RPs. In the last stage, a convolutional neural network-based classifier is used to improve the precision of the entire framework. Experiments show that the M-FCN has high precision, recall, and location accuracy. Yiding Yang, Yin Zhuang, Fukun Bi, Hao Shi 0006, Yizhuang Xie |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2017 | Harbor Water Area Extraction From Pan-Sharpened Remotely Sensed Images Based on the Definition Circle ModelabstractHarbor water area extraction is a key step in nearshore environment pollution surveillance using remote sensing image processing techniques. This letter proposes the definition circle (DC) model of color gradient to describe color fluctuations in harbor water surface areas based on pan-sharpened remote sensing images. The DC model includes two steps: center setting and radius tuning. In the center setting process, labeled training set pixels are selected in the red, green, and blue color space. Then, center setting is completed in the hue, saturation, and intensity color space using the perceptron model. In the radius tuning process, positive and negative sample pixels are used to tune the radius value. After these two steps, the DC model can describe the color gradient of a water surface area and provide accurate harbor water area extraction. A series of experiments shows that the proposed DC model is robust and performs better than other extraction methods based on pan-sharpened remote sensing images. Yin Zhuang, Penglin Wang, Yiding Yang, Hao Shi 0006, He Chen 0004, Fukun Bi |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2014 | Design of driving fatigue detection system based on hybrid measures using wavelet-packets transformabstractWith the rapid development of urbanization and motorization in China, fatigue driving has become an increasingly serious road traffic problem. Driving fatigue affects drivers' alertness, decreasing an individual's ability to operate a vehicle safely and increasing the risk of human error that could lead to fatalities, which have been widely recognized as critical safety issues that cut across all modes in the transportation industry. In this paper, firstly, with a virtual driving system we developed, driving simulation experiments were designed to collect subjects' electroencephalogram (EEG) signals and mental fatigue data. To detect drivers' mental state in real time, wavelet-packets transform (WPT) was selected to extract continuous features; then, the subjective evaluation combined with video monitoring was used to evaluate driver's mental state in experiment accurately. At last, with fatigue feature as the input and fatigue state as the output, driving fatigue detection model can be constructed by classification methods. In this paper, Support Vector Machine (SVM) was used to build driving fatigue detection model to estimate mental fatigue state of EEG signal features, and the binary classification accuracy can be achieved up to 88.6207%. Shaonan Wang, Xihui Wang, Yiding Yang |
ICRA | 5 |