EDBT 2026 Demo / reviewers in the wild / expert
Shikun Feng
dblp:26/7906
· DBLP profile ↗
30ranked-venue papers
7as first author
25since 2021 · last 2026
0000-0002-0191-4854ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 7 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Query-calibrated attention and visibility-aware anisotropic regression for small object detection
Ruxiang Duan, Qiliang Du, Lianfang Tian, Shikun Feng, Shuwei Huo |
Neurocomputing | 5 |
| 2025 | Simplifying Control Mechanism in Text-to-Image Diffusion ModelsabstractControlNet has significantly advanced controllable image generation by integrating dense conditions (such as depth and canny edges) with text-to-image diffusion models. However, ControlNet's integration requires an additional amount nearly equal to half of the base diffusion model's parameters, making it inefficient. To address this, we introduce Simple-ControlNet, an efficient and streamlined network for controllable text-to-image generation. It employs a single-scale projection layer to incorporate condition information into the denoising U-Net. It is supplemented by Low-Rank Adapter (LoRA) parameters to facilitate condition learning. Impressively, Simple-ControlNet requires fewer than 3 million parameters for the control mechanism, substantially less than the 300 million needed by ControlNet. Our extensive experiments confirm that Simple-ControlNet matches and surpasses ControlNet's performance across a broad range of tasks and base diffusion models, showcasing its utility and efficiency. Zhida Feng, Li Chen 0011, Yuenan Sun, Jiaxiang Liu 0004, Shikun Feng |
AAAI | 5 |
| 2025 | UniGEM: A Unified Approach to Generation and Property Prediction for MoleculesabstractMolecular generation and molecular property prediction are both crucial for drug discovery, but they are often developed independently. Inspired by recent studies, which demonstrate that diffusion model, a prominent generative approach, can learn meaningful data representations that enhance predictive tasks, we explore the potential for developing a unified generative model in the molecular domain that effectively addresses both molecular generation and property prediction tasks. However, the integration of these tasks is challenging due to inherent inconsistencies, making simple multi-task learning ineffective. To address this, we propose UniGEM, the first unified model to successfully integrate molecular generation and property prediction, delivering superior performance in both tasks. Our key innovation lies in a novel two-phase generative process, where predictive tasks are activated in the later stages, after the molecular scaffold is formed. We further enhance task balance through innovative training strategies. Rigorous theoretical analysis and comprehensive experiments demonstrate our significant improvements in both tasks. The principles behind UniGEM hold promise for broader applications, including natural language processing and computer vision. Shikun Feng, Yuyan Ni, Zhiming Ma, Wei-Ying Ma, Yanyan Lan |
ICLR | 1 |
| 2025 | Fine-tuned Multimodal Large Language Models are Zero-shot Learners in Image Quality AssessmentabstractImage quality assessment (IQA) has traditionally relied on task-specific models, often limiting their adaptability and generalization to diverse image content. To this end, this paper introduces LV-IQA, a fine-tuned Multimodal Large Language Model (MLLM) via Visual Grounding, demonstrating zero-shot learning in IQA. The key contributions of this paper involve the proposal of a cross-modal chain of thought rooted in hierarchical semantics and quality levels. Additionally, To enhance its adaptability, the paper incorporates a visual grounding prompt generated through segmentation masks and bounding boxes to learn the correspondence between original images and their hierarchical semantics. By observing the original images and their corresponding visual grounding information, LV-IQA is empowered to employ a sophisticated chain of thought, involving the perception and comprehension of local details and area structures, to infer the quality of the images. Through experiments, LV-IQA demonstrates reliable zero-shot capabilities in IQA. Zhida Feng, Shikun Feng |
ICME | 5 |
| 2025 | FIGRDock: Fast Interaction-Guided Regression for Flexible DockingabstractFlexible docking, which predicts the binding conformations of both proteins and small molecules by modeling their structural flexibility, plays a vital role in structure-based drug design. Although recent generative approaches, particularly diffusion-based models, have shown promising results, they require iterative sampling to generate candidate structures and depend on separate scoring functions for pose selection. This leads to an inefficient pipeline that is difficult to scale in real-world drug discovery workflows. To overcome these challenges, we introduce FIGRDock, a fast and accurate flexible docking framework that understands complicated interactions between molecules and proteins with a regression-based approach. FIGRDock leverages initial docking poses from conventional tools to distill interaction-aware distance patterns, which serve as explicit structural conditions to directly guide the prediction of the final protein-ligand complex via a regression model. This one-shot inference paradigm enables rapid and precise pose prediction without reliance on multi-step sampling or external scoring stages. Experimental results show that FIGRDock achieves up to 100× faster inference than diffusion-based docking methods, while consistently surpassing them in accuracy across standard benchmarks. These results suggest that FIGRDock has the potential to offer a scalable and efficient solution for flexible docking, advancing the pace of structure-based drug discovery. Shikun Feng, Bicheng Lin, Yuanhuan Mo, Yuyan Ni, Wenyu Zhu, Wei-Ying Ma, Yanyan Lan |
NeurIPS | 1 |
| 2025 | Straight-Line Diffusion Model for Efficient 3D Molecular GenerationabstractDiffusion-based models have shown great promise in molecular generation but often require a large number of sampling steps to generate valid samples. In this paper, we introduce a novel Straight-Line Diffusion Model (SLDM) to tackle this problem, by formulating the diffusion process to follow a linear trajectory. The proposed process aligns well with the noise sensitivity characteristic of molecular structures and uniformly distributes reconstruction effort across the generative process, thus enhancing learning efficiency and efficacy. Consequently, SLDM achieves state-of-the-art performance on 3D molecule generation benchmarks, delivering a 100-fold improvement in sampling efficiency. Yuyan Ni, Shikun Feng, Haohan Chi, Huan-ang Gao, Wei-Ying Ma, Zhiming Ma, Yanyan Lan |
NeurIPS | 2 |
| 2025 | MME-VirtualWorld: Simplifying Multimodal Assessment via Programmable Synthetic Benchmarking
Shenao Chen, Aoqi Fu, Shikun Feng |
PRCV (12) | 7 |
| 2025 | Efficient Text-Guided 3D-Aware Generation With Score Distillation on 3D DistributionabstractText-to-3D generation enables the creation of 3D content with infinite possibilities. Existing methods typically involve training 3D generative models, which suffer from poor semantic alignment due to the scarcity of paired 3D data, or optimizing a 3D representation with 2D diffusion guidance, resulting in slow inference, low diversity, and Janus problems. In this paper, we introduce InstantDreamer, a model designed for text-guided 3D-aware generation in a single forward pass without requiring paired training datasets, thereby enhancing efficiency. To accomplish this, we extend score distillation to learn a 3D-aware semantics distribution. We distill priors from diffusion models into a 3D-aware generator, amortizing the optimization time required for new prompts and eliminating the necessity of paired training data. We equip the generator with hierarchical semantics conditioning, explicitly allowing the model to perceive the correspondence between the text distribution and the 3D latent space. Our elaborate designs empower our 3D generative model with multi-view semantic consistency and feed-forward 3D generation capabilities, thus eliminating the need for score distillation-based optimization for each prompt. Both quantitative and qualitative results on the mainstream benchmarks demonstrate that our InstantDreamer generates competitive multi-view semantic consistent 3D assets compared with state-of-the-art methods. Our method outperforms previous approaches in terms of CLIP R-Precision (66.31) and FID (28.47) while also exhibiting a significant boost in generation speed. Yiji Cheng, Xiaoke Huang 0001, Jiaxiang Liu 0004, Shikun Feng, Yujiu Yang 0001, Yansong Tang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | A Multi-Node Multi-GPU Distributed GNN Training Framework for Large-Scale Online AdvertisingabstractGraph Neural Networks (GNNs) have become critical in various domains such as online advertising but face scalability challenges due to the growing size of graph data, leading to the needs for advanced distributed GPU computation strategies across multiple nodes. This paper presents PGLBox-Cluster, a robust distributed graph learning framework constructed atop the PaddlePaddle platform, implemented to efficiently process graphs comprising billions of nodes and edges. Through strategic partitioning of the model, node attributes, and graph data and leveraging industrial-grade RPC and NCCL for communication, PGLBox-Cluster facilitates effective distributed computation. The extensive experimental results confirm that PGLBox-Cluster achieves a 1.94x to 2.93x speedup over the single-node configuration, significantly elevating graph neural network scalability and efficiency by handling datasets exceeding 3 billion nodes and 120 billion edges with its novel asynchronous communication and graph partitioning techniques. The repository is released at This Link. Xuewu Jiao, Xinsheng Luo, Jiang Bian 0003, Junchao Yang 0001, Mingqing Hu, Weipeng Lu, Shikun Feng, Danlei Feng, Haoyi Xiong, Shuanglong Li |
CIKM | 9 |
| 2024 | Named Entity Driven Zero-Shot Image ManipulationabstractWe introduced StyleEntity, a zero-shot image manipulation model that utilizes named entities as proxies during its training phase. This strategy enables our model to manipulate images using unseen textual descriptions during inference, all within a single training phase. Additionally, we proposed an inference technique termed Prompt Ensemble Latent Averaging (PELA). PELA averages the manipulation directions derived from various named entities during inference, effectively eliminating the noise directions, thus achieving stable manipulation. In our experiments, StyleEntity exhibited superior performance in a zero-shot setting compared to other methods. The code, model weights, and datasets are available at https://github.com/feng-zhida/StyleEntity. Zhida Feng, Li Chen 0011, Jing Tian 0002, Jiaxiang Liu 0004, Shikun Feng |
CVPR | 5 |
| 2024 | Protein-ligand binding representation learning from fine-grained interactionsabstractThe binding between proteins and ligands plays a crucial role in the realm of drug discovery. Previous deep learning approaches have shown promising results over traditional computationally intensive methods, but resulting in poor generalization due to limited supervised data. In this paper, we propose to learn protein-ligand binding representation in a self-supervised learning manner. Different from existing pre-training approaches which treat proteins and ligands individually, we emphasize to discern the intricate binding patterns from fine-grained interactions. Specifically, this self-supervised learning problem is formulated as a prediction of the conclusive binding complex structure given a pocket and ligand with a Transformer based interaction module, which naturally emulates the binding process. To ensure the representation of rich binding information, we introduce two pre-training tasks, i.e. atomic pairwise distance map prediction and mask ligand reconstruction, which comprehensively model the fine-grained interactions from both structure and feature space. Extensive experiments have demonstrated the superiority of our method across various binding tasks, including protein-ligand affinity prediction, virtual screening and protein-ligand docking. Shikun Feng, Yinjun Jia, Wei-Ying Ma, Yanyan Lan |
ICLR | 1 |
| 2024 | Sliced Denoising: A Physics-Informed Molecular Pre-Training MethodabstractWhile molecular pre-training has shown great potential in enhancing drug discovery, the lack of a solid physical interpretation in current methods raises concerns about whether the learned representation truly captures the underlying explanatory factors in observed data, ultimately resulting in limited generalization and robustness. Although denoising methods offer a physical interpretation, their accuracy is often compromised by ad-hoc noise design, leading to inaccurate learned force fields. To address this limitation, this paper proposes a new method for molecular pre-training, called sliced denoising (SliDe), which is based on the classical mechanical intramolecular potential theory. SliDe utilizes a novel noise strategy that perturbs bond lengths, angles, and torsion angles to achieve better sampling over conformations. Additionally, it introduces a random slicing approach that circumvents the computationally expensive calculation of the Jacobian matrix, which is otherwise essential for estimating the force field. By aligning with physical principles, SliDe shows a 42\% improvement in the accuracy of estimated force fields compared to current state-of-the-art denoising methods, and thus outperforms traditional baselines on various molecular property prediction tasks. Yuyan Ni, Shikun Feng, Wei-Ying Ma, Zhiming Ma, Yanyan Lan |
ICLR | 2 |
| 2024 | Multimodal Molecular Pretraining via Modality BlendingabstractSelf-supervised learning has recently gained growing interest in molecular modeling for scientific tasks such as AI-assisted drug discovery. Current studies consider leveraging both 2D and 3D molecular structures for representation learning. However, relying on straightforward alignment strategies that treat each modality separately, these methods fail to exploit the intrinsic correlation between 2D and 3D representations that reflect the underlying structural characteristics of molecules, and only perform coarse-grained molecule-level alignment. To derive fine-grained alignment and promote structural molecule understanding, we introduce an atomic-relation level "blend-then-predict" self-supervised learning approach, MoleBLEND, which first blends atom relations represented by different modalities into one unified relation matrix for joint encoding, then recovers modality-specific information for 2D and 3D structures individually. By treating atom relationships as anchors, MoleBLEND organically aligns and integrates visually dissimilar 2D and 3D modalities of the same molecule at fine-grained atomic level, painting a more comprehensive depiction of each molecule. Extensive experiments show that MoleBLEND achieves state-of-the-art performance across major 2D/3D molecular benchmarks. We further provide theoretical insights from the perspective of mutual-information maximization, demonstrating that our method unifies contrastive, generative (cross-modality prediction) and mask-then-predict (single-modality prediction) objectives into one single cohesive framework. Qiying Yu, Yudi Zhang 0008, Yuyan Ni, Shikun Feng, Yanyan Lan |
ICLR | 4 |
| 2024 | UniCorn: A Unified Contrastive Learning Approach for Multi-view Molecular Representation LearningabstractRecently, a noticeable trend has emerged in developing pre-trained foundation models in the domains of CV and NLP. However, for molecular pre-training, there lacks a universal model capable of effectively applying to various categories of molecular tasks, since existing prevalent pre-training methods exhibit effectiveness for specific types of downstream tasks. Furthermore, the lack of profound understanding of existing pre-training methods, including 2D graph masking, 2D-3D contrastive learning, and 3D denoising, hampers the advancement of molecular foundation models. In this work, we provide a unified comprehension of existing pre-training methods through the lens of contrastive learning. Thus their distinctions lie in clustering different views of molecules, which is shown beneficial to specific downstream tasks. To achieve a complete and general-purpose molecular representation, we propose a novel pre-training framework, named UniCorn, that inherits the merits of the three methods, depicting molecular views in three different levels. SOTA performance across quantum, physicochemical, and biological tasks, along with comprehensive ablation study, validate the universality and effectiveness of UniCorn. Shikun Feng, Yuyan Ni, Yanwen Huang, Zhiming Ma, Wei-Ying Ma, Yanyan Lan |
ICML | 1 |
| 2024 | Spectral Heterogeneous Graph Convolutions via Positive Noncommutative PolynomialsabstractHeterogeneous Graph Neural Networks (HGNNs) have gained significant popularity in various heterogeneous graph learning tasks. However, most existing HGNNs rely on spatial domain-based methods to aggregate information, i.e., manually selected meta-paths or some heuristic modules, lacking theoretical guarantees. Furthermore, these methods cannot learn arbitrary valid heterogeneous graph filters within the spectral domain, which have limited expressiveness. To tackle these issues, we present a positive spectral heterogeneous graph convolution via positive noncommutative polynomials. Then, using this convolution, we propose PSHGCN, a novel Positive Spectral Heterogeneous Graph Convolutional Network. PSHGCN offers a simple yet effective method for learning valid heterogeneous graph filters. Moreover, we demonstrate the rationale of PSHGCN in the graph optimization framework. We conducted an extensive experimental study to show that PSHGCN can learn diverse heterogeneous graph filters and outperform all baselines on open benchmarks. Notably, PSHGCN exhibits remarkable scalability, efficiently handling large real-world graphs comprising millions of nodes and edges. Our codes are available at https://github.com/ivam-he/PSHGCN. Mingguo He, Zhewei Wei, Shikun Feng, Zhengjie Huang, Weibin Li 0004, Yu Sun 0029, Dianhai Yu |
WWW | 3 |
| 2023 | ERNIE-ViLG 2.0: Improving Text-to-Image Diffusion Model with Knowledge-Enhanced Mixture-of-Denoising-ExpertsabstractRecent progress in diffusion models has revolutionized the popular technology of text-to-image generation. While existing approaches could produce photorealistic high-resolution images with text conditions, there are still several open problems to be solved, which limits the further improvement of image fidelity and text relevancy. In this paper, we propose ERNIE-ViLG 2.0, a large-scale Chinese text-to-image diffusion model, to progressively upgrade the quality of generated images by: (1) incorporating fine-grained textual and visual knowledge of key elements in the scene, and (2) utilizing different denoising experts at different denoising stages. With the proposed mechanisms, ERNIE-ViLG 2.01not only achieves a new state-of-the-art on MS-COCO with zero-shot FID-30k score of 6.75, but also significantly outperforms recent models in terms of image fidelity and image-text alignment, with side-by-side human evaluation on the bilingual prompt set ViLG-300. Zhida Feng, Zhenyu Zhang 0006, Yewei Fang, Lanxin Li, Xuyi Chen, Jiaxiang Liu 0004, Weichong Yin, Shikun Feng, Yu Sun 0004, Li Chen 0011, Hao Tian 0005, Hua Wu 0003, Haifeng Wang 0001 |
CVPR | 10 |
| 2023 | ICDAR 2023 Competition on Structured Text Extraction from Visually-Rich Document Images
Wenwen Yu, Chengquan Zhang, Haoyu Cao 0001, Wei Hua 0005, Bohan Li 0010, Mingrui Chen 0001, Jianfeng Kuang, Mengjun Cheng, Yuning Du, Shikun Feng, Xiaoguang Hu, Pengyuan Lv, Yuechen Yu, Wanxiang Che, Errui Ding, Cheng-Lin Liu 0001, Jiebo Luo 0001, Shuicheng Yan, Min Zhang 0005, Dimosthenis Karatzas, Xing Sun 0001, Jingdong Wang 0001, Xiang Bai |
ICDAR (2) | 12 |
| 2023 | Fractional Denoising for 3D Molecular Pre-trainingabstractCoordinate denoising is a promising 3D molecular pre-training method, which has achieved remarkable performance in various downstream drug discovery tasks. Theoretically, the objective is equivalent to learning the force field, which is revealed helpful for downstream tasks. Nevertheless, there are two challenges for coordinate denoising to learn an effective force field, i.e. low coverage samples and isotropic force field. The underlying reason is that molecular distributions assumed by existing denoising methods fail to capture the anisotropic characteristic of molecules. To tackle these challenges, we propose a novel hybrid noise strategy, including noises on both dihedral angel and coordinate. However, denoising such hybrid noise in a traditional way is no more equivalent to learning the force field. Through theoretical deductions, we find that the problem is caused by the dependency of the input conformation for covariance. To this end, we propose to decouple the two types of noise and design a novel fractional denoising method (Frad), which only denoises the latter coordinate part. In this way, Frad enjoys both the merits of sampling more low-energy structures and the force field equivalence. Extensive experiments show the effectiveness of Frad in molecule representation, with a new state-of-the-art on 9 out of 12 tasks of QM9 and on 7 out of 8 targets of MD17. Shikun Feng, Yuyan Ni, Yanyan Lan, Zhiming Ma, Wei-Ying Ma |
ICML | 1 |
| 2023 | PGLBox: Multi-GPU Graph Learning Framework for Web-Scale RecommendationabstractWhile having been used widely for large-scale recommendation and online advertising, the Graph Neural Network (GNN) has demonstrated its representation learning capacity to extract embeddings of nodes and edges through passing, transforming, and aggregating information over the graph. In this work, we propose PGLBox1 - a multi-GPU graph learning framework based on PaddlePaddle [24], incorporating with optimized storage, computation, and communication strategies, to train deep GNNs based on web-scale graphs for the recommendation. Specifically, PGLBox adopts a hierarchical storage system with three layers to facilitate I/O, where graphs and embeddings are stored in the HBMs and SSDs, respectively, with MEMs as the cache. To fully utilize multi-GPUs and I/O bandwidth, PGLBox proposes an asynchronous pipeline with three stages - it first samples the subgraphs from the input graph, then pulls & updates embeddings and trains GNNs on the subgraph with parameters updating queued at the end of the pipeline. Thanks to the capacity of PGLBox in handling web-scale graphs, it becomes feasible to unify the view of GNN-based recommendation tasks for multiple advertising verticals and fuse all these graphs into a unified yet huge one. We evaluate PGLBox using a bucket of realistic GNN training tasks for the recommendation, and compare the performance of PGLBox on top of a multi-GPU server (Tesla A100×8) and the legacy training system based on a 40-node MPI cluster at Baidu. The overall comparisons show that PGLBox could save up to 55% monetary cost for training GNN models, and achieve up to 14× training speedup with the same accuracy as the legacy trainer. The open-source implementation of PGLBox is available at https://github.com/PaddlePaddle/PGL/tree/main/apps/PGLBox. Xuewu Jiao, Weibin Li 0004, Xinxuan Wu, Jiang Bian 0003, Siming Dai, Xinsheng Luo, Mingqing Hu, Zhengjie Huang, Danlei Feng, Junchao Yang 0001, Shikun Feng, Haoyi Xiong, Dianhai Yu, Shuanglong Li, Jingzhou He, Yanjun Ma |
KDD | 13 |
| 2023 | Label Information Enhanced Fraud Detection against Low Homophily in GraphsabstractNode classification is a substantial problem in graph-based fraud detection. Many existing works adopt Graph Neural Networks (GNNs) to enhance fraud detectors. While promising, currently most GNN-based fraud detectors fail to generalize to the low homophily setting. Besides, label utilization has been proved to be significant factor for node classification problem. But we find they are less effective in fraud detection tasks due to the low homophily in graphs. In this work, we propose GAGA, a novel Group AGgregation enhanced TrAnsformer, to tackle the above challenges. Specifically, the group aggregation provides a portable method to cope with the low homophily issue. Such an aggregation explicitly integrates the label information to generate distinguishable neighborhood information. Along with group aggregation, an attempt towards end-to-end trainable group encoding is proposed which augments the original feature space with the class labels. Meanwhile, we devise two additional learnable encodings to recognize the structural and relational context. Then, we combine the group aggregation and the learnable encodings into a Transformer encoder to capture the semantic information. Experimental results clearly show that GAGA outperforms other competitive graph-based fraud detectors by up to 24.39% on two trending public datasets and a real-world industrial dataset from Baidu. Even more, the group aggregation is demonstrated to outperform other label utilization methods (e.g., C&S, BoT/UniMP) in the low homophily setting. Jinghui Zhang 0001, Zhengjie Huang, Weibin Li 0004, Shikun Feng, Ziheng Ma, Yu Sun 0029, Dianhai Yu, Fang Dong 0001, Jiahui Jin 0001, Beilun Wang, Junzhou Luo |
WWW | 5 |
| 2022 | DuETA: Traffic Congestion Propagation Pattern Modeling via Efficient Graph Learning for ETA Prediction at Baidu MapsabstractEstimated time of arrival (ETA) prediction, also known as travel time estimation, is a fundamental task for a wide range of intelligent transportation applications, such as navigation, route planning, and ride-hailing services. To accurately predict the travel time of a route, it is essential to take into account both contextual and predictive factors, such as spatial-temporal interaction, driving behavior, and traffic congestion propagation inference. The ETA prediction models previously deployed at Baidu Maps have addressed the factors of spatial-temporal interaction (ConSTGAT) and driving behavior (SSML). In this work, we believe that modeling traffic congestion propagation patterns is of great importance toward accurately performing ETA prediction, and we focus on this factor to improve ETA performance. Traffic congestion propagation pattern modeling is challenging, and it requires accounting for impact regions over time and cumulative effect of delay variations over time caused by traffic events on the road network. In this paper, we present a practical industrial-grade ETA prediction framework named DuETA. Specifically, we construct a congestion-sensitive graph based on the correlations of traffic patterns, and we develop a route-aware graph transformer to directly learn the long-distance correlations of the road segments. This design enables DuETA to capture the interactions between the road segment pairs that are spatially distant but highly correlated with traffic conditions. Extensive experiments are conducted on large-scale, real-world datasets collected from Baidu Maps. Experimental results show that ETA prediction can significantly benefit from the learned traffic congestion propagation patterns, which demonstrates the effectiveness and practical applicability of DuETA. In addition, DuETA has already been deployed in production at Baidu Maps, serving billions of requests every day. This demonstrates that DuETA is an industrial-grade and robust solution for large-scale ETA prediction services. Jizhou Huang, Zhengjie Huang, Xiaomin Fang, Shikun Feng, Xuyi Chen, Jiaxiang Liu 0004, Haitao Yuan 0002, Haifeng Wang 0001 |
CIKM | 4 |
| 2022 | Simple and Effective Relation-based Embedding Propagation for Knowledge Representation LearningabstractRelational graph neural networks have garnered particular attention to encode graph context in knowledge graphs (KGs). Although they achieved competitive performance on small KGs, how to efficiently and effectively utilize graph context for large KGs remains an open problem. To this end, we propose the Relation-based Embedding Propagation (REP) method. It is a post-processing technique to adapt pre-trained KG embeddings with graph context. As relations in KGs are directional, we model the incoming head context and the outgoing tail context separately. Accordingly, we design relational context functions with no external parameters. Besides, we use averaging to aggregate context information, making REP more computation-efficient. We theoretically prove that such designs can avoid information distortion during propagation. Extensive experiments also demonstrate that REP has significant scalability while improving or maintaining prediction quality. Particularly, it averagely brings about 10% relative improvement to triplet-based embedding methods on OGBL-WikiKG2 and takes 5%-83% time to achieve comparable results as the state-of-the-art GC-OTE. Siming Dai, Weiyue Su, Zeyang Fang, Zhengjie Huang, Shikun Feng, Yu Sun 0029, Dianhai Yu |
IJCAI | 7 |
| 2022 | ERNIE-GeoL: A Geography-and-Language Pre-trained Model and its Applications in Baidu MapsabstractPre-trained models (PTMs) have become a fundamental backbone for downstream tasks in natural language processing and computer vision. Despite initial gains that were obtained by applying generic PTMs to geo-related tasks at Baidu Maps, a clear performance plateau over time was observed. One of the main reasons for this plateau is the lack of readily available geographic knowledge in generic PTMs. To address this problem, in this paper, we present ERNIE-GeoL, which is a geography-and-language pre-trained model designed and developed for improving the geo-related tasks at Baidu Maps. ERNIE-GeoL is elaborately designed to learn a universal representation of geography-language by pre-training on large-scale data generated from a heterogeneous graph that contains abundant geographic knowledge. Extensive quantitative and qualitative experiments conducted on large-scale real-world datasets demonstrate the superiority and effectiveness of ERNIE-GeoL. ERNIE-GeoL has already been deployed in production at Baidu Maps since April 2021, which significantly benefits the performance of various downstream tasks. This demonstrates that ERNIE-GeoL can serve as a fundamental backbone for a wide range of geo-related tasks. Jizhou Huang, Haifeng Wang 0001, Yunsheng Shi, Zhengjie Huang, An Zhuo, Shikun Feng |
KDD | 7 |
| 2022 | mmLayout: Multi-grained MultiModal Transformer for Document UnderstandingabstractRecent efforts of multimodal Transformers have improved Visually Rich Document Understanding (VrDU) tasks via incorporating visual and textual information. However, existing approaches mainly focus on fine-grained elements such as words and document image patches, making it hard for them to learn from coarse-grained elements, including natural lexical units like phrases and salient visual regions like prominent image regions. In this paper, we attach more importance to coarse-grained elements containing high-density information and consistent semantics, which are valuable for document understanding. At first, a document graph is proposed to model complex relationships among multi-grained multimodal elements, in which salient visual regions are detected by a cluster-based method. Then, a multi-grained multimodal Transformer called mmLayout is proposed to incorporate coarse-grained information into existing pre-trained fine-grained multimodal Transformers based on the graph. In mmLayout, coarse-grained information is aggregated from fine-grained, and then, after further processing, is fused back into fine-grained for final prediction. Furthermore, common sense enhancement is introduced to exploit the semantic information of natural lexical units. Experimental results on four tasks, including information extraction and document question answering, show that our method can improve the performance of multimodal Transformers based on fine-grained elements and achieve better performance with fewer parameters. Qualitative analyses show that our method can capture consistent semantics in coarse-grained elements. Wenjin Wang 0003, Zhengjie Huang, Qianglong Chen, Qiming Peng, Yinxu Pan, Weichong Yin, Shikun Feng, Yu Sun 0029, Dianhai Yu, Yin Zhang 0006 |
ACM Multimedia | 8 |
| 2021 | Masked Label Prediction: Unified Message Passing Model for Semi-Supervised ClassificationabstractGraph neural network (GNN) and label propagation algorithm (LPA) are both message passing algorithms, which have achieved superior performance in semi-supervised classification. GNN performs feature propagation by a neural network to make predictions, while LPA uses label propagation across graph adjacency matrix to get results. However, there is still no effective way to directly combine these two kinds of algorithms. To address this issue, we propose a novel Unified Message Passaging Model (UniMP) that can incorporate feature and label propagation at both training and inference time. First, UniMP adopts a Graph Transformer network, taking feature embedding and label embedding as input information for propagation. Second, to train the network without overfitting in self-loop input label information, UniMP introduces a masked label prediction strategy, in which some percentage of input label information are masked at random, and then predicted. UniMP conceptually unifies feature propagation and label propagation and is empirically powerful. It obtains new state-of-the-art semi-supervised classification results in Open Graph Benchmark (OGB). Yunsheng Shi, Zhengjie Huang, Shikun Feng, Yu Sun 0029 |
IJCAI | 3 |
| 2020 | ERNIE 2.0: A Continual Pre-Training Framework for Language UnderstandingabstractRecently pre-trained models have achieved state-of-the-art results in various language understanding tasks. Current pre-training procedures usually focus on training the model with several simple tasks to grasp the co-occurrence of words or sentences. However, besides co-occurring information, there exists other valuable lexical, syntactic and semantic information in training corpora, such as named entities, semantic closeness and discourse relations. In order to extract the lexical, syntactic and semantic information from training corpora, we propose a continual pre-training framework named ERNIE 2.0 which incrementally builds pre-training tasks and then learn pre-trained models on these constructed tasks via continual multi-task learning. Based on this framework, we construct several tasks and train the ERNIE 2.0 model to capture lexical, syntactic and semantic aspects of information in the training data. Experimental results demonstrate that ERNIE 2.0 model outperforms BERT and XLNet on 16 tasks including English tasks on GLUE benchmarks and several similar tasks in Chinese. The source codes and pre-trained models have been released at https://github.com/PaddlePaddle/ERNIE. Yu Sun 0004, Shuohuan Wang, Yu-Kun Li, Shikun Feng, Hao Tian 0005, Hua Wu 0003, Haifeng Wang 0001 |
AAAI | 4 |
| 2015 | High-Performance Video Condensation SystemabstractVideo synopsis or condensation is a smart solution for fast video browsing and storage. However, most of the existing methods work offline, where two main phases are required. The first phase is to prepare tubes and background images. The second phase is to rearrange tubes and stitch them into backgrounds. However, with a long video sequence, the first phase is memory consuming for data storage, and the second phase is computationally expensive to rearrange all tubes simultaneously. To overcome these problems, we propose a high-performance video condensation system based on an online content-aware framework. The online framework transforms the optimization problem of tube rearrangement into a stepwise optimization problem. Therefore, it can condense video with much less memory and higher speed than the offline framework. With the aid of this transformation, the proposed system can process input videos and produce condensed videos simultaneously. Thus it is suitable for real-time endless surveillance videos. Meanwhile, the online mechanism allows users to directly visit the condensation video that has been generated. Moreover, the content-aware mechanism makes the proposed system able to automatically determine the duration of a condensed video. Finally, the proposed system uses Graphic Processing Unit (GPU) and multicore techniques to improve the speed. Extensive experiments that validate the high efficiency of the system are presented. Jianqing Zhu, Shikun Feng, Dong Yi, Shengcai Liao, Zhen Lei 0001, Stan Z. Li |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2012 | Water Filling: Unsupervised People Counting via Vertical Kinect SensorabstractPeople counting is one of the key components in video surveillance applications, however, due to occlusion, illumination, color and texture variation, the problem is far from being solved. Different from traditional visible camera based systems, we construct a novel system that uses vertical Kinect sensor for people counting, where the depth information is used to remove the affect of the appearance variation. Since the head is always closer to the Kinect sensor than other parts of the body, people counting task equals to find the suitable local minimum regions. According to the particularity of the depth map, we propose a novel unsupervised water filling method that can find these regions with the property of robustness, locality and scale-invariance. Experimental comparisons with mean shift and random forest on two databases validate the superiority of our water filling algorithm in people counting. Xucong Zhang, Shikun Feng, Zhen Lei 0001, Dong Yi, Stan Z. Li |
AVSS | 3 |
| 2012 | Online content-aware video condensationabstractExplosive growth of surveillance video data presents formidable challenges to its browsing, retrieval and storage. Video synopsis, an innovation proposed by Peleg and his colleagues, is aimed for fast browsing by shortening the video into a synopsis while keeping activities in video captured by a camera. However, the current techniques are offline methods requiring that all the video data be ready for the processing, and are expensive in time and space. In this paper, we propose an online and efficient solution, and its supporting algorithms to overcome the problems. The method adopts an online content-aware approach in a step-wise manner, hence applicable to endless video, with less computational cost. Moreover, we propose a novel tracking method, called sticky tracking, to achieve high-quality visualization. The system can achieve a faster-than-real-time speed with a multi-core CPU implementation. The advantages are demonstrated by extensive experiments with a wide variety of videos. The proposed solution and algorithms could be integrated with surveillance cameras, and impact the way that surveillance videos are recorded. Shikun Feng, Zhen Lei 0001, Dong Yi, Stan Z. Li |
CVPR | 1 |
| 2010 | Online Principal Background Selection for Video SynopsisabstractVideo synopsis provides a means for fast browsing of activities in video. Principal background selection (PBS) is an important step in video synopsis. Existing methods make PBS in an offline way and at a high memory cost. In this paper we propose a novel background selection method, ``online principal background selection'' (OPBS). The OPBS selects n principal backgrounds from N backgrounds in an online fashion with a low memory cost, making it possible to build an efficient online video synopsis system. Another advantage is that, with OPBS, the selected backgrounds are related to not only background changes over time but also video activities. Experimental results demonstrate the advantages of the proposed OPBS. Shikun Feng, Shengcai Liao, Stan Z. Li |
ICPR | 1 |