VLDB 2026 Research / reviewers in the wild / expert
Weihong Lin
dblp:209/1942
· DBLP profile ↗
20ranked-venue papers
4as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 2 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DIN: Dual Impulse Network for Multi-view Representation LearningabstractMulti-view representation learning, which utilizes multiple channels to improve perceptual accuracy, is recognized for its effectiveness in the analysis of multi-view data. However, deploying these methods in real-world scenarios presents two primary challenges. 1) Lack of Variegation: Multi-view representation techniques commonly observe along a singular axis, i.e., the attribute axis; 2) Insufficient Relationship: Most multi-view models lack mechanisms for exploring potential relationships between attribute axis and channel axis. To mitigate these obstacles, we design a Dual Impulse Network framework for multi-view representation learning (DIN) to train a feature representation. In this framework, a strategy observed along the channel axis and attribute axis simultaneously is introduced, and two different representations are generated by two analogous impulse networks, which are capable of extracting information corresponding to different axes. Furthermore, we incorporate an integration network that analyzes the potential relationship between attribute axis and channel axis to generate two attention matrices. The final two feature representations derived from these attention matrices are aggregated to amplify the expression of internal information. Comprehensive experimental results support the efficacy and superiority of the proposed framework, demonstrating improvements in classification performance compared to state-of-the-art methods. Yilin Wu 0001, Weihong Lin, Renjie Lin, Zihan Fang 0002, Shide Du, Shiping Wang |
AAAI | 2 |
| 2026 | Advancing multi-omics analysis via dynamic labeling with shared-specific information
Jiecheng Wu, Zhaoliang Chen, Yali Pu, Weihong Lin, Yuanfei Dai, Genggeng Liu, Shiping Wang |
Inf. Sci. | 4 |
| 2026 | Beyond symmetric propagation: Modeling asymmetric node influences in multi-view learning
Hongyang Dong, Hongzhi He, Yilin Wu 0001, Weihong Lin, Yiqing Shi, Shiping Wang |
Knowl. Based Syst. | 4 |
| 2026 | Heterogeneous graph neural network with multi-relational structure-aware hybrid filtering
Linmin Huang, Weihong Lin, Shiping Wang |
Knowl. Based Syst. | 2 |
| 2026 | Beyond local aggregation: Global graph contrastive learning for multi-view fusion
Xueyang Min, Zihan Fang 0002, Weihong Lin, Shiping Wang |
Neural Networks | 4 |
| 2026 | Diffusion-Guided graph generation for multi-view semi-Supervised classification
Yilin Wu 0001, Weihong Lin, Hongyang Dong, Shiping Wang |
Neural Networks | 2 |
| 2025 | HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language ModelsabstractRecent Multi-modal Large Language Models (MLLMs) have made great progress in video understanding. However, their performance on videos involving human actions is still limited by the lack of high-quality data. To address this, we introduce a two-stage data annotation pipeline. First, we design strategies to accumulate videos featuring clear human actions from the Internet. Second, videos are annotated in a standardized caption format that uses human attributes to distinguish individuals and chronologically details their actions and interactions. Through this pipeline, we curate two datasets, namely HAICTrain and HAICBench. HAICTrain comprises 126K video-caption pairs generated by Gemini-Pro and verified for training purposes. Meanwhile, HAICBench includes 412 manually annotated video-caption pairs and 2,000 QA pairs, for a comprehensive evaluation of human action understanding. Experimental results demonstrate that training with HAICTrain not only significantly enhances human understanding abilities across 4 benchmarks, but can also improve text-to-video generation results. Both the HAICTrain and HAICBench will be made open-source to facilitate further research. Xiao Wang 0056, Jingyun Hua, Weihong Lin, Yuanxing Zhang, Jianlong Wu, Di Zhang 0026, Liqiang Nie |
ACL (1) | 3 |
| 2025 | Mavors: Multi-granularity Video Representation for Multimodal Large Language ModelabstractLong-context video understanding in Multimodal Large Language Models (MLLMs) faces a critical challenge: balancing computational efficiency with the retention of fine-grained spatio-temporal patterns. Existing approaches (e.g., sparse sampling, dense sampling with low resolution, and token compression) suffer from significant information loss in temporal dynamics, spatial details, or subtle interactions, particularly in videos with complex motion or varying resolutions. To address this, we propose Mavors, a novel framework that introduces Multi-granularity video representation for holistic long-video modeling. Specifically, Mavors directly encodes raw video content into latent representations through two core components: 1) an Intra-chunk Vision Encoder (IVE) that preserves high-resolution spatial features via 3D convolutions and Vision Transformers, and 2) an Inter-chunk Feature Aggregator (IFA) that establishes temporal coherence across chunks using transformer-based dependency modeling with chunk-level rotary position encodings. Moreover, the framework unifies image and video understanding by treating images as single-frame videos via sub-image decomposition. Experiments across diverse benchmarks demonstrate Mavors' superiority in maintaining both spatial fidelity and temporal continuity, significantly outperforming existing methods in tasks requiring fine-grained spatio-temporal reasoning. Yang Shi 0009, Yushuo Guan, Yuanxing Zhang, Weihong Lin, Jingyun Hua, Xinlong Chen, Bohan Zeng, Wentao Zhang 0001, Wenjing Yang 0002, Di Zhang 0026 |
ACM Multimedia | 7 |
| 2025 | Heterogeneous Graph Neural Network with Adaptive Relation Reconstruction
Weihong Lin, Zhaoliang Chen, Shiping Wang |
Neural Networks | 1 |
| 2025 | Exploring unified cross-view hypergraph generation for multi-view semi-supervised classification
Zhibin Shi, Zhenghong Lin, Weihong Lin, Shiping Wang |
Neural Networks | 3 |
| 2024 | UniVIE: A Unified Label Space Approach to Visual Information Extraction from Form-Like Documents
Jiawei Wang 0026, Weihong Lin, Zhuoyao Zhong, Lei Sun 0003, Qiang Huo |
ICDAR (6) | 3 |
| 2023 | A Question-Answering Approach to Key Value Pair Extraction from Form-Like Document ImagesabstractIn this paper, we present a new question-answering (QA) based key-value pair extraction approach, called KVPFormer, to robustly extracting key-value relationships between entities from form-like document images. Specifically, KVPFormer first identifies key entities from all entities in an image with a Transformer encoder, then takes these key entities as questions and feeds them into a Transformer decoder to predict their corresponding answers (i.e., value entities) in parallel. To achieve higher answer prediction accuracy, we propose a coarse-to-fine answer prediction approach further, which first extracts multiple answer candidates for each identified question in the coarse stage and then selects the most likely one among these candidates in the fine stage. In this way, the learning difficulty of answer prediction can be effectively reduced so that the prediction accuracy can be improved. Moreover, we introduce a spatial compatibility attention bias into the self-attention/cross-attention mechanism for KVPFormer to better model the spatial interactions between entities. With these new techniques, our proposed KVPFormer achieves state-of-the-art results on FUNSD and XFUND datasets, outperforming the previous best-performing method by 7.2% and 13.2% in F1 score, respectively. Zhuoyuan Wu, Zhuoyao Zhong, Weihong Lin, Lei Sun 0003, Qiang Huo |
AAAI | 4 |
| 2023 | DETRs with Hybrid MatchingabstractOne-to-one set matching is a key design for DETR to establish its end-to-end capability, so that object detection does not require a hand-crafted NMS (non-maximum suppression) to remove duplicate detections. This end-to-end signature is important for the versatility of DETR, and it has been generalized to broader vision tasks. However, we note that there are few queries assigned as positive samples and the one-to-one set matching significantly reduces the training efficacy of positive samples. We propose a simple yet effective method based on a hybrid matching scheme that combines the original one-to-one matching branch with an auxiliary one-to-many matching branch during training. Our hybrid strategy has been shown to significantly improve accuracy. In inference, only the original one-to-one match branch is used, thus maintaining the end-to-end merit and the same inference efficiency of DETR. The method is namedℋ-DETR, and it shows that a wide range of representative DETR methods can be consistently improved across a wide range of visual tasks, including Deformable-DETR, PETRv2, PETR, and TransTrack, among others. Code is available at: https://github.com/HDETR. Ding Jia, Yuhui Yuan, Haodi He, Xiaopei Wu, Haojun Yu, Weihong Lin, Lei Sun 0003, Chao Zhang 0001, Han Hu 0001 |
CVPR | 6 |
| 2023 | Robust Table Detection and Structure Recognition from Heterogeneous Document ImagesabstractWe introduce a new table detection and structure recognition approach named RobusTabNet to detect the boundaries of tables and reconstruct the cellular structure of each table from heterogeneous document images. For table detection, we propose to use CornerNet as a new region proposal network to generate higher quality table proposals for Faster R-CNN, which has significantly improved the localization accuracy of Faster R-CNN for table detection. Consequently, our table detection approach achieves state-of-the-art performance on three public table detection benchmarks, namely cTDaR TrackA, PubLayNet and IIIT-AR-13K, by only using a lightweight ResNet-18 backbone network. Furthermore, we propose a new split-and-merge based table structure recognition approach, in which a novel spatial CNN based separation line prediction module is proposed to split each detected table into a grid of cells, and a Grid CNN based cell merging module is applied to recover the spanning cells. As the spatial CNN module can effectively propagate contextual information across the whole table image, our table structure recognizer can robustly recognize tables with large blank spaces and geometrically distorted (even curved) tables. Thanks to these two techniques, our table structure recognition approach achieves state-of-the-art performance on three public benchmarks, including SciTSR, PubTabNet and cTDaR TrackB2-Modern. Moreover, we have further demonstrated the advantages of our approach in recognizing tables with complex structures, large blank spaces, as well as geometrically distorted or even curved shapes on a more challenging in-house dataset. Chixiang Ma, Weihong Lin, Lei Sun 0003, Qiang Huo |
Pattern Recognit. | 2 |
| 2023 | Robust table structure recognition with dynamic queries enhanced detection transformer
Jiawei Wang 0026, Weihong Lin, Chixiang Ma, Lei Sun 0003, Qiang Huo |
Pattern Recognit. | 2 |
| 2022 | TSRFormer: Table Structure Recognition with TransformersabstractWe present a new table structure recognition (TSR) approach, called TSRFormer, to robustly recognizing the structures of complex tables with geometrical distortions from various table images. Unlike previous methods, we formulate table separation line prediction as a line regression problem instead of an image segmentation problem and propose a new two-stage DETR based separator prediction approach, dubbed Sep arator RE gression TR ansformer (SepRETR), to predict separation lines from table images directly. To make the two-stage DETR framework work efficiently and effectively for the separation line prediction task, we propose two improvements: 1) A prior-enhanced matching strategy to solve the slow convergence issue of DETR; 2) A new cross attention module to sample features from a high-resolution convolutional feature map directly so that high localization accuracy is achieved with low computational cost. After separation line prediction, a simple relation network based cell merging module is used to recover spanning cells. With these new techniques, our TSRFormer achieves state-of-the-art performance on several benchmark datasets, including SciTSR, PubTabNet and WTW. Furthermore, we have validated the robustness of our approach to tables with complex structures, borderless cells, large blank spaces, empty or spanning cells as well as distorted or even curved shapes on a more challenging real-world in-house dataset. Weihong Lin, Chixiang Ma, Jiawei Wang 0026, Lei Sun 0003, Qiang Huo |
ACM Multimedia | 1 |
| 2022 | Expediting Large-Scale Vision Transformer for Dense Prediction without Fine-tuningabstractVision transformers have recently achieved competitive results across various vision tasks but still suffer from heavy computation costs when processing a large number of tokens. Many advanced approaches have been developed to reduce the total number of tokens in the large-scale vision transformers, especially for image classification tasks. Typically, they select a small group of essential tokens according to their relevance with the [\texttt{class}] token, then fine-tune the weights of the vision transformer. Such fine-tuning is less practical for dense prediction due to the much heavier computation and GPU memory cost than image classification.In this paper, we focus on a more challenging problem, \ie, accelerating large-scale vision transformers for dense prediction without any additional re-training or fine-tuning. In response to the fact that high-resolution representations are necessary for dense prediction, we present two non-parametric operators, a \emph{token clustering layer} to decrease the number of tokens and a \emph{token reconstruction layer} to increase the number of tokens. The following steps are performed to achieve this: (i) we use the token clustering layer to cluster the neighboring tokens together, resulting in low-resolution representations that maintain the spatial structures; (ii) we apply the following transformer layers only to these low-resolution representations or clustered tokens; and (iii) we use the token reconstruction layer to re-create the high-resolution representations from the refined low-resolution representations. The results obtained by our method are promising on five dense prediction tasks including object detection, semantic segmentation, panoptic segmentation, instance segmentation, and depth estimation. Accordingly, our method accelerates $40\%\uparrow$ FPS and saves $30\%\downarrow$ GFLOPs of ``Segmenter+ViT-L/$16$'' while maintaining $99.5\%$ of the performance on ADE$20$K without fine-tuning the official weights. Weicong Liang, Yuhui Yuan, Henghui Ding, Xiao Luo 0001, Weihong Lin, Ding Jia, Zheng Zhang 0022, Chao Zhang 0001, Han Hu 0001 |
NeurIPS | 5 |
| 2021 | ViBERTgrid: A Jointly Trained Multi-modal 2D Document Representation for Key Information Extraction from Documents
Weihong Lin, Qifang Gao, Lei Sun 0003, Zhuoyao Zhong, Qin Ren 0003, Qiang Huo |
ICDAR (1) | 1 |
| 2021 | HRFormer: High-Resolution Vision Transformer for Dense PredictabstractWe present a High-Resolution Transformer (HRFormer) that learns high-resolution representations for dense prediction tasks, in contrast to the original Vision Transformer that produces low-resolution representations and has high memory and computational cost. We take advantage of the multi-resolution parallel design introduced in high-resolution convolutional networks (HRNet [45]), along with local-window self-attention that performs self-attention over small non-overlapping image windows [21], for improving the memory and computation efficiency. In addition, we introduce a convolution into the FFN to exchange information across the disconnected image windows. We demonstrate the effectiveness of the HighResolution Transformer on both human pose estimation and semantic segmentation tasks, e.g., HRFormer outperforms Swin transformer [27] by 1.3 AP on COCO pose estimation with 50% fewer parameters and 30% fewer FLOPs. Code is available at: https://github.com/HRNet/HRFormer Yuhui Yuan, Lang Huang 0001, Weihong Lin, Chao Zhang 0001, Xilin Chen 0001, Jingdong Wang 0001 |
NeurIPS | 4 |
| 2017 | A Caching Miss Ratio Aware Path Selection Algorithm for Information-Centric NetworksabstractIn Information-Centric Networks (ICN), contents are cached on some intermediary routers. This creates thus a new situation which is totally different from the traditional path-selection paradigm: the source/destination paradigm no longer exists; instead, the new paradigm is how to find a path through a selected group of caches, so that the content is delivered via the shortest way. This paper addresses this issue and proposes a path-selection algorithm taking into account both the caching capability of router and the more traditional link cost between routers. We formulated the problem as a convex optimization problem (named ESP) which aims to get expected shortest path (ESP) by minimizing the transportation cost. By applying the Lagrangian dual theorem, we solved the ESP problem and obtained a criterion for request (and reversely, data) routing. Based on this path-selection criterion, we provide a fully distributed distance-based ESP algorithm that enables routers maintain routes to nearest content, without knowing a network topology and the caching miss ratio of content at other routers. Simulations confirm the efficiency of our approach versus the traditional shortest path algorithm. Weihong Lin, Xinggong Zhang, Yu Guan 0005, Zongming Guo |
LCN | 1 |