VLDB 2026 Research / reviewers in the wild / expert
Hongchao Lu
dblp:00/6909
· DBLP profile ↗
13ranked-venue papers
6as first author
6since 2021 · last 2026
0000-0002-6740-7927ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorSystems, architecture and hardware · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Reinforcement learning · 26% Robot manipulation · 22% Motion planning and robot control · 12% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
offline reinforcement learning |
1.9 | 2 | 2026 | Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models · AAAI 2026 ReinboT: Amplifying Robot Visual-Language Manipulation with Reinforcement Learning · ICML 2025 |
Robotics › Robot manipulation › embodied foundation models
vision-language-action model |
1.9 | 2 | 2026 | Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models · AAAI 2026 ReinboT: Amplifying Robot Visual-Language Manipulation with Reinforcement Learning · ICML 2025 |
Robotics › Motion planning and robot control › robot learning › robot policy learning
flow matching policies |
1.0 | 1 | 2026 | Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models · AAAI 2026 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.5 | 1 | 2021 | Feature Enhanced Projection Network for Zero-shot Semantic Segmentation · ICRA 2021 |
Natural language and speech › Language models and text generation › text representation
semantic word embeddings |
0.5 | 1 | 2021 | Feature Enhanced Projection Network for Zero-shot Semantic Segmentation · ICRA 2021 |
Machine learning › Transfer learning and domain adaptation
zero-shot learning |
0.5 | 1 | 2021 | Feature Enhanced Projection Network for Zero-shot Semantic Segmentation · ICRA 2021 |
Computer vision › Segmentation and scene understanding › semantic segmentation › open-vocabulary segmentation
zero-shot semantic segmentation |
0.5 | 1 | 2021 | Feature Enhanced Projection Network for Zero-shot Semantic Segmentation · ICRA 2021 |
Machine learning › Efficient and distributed learning
inference efficiency |
0.4 | 1 | 2019 | Fast Object Detection in Compressed Video · ICCV 2019 |
Computer vision › Image recognition and object detection
object detection |
0.4 | 1 | 2019 | Fast Object Detection in Compressed Video · ICCV 2019 |
Computer vision › Video understanding and tracking
video object detection |
0.4 | 1 | 2019 | Fast Object Detection in Compressed Video · ICCV 2019 |
Machine learning › Reinforcement learning
imitation learning |
0.3 | 1 | 2026 | Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models · AAAI 2026 |
Robotics › Autonomous driving › perception
environment perception |
0.1 | 1 | 2021 | Feature Enhanced Projection Network for Zero-shot Semantic Segmentation · ICRA 2021 |
Robotics › Autonomous driving
perception |
0.1 | 1 | 2021 | Feature Enhanced Projection Network for Zero-shot Semantic Segmentation · ICRA 2021 |
Methods — techniques the papers use, named apart from their topics
offline reinforcement learning · 1.9flow matching · 1.0bias-variance trade-off · 1.0imitation learning · 0.9dense return prediction · 0.9projection network · 0.5knowledge transfer · 0.5residual error · 0.4motion vector · 0.4convolutional neural network · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow ModelsabstractVision-Language-Action (VLA) models based on flow matching have shown excellent performance in general-purpose robotic manipulation tasks. However, the action accuracy of these models on complex downstream tasks is unsatisfactory. One important reason is that these models rely solely on the post-training paradigm of imitation learning, which makes it difficult to have a deeper understanding of the distribution properties of data quality, which is exactly what Reinforcement Learning (RL) excels at. In this paper, we theoretically propose an offline RL post-training objective for VLA flow models and induce an efficient and feasible offline RL fine-tuning algorithm −− Adaptive Reinforced Flow Matching (ARFM). By introducing an adaptively adjusted scaling factor in the VLA flow model loss, we construct a principled bias-variance trade-off objective function to optimally control the impact of RL signal on flow loss. ARFM adaptively balances RL advantage preservation and flow loss gradient variance control, resulting in a more stable and efficient fine-tuning process. Extensive simulation and real-world experimental results show that ARFM exhibits excellent generalization, robustness, few-shot learning, and continuous learning performance. Hongyin Zhang 0001, Junxi Jin, Qixin Zeng, Hongchao Lu |
AAAI | 6 |
| 2025 | ReinboT: Amplifying Robot Visual-Language Manipulation with Reinforcement LearningabstractVision-Language-Action (VLA) models have shown great potential in general robotic decision-making tasks via imitation learning. However, the variable quality of training data often constrains the performance of these models. On the other hand, offline Reinforcement Learning (RL) excels at learning robust policy models from mixed-quality data. In this paper, we introduce Reinforced robot GPT (ReinboT), a novel end-to-end VLA model that integrates the RL principle of maximizing cumulative reward. ReinboT achieves a deeper understanding of the data quality distribution by predicting dense returns that capture the nuances of manipulation tasks. The dense return prediction capability enables the robot to generate more robust decision-making actions, oriented towards maximizing future benefits. Extensive experiments show that ReinboT achieves state-of-the-art performance on the CALVIN mixed-quality dataset and exhibits superior few-shot learning and out-of-distribution generalization capabilities in real-world tasks. Hongyin Zhang 0001, Zifeng Zhuang, Han Zhao 0008, Pengxiang Ding, Hongchao Lu |
ICML | 5 |
| 2021 | Learning With Memory For Few-Shot Semantic SegmentationabstractDespite great progress made in the few-shot semantic segmentation task, the existing works still suffer from problems of incompleteness and inconsistency of segmentation. In this paper, a novel attention-aided LSTM optimization network called LONet is proposed, which optimizes predictions without forgetting useful inner cues. Particularly, we calculate an attention map to align and match possible locations with query features to deal with incomplete segmentation. Then, an LSTM-based module is designed to overcome the segmentation inconsistency by memorizing and updating useful cues iteratively. Extensive experiments are conducted on two popular few-shot segmentation datasets including PASCAL-5iand FSS-1000. The experimental results on the FSS-1000 dataset demonstrate that our LONet exceeds the state-of-the-art results by 2.1% and 2.3%, respectively. Hongchao Lu, Zhidong Deng |
ICIP | 1 |
| 2021 | Feature Enhanced Projection Network for Zero-shot Semantic SegmentationabstractIn environmental perception of autonomous driving, zero-shot semantic segmentation that can make prediction of new categories without using any labeled training samples is considered as a challenging task. One key step in this task is to transfer knowledge across categories via auxiliary semantic word embeddings. In this paper, we propose a feature enhanced projection network (FEPNet) that takes full advantage of transferred knowledge to enrich semantic representations. In FEPNet, two projection layers are added to a segmentation network so as to map features into seen (S) and unseen (U) category spaces, respectively. During training, U-space features are transferred to S-space using similarity relations to enhance the representation of seen categories. In the inference stage, the representation of unseen categories is also strengthened by incorporating features transferred from S-space. Moreover, a novel strategy is proposed to effectively alleviate prediction bias by performing segmentation independently in separate areas that contain seen and unseen categories. We conduct extensive experiments on three benchmark datasets. The experimental results show that our FEPNet achieves new state-of-the-art results compared to existing approaches. Hongchao Lu, Longwei Fang, Matthieu Lin, Zhidong Deng |
ICRA | 1 |
| 2021 | Lane Intrusion Behaviors Dataset: Action Recognition in Real-world Highway Scenarios for Self-drivingabstractIt is necessary for the development of self-driving to fulfill the requirements for safety, stability, and intelligence, especially in high-speed conditions. Therefore, the detection of pedestrians that may occur in highway scenarios during driving and understanding the meaning of their behaviors in advance are significantly important for the self-driving vehicle to make correct decisions. However, no existing datasets are available for behavior recognition in self-driving scenarios. In order to advance the task of interactive cognition between the vehicle and pedestrians, in this paper, we present a new dataset, called THU-IntrudBehavior, that collects lane intrusion behaviors of pedestrians that can be applied in real world highway scenarios. The dataset contains diverse behaviors of single or multiple pedestrians/cyclists that are simulated in different urban roads under various weather conditions. We describe annotations of each video and report several experimental results of baseline methods on our self-collected dataset. Our THU-IntrudBehavior dataset provides new support for behavior recognition in high-speed conditions for self-driving. Hongchao Lu, Zhidong Deng, Ruiwen Zhang |
IJCNN | 1 |
| 2021 | Phase Space Reconstruction Network for Lane Intrusion Action RecognitionabstractIn a complex road traffic scene, illegal lane intrusion of pedestrians or cyclists constitutes one of the main safety challenges in autonomous driving application. In this paper, we propose a novel object-level phase space reconstruction network (PSRNet) for motion time series classification, aiming to recognize lane intrusion actions that occur 150m ahead through a monocular camera fixed on moving vehicle. In the PSRNet, the movement of pedestrians and cyclists, specifically viewed as an observable object-level dynamic process, can be reconstructed as trajectories of state vectors in a latent phase space and further characterized by a learnable Lyapunov exponent-like classifier that indicates discrimination in terms of average exponential divergence of state trajectories. Additionally, in order to first transform video inputs into one-dimensional motion time series of each object, a lane width normalization based on visual object tracking-by-detection is presented. Extensive experiments are conducted on the THU-IntrudBehavior dataset collected from real urban roads. The results show that our PSRNet could reach the best accuracy of 98.0%, which remarkably exceeds existing action recognition approaches by more than 30%. Ruiwen Zhang, Zhidong Deng, Hongsen Lin, Hongchao Lu |
IJCNN | 4 |
| 2020 | A Boundary-aware Distillation Network for Compressed Video Semantic SegmentationabstractIn recent years optical flow is often estimated to reuse features so as to accelerate video semantic segmentation. With addition of optical flow network, however, extra cost may incur and accuracy may thus be degraded because of repeated warping operation. In this paper, we propose a boundary-aware distillation network (BDNet) that replaces optical flow network with block motion vectors encoded in compressed video, resulting in negligible computational complexity. In order to make salient features, an auxiliary boundary-aware stream is added to the main stream to jointly estimate silhouette and segmentation of objects. To further correct warped features, a well-trained teacher network is employed to transfer knowledge to the main stream. Both boundary-aware stream and the teacher network are neglected during inference stage, so that video segmentation network enables to get faster without increasing any computational burden. By splitting the task into three components, our BDNet shows almost 10% time saving as well as 1.6% accuracy improvement over baseline on the Cityscapes dataset. Hongchao Lu, Zhidong Deng |
ICPR | 1 |
| 2019 | Fast Object Detection in Compressed VideoabstractObject detection in videos has drawn increasing attention since it is more practical in real scenarios. Most of the deep learning methods use CNNs to process each decoded frame in a video stream individually. However, the free of charge yet valuable motion information already embedded in the video compression format is usually overlooked. In this paper, we propose a fast object detection method by taking advantage of this with a novel Motion aided Memory Network (MMNet). The MMNet has two major advantages: 1) It significantly accelerates the procedure of feature extraction for compressed videos. It only need to run a complete recognition network for I-frames, i.e. a few reference frames in a video, and it produces the features for the following P frames (predictive frames) with a light weight memory network, which runs fast; 2) Unlike existing methods that establish an additional network to model motion of frames, we take full advantage of both motion vectors and residual errors that are freely available in video streams. To our best knowledge, the MMNet is the first work that investigates a deep convolutional detector on compressed videos. Our method is evaluated on the large-scale ImageNet VID dataset, and the results show that it is 3× times faster than single image detector R-FCN and 10× times faster than high-performance detector MANet at a minor accuracy loss. Shiyao Wang 0001, Hongchao Lu, Zhidong Deng |
ICCV | 2 |
| 2018 | SRC-Disp: Synthetic-Realistic Collaborative Disparity Learning for Stereo Matching
Guorun Yang, Zhidong Deng, Hongchao Lu, Zeping Li |
ACCV (5) | 3 |
| 2018 | Semantic Image Segmentation Based on Attentions to Intra Scales and Inner ChannelsabstractMulti-scale features provide different context information of objects, which is significant for better performance in semantic segmentation tasks. But different scale features contribute equally to final predicitons. In this paper, we propose a new attention mechanism that not only learns weights between different scales but also allocates importance to subregions of inner channels. The network architecture is built on a state-of-the-art feedforward network to generate strong semantic feature maps, accompanying with a top-down pathway that incorporates larger scale feature maps through lateral connections. The proposed intra-scale attention module softly weights features between scales pixel by pixel. To enhance impact of each feature map in intermediate layers on performance, we further present an inner-channel attention module to pay attention to subregions within each channel. Moreover, an extra supervision is presented to achieve excellent performance. Importantly, the inner-channel attention module adaptively changes features as the layer goes deeper and could be inserted into any other layers. Extensive experiments are conducted on PASCAL VOC2012 to verify the network effectiveness. The experimental results show that both intra-scale and inner-channel attention modules could yield better performance. Hongchao Lu, Zhidong Deng, Xiaolong Liu 0010 |
IJCNN | 1 |
| 2009 | Predicting Functional Alternative Splicing by Measuring RNA Selection Pressure from Multigenome AlignmentsabstractHigh-throughput methods such as EST sequencing, microarrays and deep sequencing have identified large numbers of alternative splicing (AS) events, but studies have shown that only a subset of these may be functional. Here we report a sensitive bioinformatics approach that identifies exons with evidence of a strong RNA selection pressure ratio (RSPR)--i.e., evolutionary selection against mutations that change only the mRNA sequence while leaving the protein sequence unchanged--measured across an entire evolutionary family, which greatly amplifies its predictive power. Using the UCSC 28 vertebrate genome alignment, this approach correctly predicted half to three-quarters of AS exons that are known binding targets of the NOVA splicing regulatory factor, and predicted 345 strongly selected alternative splicing events in human, and 262 in mouse. These predictions were strongly validated by several experimental criteria of functional AS such as independent detection of the same AS event in other species, reading frame-preservation, and experimental evidence of tissue-specific regulation: 75% (15/20) of a sample of high-RSPR exons displayed tissue specific regulation in a panel of ten tissues, vs. only 20% (4/20) among a sample of low-RSPR exons. These data suggest that RSPR can identify exons with functionally important splicing regulation, and provides biologists with a dataset of over 600 such exons. We present several case studies, including both well-studied examples (GRIN1) and novel examples (EXOC7). These data also show that RSPR strongly outperforms other approaches such as standard sequence conservation (which fails to distinguish amino acid selection pressure from RNA selection pressure), or pairwise genome comparison (which lacks adequate statistical power for predicting individual exons). Hongchao Lu, Seiko Sato, Yi Xing, Christopher J. Lee |
PLoS Comput. Biol. | 1 |
| 2006 | Phylophenetic properties of metabolic pathway topologies as revealed by global analysisabstractBACKGROUND: As phenotypic features derived from heritable characters, the topologies of metabolic pathways contain both phylogenetic and phenetic components. In the post-genomic era, it is possible to measure the "phylophenetic" contents of different pathways topologies from a global perspective. RESULTS: We reconstructed phylophenetic trees for all available metabolic pathways based on topological similarities, and compared them to the corresponding 16S rRNA-based trees. Similarity values for each pair of trees ranged from 0.044 to 0.297. Using the quartet method, single pathways trees were merged into a comprehensive tree containing information from a large part of the entire metabolic networks. This tree showed considerably higher similarity (0.386) to the corresponding 16S rRNA-based tree than any tree based on a single pathway, but was, on the other hand, sufficiently distinct to preserve unique phylogenetic information not reflected by the 16S rRNA tree. CONCLUSION: We observed that the topology of different metabolic pathways provided different phylogenetic and phenetic information, depicting the compromise between phylogenetic information and varying evolutionary pressures forming metabolic pathway topologies in different organisms. The phylogenetic information content of the comprehensive tree is substantially higher than that of any tree based on a single pathway, which also gave clues to constraints working on the topology of the global metabolic networks, information that is only partly reflected by the topologies of individual metabolic pathways. Yong Zhang 0006, Shaojuan Li, Geir Skogerbø, Shiwei Sun, Hongchao Lu, Baochen Shi, Runsheng Chen |
BMC Bioinform. | 8 |
| 2006 | Dynamic Changes in Subgraph Preference Profiles of Crucial Transcription FactorsabstractTranscription factors with a large number of target genes--transcription hub(s), or THub(s)--are usually crucial components of the regulatory system of a cell, and the different patterns through which they transfer the transcriptional signal to downstream cascades are of great interest. By profiling normalized abundances (A(N)) of basic regulatory patterns of individual THubs in the yeast Saccharomyces cerevisiae transcriptional regulation network under five different cellular states and environmental conditions, we have investigated their preferences for different basic regulatory patterns. Subgraph-normalized abundances downstream of individual THubs often differ significantly from that of the network as a whole, and conversely, certain over-represented subgraphs are not preferred by any THub. The THub preferences changed substantially when the cellular or environmental conditions changed. This switching of regulatory pattern preferences suggests that a change in conditions does not only elicit a change in response by the regulatory network, but also a change in the mechanisms by which the response is mediated. The THub subgraph preference profile thus provides a novel tool for description of the structure and organization between the large-scale exponents and local regulatory patterns. Changning Liu, Geir Skogerbø, Hongchao Lu, Baochen Shi, Yong Zhang 0006, Tao Wu 0002, Runsheng Chen |
PLoS Comput. Biol. | 5 |