VLDB 2026 Research / reviewers in the wild / expert
Linlin Ou
dblp:17/7817
· DBLP profile ↗
22ranked-venue papers
1as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 1 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 5 since 2021Systems, architecture and hardware · 6 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Retrieval-Enhanced Visual Prompt Learning for Few-Shot ClassificationabstractThe Contrastive Language-Image Pretraining (CLIP) model has been widely used in various downstream vision tasks. The few-shot learning paradigm has been widely adopted to augment its capacity for these tasks. However, current paradigms may struggle with fine-grained classification, such as satellite image recognition, due to widening domain gaps. To address this limitation, we propose retrieval-enhanced visual prompt learning (RePrompt), which introduces retrieval mechanisms to cache and reuse the knowledge of downstream tasks. RePrompt constructs a retrieval database from either training examples or external data if available, and uses a retrieval mechanism to enhance multiple stages of a simple prompt learning baseline, thus narrowing the domain gap. During inference, our enhanced model can reference similar samples brought by retrieval to make more accurate predictions. A detailed analysis reveals that retrieval helps to improve the distribution of late features, thus, improving generalization for downstream tasks. RePrompt attains state-of-the-art performance on a wide range of vision datasets, including 11 image datasets, 3 video datasets, 1 multi-view dataset, and 4 domain generalization benchmarks. Jintao Rong 0001, Hao Chen 0041, Linlin Ou, Tianxiao Chen, Yifan Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Channel Merging: Preserving Specialization for Merged ExpertsabstractLately, the practice of utilizing task-specific fine-tuning has been implemented to improve the performance of large language models (LLM) in subsequent tasks. Through the integration of diverse LLMs, the overall competency of LLMs is significantly boosted. Nevertheless, traditional ensemble methods are notably memory-intensive, necessitating the simultaneous loading of all specialized models into GPU memory. To address the inefficiency, model merging strategies have emerged, merging all LLMs into one model to reduce the memory footprint during inference. Despite these advances, model merging often leads to parameter conflicts and performance decline as the number of experts increases. Previous methods to mitigate these conflicts include post-pruning and partial merging. However, both approaches have limitations, particularly in terms of performance and storage efficiency when merged experts increase. To address these challenges, we introduce Channel Merging, a novel strategy designed to minimize parameter conflicts while enhancing storage efficiency. This method initially clusters and merges channel parameters based on their similarity to form several groups offline. By ensuring that only highly similar parameters are merged within each group, it significantly reduces parameter conflicts. During inference, we can instantly look up the expert parameters from the merged groups, preserving specialized knowledge. Our experiments demonstrate that Channel Merging consistently delivers high performance, matching unmerged models in tasks like English and Chinese reasoning, mathematical reasoning, and code generation. Moreover, it obtains results comparable to model ensemble with just 53% parameters when used with a task-specific router. Mingyang Zhang 0007, Jing Liu 0048, Ganggui Ding, Linlin Ou, Bohan Zhuang |
AAAI | 4 |
| 2025 | Towards Robust Category-level Articulation Pose Estimation via Integrated Differentiable RenderingabstractAccurate object pose estimation is crucial for embodied intelligence tasks such as manipulation, grasping, and human-robot interaction. However, due to the inherent characteristics of articulated objects, such as kinematic constraints and self-occlusion, pose estimation for articulated objects has remained a significant challenge. To address these issues, this paper proposes CAPED, an end-to-end robust Category-level Articulated object Pose Estimator integrated differentiable rendering. Given partial point cloud as input, CAPED outputs the per-part 6D pose for articulation. Specifically, with the proposed joint-centric modeling manner, CAPED firstly estimates the pose for the free part. Afterward, we canonicalize the input point cloud to estimate constrained parts’ poses by predicting the joint parameters and states as replacements. For further refinement, we propose a differentiable rendering scheme for pose optimization. Evaluations of the ArtImage and RobotArm datasets demonstrate that CAPED exhibits outstanding effectiveness and generalization in tasks ranging from synthetic data to real-world scenarios. We will publicly release the code. Li Zhang 0104, Yukang Huo, Lin Wu 0001, Yanyan Wei, Harshal Suresh Shende, Liu Liu 0012, Linlin Ou |
ICASSP | 10 |
| 2025 | A Novel Hybrid Hysteresis Modeling Method for Multiloop-Asymmetry Hysteresis Behavior of Nonlinear Compliant ActuatorsabstractNonlinear compliant actuators are being increasingly used in human-robot interaction scenarios due to their inherent flexibility. However, a limitation is that nonlinear hysteresis exists, which will degrade the force/torque tracking performance if the hysteresis is not modeled accurately. Moreover, the existing methods are difficult to deal with the multi-loop asymmetry hysteresis. In this work, we present a novel modeling method, in which the hysteresis curves are decoupled into nonlinear reference lines and symmetrical hysteresis loops. A hybrid hysteresis model based on power function and Maxwellslip model is then developed to fit the nonlinear reference lines and symmetrical hysteresis loops respectively. Experiments were conducted on a nonlinear compliant actuator and the results show that the root-mean-square-errors (RMSE) of the hysteresis model decreases by 24.4% when compared with the Maxwellslip based hysteresis model. Lingpeng Xu, Linlin Ou, Yalei Feng, Shaoping Bai |
ICRA | 3 |
| 2024 | Improving Neural Indoor Surface Reconstruction with Mask-Guided Adaptive Consistency Constraintsabstract3D scene reconstruction from 2D images has been a long-standing task. Instead of estimating per-frame depth maps and fusing them in 3D, recent researches leverage the neural implicit surface as a global representation for 3D reconstruction. Equipped with data-driven pre-trained geometric cues, these methods have demonstrated promising performance. However, the inevitable inaccurate estimation of priors can lead to suboptimal reconstruction quality, particularly in some geometrically complex regions. In this paper, we propose a two-stage training process to further improve the reconstruction quality. It decouples the view-dependent and view-independent colors, and leverages two novel consistency constraints to enhance detail reconstruction performance without requiring extra priors. Additionally, we introduce an essential mask scheme to adaptively influence the selection of supervision constraints, thereby improving performance in a self-supervised paradigm. Experiments on synthetic and real-world datasets show the capability of reducing the side effects of inaccurately estimated priors and achieving high-quality scene reconstruction with rich geometric details. Liqin Lu, Jintao Rong 0001, Guangkai Xu, Linlin Ou |
ICRA | 5 |
| 2024 | EfficientCAPER: An End-to-End Framework for Fast and Robust Category-Level Articulated Object Pose EstimationabstractHuman life is populated with articulated objects. Pose estimation for category-level articulated objects is a significant challenge due to their inherent complexity and diverse kinematic structures. Current methods for this task usually meet the problems of insufficient consideration of kinematic constraints, self-occlusion, and optimization requirements. In this paper, we propose EfficientCAPER, an end-to-end Category-level Articulated object Pose EstimatoR, eliminating the need for optimization functions as post-processing and utilizing the kinematic structure for joint-centric pose modeling, thus enhancing the efficiency and applicability. Given a partial point cloud as input, the EfficientCAPER firstly estimates the pose for the free part of an articulated object using decoupled rotation representation. Next, we canonicalize the input point cloud to estimate constrained parts' poses by predicting the joint parameters and states as replacements. Evaluations on three diverse datasets, ArtImage, ReArtMix, and RobotArm, show EfficientCAPER's effectiveness and generalization ability to real-world scenarios. The framework exhibits excellent static pose estimation performance for articulated objects, contributing to the advancement of category-level pose estimation. Codes will be made publicly available. Li Zhang 0104, Lin Wu 0001, Linlin Ou, Liu Liu 0012 |
NeurIPS | 5 |
| 2024 | DMFusion: LiDAR-camera fusion framework with depth merging and temporal aggregation
Linlin Ou |
Appl. Intell. | 4 |
| 2024 | Human-robot collaborative interaction with human perception and action recognitionabstractThis paper presents a human–robot interaction system (HRIS) that utilizes human perception and action recognition to enable the robot to understand human intentions and flexibly interact with humans. A monocular multi-person three-dimensional (3D) pose estimation method is first proposed to perceive multi-person two-dimensional (2D) and 3D poses in interaction scenarios. Furthermore, a 3D skeleton poses tracking approach is adopted to locate the identity of each person in consecutive frames and enhance interactive stability. Then, an action recognition model is developed, which exploits tracked pose features to recognize the intentions of humans. An action-controlled interaction system is built with a modular approach to ensure flexibility in meeting multiple task requirements and facilitating flexible interaction. In the system, a distance-based safety solution is designed to avoid collisions between humans and robots. Finally, experimental results are presented to demonstrate the feasibility and effectiveness of the proposed methods and system. Chengjun Xu, Linlin Ou |
Neurocomputing | 4 |
| 2024 | Asymmetric time-varying integral barrier Lyapunov function based adaptive optimal control for nonlinear systems with dynamic state constraintsabstractThis paper investigates the issue of adaptive optimal tracking control for nonlinear systems with dynamic state constraints. An asymmetric time-varying integral barrier Lyapunov function (ATIBLF) based integral reinforcement learning (IRL) control algorithm with an actor–critic structure is first proposed. The ATIBLF items are appropriately arranged in every step of the optimized backstepping control design to ensure that the dynamic full-state constraints are never violated. Thus, optimal virtual/actual control in every backstepping subsystem is decomposed with ATIBLF items and also with an adaptive optimized item. Meanwhile, neural networks are used to approximate the gradient value functions. According to the Lyapunov stability theorem, the boundedness of all signals of the closed-loop system is proved, and the proposed control scheme ensures that the system states are within predefined compact sets. Finally, the effectiveness of the proposed control approach is validated by simulations. Mingshuang Hao, Linlin Ou |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2024 | DynamicAug: Enhancing Transfer Learning Through Dynamic Data Augmentation Strategies Based on Model StateabstractAbstract Transfer learning has made significant advancements, however, the issue of overfitting continues to pose a major challenge. Data augmentation has emerged as a highly promising technique to counteract this challenge. Current data augmentation methods are fixed in nature, requiring manual determination of the appropriate intensity prior to the training process. However, this entails substantial computational costs. Additionally, as the model approaches convergence, static data augmentation strategies can become suboptimal. In this paper, we introduce the concept of Dynamic Data Augmentation (DynamicAug), a method that autonomously adjusts the intensity of data augmentation, taking into account the convergence state of the model. During each iteration of the model’s forward pass, we utilize a Gaussian distribution based sampler to stochastically sample the current intensity of data augmentation. To ensure that the sampled intensity is aligned with the convergence state of the model, we introduce a learnable expectation to the sampler and update the expectation iteratively. In order to assess the convergence status of the model, we introduce a novel loss function called the convergence loss. Through extensive experiments conducted over 27 vision datasets, we have demonstrated that DynamicAug can significantly enhance the performance of existing transfer learning methods. Haodong Zhao, Mingyang Zhang 0007, Linlin Ou |
Neural Process. Lett. | 6 |
| 2023 | ShiftNAS: Improving One-shot NAS via Probability ShiftabstractOne-shot Neural architecture search (One-shot NAS) has been proposed as a time-efficient approach to obtain optimal subnet architectures and weights under different complexity cases by training only once. However, the subnet performance obtained by weight sharing is often inferior to the performance achieved by retraining. In this paper, we investigate the performance gap and attribute it to the use of uniform sampling, which is a common approach in supernet training. Uniform sampling concentrates training resources on subnets with intermediate computational resources, which are sampled with high probability. However, subnets with different complexity regions require different optimal training strategies for optimal performance.To address the problem of uniform sampling, we propose ShiftNAS, a method that can adjust the sampling probability based on the complexity of subnets. We achieve this by evaluating the performance variation of subnets with different complexity and designing an architecture generator that can accurately and efficiently provide subnets with the desired complexity. Both the sampling probability and the architecture generator can be trained end-to-end in a gradient-based manner. With ShiftNAS, we can directly obtain the optimal model architecture and parameters for a given computational complexity. We evaluate our approach on multiple visual network models, including convolutional neural networks (CNNs) and vision transformers (ViTs), and demonstrate that ShiftNAS is model-agnostic. Experimental results on ImageNet show that ShiftNAS can improve the performance of one-shot NAS without additional consumption. Source codes are available at GitHub. Mingyang Zhang 0007, Haodong Zhao, Linlin Ou |
ICCV | 4 |
| 2023 | Repnas: Searching for Efficient Re-Parameterizing BlocksabstractIn the past years, significant improvements in the field of neural architecture search(NAS) have been made. However, in order to improve the performance of the model, many search spaces contain multi-branch architectures, which leads to the gap between the searched constraint and real inference time. In this work, we propose a re-parameterization (Rep) search space based on structural Rep techniques. In the Rep search space, the subnets are multi-branch in training and single-path in inference. Furthermore, RepNAS, a one-stage NAS approach, is present to efficiently search the optimal diverse branch block (ODBB) for each layer under the branch number constraint. Our experimental results show the searched ODBB can easily surpass the manual diverse branch block (DBB) with efficient training. Mingyang Zhang 0007, Jintao Rong 0001, Linlin Ou |
ICME | 4 |
| 2023 | Conditional generative data-free knowledge distillation
Linlin Ou |
Image Vis. Comput. | 5 |
| 2023 | Def-Grasp: A Robot Grasping Detection Method for Deformable Objects Without Force Sensor
Chongliang Zhao, Linlin Ou |
Neural Process. Lett. | 5 |
| 2023 | Efficient Re-parameterization Operations Search for Easy-to-Deploy Network Based on Directional Evolutionary Strategy
Jintao Rong 0001, Mingyang Zhang 0007, Linlin Ou |
Neural Process. Lett. | 5 |
| 2022 | Graph pruning for model compression
Mingyang Zhang 0007, Jintao Rong 0001, Linlin Ou |
Appl. Intell. | 4 |
| 2020 | 3D Semantic Map Construction System Based on Visual SLAM and CNNsabstractTraditional approaches to simultaneous localization and mapping (SLAM) are unable to extract semantic information from the scene or meet the high-level task requirements of robots, and the efficiency of 3D map construction is low. To solve this problem, a 3D semantic map construction system is proposed in the paper to build a 3D semantic map. Firstly, the current position of the camera is estimated and optimized based on ORB-SLAM algorithm, and the globally consistent trajectory and pose are obtained. Then the semantic segmentation network is designed to predict semantic category of every pixel. The 3D semantic point cloud information is generated by combing the semantic information and the object point cloud. The global consistent camera pose estimated by visual SLAM algorithm is integrated with the semantic point cloud information to generate a 3D semantic map. Finally, we use an Octomap which is applied to navigation projects for map storage to reduce the storage capacity of the map. The experimental result verifies the accuracy and efficiency of this method. Lei Lai, Xuecheng Qian, Linlin Ou |
IECON | 4 |
| 2020 | Soft Taylor Pruning for Accelerating Deep Convolutional Neural NetworksabstractNetworking pruning is widely utilized for accelerating the inference procedure of deep models in low-resource settings. In this paper, a novel Gradient-based method, Soft Taylor Pruning(STP), are proposed to reduce the network complexity in dynamic way. Given a global compression rate, all filters are categorized into two parts by the gradient-based evaluation criterion in the pruning process. Then, two types of filters are remixed and updated in the training epoch. When a pruning error occurs, the model can correct the pruning error by re-active pruned filters in the next training epoch. In this way, the capacity of the model channel space remains the same until the network structure converges. In order to reduce the impact of large-weighted filters on criterion, We take the absolute value of the product of the feature map and the gradient as the evaluation criterion. So as to reduce the model pruning time, STP allows simultaneous pruning on multiple layers by controlling the opening and closing of multiple mask layers. Moreover, STP can be applied to various advanced CNNs, such as MobileNet. The features of our method are: 1) maintain the integrity of the model channel space; 2) less time cost of model compression; 3)less dependence on the pretrained model. Jintao Rong 0001, Xiyi Yu, Mingyang Zhang 0007, Linlin Ou |
IECON | 4 |
| 2020 | Construction and Optimization of Digital Twin Model for Hardware Production LineabstractA digital twin system about hardware production line is designed in this paper. Based on 3D modeling and visualization technology, a model database is established to meet the needs of production line. Genetic Algorithm is used to improve the process layout, which raises the equipment utilization. In order to achieve better obstacle avoidance, the robot path planning is carried out to avoid interference. By combining the above results, the digital model is built in Demo3D software to realize the simulate and virtual debugging. The result shows that this system of hardware production line can complete the debugging of real physical equipment in virtual environment. The feasibility and practicability of this method are verified. Liting Yuan, Linlin Ou |
IECON | 4 |
| 2020 | Multi-View Human Pose Estimation in Human-Robot InteractionabstractThis paper contributes a real-time human-robot interaction system based on multi-human poses capture. First, we propose an iterative method for multi-human 3D poses estimation from multi-view. There are three mainly steps: (1) all independent 2D poses under each view are obtained by using the pose detector; (2) 2D poses from different perspectives are associated based on a greedy algorithm; (3) 3D skeletons are generated with the 2D poses that have been correctly associated. In the process of each iteration, the poses with the highest correspondence are always paired priority to ensure the accuracy of the results. Due to the occlusion in each view, previous methods that rely on the appearance of the human body and wearable devices are not suitable for multi-human 3D poses estimation, especially in crowded scenes. The proposed method does not required any markers or devices, which has good robustness even under occlusion or crowded. In addition, we also designed three human-robot interaction modes, including action following, target specification and dynamic obstacle avoidance. For each mode, target generation and correction methods are present. Combining human body pose estimation and robot motion, design experiments to verify the smoothness of the robot control method for dynamic target tracking and the effectiveness in different interaction modes. Chengjun Xu, Zhengan Wang, Linlin Ou |
IECON | 4 |
| 2012 | Decentralized PID controller design for the cooperative control of networked multi-agent systemsabstractFor the networked multi-agent system with arbitrary-order time-delayed agent dynamics, the parametric H∞design method of the decentralized PID controller is proposed in this paper. The closed-loop framework representation is first given for the multi-agent system with the decentralized PID controller imposed on each agent. Based on this close-loop framework, the H∞performance criterion of the entire system is transformed into several local H∞performance constraints of the subsystem which is related to the eigenvalues of the Laplacian matrix. Thus, the design problem of the decentralized H∞PID controller is converted to the stabilization problem of the PID controller simultaneously for a family of complex quasipolynomials. Then, two parametric approaches are given to determine the region of the PID control parameters that can guarantee the stability of the complex quasipolynomial. Finally, the decentralized H∞PID controller is derived by finding the intersection of the stabilizing PID regions for all resultant quasipolynomials. Linlin Ou, Qike Shao, Yuan Su, Li Yu 0001 |
ICARCV | 1 |
| 2012 | Stability region of fractional-order PI λDμ controller for fractional-order systems with time delayabstractA simple and effective method to determine the region of fractional-order PIλDμcontrollers that can stabilize a given fractional-order system with time delay is proposed in this paper. For each known proportional, integral or derivative gain in the PIλDμcontrollers, the stability region with respect to the other two control gains is derived. Firstly, the boundaries of the fractional-order PIλDμcontrollers are determined by using the D-decomposition method. Then, an analytical approach is presented to judge which region is the stability one among a lot of areas divided by the resultant boundaries. In comparison with other relevant methods, the main advantage of the proposed method lies in that it can effectively avoid choosing one point from each divided area and finding the stability region of the fractional-order PID controller by testing the system stability corresponding to each chosen point. Moreover, a special phenomenon is revealed: if λ + μ ≠ 2, the boundaries of the stability region in ki-kdplane are the curves for a given k value; otherwise, the stability regions in ki-kdplane are convex polygons. A numerical example is presented to check the validity of the proposed method. The proposed method can be applied to the fractional-order system free of the detailed model and only the frequency response data of the fractional-order system is required. Qunhong Wu, Linlin Ou, Hongjie Ni, Weidong Zhang 0004 |
ICARCV | 2 |