VLDB 2026 Research / reviewers in the wild / expert
Qihao Liu
dblp:158/2755
· DBLP profile ↗
38ranked-venue papers
11as first author
38since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 11 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Integrated multi-robot scheduling for collaborative processing and autonomous mobility: A knowledge-guided spatiotemporal evolutionary approach
Qingsong Fan, Xinyu Li 0001, Chunjiang Zhang, Qihao Liu, Liang Gao 0001 |
Expert Syst. Appl. | 4 |
| 2026 | 4D-Animal: Freely Reconstructing Animatable 3D Animals from VideosabstractReconstructing animatable 3D animals from videos traditionally depends on sparse semantic keypoints to fit parametric models. Acquiring these keypoints is labor-intensive, and detectors trained on limited animal datasets are often unreliable. We propose 4D-Animal, a keypoint-free framework that reconstructs animatable 3D animals directly from videos. Our method employs a dense feature network to map 2D image representations to SMAL parameters, improving both efficiency and stability. Additionally, we introduce a hierarchical alignment strategy that leverages silhouette, part-level, pixel-level, and temporal cues from pretrained 2D models, ensuring accurate and temporally coherent reconstructions. Extensive experiments demonstrate that 4D-Animal outperforms both model-based and model-free baselines on dog dataset. Moreover, the high-quality 3D assets generated by our method can benefit other 3D tasks, underscoring its potential for large-scale applications. The code is released at https://github.com/zhongshsh/4D-Animal. Shanshan Zhong, Zehan Zheng, Zhongzhan Huang, Wufei Ma, Guofeng Zhang 0025, Qihao Liu, Alan L. Yuille, Jieneng Chen |
WACV | 7 |
| 2026 | Adaptive quantum differential evolution with experience-guided learning for integrated production and collaborative mobile robot scheduling
Qingsong Fan, Liang Gao 0001, Xinyu Li 0001, Chunjiang Zhang, Qihao Liu |
Expert Syst. Appl. | 5 |
| 2026 | Feedback-Driven Population Self-Evolution Framework for Dispatching Rule Generation in Dynamic Job Shop via Knowledge Distillation
Zhengqi Shi, Qihao Liu, Xinyu Li 0001, Liang Gao 0001 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2026 | Constraint Programming for AGV and Machine Integrated Scheduling Problem in Flexible Manufacturing SystemabstractThe finite resources of automated guided vehicles (AGVs) and machines in a flexible manufacturing system necessitate the integrated scheduling of production and transportation tasks to minimize delays in the production process. Constraint programming (CP) has demonstrated strong solving capabilities in complex shop scheduling problems. However, existing CP models exhibit significant limitations, typically yielding suboptimal solutions in specific scenarios. To address these challenges, this paper introduces a novel CP model that consistently delivers correct optimal solutions across all scenarios. First, the interrelationships among the four key decision sub-problems in AGV and machine integrated scheduling for flexible manufacturing systems are thoroughly analyzed. Next, based on the above analysis and leveraging the presence of transportation tasks, a new CP model is proposed to efficiently handle special cases where jobs do not require transportation. Finally, the model is benchmarked against state-of-the-art methods across three benchmarks and validated through a real-world case study. The results show that the proposed model outperforms existing approaches in both solution quality and efficiency. Notably, the proposed model updates the best-known solutions for the EX72 and EX84 instances, and proves the optimality of all EX instances for the first time. You-Jie Yao 0001, Qihao Liu, Xinyu Li 0001, Liang Gao 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2026 | A Knowledge-Enhanced Evolutionary Multitasking Memetic Algorithm for Multimodal Multiobjective Flexible Job Shop Scheduling Considering SpeedabstractMost research on flexible job shop scheduling assumes constant processing speeds. However, in real production, machines need to operate at variable speeds to achieve energy-efficient scheduling, which requires balancing multiobjective between production efficiency and green development. Such tradeoffs thus trigger the phenomenon in which massive solutions converge to identical objective values (i.e., the multimodal property), which is often neglected in scheduling problems. To address the above challenges, this work introduces a knowledge-enhanced evolutionary multitasking memetic algorithm (KEMMA) to solve the multimodal multiobjective flexible job shop scheduling problem considering speed (MMFJSP-S). First, self-paced learning motivated us to construct a simple auxiliary task and employ an evolutionary multitasking (EMT) framework to tackle the complex MMFJSP-S. Moreover, a knowledge enhancement and explicit transfer strategy is designed to reduce the effects of negative transfer by reinforcing and sharing beneficial knowledge across tasks. Finally, a mapping transformation mechanism is proposed to handle the multimodal property of the MMFJSP-S in the decision space. By comparing with ten advanced algorithms, the experimental results verify the remarkable superiority of the proposed KEMMA in solving MMFJSP-S and reveal the significance of studying the multimodal property. Xinyu Li 0001, Liang Gao 0001, Qihao Liu, Qingsong Fan |
IEEE Trans. Cybern. | 4 |
| 2026 | Automatic Programming via Large Language Models With Population Self-Evolution for Dynamic Fuzzy Job Shop Scheduling ProblemabstractHeuristic dispatching rules (HDRs) are widely used for solving the dynamic fuzzy job shop scheduling problem (DFJSSP). However, their performance is highly sensitive to specific scenarios and often necessitates expert customization. To overcome this, automated design methods like genetic programming (GP) and gene expression programming (GEP) have been proposed. Despite their success, these methods face challenges, such as high randomness in the search process. Recently, the combination of large language models (LLMs) with evolutionary algorithms has opened new possibilities for prompt engineering and automated algorithm design. To improve the ability of LLMs in automatic HDR design, this paper introduces a novel population self-evolutionary (SeEvo) framework, which draws inspiration from the self-reflective design strategies employed by human experts. Notably, this framework employs a novel teacher-student learning mechanism, allowing the LLM (student) to generate robust HDRs. Guided by a teacher model with complete knowledge of actual processing times, the student learns to infer fuzzy uncertainties from historical deviations, enabling it to effectively anticipate and adapt to fuzzy impacts. Experimental results demonstrate that SeEvo significantly outperforms GP, GEP, deep reinforcement learning (DRL) methods, and more than ten commonly used HDRs from the literature, particularly in previously unseen and dynamic scenarios. Qihao Liu, Xinyu Li 0001, Liang Gao 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2025 | A Discrete Grey Wolf Optimizer with an Active-Decoding Strategy for Reconfigurable Manufacturing System Scheduling ProblemabstractReconfigurable manufacturing systems offer enhanced flexibility to adapt to rapidly changing market demands. However, the reconfigurability of equipment introduces significant challenges to production scheduling, complicating optimization. This paper addresses the scheduling problem in reconfigurable manufacturing systems and proposes a discrete grey wolf optimizer algorithm with an active-decoding strategy (DGWO). A novel operation-configuration encoding scheme is proposed to comprehensively represent the solution space, accompanied by an active-decoding strategy that maximizes solution exploration and minimizes idle time. In the GWO, two crossover operators are introduced to enhance the search space, while the random walk strategy is introduced to prevent the algorithm from falling into premature convergence. Additionally, four neighborhood structures are defined based on the encoding space, and an efficient randomized enhanced local search is developed based on these structures to improve the algorithm's exploitation capability. The proposed DGWO algorithm is evaluated on 60 benchmark instances and compared with several related algorithms, demonstrating superior effectiveness and convergence performance. Cuiyu Wang, Xinyu Li 0001, Qihao Liu, Yiping Gao, Liang Gao 0001 |
CSCWD | 4 |
| 2025 | A Dense Pixel-Based Genetic Algorithm for Additive Manufacturing Scheduling ProblemabstractAdditive manufacturing has revolutionized the way to design and manufacture products by enabling complex geometries and on-demand production. However, 3D printing without scheduling is time-consuming and space-inefficient. To address these issues, a dense pixel-based genetic algorithm (DPGA) is proposed. To fast characterize 3D parts, a tolerant pixel matrix (TPM) is adopted for abstraction. Based on the TPM, a double-layer encoding scheme is designed to represent the solutions. An active decoding strategy is designed to maximize space utilization, which can improve the quality of schemes decoded with the same encoding. In the section of operator design, an initialization strategy based on load balancing is developed, which can effectively improve the quality of initial population. Additionally, a population rebirth mechanism is designed to efficiently escape from local optima. Computational experiments demonstrate the effectiveness of DPGA in solving additive manufacturing scheduling problems with varying sizes and complexities. The proposed DPGA outperforms traditional genetic algorithm and other state-of-the-art methods in terms of processing time and packing density. Zipeng Yang, Xinyu Li 0001, Qihao Liu, Chunjiang Zhang, Liang Gao 0001 |
CSCWD | 3 |
| 2025 | Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality EvolutionabstractDiffusion models, and their generalization, flow matching, have had a remarkable impact on the field of media generation. Here, the conventional approach is to learn the complex mapping from a simple source distribution of Gaussian noise to the target media distribution. For cross-modal tasks such as text-to-image generation, this same mapping from noise to image is learnt whilst including a conditioning mechanism in the model. One key and thus far relatively unexplored feature of flow matching is that, unlike Diffusion models, they are not constrained for the source distribution to be noise. Hence, in this paper, we propose a paradigm shift, and ask the question of whether we can instead train flow matching models to learn a direct mapping from the distribution of one modality to the distribution of another, thus obviating the need for both the noise distribution and conditioning mechanism. We present a general and simple framework, CrossFlow, for cross-modal flow matching. We show the importance of applying Variational Encoders to the input data, and introduce a method to enable Classifier-free guidance. Surprisingly, for text-to-image, CrossFlow with a vanilla transformer without cross attention slightly outperforms standard flow matching, and we show that it scales better with training steps and model size, while also allowing for interesting latent arithmetic which results in semantically meaningful edits in the output space. To demonstrate the generalizability of our approach, we also show that CrossFlow is on par with or outperforms the state-of-the-art for various cross-modal / intra-modal mapping tasks, viz. image captioning, depth estimation, and image super-resolution. We hope this paper contributes to accelerating progress in cross-modal media generation. Qihao Liu, Xi Yin 0001, Alan L. Yuille, Mannat Singh |
CVPR | 1 |
| 2025 | FlowTok: Flowing Seamlessly Across Text and Image TokensabstractBridging different modalities lies at the heart of cross-modality generation. While conventional approaches treat the text modality as a conditioning signal that gradually guides the denoising process from Gaussian noise to the target image modality, we explore a much simpler paradigm-directly evolving between text and image modalities through flow matching. This requires projecting both modalities into a shared latent space, which poses a significant challenge due to their inherently different representations: text is highly semantic and encoded as 1D tokens, whereas images are spatially redundant and represented as 2D latent embeddings. To address this, we introduce FlowTok, a minimal framework that seamlessly flows across text and images by encoding images into a compact 1D token representation. Compared to prior methods, this design reduces the latent space size by 3.3x at an image resolution of 256, eliminating the need for complex conditioning mechanisms or noise scheduling. Moreover, FlowTok naturally extends to image-to-text generation under the same formulation. With its streamlined architecture centered around compact 1D tokens, FlowTok is highly memory-efficient, requires significantly fewer training resources, and achieves much faster sampling speeds-all while delivering performance comparable to state-of-the-art models. Code is available at https://github.com/TACJu/FlowTok. Ju He, Qihang Yu, Qihao Liu, Liang-Chieh Chen |
ICCV | 3 |
| 2025 | Automatic MILP Model Construction for Multi-Robot Task Allocation and Scheduling Based on Large Language ModelsabstractWith the accelerated development of Industry 4.0, intelligent manufacturing systems increasingly require efficient task allocation and scheduling in multi-robot systems. However, existing methods rely on domain expertise and face challenges in adapting to dynamic production constraints. Additionally, enterprises have high privacy requirements for production scheduling data, which prevents the use of cloud-based large language models (LLMs) for solution development. To address these challenges, there is an urgent need for an automated modeling solution that meets data privacy requirements. This study proposes a knowledge-augmented mixed integer linear programming (MILP) automated formulation framework, integrating local LLMs with domain-specific knowledge bases to generate executable code from natural language descriptions automatically. The framework employs a knowledge-guided DeepSeek-R1-Distill-Qwen-32B model to extract complex spatiotemporal constraints (82% average accuracy) and leverages a supervised fine-tuned Qwen2.5-Coder-7B-Instruct model for efficient MILP code generation (90% average accuracy). Experimental results demonstrate that the framework successfully achieves automatic modeling in the aircraft skin manufacturing case while ensuring data privacy and computational efficiency. This research provides a low-barrier and highly reliable technical path for modeling in complex industrial scenarios. Mingming Peng, Zhengqi Shi, Qihao Liu, Xinyu Li 0001, Liang Gao 0001 |
IROS | 6 |
| 2025 | SpatialReasoner: Towards Explicit and Generalizable 3D Spatial ReasoningabstractDespite recent advances on multi-modal models, 3D spatial reasoning remains a challenging task for state-of-the-art open-source and proprietary models. Recent studies explore data-driven approaches and achieve enhanced spatial reasoning performance by fine-tuning models on 3D-related visual question-answering data. However, these methods typically perform spatial reasoning in an implicit manner and often fail on questions that are trivial to humans, even with long chain-of-thought reasoning. In this work, we introduce SpatialReasoner, a novel large vision-language model (LVLM) that addresses 3D spatial reasoning with explicit 3D representations shared between multiple stages--3D perception, computation, and reasoning. Explicit 3D representations provide a coherent interface that supports advanced 3D spatial reasoning and improves the generalization ability to novel question types. Furthermore, by analyzing the explicit 3D representations in multi-step reasoning traces of SpatialReasoner, we study the factual errors and identify key shortcomings of current LVLMs. Results show that our SpatialReasoner achieves improved performance on a variety of spatial reasoning benchmarks, outperforming Gemini 2.0 by 9.2% on 3DSRBench, and generalizes better when evaluating on novel 3D spatial reasoning questions. Our study bridges the 3D parsing capabilities of prior visual foundation models with the powerful reasoning abilities of large language models, opening new directions for 3D spatial reasoning. Wufei Ma, Yu-Cheng Chou, Qihao Liu, Xingrui Wang, Celso de Melo, Jianwen Xie, Alan L. Yuille |
NeurIPS | 3 |
| 2025 | Constraint programming-based layered method for integrated process planning and scheduling in extensive flexible manufacturing
Xinyu Li 0001, Liang Gao 0001, Qihao Liu |
Adv. Eng. Informatics | 4 |
| 2025 | Tackling dual-resource flexible job shop scheduling problem in the production line reconfiguration scenario: An efficient meta-heuristic with critical path-based neighborhood search
Xinyu Li 0001, Liang Gao 0001, Qihao Liu |
Adv. Eng. Informatics | 4 |
| 2025 | A heterogeneous graph attention-enhanced deep reinforcement learning framework for flexible job shop scheduling problem with variable sublotsabstractVariable lot-sizing is an effective approach to improve production efficiency by splitting an operation into several sublots, which has been widely applied in flexible manufacturing systems. However, the flexibility of lot-sizing will dramatically expand the solution space, leading to excessive computation time in converging to the relative optimum. To address this challenge, this paper introduces an end-to-end deep reinforcement learning framework based on heterogeneous graph attention mechanisms (HGADRL) for flexible job shop scheduling problem with variable sublots. Unlike traditional heuristic and rule-based methods, HGADRL dynamically learns the high-dimensional nature, providing a more generalizable solution in a very short time. In HGADRL, a modified heterogeneous disjunctive graph is designed to represent the dynamic scheduling status, including operation selection and sublot division. A dual-scale graph attention network combined with two interconnected attention modules is developed, enabling the precise capture of complex interdependencies between heterogeneous vertices. This approach can significantly enhance the agent's ability to self-learn and evolve optimal policies. By leveraging local and global features extracted through the graph attention network, an actor-critic network is employed for high-quality scheduling in different states. Experimental results demonstrate that the proposed method outperforms the 12 mixed priority dispatching rules, two meta-heuristic methods and two deep reinforcement learning methods in all 500 synthetic instances. Additionally, the proposed method outperforms all compared methods across 16 unseen scales of instances and four real-world instances, demonstrating its strong generalization capabilities. Zipeng Yang, Xinyu Li 0001, Liang Gao 0001, Qihao Liu |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | A Lightweight Multi-Scale Feature Enhancement Network for Person Re- IDabstractABSTRACT Recent advances in deep learning have significantly propelled progress in person re‐identification. However, many current solutions often prioritize architectural optimization. Although this approach has led to considerable performance improvements, it may inadvertently overlook the critical challenge of suppressing interference from complex background noise. To explore a path toward addressing this aspect, we propose a Multi‐scale Feature Enhancement Network (MSFENet). Our approach includes a spatial‐frequency fusion module designed to guide the network's attention toward pedestrian‐specific regions. Furthermore, the incorporation of frequency‐domain cues is intended to facilitate the capture of fine‐grained details, thereby potentially enhancing robustness. We also design a Multi‐Granularity Fusion (MGFusion) module to help alleviate overfitting and information loss during feature interaction. Experimental results indicate that MSFENet achieves competitive performance across evaluated tasks on the Market1501 and MSMT17 datasets, as well as under cross‐domain settings (MSMT17 → Market1501). Qihao Liu, Pengyuan Shen, Tiancun Guo |
Expert Syst. J. Knowl. Eng. | 1 |
| 2025 | A Flexible Job Shop Scheduling Problem Considering On-Site Machining Fixtures: A Case Study From Customized Manufacturing EnterpriseabstractThe joint optimization of production scheduling and resource constraints is critical to modern manufacturing systems. The number of auxiliary resources (fixtures) is usually insufficient in customized manufacturing. Thus, on-site machining fixtures (Type II fixtures) should be prepared in the workshop to reduce the shortage. In this way, Type II fixtures are production tasks and resource constraints, while Type I fixtures are only resource constraints. The existing studies mainly concentrate on Type I fixtures, whereas the research on Type II fixtures is limited. Therefore, this paper focuses on a flexible job shop with on-site machining fixtures (FJSP-F). Firstly, a mathematical model is developed to minimize total weight tardiness (TWT). Secondly, a job-fixture-machine (JFM) encoding and novel decoding methods are presented to obtain a feasible schedule solution. Thirdly, an improved genetic algorithm (IGA4F) with problem-specific variable neighborhood search (PVNS) is proposed to balance the exploration and exploitation. Finally, the proposed algorithm is tested on 20 instances with comparison algorithms. The results demonstrate that IGA4F is a competitive algorithm in large-scale instances. From the case study results, the performance gains of the TWT and makespan obtained by IGA4F are 49.27% and 28.94% compared to the original schedule solution. Note to Practitioners—The integrated problem of fixture allocation and production scheduling is widespread in highly customized manufacturing enterprises, such as aerospace and shipbuilding. A well-balanced allocation between fixtures and machines can facilitate productivity and resource utilization. In general, Type I fixtures can be used directly if they are idle, and these fixtures are treated as resource constraints. However, due to the limited number of Type II fixtures, they are only available when finished in the workshop. Hence, Type II fixtures are considered production tasks and resource constraints, and the number of these fixtures is dynamic during the production cycle. Therefore, it is necessary for enterprise managers to investigate the effect of Type II fixtures on production scheduling. This paper proposes novel encoding and decoding methods to represent the solution and objective spaces. The evolutionary-based algorithm is proposed to solve the daily order of a real-world enterprise. The obtained results from the proposed algorithm can guide the managers to promote the workshop’s productivity. Jiahang Li 0003, Xinyu Li 0001, Liang Gao 0001, Qihao Liu |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | Integrated Nesting and Scheduling for SLA 3D Printing: A Pixel-Based Evolutionary Algorithm With Convolutional AccelerationabstractAdditive manufacturing (AM) constructs complex products through the layer-by-layer deposition of materials. AM enables complex geometry fabrication but faces spatiotemporal optimization challenges: maximizing space utilization and minimizing printing time, requiring intelligent nesting of products and scheduling of resources. The integration of nesting and scheduling further expands the solution space, but existing methods frequently overlook essential geometric intricacies in irregular products. This paper proposes a novel pixel-based grey wolf optimizer algorithm (PGWO) to improve packing density and reduce time cost in stereolithography (SLA) printing. In the proposed PGWO, point-cloud pixelization is employed to simplify 3D irregular parts. Each part independently performs autonomous orientation to minimize local space occupancy at a low cost. Based on the time-frequency domain conversion, a convolutionally accelerated localization strategy (Cals) is proposed to improve the speed of nesting. For efficient scheduling, a bottleneck-balanced local search phase is designed with two operators. By targeting critical bottlenecks in printing, two operators can effectively reduce the frequency of layer changes, further optimizing local optimal solutions. PGWO is evaluated on 70 instances with diverse scales, and shows significant superiority over other state-of-the-art methods in over 88% of instances. The results demonstrate its superior performance in reducing printing time and enhancing packing density. Zipeng Yang, Xinyu Li 0001, Qihao Liu, Liang Gao 0001 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | A Novel Mathematical Model for the Flexible Job-Shop Scheduling Problem With Limited Automated Guided VehiclesabstractAutomated Guided Vehicles (AGVs) have found widespread application in discrete manufacturing systems. In flexible job-shop environments, the integrated scheduling of machines and AGVs is a significant research direction to improve the productivity. However, the existing mathematical model assigns non-existent transport tasks to the corresponding AGVs, resulting in poor performance. To tackle this weakness, this paper proposes a novel mixed integer linear programming (MILP) model. Firstly, the flexible job-shop scheduling problem with limited AGVs (FJSPLA) is decomposed into four sub-problems, and the interactions and dependencies between the sub-problems are elaborated. Secondly, the existence of transport tasks is explained in detail based on the disjunctive graph model. Subsequently, a more efficient MILP model is proposed, leveraging insights from the four sub-problems and the disjunctive graph model. Finally, comparison experiments are conducted, encompassing two benchmarks (FJSPT and EX), along with a real-world case. The proposed model exhibits a more streamlined formulation with fewer decision variables and constraints in comparison to existing models. It successfully proves optimality for the most challenging instance FJSPT7 as well as 15 instances in EX benchmark. Compared with the existing model, the experimental results not only demonstrate the effectiveness and superior performance of the proposed model but also show the practicality in addressing real workshop problems.Note to Practitioners—Automated guided vehicles (AGVs) have been extensive application in various industries, prompting practitioners to integrate the scheduling of machines and AGVs during production planning. To address this realistic production problem, this study develops a novel MILP model. Through comprehensive analyses, integrated scheduling is decomposed into four sub-problems and the correlations between the four sub-problems are accurately presented. For each sub-problem, we establish the corresponding mathematical formulations. Practitioners can use the work in this paper to clearly understand the integrated scheduling problem, and can easily use the optimization software to solve the model. As in our case study, practitioners collate the production information according to their workshop, and the model can give the optimal solution for integrated scheduling in an acceptable time. The optimal solution obtained from the proposed model can guide the practitioners to maximize the productivity of the workshop. You-Jie Yao 0001, Qihao Liu, Xinyu Li 0001, Yanbin Yu, Liang Gao 0001, Wei Zhou 0070 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | A Dual-Space Artificial Bee Colony Algorithm Integrating Configuration-Coupled Heterogeneous Disjunctive Graph for Scheduling Problem in Reconfigurable Manufacturing SystemsabstractReconfigurable manufacturing systems root mean square (rms) offer high flexibility, enabling efficient adaptation to changing market demands. However, this reconfigurability significantly increases the complexity of production scheduling. This article addresses the rms scheduling problem (RMSSP) to minimize the makespan. A configuration-coupled heterogeneous disjunctive graph (CHDG) model is proposed to represent feasible solutions by incorporating machine-configuration arcs and reconfiguration nodes, capturing reconfiguration processes and operation statuses. Feasibility theorems for intramachine and intermachine movements are developed, and six solution-space clipping strategies are introduced to reduce invalid searches. Based on these, a dual-space artificial bee colony (DABC) algorithm is proposed, featuring a novel operation-configuration encoding scheme and configuration-associated active-decoding strategy to maximize the potential of encoding. A hierarchical crossover operator and CHDG-based neighborhood search operators collaboratively explore the encoding and disjunctive graph (DG) spaces for efficient optimization. Numerical experiment results on 60 benchmark instances show that integrating CHDG significantly improves DABC’s performance in solving RMSSP. In addition, the six clipping theorems reduce invalid intramachine neighborhood searches by 55.3%. Qihao Liu, Zipeng Yang, Liang Gao 0001, Xinyu Li 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2025 | Graph-Based Dual-Agent Deep Reinforcement Learning for Dynamic Human-Machine Hybrid Reconfiguration Manufacturing SchedulingabstractHuman–machine hybrid reconfiguration manufacturing is an emerging paradigm in the field of precision equipment production and can greatly improve the production capability of the workshop. However, numerous complex constraints and a dynamic environment make reasonable scheduling very difficult. To this end, this article studies the dynamic human–machine hybrid reconfiguration manufacturing scheduling problem (DHMRSP) and proposes a novel deep reinforcement learning (DRL) scheduling method. Specifically, a dual-agent Markov decision process (MDP) is established, which can handle seven complex constraints and three disturbance events. Then, a heterogeneous competition graph attention network (HCGAN) is designed, where the meta-path-based subgraph conversion reflects the resource-operation competition, and three modules use node-level attention and semantic-level attention to realize important information embedding. Afterward, a dual proximal policy optimization (PPO) algorithm with HCGAN and mixed action space (HM-DPPO) is proposed, where the allocation agent and reconfiguration agent achieve collaborative learning by taking joint action and sharing graph embeddings and reward. Experimental results prove that the proposed approach outperforms rules, genetic programming (GP), and three DRL methods on different instances and can effectively handle various disturbance events. Qihao Liu, Chunjiang Zhang, Xinyu Li 0001, Liang Gao 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2024 | DIRECT-3D: Learning Direct Text-to-3D Generation on Massive Noisy 3D DataabstractWe present DIRECT-3D, a diffusion-based 3D generative model for creating high-quality 3D assets (represented by Neural Radiance Fields) from text prompts. Unlike recent 3D generative models that rely on clean and well-aligned 3D data, limiting them to single or few-class generation, our model is directly trained on extensive noisy and unaligned ‘in-the-wild’ 3D assets, mitigating the key challenge (i.e., data scarcity) in large-scale 3D generation. In particular, DIRECT-3D is a tri-plane diffusion model that integrates two innovations: 1) A novel learning framework where noisy data are filtered and aligned automatically during the training process. Specifically, after an initial warm-up phase using a small set of clean data, an iterative optimization is introduced in the diffusion process to explicitly estimate the 3D pose of objects and select beneficial data based on conditional density. 2) An efficient 3D representation that is achieved by disentangling object geometry and color features with two separate conditional diffusion models that are optimized hierarchically. Given a prompt input, our model generates high-quality, high-resolution, realistic, and complex 3D objects with accurate geometric details in seconds. We achieve state-of-the-art performance in both single-class generation and text-to-3D generation. We also demonstrate that DIRECT-3D can serve as a useful 3D geometric prior of objects, for example to alleviate the well-known Janus problem in 2D-lifting methods such as DreamFusion. The code and models are available for research purposes at: https://github.com/qihao067/direct3d. Qihao Liu, Yi Zhang 0099, Song Bai 0001, Adam Kortylewski, Alan L. Yuille |
CVPR | 1 |
| 2024 | General Object Foundation Model for Images and Videos at ScaleabstractWe present GLEE in this work, an object-level foundation model for locating and identifying objects in images and videos. Through a unified framework, GLEE accomplishes detection, segmentation, tracking, grounding, and identification of arbitrary objects in the open world scenario for various object perception tasks. Adopting a cohesive learning strategy, GLEE acquires knowledge from diverse data sources with varying supervision levels to formu-late general object representations, excelling in zero-shot transfer to new data and tasks. Specifically, we employ an image encoder, text encoder, and visual prompter to handle multimodal inputs, enabling to simultaneously solve various object-centric downstream tasks while maintaining state-of-the-art performance. Demonstrated through extensive training on over five million images from diverse benchmarks, GLEE exhibits remarkable versatility and improved generalization performance, efficiently tack-ling downstream tasks without the need for task-specific adaptation. By integrating large volumes of automatically labeled data, we further enhance its zero-shot generalization capabilities. Additionally, GLEE is capable of being integrated into Large Language Models, serving as a foun-dational model to provide universal object-level information for multimodal tasks. We hope that the versatility and universality of our method will mark a significant step in the development of efficient visual foundation models for AGI systems. The models and code are released at https://github.com/FoundationVision/GLEE. Junfeng Wu 0003, Yi Jiang 0009, Qihao Liu, Zehuan Yuan, Xiang Bai, Song Bai 0001 |
CVPR | 3 |
| 2024 | Rethinking Video-Text Understanding: Retrieval from Counterfactually Augmented Data
Wufei Ma, Kai Li 0016, Zhongshi Jiang, Moustafa Meshry, Qihao Liu, Christian Häne, Alan L. Yuille |
ECCV (13) | 5 |
| 2024 | Discovering Failure Modes of Text-guided Diffusion Models via Adversarial SearchabstractText-guided diffusion models (TDMs) are widely applied but can fail unexpectedly. Common failures include: _(i)_ natural-looking text prompts generating images with the wrong content, or _(ii)_ different random samples of the latent variables that generate vastly different, and even unrelated, outputs despite being conditioned on the same text prompt. In this work, we aim to study and understand the failure modes of TDMs in more detail. To achieve this, we propose SAGE, the first adversarial search method on TDMs that systematically explores the discrete prompt space and the high-dimensional latent space, to automatically discover undesirable behaviors and failure cases in image generation. We use image classifiers as surrogate loss functions during searching, and employ human inspections to validate the identified failures. For the first time, our method enables efficient exploration of both the discrete and intricate human language space and the challenging latent space, overcoming the gradient vanishing problem. Then, we demonstrate the effectiveness of SAGE on five widely used generative models and reveal four typical failure modes that have not been systematically studied before: (1) We find a variety of natural text prompts that generate images failing to capture the semantics of input texts. We further discuss the underlying causes and potential solutions based on the results. (2) We find regions in the latent space that lead to distorted images independent of the text prompt, suggesting that parts of the latent space are not well-structured. (3) We also find latent samples that result in natural-looking images unrelated to the text prompt, implying a possible misalignment between the latent and prompt spaces. (4) By appending a single adversarial token embedding to any input prompts, we can generate a variety of specified target objects, with minimal impact on CLIP scores, demonstrating the fragility of language representations. Qihao Liu, Adam Kortylewski, Yutong Bai, Song Bai 0001, Alan L. Yuille |
ICLR | 1 |
| 2024 | Generating Images with 3D Annotations Using Diffusion ModelsabstractDiffusion models have emerged as a powerful generative method, capable of producing stunning photo-realistic images from natural language descriptions. However, these models lack explicit control over the 3D structure in the generated images. Consequently, this hinders our ability to obtain detailed 3D annotations for the generated images or to craft instances with specific poses and distances. In this paper, we propose 3D Diffusion Style Transfer (3D-DST), which incorporates 3D geometry control into diffusion models. Our method exploits ControlNet, which extends diffusion models by using visual prompts in addition to text prompts. We generate images of the 3D objects taken from 3D shape repositories~(e.g., ShapeNet and Objaverse), render them from a variety of poses and viewing directions, compute the edge maps of the rendered images, and use these edge maps as visual prompts to generate realistic images. With explicit 3D geometry control, we can easily change the 3D structures of the objects in the generated images and obtain ground-truth 3D annotations automatically. This allows us to improve a wide range of vision tasks, e.g., classification and 3D pose estimation, in both in-distribution (ID) and out-of-distribution (OOD) settings. We demonstrate the effectiveness of our method through extensive experiments on ImageNet-100/200, ImageNet-R, PASCAL3D+, ObjectNet3D, and OOD-CV. The results show that our method significantly outperforms existing methods, e.g., 3.8 percentage points on ImageNet-100 using DeiT-B. Our code is available at <https://ccvl.jhu.edu/3D-DST/> Wufei Ma, Qihao Liu, Jiahao Wang 0001, Angtian Wang, Xiaoding Yuan, Yi Zhang 0099, Zihao Xiao 0001, Guofeng Zhang 0020, Beijia Lu, Ruxiao Duan, Yongrui Qi, Adam Kortylewski, Yaoyao Liu 0001, Alan L. Yuille |
ICLR | 2 |
| 2024 | Alleviating Distortion in Image Generation via Multi-Resolution Diffusion Models and Time-Dependent Layer NormalizationabstractThis paper presents innovative enhancements to diffusion models by integrating a novel multi-resolution network and time-dependent layer normalization.
Diffusion models have gained prominence for their effectiveness in high-fidelity image generation.
While conventional approaches rely on convolutional U-Net architectures, recent Transformer-based designs have demonstrated superior performance and scalability.
However, Transformer architectures, which tokenize input data (via "patchification"), face a trade-off between visual fidelity and computational complexity due to the quadratic nature of self-attention operations concerning token length.
While larger patch sizes enable attention computation efficiency, they struggle to capture fine-grained visual details, leading to image distortions.
To address this challenge, we propose augmenting the **Di**ffusion model with the **M**ulti-**R**esolution network (DiMR), a framework that refines features across multiple resolutions, progressively enhancing detail from low to high resolution.
Additionally, we introduce Time-Dependent Layer Normalization (TD-LN), a parameter-efficient approach that incorporates time-dependent parameters into layer normalization to inject time information and achieve superior performance.
Our method's efficacy is demonstrated on the class-conditional ImageNet generation benchmark, where DiMR-XL variants surpass previous diffusion models, achieving FID scores of 1.70 on ImageNet $256 \times 256$ and 2.89 on ImageNet $512 \times 512$. Our best variant, DiMR-G, further establishes a state-of-the-art 1.63 FID on ImageNet $256 \times 256$. Qihao Liu, Zhanpeng Zeng, Ju He, Qihang Yu, Xiaohui Shen, Liang-Chieh Chen |
NeurIPS | 1 |
| 2024 | ImageNet3D: Towards General-Purpose Object-Level 3D UnderstandingabstractA vision model with general-purpose object-level 3D understanding should be capable of inferring both 2D (e.g., class name and bounding box) and 3D information (e.g., 3D location and 3D viewpoint) for arbitrary rigid objects in natural images. This is a challenging task, as it involves inferring 3D information from 2D signals and most importantly, generalizing to rigid objects from unseen categories. However, existing datasets with object-level 3D annotations are often limited by the number of categories or the quality of annotations. Models developed on these datasets become specialists for certain categories or domains, and fail to generalize. In this work, we present ImageNet3D, a large dataset for general-purpose object-level 3D understanding. ImageNet3D augments 200 categories from the ImageNet dataset with 2D bounding box, 3D pose, 3D location annotations, and image captions interleaved with 3D information. With the new annotations available in ImageNet3D, we could (i) analyze the object-level 3D awareness of visual foundation models, and (ii) study and develop general-purpose models that infer both 2D and 3D information for arbitrary rigid objects in natural images, and (iii) integrate unified 3D models with large language models for 3D-related reasoning. We consider two new tasks, probing of object-level 3D awareness and open vocabulary pose estimation, besides standard classification and pose estimation. Experimental results on ImageNet3D demonstrate the potential of our dataset in building vision models with stronger general-purpose object-level 3D understanding. Our dataset and project page are available here: https://imagenet3d.github.io. Wufei Ma, Guofeng Zhang 0020, Qihao Liu, Guanning Zeng, Adam Kortylewski, Yaoyao Liu 0001, Alan L. Yuille |
NeurIPS | 3 |
| 2023 | PoseExaminer: Automated Testing of Out-of-Distribution Robustness in Human Pose and Shape EstimationabstractHuman pose and shape (HPS) estimation methods achieve remarkable results. However, current HPS bench-marks are mostly designed to test models in scenarios that are similar to the training data. This can lead to critical situations in real-world applications when the observed data differs significantly from the training data and hence is out-of-distribution (OOD). It is therefore important to test and improve the OOD robustness of HPS methods. To address this fundamental problem, we develop a simulator that can be controlled in a fine-grained manner using interpretable parameters to explore the manifold of images of human pose, e.g. by varying poses, shapes, and clothes. We introduce a learning-based testing method, termed PoseExaminer, that automatically diagnoses HPS algorithms by searching over the parameter space of human pose images to find the failure modes. Our strategy for exploring this high-dimensional parameter space is a multiagent reinforcement learning system, in which the agents collaborate to explore different parts of the parameter space. We show that our PoseExaminer discovers a variety of limitations in current state-of-the-art models that are relevant in real-world scenarios but are missed by current benchmarks. For example, it finds large regions of realistic human poses that are not predicted correctly, as well as reduced performance for humans with skinny and corpulent body shapes. In addition, we show that fine-tuning HPS methods by exploiting the failure modes found by PoseExaminer improve their robustness and even their performance on standard benchmarks by a significant margin. The code are available for research purposes at https://github.com/qihao0671PoseExaminer. Qihao Liu, Adam Kortylewski, Alan L. Yuille |
CVPR | 1 |
| 2023 | InstMove: Instance Motion for Object-centric Video SegmentationabstractDespite significant efforts, cutting-edge video segmentation methods still remain sensitive to occlusion and rapid movement, due to their reliance on the appearance of objects in the form of object embeddings, which are vulnerable to these disturbances. A common solution is to use optical flow to provide motion information, but essentially it only considers pixel-level motion, which still relies on appearance similarity and hence is often inaccurate under occlusion and fast movement. In this work, we study the instance-level motion and present InstMove, which stands for Instance Motion for Object-centric Video Segmentation. In comparison to pixel-wise motion, Inst-Move mainly relies on instance-level motion information that is free from image feature embeddings, and features physical interpretations, making it more accurate and robust toward occlusion and fast-moving objects. To better fit in with the video segmentation tasks, InstMove uses instance masks to model the physical presence of an object and learns the dynamic model through a memory network to predict its position and shape in the next frame. With only a few lines of code, InstMove can be integrated into current SOTA methods for three different video segmentation tasks and boost their performance. Specifically, we improve the previous arts by 1.5 AP on OVIS dataset, which features heavy occlusions, and 4.9 AP on YouTube VIS-Long dataset, which mainly contains fast moving objects. These results suggest that instance-level motion is robust and accurate, and hence serving as a powerful solution in complex scenarios for object-centric video segmentation. Qihao Liu, Junfeng Wu 0003, Yi Jiang 0009, Xiang Bai, Alan L. Yuille, Song Bai 0001 |
CVPR | 1 |
| 2023 | Animal3D: A Comprehensive Dataset of 3D Animal Pose and ShapeabstractAccurately estimating the 3D pose and shape is an essential step towards understanding animal behavior, and can potentially benefit many downstream applications, such as wildlife conservation. However, research in this area is held back by the lack of a comprehensive and diverse dataset with high-quality 3D pose and shape annotations. In this paper, we propose Animal3D, the first comprehensive dataset for mammal animal 3D pose and shape estimation. Animal3D consists of 3379 images collected from 40 mammal species, high-quality annotations of 26 key-points, and importantly the pose and shape parameters of the SMAL [50] model. All annotations were labeled and checked manually in a multi-stage process to ensure highest quality results. Based on the Animal3D dataset, we benchmark representative shape and pose estimation models at: (1) supervised learning from only the Animal3D data, (2) synthetic to real transfer from synthetically generated images, and (3) fine-tuning human pose and shape estimation models. Our experimental results demonstrate that predicting the 3D shape and pose of animals across species remains a very challenging task, despite significant advances in human pose estimation. Our results further demonstrate that synthetic pre-training is a viable strategy to boost the model performance. Overall, Animal3D opens new directions for facilitating future research in animal 3D pose and shape estimation, and is publicly available. Jiacong Xu, Yi Zhang 0099, Wufei Ma, Artur Jesslen, Pengliang Ji, Qixin Hu, Qihao Liu, Jiahao Wang 0001, Wei Ji 0011, Chen Wang 0049, Xiaoding Yuan, Prakhar Kaushik, Guofeng Zhang 0020, Jie Liu 0044, Yushan Xie, Yawen Cui, Alan L. Yuille, Adam Kortylewski |
ICCV | 9 |
| 2023 | A multi-population co-evolutionary algorithm for green integrated process planning and scheduling considering logistics system
Qihao Liu, Cuiyu Wang, Xinyu Li 0001, Liang Gao 0001 |
Eng. Appl. Artif. Intell. | 1 |
| 2022 | Learning Part Segmentation through Unsupervised Domain Adaptation from Synthetic VehiclesabstractPart segmentations provide a rich and detailed part-level description of objects. However, their annotation requires an enormous amount of work, which makes it difficult to apply standard deep learning methods. In this paper, we propose the idea of learning part segmentation through unsupervised domain adaptation (UDA) from synthetic data. We first introduce UDA-Part, a comprehensive part segmentation dataset for vehicles that can serve as an adequate benchmark for UDA11https://qliu24.github.io/udapart/. In UDA-Part, we label parts on 3D CAD models which enables us to generate a large set of annotated synthetic images. We also annotate parts on a number of real images to provide a real test set. Secondly, to advance the adaptation of part models trained from the synthetic data to the real images, we introduce a new UDA algorithm that leverages the object's spatial structure to guide the adaptation process. Our experimental results on two real test datasets confirm the superiority of our approach over existing works, and demonstrate the promise of learning part segmentation for general objects from synthetic data. We believe our dataset provides a rich testbed to study UDA for part segmentation and will help to significantly push forward research in this area. Qing Liu 0017, Adam Kortylewski, Zhishuai Zhang, Zizhang Li, Mengqi Guo, Qihao Liu, Xiaoding Yuan, Jiteng Mu, Weichao Qiu, Alan L. Yuille |
CVPR | 6 |
| 2022 | Explicit Occlusion Reasoning for Multi-person 3D Human Pose Estimation
Qihao Liu, Yi Zhang 0099, Song Bai 0001, Alan L. Yuille |
ECCV (5) | 1 |
| 2022 | In Defense of Online Models for Video Instance Segmentation
Junfeng Wu 0003, Qihao Liu, Yi Jiang 0009, Song Bai 0001, Alan L. Yuille, Xiang Bai |
ECCV (28) | 2 |
| 2021 | PNS: Population-Guided Novelty Search for Reinforcement Learning in Hard Exploration EnvironmentsabstractReinforcement Learning (RL) has made remarkable achievements, but it still suffers from inadequate exploration strategies, sparse reward signals, and deceptive reward functions. To alleviate these problems, a Population-guided Novelty Search (PNS) parallel learning method is proposed in this paper. In PNS, the population is divided into multiple sub-populations, each of which has one chief agent and several exploring agents. The chief agent evaluates the policies learned by exploring agents and shares the optimal policy with all sub-populations. The exploring agents learn their policies in collaboration with the guidance of the optimal policy and, simultaneously, upload their policies to the chief agent. To balance exploration and exploitation, the Novelty Search (NS) is employed in every chief agent to encourage policies with high novelty while maximizing per-episode performance. We apply PNS to the twin delayed deep deterministic (TD3) policy gradient algorithm. The effectiveness of PNS to promote exploration and improve performance in continuous control domains is demonstrated in the experimental section. Notably, PNS-TD3 achieves rewards that far exceed the SOTA methods in environments with sparse or delayed reward signals. We also demonstrate that PNS enables robotic agents to learn control policies directly from pixels for sparse-reward manipulation in both simulated and real-world settings. Qihao Liu |
IROS | 1 |
| 2021 | A Modified Genetic Algorithm With New Encoding and Decoding Methods for Integrated Process Planning and Scheduling ProblemabstractDue to the complementarity of the process planning and shop scheduling, their integration can greatly facilitate the development of the intelligent manufacturing system. In the last decade, the integrated process planning and scheduling (IPPS) problem has become a research hotspot in the manufacturing system area. It is an NP-hard problem and is more complicated than the job shop scheduling problem. Although some progress has been obtained in the IPPS field, there are still many unsolved open problems. In this article, the novel integrated encoding and decoding methods are proposed by considering the OR-node of the process network graph. Moreover, a modified genetic algorithm (MGA) is designed based on the proposed coding methods. The process planning and the scheduling parts can be represented simultaneously in one individual. As for the precedence constraints between operations, the specifically designed operators are able to guarantee the feasibility of the operation sequence during the searching procedure. Then, the superiority of MGA is verified by updating nine new records on 37 well-known open problems, four of them reach their lower bounds. In addition, the proposed algorithm is also tested on a real-world case from a nonstandard equipment workshop in a Chinese machine tool company, which produces a common module of a packaging machine. The results show that the proposed MGA can solve the real-world case better than the comparative algorithms. Qihao Liu, Xinyu Li 0001, Liang Gao 0001, Yingli Li |
IEEE Trans. Cybern. | 1 |