Shuqi Zhao

dblp:45/6764 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Systems, architecture and hardware · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 X-Drive: Cross-modality Consistent Multi-Sensor Data Synthesis for Driving Scenarios
abstract
Recent advancements have exploited diffusion models for the synthesis of either LiDAR point clouds or camera image data in driving scenarios. Despite their success in modeling single-modality data marginal distribution, there is an under- exploration in the mutual reliance between different modalities to describe com- plex driving scenes. To fill in this gap, we propose a novel framework, X-DRIVE, to model the joint distribution of point clouds and multi-view images via a dual- branch latent diffusion model architecture. Considering the distinct geometrical spaces of the two modalities, X-DRIVE conditions the synthesis of each modality on the corresponding local regions from the other modality, ensuring better alignment and realism. To further handle the spatial ambiguity during denoising, we design the cross-modality condition module based on epipolar lines to adaptively learn the cross-modality local correspondence. Besides, X-DRIVE allows for controllable generation through multi-level input conditions, including text, bounding box, image, and point clouds. Extensive results demonstrate the high-fidelity synthetic results of X-DRIVE for both point clouds and multi-view images, adhering to input conditions while ensuring reliable cross-modality consistency. Our code will be made publicly available at https://github.com/yichen928/X-Drive.
Yichen Xie 0002, Chenfeng Xu, Chensheng Peng, Shuqi Zhao, Nhat Ho, Alexander T. Pham, Mingyu Ding, Masayoshi Tomizuka
ICLR4
2025 PhyGrasp: Generalizing Robotic Grasping with Physics-informed Large Multimodal Models
abstract
Robotic grasping, crucial for robot interaction with objects, still struggles with counter-intuitive or long-tailed scenarios like uncommon materials and shapes. Humans, however, intuitively adjust grasps with their physics-informed interpretations of the object, using visual and linguistic cues. This work introduces PhyGrasp, a large multimodal model and dataset that enhance robotic manipulation by combining natural language and 3D point clouds using a bridge module to integrate these inputs. The language modality exhibits robust reasoning capabilities concerning the impacts of diverse physical properties on grasping, while the 3D modality comprehends object shapes and parts. With these two capabilities, PhyGrasp is able to accurately assess the physical properties of object parts and determine optimal grasping poses. Additionally, the model’s language comprehension enables human instruction interpretation, generating grasping poses that align with human preferences. To train PhyGrasp, we construct a dataset PhyPartNet with 195K object instances with varying physical properties and human preferences, alongside their corresponding language descriptions. Extensive experiments conducted in the simulation and on the real robots demonstrate that PhyGrasp achieves state-of-the-art performance, particularly in long-tailed cases, e.g., about 10% improvement in success rate over GraspNet. More demos and information are available on https://sites.google.com/view/phygrasp.
Dingkun Guo, Yuqi Xiang, Shuqi Zhao, Xinghao Zhu, Masayoshi Tomizuka, Mingyu Ding
IROS3
2025 Spatial-Spectral Dual Guided Network With Joint Attention for Pansharpening
abstract
Deep learning (DL) has been widely recognized for its strong feature representation capability, making it a promising technique to improve pansharpening methods. However, existing DL-based methods commonly extract the spectral information and spatial information from the high-resolution panchromatic (PAN) images and low-resolution multispectral (MS) images, respectively. This separation limits the effective extraction and integration of potential spectral-spatial information, ultimately reducing the quality of the generated high-resolution multispectral (HRMS) images. In this paper, we propose a novel spatialspectral dual guided network (SSDGN), aiming to fully capitalize on the spectral and spatial information contained in both the PAN and MS images. Firstly, to enhance feature extraction, we introduce two subnetworks: the progressive spectral feature extraction (PSpeFE) subnetwork and progressive spatial feature extraction (PSpaFE) for spatial features. These extract information from the PAN and MS images. Additionally, features from the frequency domain (FD) and intensity domain (ID) of both image types are leveraged to guide and enhance the efficiency of feature extraction. Then, a joint spatial-spectral attention feature fusion module and a multi-stage residual reconstruction module are devised to efficiently harness the extracted spatial and spectral information. Finally, extensive experiments are conducted to evaluate the performance and effectiveness of the proposed SSDGN. Compared to the second-best methods, our approach achieves an average reduction of 12.3% in ERGAS across three satellite datasets. The QNR metric improves by up to 2.1% and averages a 0.8% gain, demonstrating consistent advantages in spectral fidelity and fusion quality.
Shuyin Zhang, Laituan Qiao, Fan Zhang 0041, Chao Xu 0007, Shuqi Zhao, Quanwei Gao
IEEE Trans. Geosci. Remote. Sens.5
2023 More Than Capacity: Performance-oriented Evolution of Pangu in Alibaba
Qiang Li 0045, Qiao Xiang, Yuxin Wang 0003, Ridi Wen, Wenhui Yao, Shuqi Zhao, Zhaosheng Zhu, Huayong Wang, Shanyang Liu, Lulu Chen, Zhiwu Wu, Haonan Qiu, Derui Liu, Gexiao Tian, Shaozong Liu, Yaohui Wu, Zicheng Luo, Yuchao Shao, Junping Wu, Zheng Cao 0003, Zhongjie Wu, Jiaji Zhu, Jiwu Shu, Jiesheng Wu
FAST8
2023 Failure-aware Policy Learning for Self-assessable Robotics Tasks
abstract
Self-assessment rules play an essential role in safe and effective real-world robotic applications, which verify the feasibility of the selected action before actual execution. But how to utilize the self-assessment results to re-choose actions remains a challenge. Previous methods eliminate the selected action evaluated as failed by the self-assessment rules, and re-choose one with the next-highest affordance (i.e. process-of-elimination strategy [1]), which ignores the dependency between the self-assessment results and the remaining untried actions. However, this dependency is important since the previous failures might help trim the remaining over-estimated actions. In this paper, we set to investigate this dependency by learning a failure-aware policy. We propose two architectures for the failure-aware policy by representing the self-assessment results of previous failures as the variable state, and leveraging recurrent neural networks to implicitly memorize the previous failures. Experiments conducted on three tasks demonstrate that our method can achieve better performances with higher task success rates by less trials. Moreover, when the actions are correlated, learning a failure-aware policy can achieve better performance than the process-of-elimination strategy.
Kechun Xu, Runjian Chen, Shuqi Zhao, Zizhang Li, Hongxiang Yu, Ci Chen 0004, Yue Wang 0020, Rong Xiong
ICRA3
2023 A Joint Modeling of Vision-Language-Action for Target-oriented Grasping in Clutter
abstract
We focus on the task of language-conditioned grasping in clutter, in which a robot is supposed to grasp the target object based on a language instruction. Previous works separately conduct visual grounding to localize the target object, and generate a grasp for that object. However, these works require object labels or visual attributes for grounding, which calls for handcrafted rules in planner and restricts the range of language instructions. In this paper, we propose to jointly model vision, language and action with object-centric representation. Our method is applicable under more flexible language instructions, and not limited by visual grounding error. Besides, by utilizing the powerful priors from the pre-trained multi-modal model and grasp model, sample efficiency is effectively improved and the sim2real problem is relived without additional data for transfer. A series of experiments carried out in simulation and real world indicate that our method can achieve better task success rate by less times of motion under more flexible language instructions. Moreover, our method is capable of generalizing better to scenarios with unseen objects and language instructions.
Kechun Xu, Shuqi Zhao, Zhongxiang Zhou, Zizhang Li, Huaijin Pi, Yue Wang 0020, Rong Xiong
ICRA2
2012 ABrowse - a customizable next-generation genome browser framework
abstract
BACKGROUND: With the rapid growth of genome sequencing projects, genome browser is becoming indispensable, not only as a visualization system but also as an interactive platform to support open data access and collaborative work. Thus a customizable genome browser framework with rich functions and flexible configuration is needed to facilitate various genome research projects. RESULTS: Based on next-generation web technologies, we have developed a general-purpose genome browser framework ABrowse which provides interactive browsing experience, open data access and collaborative work support. By supporting Google-map-like smooth navigation, ABrowse offers end users highly interactive browsing experience. To facilitate further data analysis, multiple data access approaches are supported for external platforms to retrieve data from ABrowse. To promote collaborative work, an online user-space is provided for end users to create, store and share comments, annotations and landmarks. For data providers, ABrowse is highly customizable and configurable. The framework provides a set of utilities to import annotation data conveniently. To build ABrowse on existing annotation databases, data providers could specify SQL statements according to database schema. And customized pages for detailed information display of annotation entries could be easily plugged in. For developers, new drawing strategies could be integrated into ABrowse for new types of annotation data. In addition, standard web service is provided for data retrieval remotely, providing underlying machine-oriented programming interface for open data access. CONCLUSIONS: ABrowse framework is valuable for end users, data providers and developers by providing rich user functions and flexible customization approaches. The source code is published under GNU Lesser General Public License v3.0 and is accessible at http://www.abrowse.org/. To demonstrate all the features of ABrowse, a live demo for Arabidopsis thaliana genome has been built at http://arabidopsis.cbi.edu.cn/.
Jun Wang 0060, Shuqi Zhao, Xiaocheng Gu, Jingchu Luo, Ge Gao 0004
BMC Bioinform.3
2007 ABCGrid: Application for Bioinformatics Computing Grid
abstract
UNLABELLED: We have developed a package named Application for Bioinformatics Computing Grid (ABCGrid). ABCGrid was designed for biology laboratories to use heterogeneous computing resources and access bioinformatics applications from one master node. ABCGrid is very easy to install and maintain at the premise of robustness and high performance. We implement a mechanism to install and update all applications and databases in worker nodes automatically to reduce the workload of manual maintenance. We use a backup task method and self-adaptive job dispatch approach to improve performance. Currently, ABCGrid integrates NCBI_BLAST, Hmmpfam and CE, running on a number of computing platforms including UNIX/Linux, Windows and Mac OS X. AVAILABILITY: The source code, executables and documents can be downloaded from http://abcgrid.cbi.pku.edu.cn
Shuqi Zhao, Huashan Yu, Ge Gao 0004, Jingchu Luo
Bioinform.2
2007 Finding new structural and sequence attributes to predict possible disease association of single amino acid polymorphism (SAP)
abstract
MOTIVATION: The rapid accumulation of single amino acid polymorphisms (SAPs), also known as non-synonymous single nucleotide polymorphisms (nsSNPs), brings the opportunities and needs to understand and predict their disease association. Currently published attributes are limited, the detailed mechanisms governing the disease association of a SAP remain unclear and thus, further investigation of new attributes and improvement of the prediction are desired. RESULTS: A SAP dataset was compiled from the Swiss-Prot variant pages. We extracted and demonstrated the effectiveness of several new biologically informative attributes including the structural neighbor profiles that describe the SAP's microenvironment, nearby functional sites that measure the structure-based and sequence-based distances between the SAP site and its nearby functional sites, aggregation properties that measure the likelihood of protein aggregation and disordered regions that consider whether the SAP is located in structurally disordered regions. The new attributes provided insights into the mechanisms of the disease association of SAPs. We built a support vector machines (SVMs) classifier employing a carefully selected set of new and previously published attributes. Through a strict protein-level 5-fold cross-validation, we attained an overall accuracy of 82.61%, and an MCC of 0.60. Moreover, a web server was developed to provide a user-friendly interface for biologists. AVAILABILITY: The web server is available at http://sapred.cbi.pku.edu.cn/
Zhi-Qiang Ye, Shuqi Zhao, Xiao-Qiao Liu, Robert E. Langlois, Hui Lu 0004, Liping Wei
Bioinform.2