Zongyi Li

dblp:209/9986 · DBLP profile ↗
← Back
44ranked-venue papers
15as first author
39since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 11 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 6 first-author · 17 since 2021Computer networks · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Dual-stream Relation-modeling Disentanglement for Cloth-Changing Person Re-Identification
abstract
Cloth-changing person re-identification (CC-ReID) aims to identify individuals across non-overlapping cameras despite clothing variations. Existing methods are often constrained by two primary limitations: approaches using auxiliary modalities typically rely on a single specific cue, limiting their robustness, while feature disentanglement methods struggle with discrete labels that create inconsistencies between ground truth labels and modality semantic similarity. To overcome these limitations, we propose DRDnet, a unified framework that synergistically integrates dual auxiliary cues and advanced relation modeling. Specifically, our Dual-Stream Disentanglement (DSD) module leverages textual descriptions and parsing images to decouple clothing factors through high-level semantic supervision and pixel-level operations, yielding robust clothing-agnostic features. Simultaneously, our Modal Relation Modeling (MRM) module constructs feature memory banks and employs adaptive soft label smoothing, effectively enhancing image-text semantic alignment and reinforcing identity consistency across clothing changes. We evaluate DRDnet on several CC-ReID benchmarks to demonstrate its effectiveness and provide state-of-the-art performance across all benchmarks.
Shijuan Huang, Zongyi Li, Zhao Lv
AAAI3
2026 Performance Analysis of RIS-Aided MISO URLLC Systems with Channel Aging and EMI
Zongyi Li, Jiayi Zhang 0001, Yu Lu 0011, Ziling Xu, Shuxian Wen
ICC1
2025 Exploring the Potential of Large Vision-Language Models for Unsupervised Text-Based Person Retrieval
abstract
The aim of text-based person retrieval is to identify pedestrians using natural language descriptions within a large-scale image gallery. Traditional methods rely heavily on manually annotated image-text pairs, which are resource-intensive to obtain. With the emergence of Large Vision-Language Models (LVLMs), the advanced capabilities of contemporary models in image understanding have led to the generation of highly accurate captions. Therefore, this paper explores the potential of employing Large Vision-Language Models for unsupervised text-based pedestrian image retrieval and proposes a Multi-grained Uncertainty Modeling and Alignment framework (MUMA). Initially, multiple Large Vision-Language Models are employed to generate diverse and hierarchically structured pedestrian descriptions across different styles and granularities. However, the generated captions inevitably introduce noise. To address this issue, an uncertainty-guided sample filtration module is proposed to estimate and filter out unreliable image-text pairs. Additionally, to simulate the diversity of styles and granularities in captions, a multi-grained uncertainty modeling approach is applied to model the distributions of captions, with each caption represented as a multivariate Gaussian distribution. Finally, a multi-level consistency distillation loss is employed to integrate and align the multi-grained captions, aiming to transfer knowledge across different granularities. Experimental evaluations conducted on three widely-used datasets demonstrate the significant advancements achieved by our approach.
Zongyi Li, Yuxuan Shi 0002, Jiazhong Chen, Shijuan Huang, Linnan Tu, Fei Shen 0004
AAAI1
2025 AD2T: Adversarial Distortion Domain Translation for Robust Watermarking against Non-differentiable Distortions
abstract
Deep watermarking models optimize robustness by incorporating distortions between the encoder and decoder. To tackle non-differentiable distortions, current methods only train the decoder with distorted images, which breaks the joint optimization of the encoder-decoder, resulting in suboptimal performance. To address this problem, we propose an Adversarial Distortion Domain Translation (AD2T) method by treating the distortion as an image-to-image translation task. AD2T adopts conditional GANs to learn the non-differentiable distortion mappings. It employs generators to transform the encoded image into the distorted one to bridge the encoder-decoder for joint optimization. We also supervise the GANs to generate challenging distorted samples to augment the watermarking model via adversarial training. This further improves the model robustness by minimizing the maximum decoding loss. Extensive experiments demonstrate the superiority of our method when tested on non-differentiable distortions, including lossy compression and style transfers. Codes are released here: https://github.com/zcx-language/AdversarialDistortionDomainTranslation.
Chengxin Zhao, Jiazhong Chen, Han Fang 0004, Zongyi Li, Sijing Xie
ICASSP5
2025 ARLON: Boosting Diffusion Transformers with Autoregressive Models for Long Video Generation
abstract
Text-to-video (T2V) models have recently undergone rapid and substantial advancements. Nevertheless, due to limitations in data and computational resources, achieving efficient generation of long videos with rich motion dynamics remains a significant challenge. To generate high-quality, dynamic, and temporally consistent long videos, this paper presents ARLON, a novel framework that boosts diffusion Transformers with autoregressive (\textbf{AR}) models for long (\textbf{LON}) video generation, by integrating the coarse spatial and long-range temporal information provided by the AR model to guide the DiT model effectively. Specifically, ARLON incorporates several key innovations: 1) A latent Vector Quantized Variational Autoencoder (VQ-VAE) compresses the input latent space of the DiT model into compact and highly quantized visual tokens, bridging the AR and DiT models and balancing the learning complexity and information density; 2) An adaptive norm-based semantic injection module integrates the coarse discrete visual units from the AR model into the DiT model, ensuring effective guidance during video generation; 3) To enhance the tolerance capability of noise introduced from the AR inference, the DiT model is trained with coarser visual latent tokens incorporated with an uncertainty sampling module. Experimental results demonstrate that ARLON significantly outperforms the baseline OpenSora-V1.2 on eight out of eleven metrics selected from VBench, with notable improvements in dynamic degree and aesthetic quality, while delivering competitive results on the remaining three and simultaneously accelerating the generation process. In addition, ARLON achieves state-of-the-art performance in long video generation, outperforming other open-source models in this domain. Detailed analyses of the improvements in inference efficiency are presented, alongside a practical application that demonstrates the generation of long videos using progressive text prompts. Project page: \url{http://aka.ms/arlon}.
Zongyi Li, Shujie Hu, Shujie Liu 0001, Jeongsoo Choi, Lingwei Meng, Jinyu Li 0001, Furu Wei
ICLR1
2025 Harnessing Multimodal Large Language Models for Personalized Product Search with Query-aware Refinement
abstract
Personalized product search (PPS) aims to retrieve products relevant to the given query considering user preferences within their purchase histories. Since large language models (LLM) exhibit impressive potential in content understanding and reasoning, current methods explore to leverage LLM to comprehend the complicated relationships among user, query and product to improve the search performance of PPS. Despite the progress, LLM-based PPS solutions merely take textual contents into consideration, neglecting multimodal contents which play a critical role for product search. Motivated by this, we propose a novel framework, HMPPS, for Harnessing Multimodal large language models (MLLM) to deal with Personalized Product Search based on multimodal contents. Nevertheless, the redundancy and noise in PPS input stand for a great challenge to apply MLLM for PPS, which not only misleads MLLM to generate inaccurate search results but also increases the computation expense of MLLM. To deal with this problem, we additionally design two query-aware refinement modules for HMPPS: 1) a perspective-guided summarization module that generates refined product descriptions around core perspectives relevant to search query, reducing noise and redundancy within textual contents; and 2) a two-stage training paradigm that introduces search query for user history filtering based on multimodal representations, capturing precise user preferences and decreasing the inference cost. Extensive experiments are conducted on four public datasets to demonstrate the effectiveness of HMPPS. Furthermore, HMPPS is deployed on an online search system with billion-level daily active users and achieves an evident gain in A/B testing.
Beibei Zhang 0005, Yanan Lu, Ruobing Xie, Zongyi Li, Tongwei Ren, Fen Lin 0002
ACM Multimedia4
2025 Incentivizing Reasoning for Advanced Instruction-Following of Large Language Models
abstract
Existing large language models (LLMs) face challenges of following complex instructions, especially when multiple constraints are present and organized in paralleling, chaining, and branching structures. One intuitive solution, namely chain-of-thought (CoT), is expected to universally improve capabilities of LLMs. However, we find that the vanilla CoT exerts a negative impact on performance due to its superficial reasoning pattern of simply paraphrasing the instructions. It fails to peel back the compositions of constraints for identifying their relationship across hierarchies of types and dimensions. To this end, we propose RAIF, a systematic method to boost LLMs in dealing with complex instructions via incentivizing reasoning for test-time compute scaling. First, we stem from the decomposition of complex instructions under existing taxonomies and propose a reproducible data acquisition method. Second, we exploit reinforcement learning (RL) with verifiable rule-centric reward signals to cultivate reasoning specifically for instruction following. We address the shallow, non-essential nature of reasoning under complex instructions via sample-wise contrast for superior CoT enforcement. We also exploit behavior cloning of experts to facilitate steady distribution shift from fast-thinking LLMs to skillful reasoners. Extensive evaluations on seven comprehensive benchmarks confirm the validity of the proposed method, where a 1.5B LLM achieves 11.74% gains with performance comparable to a 8B LLM. Evaluation on OOD constraints also confirms the generalizability of our RAIF.
Yulei Qin, Zongyi Li, Zhekai Lin, Ke Li 0015, Xing Sun 0001
NeurIPS3
2025 Autoregressive Motion Generation with Gaussian Mixture-Guided Latent Sampling
abstract
Existing efforts in motion synthesis typically utilize either generative transformers with discrete representations or diffusion models with continuous representations. However, the discretization process in generative transformers can introduce motion errors, while the sampling process in diffusion models tends to be slow. In this paper, we propose a novel text-to-motion synthesis method GMMotion that combines a continuous motion representation with an autoregressive model, using the Gaussian mixture model (GMM) to represent the conditional probability distribution. Unlike autoregressive approaches relying on residual vector quantization, our model employs continuous motion representations derived from the VAE's latent space. This choice streamlines both the training and the inference processes. Specifically, we utilize a causal transformer to learn the distributions of continuous motion representations, which are modeled with a learnable Gaussian mixture model. Extensive experiments demonstrate that our model surpasses existing state-of-the-art models in the motion synthesis task.
Linnan Tu, Lingwei Meng, Zongyi Li, Shijuan Huang
NeurIPS3
2025 TSAD: Temporal-spatial association differences-based unsupervised anomaly detection for multivariate time-series
Hanbing Zhu, Zongyi Li, Yuxuan Shi 0002, Chuang Zhao 0001, Hongxu Ji, Ping Li 0021
Neurocomputing4
2025 Geometric Operator Learning with Optimal Transport
abstract
We propose integrating optimal transport (OT) into operator learning for partial differential equations (PDEs) on complex geometries. Classical geometric learning methods typically represent domains as meshes, graphs, or point clouds. Our approach generalizes discretized meshes to mesh density functions, formulating geometry embedding as an OT problem that maps these functions to a uniform density in a reference space. Compared to previous methods relying on interpolation or shared deformation, our OT-based method employs instance-dependent deformation, offering enhanced flexibility and effectiveness. For 3D simulations focused on surfaces, our OT-based neural operator embeds the surface geometry into a 2D parameterized latent space. By performing computations directly on this 2D representation of the surface manifold, it achieves significant computational efficiency gains compared to volumetric simulation. Experiments with Reynolds-averaged Navier-Stokes equations (RANS) on the ShapeNet-Car and DrivAerNet-Car datasets show that our method achieves better accuracy and also reduces computational expenses in terms of both time and memory usage compared to existing machine learning models. Additionally, our model demonstrates significantly improved accuracy on the FlowBench dataset, underscoring the benefits of employing instance-dependent deformation for datasets with highly variable geometries.
Zongyi Li, Nikola B. Kovachki, Anima Anandkumar
J. Mach. Learn. Res.2
2025 Cross-Modality Relation and Uncertainty Exploration for Text-Based Person Search
abstract
Text-based person search aims to retrieve specific individuals from an extensive image gallery using textual queries. Recent approaches have delved into aligning global and part features in both text and image modalities, yielding substantial improvements. However, these methods often overlook intra-modality instance relations and uncertainties inherent in text-based person search. In response to these challenges, we propose the Cross-Modality Relation and Uncertainty Exploration (CRUE) method to model the relations and uncertainties in the matching procedure. To alleviate the strict alignment issues arising from hard labels in the original contrastive loss, an Intra-Modality Relation Exploration (IRE) module is introduced. This module is specifically designed to smooth hard-matching relations by modeling intra-modality similarity. Additionally, to address uncertain matching problems stemming from many-to-many relations, we propose a novel Uncertainty-Guided Modeling (UGM) module. This module is specifically designed to handle weak and noise-matched image–text pairs by modeling features as distributions, thereby alleviating instability and noise. Both the IRE and UGM modules effectively consider genuine intra-modality similarities and reduce the negative impact of uncertainties. Experimental results demonstrate significant improvements across three widely used person search datasets, thereby validating the efficacy of the CRUE method in enhancing text-based person search. Our code will be available on GitHub at https://github.com/ShijuanHuang/CRUE .
Shijuan Huang, Zongyi Li
ACM Trans. Multim. Comput. Commun. Appl.2
2025 GSyncCode: Geometry Synchronous Hidden Code for One-step Photography Decoding
abstract
Invisible hyperlinks and hidden barcodes have recently emerged as a hot topic in offline-to-online messaging, where an invisible message or barcode is embedded in an image and can be decoded via camera shooting. Current schemes involve a two-step decoding process: starting with vertex localization of the embedded region to correct the perspective distortion introduced by shooting, followed by decoding the message from the corrected region. However, vertex localization can be complex and time-consuming, which affects the efficiency and accuracy of message decoding. To address this issue, this article proposes a geometry synchronous decoding scheme called GSyncCode, allowing for one-step extraction of a Data Matrix code from the photograph. Instead of correction before decoding, GSyncCode directly decodes a geometry-transformed Data Matrix that is synchronized with the embedded region. A barcode scanner is then used to efficiently retrieve messages. We design a Haar transform-based encoder HaarUNet and a HaarLoss visual function to select the key component of the Data Matrix for embedding. They improve the visual quality of the embedded image by reducing redundant embedding signals. Extensive simulated and real-world experiments demonstrate the superiority of GSyncCode in both decoding efficiency and accuracy. Our codes are published at: https://github.com/zcx-language/GSyncCode .
Chengxin Zhao, Jialie Shen 0001, Han Fang 0004, Sijing Xie, Yaokun Fang, Zongyi Li, Ping Li 0021
ACM Trans. Multim. Comput. Commun. Appl.7
2024 Uncertainty-Guided Person Search Model with Auxiliary Shallow Feature Exploration
abstract
Person search is a unified system aimed at jointly localizing and identifying a person of interest from a gallery of whole scene images. Due to the inherent properties of the person search, it faces significant challenges of large-scale variations, inaccurate detection boxes, and crowded scenes. To address these issues, we proposed an uncertainty-guided framework coupled with auxiliary shallow feature exploration, which includes a shallow feature fusion module and an uncertainty-guided module. Firstly, considering the scales of the person are varied due to various scenes and their relative positions to the camera, a shallow feature fusion module is designed to extract multi-scale features to assist the re-id sub-task. Additionally, a self-distillation loss is proposed to align features across different scales. Furthermore, to alleviate the problem that the model can be easily affected by coarse samples resulting from crowded scenes and inaccurate detection boxes, we introduce an uncertainty guidance module to reduce the negative impact of these coarse targets. The experimental results demonstrate the effectiveness of our proposed methods on two benchmarks (i.e., CUHK-SYSU, and PRW).
Zongyi Li, Yuxuan Shi 0002, Jiazhong Chen, Runsheng Wang, Ping Li 0021
ICASSP1
2024 Training-Free Diffusion Models for Content-Style Synthesis
Ruipeng Xu, Fei Shen 0004, Zongyi Li
ICIC (10)4
2024 Enhancing Landslide Segmentation with Guide Attention Mechanism and Fast Fourier Transformer
Kaiyu Yan, Fei Shen 0004, Zongyi Li
ICIC (10)3
2024 Cross-modal Generation and Alignment via Attribute-guided Prompt for Unsupervised Text-based Person Retrieval
Zongyi Li, Yuxuan Shi 0002, Jiazhong Chen, Runsheng Wang, Shijuan Huang
IJCAI1
2024 DBDH: A Dual-Branch Dual-Head Neural Network for Invisible Embedded Regions Localization
abstract
Embedding invisible hyperlinks or hidden codes in images to replace QR codes has become a hot topic recently. This technology requires first localizing the embedded region in the captured photos before decoding. Existing methods that train models to find the invisible embedded region struggle to obtain accurate localization results, leading to degraded decoding accuracy. This limitation is primarily because the CNN network is sensitive to low-frequency signals, while the embedded signal is typically in the high-frequency form. Based on this, this paper proposes a Dual-Branch Dual-Head (DBDH) neural network tailored for the precise localization of invisible embedded regions. Specifically, DBDH uses a low-level texture branch containing 62 high-pass filters to capture the high-frequency signals induced by embedding. A high-level context branch is used to extract discriminative features between the embedded and normal regions. DBDH employs a detection head to directly detect the four vertices of the embedding region. In addition, we introduce an extra segmentation head to segment the mask of the embedding region during training. The segmentation head provides pixel-level supervision for model learning, facilitating better learning of the embedded signals. Based on two state-of-the-art invisible offline-to-online messaging methods, we construct two datasets and augmentation strategies for training and testing localization models. Extensive experiments demonstrate the superior performance of the proposed DBDH over existing methods.
Chengxin Zhao, Sijing Xie, Zongyi Li, Yuxuan Shi 0002, Jiazhong Chen
IJCNN5
2024 Pretraining Codomain Attention Neural Operators for Solving Multiphysics PDEs
abstract
Existing neural operator architectures face challenges when solving multiphysics problems with coupled partial differential equations (PDEs) due to complex geometries, interactions between physical variables, and the limited amounts of high-resolution training data. To address these issues, we propose *Codomain Attention Neural Operator* (CoDA-NO), which tokenizes functions along the codomain or channel space, enabling self-supervised learning or pretraining of multiple PDE systems. Specifically, we extend positional encoding, self-attention, and normalization layers to function spaces. CoDA-NO can learn representations of different PDE systems with a single model. We evaluate CoDA-NO's potential as a backbone for learning multiphysics PDEs over multiple systems by considering few-shot learning settings. On complex downstream tasks with limited data, such as fluid flow simulations, fluid-structure interactions, and Rayleigh-Bénard convection, we found CoDA-NO to outperform existing methods by over 36%.
Robert Joseph George, Mogab Elleithy, Daniel V. Leibovici, Zongyi Li, Boris Bonev, Colin White, Julius Berner, Raymond A. Yeh, Jean Kossaifi, Kamyar Azizzadenesheli, Anima Anandkumar
NeurIPS5
2024 Improving the Transferability of Adversarial Attacks on Face Recognition With Beneficial Perturbation Feature Augmentation
abstract
Face recognition (FR) models can be easily fooled by adversarial examples, which are crafted by adding imperceptible perturbations on benign face images. The existence of adversarial face examples poses a great threat to the security of society. To build a more sustainable digital nation, in this article, we improve the transferability of adversarial face examples to expose more blind spots of the existing FR models. Though generating hard samples has shown its effectiveness in improving the generalization of models in training tasks, the effectiveness of using this idea to improve the transferability of adversarial face examples remains unexplored. To this end, based on the property of hard samples and the symmetry between training tasks and adversarial attack tasks, we propose the concept of hard models, which have similar effects as hard samples for adversarial attack tasks. Using the concept of hard models, we propose a novel attack method called beneficial perturbation feature augmentation attack (BPFA), which reduces the overfitting of adversarial examples to surrogate FR models by constantly generating new hard models to craft the adversarial examples. Specifically, in the backpropagation, BPFA records the gradients on preselected feature maps and uses the gradient on the input image to craft the adversarial example. In the next forward propagation, BPFA leverages the recorded gradients to add beneficial perturbations on their corresponding feature maps to increase the loss. Extensive experiments demonstrate that BPFA can significantly boost the transferability of adversarial attacks on FR.
Fengfan Zhou, Yuxuan Shi 0002, Jiazhong Chen, Zongyi Li, Ping Li 0021
IEEE Trans. Comput. Soc. Syst.5
2024 Knowledge Consistency Distillation for Weakly Supervised One Step Person Search
abstract
Weakly supervised person search targets to detect and identify a person with only bounding box annotations. Recent approaches have focused on learning person relations in a single model, ignoring the conflicts between the detection and Re-ID heads, along with the influence of background elements, which may lead to noisy pseudo labels and inaccurate Re-ID features. To address this challenge, we introduce a novel framework named Knowledge Consistency Distillation (KCD) for weakly supervised person search, which explores the capabilities of an advanced unsupervised person re-identification (Re-ID) model to mitigate the conflicts and background influences. We propose hierarchical consistency alignments, including feature-level, cluster-level, and instance-level consistency alignment, to synchronize the knowledge from the state-of-the-art unsupervised Re-ID model. Specifically, the feature-level consistency aligns the feature through both context and relation alignment. The cluster-level consistency aligns the teacher cluster information by reusing its OIM module. To tackle the inconsistency problem between student instances and teacher cluster centroids, we incorporate pseudo-label refinement to assist the student model in comprehending the teacher’s knowledge at cluster-level while mitigating the negative effects of noisy labels. Finally, an instance-level consistency loss weighted by the similarity between the instance and its corresponding cluster is proposed to align the positive instance correlations. Our approach aims to train a one-step weakly supervised model for person search by exploiting the characteristics of unsupervised person Re-ID. Extensive experiments illustrate that our method achieves state-of-the-art performance on two widely-used person search datasets, CUHK-SYSU and PRW. Our code will be available on GitHub athttps://github.com/zongyi1999/KCD.
Zongyi Li, Yuxuan Shi 0002, Jiazhong Chen, Runsheng Wang, Chengxin Zhao, Qian Wang 0001, Shijuan Huang
IEEE Trans. Circuits Syst. Video Technol.1
2024 Gait Recognition With Multi-Level Skeleton-Guided Refinement
abstract
Existing methods combining skeleton and silhouette representations demonstrate explicit effectiveness for gait recognition. However, current related methods simply combine the video-level representations of model-based skeleton data and gait silhouettes for retrieval. Therefore, diverse skeleton information is not fully exploited in existing related works: Firstly, the position and movement of bones are not clear from individual silhouettes. This indicates that the frame-level interaction between features of skeletons and silhouettes is critical, which is ignored by previous methods. Secondly, diverse part-level skeleton-guided gait features are not fully captured in existing related approaches. To solve the above issues, we present a novel framework with multi-level skeleton-guided refinement, including frame-level, part-level, and video-level skeleton-guided refinement, for comprehensive skeleton-aided gait representation learning. First, two modules are proposed for frame-level skeleton-guided refinement. Specifically, Visual Skeleton Enhanced Backbone (VSEB) is proposed to visually highlight the global and part-level skeleton regions for the feature of each silhouette frame. Moreover, Cross-Visual-Model Frame-level Interaction (CVMFI) is proposed to further transfer the model-based skeleton information to features of the visual modalities. Secondly, part-level visual and model-based skeleton features are utilized to refine the final gait representation. Concretely, in VSEB, Part Skeleton Enhance Network (PSEN) is proposed to visually enhance the position and movement of part-level skeletons. In addition, Semantic Part Pooling (SPP) is proposed for capturing the model-based skeleton features of different semantic parts. Finally, as the video-level skeleton-guided refinement, multimodal video-level features are combined to boost the final recognition performance. Extensive experimental results on prevailing datasets demonstrate that our approach outperforms most existing methods, including the skeleton-aided multi-modal methods. With the multi-level refinement guided by the skeleton modalities, the framework is expected to provide a deeper understanding of skeleton-aided gait recognition.
Runsheng Wang, Yuxuan Shi 0002, Zongyi Li, Chengxin Zhao, Bohao Wei, He Li 0052, Ping Li 0021
IEEE Trans. Multim.4
2024 Viewpoint Disentangling and Generation for Unsupervised Object Re-ID
abstract
Unsupervised object Re-ID aims to learn discriminative identity features from a fully unlabeled dataset to solve the open-class re-identification problem. Satisfying results have been achieved in existing unsupervised Re-ID methods, primarily trained with pseudo-labels created by feature clustering. However, the viewpoint variation of objects is the key challenge, introducing noisy labels in the clustering process. To address this problem, a novel viewpoint disentangling and generation framework (VDG) is proposed to learn viewpoint-invariant ID features, including a disentangling and generation module, as well as a contrastive learning module. First, we design an ID encoder to map the viewpoint and identity features into the latent space. Second, a generator is used to disentangle view features and synthesize images with different orientations. Especially, the well-trained encoder serves as a pre-trained feature extractor in the contrastive learning module. Third, a viewpoint-aware loss and a class-level loss are integrated to facilitate contrastive learning between original and novel views. The generation of novel view images and the application of viewpoint-aware contrastive loss mutually assist model learning viewpoint-invariant ID features. Extensive experiments on Market-1501, DukeMTMC, MSMT17, and VeRi-776 demonstrate the effectiveness of the proposed VDG framework, as well as its superiority over the existing state-of-the-art approaches. The VDG model also demonstrates high quality in the image generation tasks.
Zongyi Li, Yuxuan Shi 0002, Jiazhong Chen, Boyuan Liu, Runsheng Wang, Chengxin Zhao
ACM Trans. Multim. Comput. Commun. Appl.1
2023 FusionU-Net: U-Net with Enhanced Skip Connection for Pathology Image Segmentation
Zongyi Li, Hongbing Lyu
ACML1
2023 Mutual Relative Position Learning Transformer for Cross-View Geo-Localization
abstract
Cross-view geo-localization refers to matching ground images with geo-tagged satellite imagery. Existing methods are mainly two-stage, applying a polar transform to roughly eliminate the gap between these two domains, but this might introduce distortions and reduce the discriminativeness of features. In this work, we propose a transformer-based one-stage approach, which unifies gap elimination and feature extraction. The relative position among objects provides critical clues for this task and has strong spatial correspondences between the two views. Firstly, we form the relative position by selecting representative tokens from different regions. Then the relative positions of the two views predict each other and eliminate the gap through mutual learning. Finally, we introduce a novel consistency loss to enhance feature learning by mutual transfer of relational knowledge among samples. Extensive experiments demonstrate that our method achieves state-of-the-art results on both standard and fine-grained datasets.1
Yuxuan Shi 0002, Zongyi Li, Chuang Zhao 0001, Ping Li 0021
ICIP4
2023 MEGL: Multi-Experts Guided Learning Network for Single Camera Training Person Re-Identification
abstract
The time-saving single-camera training(SCT) person re-identification aims to learn camera-invariant information without cross-camera pedestrian annotations. To address this challenging task, we propose a novel approach called Multi-Experts Guided Learning Network (MEGL-Net) for SCT-ReID that can obtain features not influenced by camera views at the global and local levels under the guidance of multi-camera experts. Firstly, to obtain camera-invariant features, an adaptive feature integration module (AFI) is introduced to adaptively integrate expert-guided features from different camera branches. Then, the proposed camera-local interactive module (CLI) facilitates interaction between the local branch and the camera experts branch for automatically extracting discriminative, domain-invariant features at a fine-grained level. Finally, our framework aggregates expert-guided features with global features and enhanced local features in the testing stage for pedestrian retrieval. Under the Market-SCT and Duke-SCT datasets, experimental results demonstrate that our approach significantly improves ReID performance and outperforms existing state-of-the-art (SOTA) methods.
He Li 0052, Yuxuan Shi 0002, Zongyi Li, Runsheng Wang, Chengxin Zhao, Ping Li 0021
ICIP4
2023 Geometry-Informed Neural Operator for Large-Scale 3D PDEs
abstract
We propose the geometry-informed neural operator (GINO), a highly efficient approach for learning the solution operator of large-scale partial differential equations with varying geometries. GINO uses a signed distance function (SDF) representation of the input shape and neural operators based on graph and Fourier architectures to learn the solution operator. The graph neural operator handles irregular grids and transforms them into and from regular latent grids on which Fourier neural operator can be efficiently applied. We provide an efficient implementation of GINO using an optimized hashing approach, which allows efficient learning in a shared, compressed latent space with reduced computation and memory costs. GINO is discretization-invariant, meaning the trained model can be applied to arbitrary discretizations of the continuous domain and applies to any shape or resolution. To empirically validate the performance of our method on large-scale simulation, we generate the industry-standard aerodynamics dataset of 3D vehicle geometries with Reynolds numbers as high as five million. For this large-scale 3D fluid simulation, numerical methods are expensive to compute surface pressure. We successfully trained GINO to predict the pressure on car surfaces using only five hundred data points. The cost-accuracy experiments show a 26,000x speed-up compared to optimized GPU-based computational fluid dynamics (CFD) simulators on computing the drag coefficient. When tested on new combinations of geometries and boundary conditions (inlet velocities), GINO obtains a one-fourth reduction in error rate compared to deep neural network approaches.
Zongyi Li, Nikola B. Kovachki, Christopher B. Choy, Jean Kossaifi, Shourya Prakash Otta, Mohammad Amin Nabian, Maximilian Stadler, Christian Hundt 0002, Kamyar Azizzadenesheli, Anima Anandkumar
NeurIPS1
2023 Neural Operator: Learning Maps Between Function Spaces With Applications to PDEs
abstract
The classical development of neural networks has primarily focused on learning mappings between finite dimensional Euclidean spaces or finite sets. We propose a generalization of neural networks to learn operators, termed neural operators, that map between infinite dimensional function spaces. We formulate the neural operator as a composition of linear integral operators and nonlinear activation functions. We prove a universal approximation theorem for our proposed neural operator, showing that it can approximate any given nonlinear continuous operator. The proposed neural operators are also discretization-invariant, i.e., they share the same model parameters among different discretization of the underlying function spaces. Furthermore, we introduce four classes of efficient parameterization, viz., graph neural operators, multi-pole graph neural operators, low-rank neural operators, and Fourier neural operators. An important application for neural operators is learning surrogate maps for the solution operators of partial differential equations (PDEs). We consider standard PDEs such as the Burgers, Darcy subsurface flow, and the Navier-Stokes equations, and show that the proposed neural operators have superior performance compared to existing machine learning based methodologies, while being several orders of magnitude faster than conventional PDE solvers.
Nikola B. Kovachki, Zongyi Li, Burigede Liu, Kamyar Azizzadenesheli, Kaushik Bhattacharya, Andrew M. Stuart, Anima Anandkumar
J. Mach. Learn. Res.2
2023 Fourier Neural Operator with Learned Deformations for PDEs on General Geometries
abstract
Deep learning surrogate models have shown promise in solving partial differential equations (PDEs). Among them, the Fourier neural operator (FNO) achieves good accuracy, and is significantly faster compared to numerical solvers, on a variety of PDEs, such as fluid flows. However, the FNO uses the Fast Fourier transform (FFT), which is limited to rectangular domains with uniform grids. In this work, we propose a new framework, viz., Geo-FNO, to solve PDEs on arbitrary geometries. Geo-FNO learns to deform the input (physical) domain, which may be irregular, into a latent space with a uniform grid. The FNO model with the FFT is applied in the latent space. The resulting Geo-FNO model has both the computation efficiency of FFT and the flexibility of handling arbitrary geometries. Our Geo-FNO is also flexible in terms of its input formats, viz., point clouds, meshes, and design parameters are all valid inputs. We consider a variety of PDEs such as the Elasticity, Plasticity, Euler's, and Navier-Stokes equations, and both forward modeling and inverse design problems. Comprehensive cost-accuracy experiments show that Geo-FNO is $10^5$ times faster than the standard numerical solvers and twice more accurate compared to direct interpolation on existing ML-based PDE solvers such as the standard FNO.
Zongyi Li, Daniel Zhengyu Huang, Burigede Liu, Anima Anandkumar
J. Mach. Learn. Res.1
2022 Reliability Exploration with Self-Ensemble Learning for Domain Adaptive Person Re-identification
abstract
Person re-identifcation (Re-ID) based on unsupervised domain adaptation (UDA) aims to transfer the pre-trained model from one labeled source domain to an unlabeled target domain. Existing methods tackle this problem by using clustering methods to generate pseudo labels. However, pseudo labels produced by these techniques may be unstable and noisy, substantially deteriorating models’ performance. In this paper, we propose a Reliability Exploration with Self-ensemble Learning (RESL) framework for domain adaptive person ReID. First, to increase the feature diversity, multiple branches are presented to extract features from different data augmentations. Taking the temporally average model as a mean teacher model, online label refning is conducted by using its dynamic ensemble predictions from different branches as soft labels. Second, to combat the adverse effects of unreliable samples in clusters, sample reliability is estimated by evaluating the consistency of different clusters’ results, followed by selecting reliable instances for training and re-weighting sample contribution within Re-ID losses. A contrastive loss is also utilized with cluster-level memory features which are updated by the mean feature. The experiments demonstrate that our method can signifcantly surpass the state-of-the-art performance on the unsupervised domain adaptive person ReID.
Zongyi Li, Yuxuan Shi 0002, Jiazhong Chen, Qian Wang 0001, Fengfan Zhou
AAAI1
2022 Generic lithography modeling with dual-band optics-inspired neural networks
abstract
Lithography simulation is a critical step in VLSI design and optimization for manufacturability. Existing solutions for highly accurate lithography simulation with rigorous models are computationally expensive and slow, even when equipped with various approximation techniques. Recently, machine learning has provided alternative solutions for lithography simulation tasks such as coarse-grained edge placement error regression and complete contour prediction. However, the impact of these learning-based methods has been limited due to restrictive usage scenarios or low simulation accuracy. To tackle these concerns, we introduce an dual-band optics-inspired neural network design that considers the optical physics underlying lithography. To the best of our knowledge, our approach yields the first published via/metal layer contour simulation at 1nm2/pixel resolution with any tile size. Compared to previous machine learning based solutions, we demonstrate that our framework can be trained much faster and offers a significant improvement on efficiency and image quality with 20× smaller model size. We also achieve 85× simulation speedup over traditional lithography simulator with ~ 1% accuracy loss.
Zongyi Li, Kumara Sastry, Saumyadip Mukhopadhyay, Mark Kilgard, Anima Anandkumar, Brucek Khailany, Haoxing Ren
DAC2
2022 Efficient Token Mixing for Transformers via Adaptive Fourier Neural Operators
John Guibas, Morteza Mardani, Zongyi Li, Andrew Tao, Anima Anandkumar, Bryan Catanzaro
ICLR3
2022 Learning Chaotic Dynamics in Dissipative Systems
abstract
Chaotic systems are notoriously challenging to predict because of their sensitivity to perturbations and errors due to time stepping. Despite this unpredictable behavior, for many dissipative systems the statistics of the long term trajectories are governed by an invariant measure supported on a set, known as the global attractor; for many problems this set is finite dimensional, even if the state space is infinite dimensional. For Markovian systems, the statistical properties of long-term trajectories are uniquely determined by the solution operator that maps the evolution of the system over arbitrary positive time increments. In this work, we propose a machine learning framework to learn the underlying solution operator for dissipative chaotic systems, showing that the resulting learned operator accurately captures short-time trajectories and long-time statistical behavior. Using this framework, we are able to predict various statistics of the invariant measure for the turbulent Kolmogorov Flow dynamics with Reynolds numbers up to $5000$.
Zongyi Li, Miguel Liu-Schiaffini, Nikola B. Kovachki, Kamyar Azizzadenesheli, Burigede Liu, Kaushik Bhattacharya, Andrew M. Stuart, Anima Anandkumar
NeurIPS1
2022 Spatial-wise and channel-wise feature uncertainty for occluded person re-identification
Yuxuan Shi 0002, Weiyi Tian, Zongyi Li, Ping Li 0021
Neurocomputing4
2021 Searching for an Effective Defender: Benchmarking Defense against Adversarial Word Substitution
abstract
Recent studies have shown that deep neural network-based models are vulnerable to intentionally crafted adversarial examples, and various methods have been proposed to defend against adversarial word-substitution attacks for neural NLP models.However, there is a lack of systematic study on comparing different defense approaches under the same attacking setting.In this paper, we seek to fill the gap through comprehensive studies on the behavior of neural text classifiers trained with various defense methods against representative adversarial attacks.In addition, we propose an effective method to further improve the robustness of neural text classifiers against such attacks, and achieved the highest accuracy on both clean and adversarial examples on AGNEWS and IMDB datasets, outperforming existing methods by a significant margin.We hope this study could provide useful clues for future research on text adversarial defense.Codes are available at https:// github.com/RockyLzy/TextDefender.
Zongyi Li, Jianhan Xu, Jiehang Zeng, Linyang Li, Xiaoqing Zheng, Qi Zhang 0001, Kai-Wei Chang 0001, Cho-Jui Hsieh
EMNLP (1)1
2021 Fourier Neural Operator for Parametric Partial Differential Equations
Zongyi Li, Nikola B. Kovachki, Kamyar Azizzadenesheli, Burigede Liu, Kaushik Bhattacharya, Andrew M. Stuart, Anima Anandkumar
ICLR1
2021 Video Saliency Prediction via Deep Eye Movement Learning
abstract
Existing methods often utilize temporal motion information and spatial layout information in video to predict video saliency. However, the fixations are not always consistent with the moving object of interest, because human eye fixations are determined not only by the spatio-temporal information, but also by the velocity of eye movement. To address this issue, a new saliency prediction method via deep eye movement learning (EML) is proposed in this paper. Compared with previous methods that use human fixations as ground truth, our method uses the optical flow of fixations between successive frames as an extra ground truth for the purpose of eye movement learning. Experimental results on DHF1K, Hollywood2, and UCF-sports datasets show the proposed EML model achieves a promising result across a wide of metrics.
Jiazhong Chen, Jie Chen 0058, Dakai Ren, Shiqi Zhang 0002, Zongyi Li
MMAsia6
2021 Video saliency prediction via spatio-temporal reasoning
Jiazhong Chen, Zongyi Li, Yi Jin 0001, Dakai Ren
Neurocomputing2
2021 Gaze estimation via bilinear pooling-based attention networks
Dakai Ren, Jiazhong Chen, Zhaoming Lu, Zongyi Li
J. Vis. Commun. Image Represent.6
2021 Saliency detection via cross-scale deep inference
Dakai Ren, Xiangming Wen, Jiazhong Chen, Zongyi Li
J. Vis. Commun. Image Represent.5
2020 Conditional Linear Regression
abstract
Work in machine learning and statistics commonly focuses on building models that capture the vast majority of data, possibly ignoring a segment of the population as outliers. However, there may not exist a good, simple model for the distribution, so we seek to find a small subset where there exists such a model. We give a computationally efficient algorithm with theoretical analysis for the conditional linear regression task, which is the joint task of identifying a significant portion of the data distribution, described by a k-DNF, along with a linear predictor on that portion with a small loss. In contrast to work in robust statistics on small subsets, our loss bounds do not feature a dependence on the density of the portion we fit, and compared to previous work on conditional linear regression, our algorithm’s running time scales polynomially with the sparsity of the linear predictor. We also demonstrate empirically that our algorithm can leverage this advantage to obtain a k-DNF with a better linear predictor in practice.
Diego Calderon, Brendan Juba, Zongyi Li, Lisa Ruan
AISTATS4
2020 Multipole Graph Neural Operator for Parametric Partial Differential Equations
abstract
One of the main challenges in using deep learning-based methods for simulating physical systems and solving partial differential equations (PDEs) is formulating physics-based data in the desired structure for neural networks. Graph neural networks (GNNs) have gained popularity in this area since graphs offer a natural way of modeling particle interactions and provide a clear way of discretizing the continuum models. However, the graphs constructed for approximating such tasks usually ignore long-range interactions due to unfavorable scaling of the computational complexity with respect to the number of nodes. The errors due to these approximations scale with the discretization of the system, thereby not allowing for generalization under mesh-refinement. Inspired by the classical multipole methods, we purpose a novel multi-level graph neural network framework that captures interaction at all ranges with only linear complexity. Our multi-level formulation is equivalent to recursively adding inducing points to the kernel matrix, unifying GNNs with multi-resolution matrix factorization of the kernel. Experiments confirm our multi-graph network learns discretization-invariant solution operators to PDEs and can be evaluated in linear time.
Zongyi Li, Nikola B. Kovachki, Kamyar Azizzadenesheli, Burigede Liu, Andrew M. Stuart, Kaushik Bhattacharya, Anima Anandkumar
NeurIPS1
2018 Conditional Linear Regression
Diego Calderon, Brendan Juba, Zongyi Li, Lisa Ruan
AAAI3
2018 Learning Abduction Using Partial Observability
Brendan Juba, Zongyi Li, Evan Miller
AAAI2
2018 Learning Abduction Under Partial Observability
abstract
Our work extends Juba’s formulation of learning abductive reasoning from examples, in which both the relative plausibility of various explanations, as well as which explanations are valid, are learned directly from data. We extend the formulation to consider partially observed examples, along with declarative background knowledge about the missing data. We show that it is possible to use implicitly learned rules together with the explicitly given declarative knowledge to support hypotheses in the course of abduction. We observe that when a small explanation exists, it is possible to obtain a much-improved guarantee in the challenging exception-tolerant setting.
Brendan Juba, Zongyi Li, Evan Miller
AAAI2