Yuwei Zhou

dblp:124/2955 · DBLP profile ↗
← Back
27ranked-venue papers
10as first author
26since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding
abstract
Long video understanding (LVU) is challenging due to rich and complicated multimodal clues in long temporal range. Current methods adopt reasoning to improve the model's ability to analyze complex video clues in long videos via text-form reasoning. However, the existing literature suffers from the fact that the text-only reasoning under fixed video context may exacerbate hallucinations since detailed crucial clues are often ignored under limited video context length due to the temporal redundancy of long videos. To address this gap, we propose Video-TwG, a curriculum reinforced framework that employs a novel Think-with-Grounding paradigm, enabling video LLMs to actively decide when to perform on-demand grounding during interleaved text–video reasoning, selectively zooming into question-relevant clips only when necessary. Video-TwG can be trained end-to-end in a straightforward manner, without relying on complex auxiliary modules or heavily annotated reasoning traces. In detail, we design a Two-stage Reinforced Curriculum Strategy, where the model first learns think-with-grounding behavior on a small short-video GQA dataset with grounding labels, and then scales to diverse general QA data with videos of diverse domains to encourage generalization. Further, to handle complex think-with-grounding reasoning for various kinds of data, we propose the TwG-GRPO algorithm, which features the fine-grained grounding reward, self-confirmed pseudo reward, and accuracy-gated mechanism. Finally, we propose to construct a new TwG-51K dataset that facilitates training. Experiments on Video-MME, LongVideoBench, and MLVU show that Video-TwG consistently outperforms strong LVU baselines. Further ablation validates the necessity of our Two-stage Reinforced Curriculum Strategy and shows our TwG-GRPO better leverages diverse unlabeled data to improve grounding quality and reduce redundant groundings without sacrificing QA performance. https://github.com/hlchen23/Video-TwG
Houlun Chen, Xin Wang 0019, Guangyao Li 0001, Yuwei Zhou, Jia Jia 0001, Wenwu Zhu 0001
SIGIR4
2026 Multi-Modal Generative AI: Multi-Modal LLMs, Diffusions, and the Unification
abstract
Multi-modal generative AI (Artificial Intelligence) has attracted increasing attention from both academia and industry. Particularly, two dominant families of techniques have emerged: i) Multi-modal large language models (LLMs) demonstrate impressive ability formulti-modal understanding; and ii) Diffusion models exhibit remarkable multi-modal powers in terms ofmulti-modal generation. Therefore, this paper provides a comprehensive overview of multi-modal generative AI, including multi-modal LLMs, diffusions, and the unification for understanding and generation. To lay a solid foundation for unified models, we first provide a detailed review of both multi-modal LLMs and diffusion models, respectively, including their probabilistic modeling procedure, multi-modal architecture design, and advanced applications to image/video LLMs as well as text-to-image/video generation. Furthermore, we explore the emerging efforts toward unified models for understanding and generation. To achieve the unification of understanding and generation, we investigate key designs including autoregressive-based and diffusion-based modeling, as well as dense and Mixture-of-Experts (MoE) architectures. We then introduce several strategies for unified models, analyzing their potential advantages and disadvantages. In addition, we summarize the common datasets widely used for multi-modal generative AI pretraining. Last but not least, we present several challenging future research directions that may contribute to the ongoing advancement of multi-modal generative AI.
Xin Wang 0019, Yuwei Zhou, Bin Huang 0004, Hong Chen 0011, Wenwu Zhu 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 Depth Estimation Based on Fisheye Cameras
abstract
Fisheye cameras, with their ultra-wide field of view, offer significant benefits for depth estimation in applications such as autonomous navigation, robotics, and immersive imaging by capturing more scene content from a single viewpoint. However, their strong radial distortion and varying spatial resolution across the image pose substantial challenges for accurate depth prediction. We present a deep learning–based framework for fisheye depth estimation that addresses these challenges while leveraging the wide coverage advantage. During training, rectified and synchronized stereo image pairs are used, with the right image and an estimated initial depth map reconstructing the left image. A refined spatial consistency loss is formulated by combining Structural Similarity Index Measure (SSIM) and L1 loss, with gradient-based weighting to emphasize disparity edges. To overcome the limitations of photometric loss in disparity learning, we normalize pixel intensities to better correlate disparity with appearance features. A fisheye-specific depth refinement module incorporates an uncertainty map derived from an inconsistency mask and a distortion distribution map, mitigating the effects of occlusion and high-distortion regions. This uncertainty map is used to weight the temporal warping loss, enhancing robustness against distortion-prone areas. During inference, only a single fisheye image is required to produce an accurate depth map. Experimental results demonstrate that our method improves reconstruction fidelity and robustness, making it well-suited for real-world fisheye-based depth estimation tasks.
Yuwei Zhou, Guoyu Lu 0001
IROS1
2025 ModuleTeam: Open-Set Multi-Conditional Image Generation with Training-Free Latent Mixture of Any Control Module
abstract
Multi-conditional image generation aims to create customized images that align with multiple specified conditions. Existing methods, whether through end-to-end training or by fine-tuning adapters to integrate pre-trained control modules of the same category (e.g., LoRA, IP-Adapter, ControlNet, T2I-Adapter), are restricted to a closed set of predefined input conditions. To overcome this limitation, we propose ModuleTeam, a training-free method for latent mixture of arbitrary control modules, capable of handling open-set conditions by incorporating the corresponding modules. The design of ModuleTeam is rooted in two key findings: (i) modules interfere with each other at the level of model parameters, and (ii) module weights contribute to the generated images by affecting the noise predictions within the diffusion process in an approximately linear manner. The first finding motivates our latent mixture approach, which mixes the control modules by aggregating their latent variables between diffusion model blocks. The second finding enables a multi-inference module reweighting strategy that balances module contributions to generation, requiring no additional training or fine-tuning overhead. Extensive results demonstrate that ModuleTeam not only outperforms existing methods but also provides flexibility in the types of conditions and scalability in their number.
Yuwei Zhou, Xin Wang 0019, Hong Chen 0011, Yipeng Zhang 0003, Zeyang Zhang 0001, Wenwu Zhu 0001
ACM Multimedia1
2025 High linearity GaN HEMT by optimized three-dimensional-gated modulation via top-MIS-gate nanowire channel structure
Can Gong, Minhan Mi, Yuwei Zhou, Hanzhen Li, Xinyi Wen, Sirui An, Xiang Du, Qing Zhu 0013, Xiaohua Ma 0001, Yue Hao 0001
Sci. China Inf. Sci.3
2025 VideoDreamer: Customized Multi-Subject Text-to-Video Generation With Disen-Mix Finetuning on Language-Video Foundation Models
abstract
Customized text-to-video generation aims to generate text-guided videos with user-given subjects, which has gained increasing attention. However, existing works are primarily limited to single-subject oriented text-to-video generation, leaving the more challenging problem of customized multi-subject generation unexplored. In this paper, we fill this gap and propose a novel VideoDreamer framework, which can generate temporally consistent text-guided videos that faithfully preserve the visual features of the given multiple subjects. Specifically, VideoDreamer adopts the pretrained Stable Diffusion with temporal modules as its base video generator, taking the power of the text-to-image model to generate diversified content. The video generator is further customized for multi-subjects, which leverages the proposed Disen-Mix Finetuning and Human-in-the-Loop Re-finetuning strategy, to tackle the attribute binding problem of multi-subject generation. Additionally, we present a disentangled motion customization strategy to finetune the temporal modules so that we can generate videos with both customized subjects and motions. To evaluate the performance of customized multi-subject text-to-video generation, we introduce the MultiStudioBench benchmark. Extensive experiments demonstrate the remarkable ability of VideoDreamer to generate videos with new content such as new events and backgrounds, tailored to the customized multiple subjects.
Hong Chen 0011, Xin Wang 0019, Guanning Zeng, Yipeng Zhang 0003, Yuwei Zhou, Feilin Han, Yaofei Wu, Wenwu Zhu 0001
IEEE Trans. Multim.5
2025 Automated Disentangled Sequential Recommendation with Large Language Models
abstract
Sequential recommendation aims to recommend the next items that a target user may have interest in based on the user’s sequence of past behaviors, which has become a hot research topic in both academia and industry. In the literature, sequential recommendation adopts a Sequence-to-Item or Sequence-to-Sequence training strategy, which supervises a sequential model with a user’s next one or more behaviors as the labels and the sequence of the past behaviors as the input. However, existing powerful sequential recommendation approaches employ more and more complex deep structures such as Transformer in order to accurately capture the sequential patterns, which heavily rely on hand-crafted designs on key attention mechanism to achieve state-of-the-art performance, thus failing to automatically obtain the optimal design of attention representation architectures in various scenarios with different data. Other works on classic automated deep recommender systems only focus on traditional settings, ignoring the problem of sequential scenarios. In this article, we study the problem of automated sequential recommendation, which faces two main challenges: (1) How can we design a proper search space tailored for attention automation in sequential recommendation, and (2) How can we accurately search effective attention representation architectures considering multiple user interests reflected in the sequential behavior. To tackle these challenges, we propose an automated disentangled sequential recommendation (AutoDisenSeq) model. In particular, we employ neural architecture search (NAS) and design a search space tailored for automated attention representation in attentive intention-disentangled sequential recommendation with an expressive and efficient space complexity of \(O(n^{2})\) given \(n\) as the number of layers. We further propose a context-aware parameter sharing mechanism taking characteristics of each sub-architecture into account to enable accurate architecture performance estimations and great flexibility for disentanglement of latent intention representation. Moreover, we propose AutoDisenSeq-large language model (LLM), which utilizes the textual understanding power of LLM as a guidance to refine the candidate list for recommendation from AutoDisenSeq. We conduct extensive experiments to show that our proposed AutoDisenSeq model and AutoDisenSeq-LLM model outperform existing baseline methods on four real-world datasets in both overall recommendation and cold-start recommendation scenarios.
Xin Wang 0019, Hong Chen 0011, Zirui Pan, Yuwei Zhou, Chaoyu Guan, Lifeng Sun, Wenwu Zhu 0001
ACM Trans. Inf. Syst.4
2024 DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation
abstract
Subject-driven text-to-image generation aims to generate customized images of the given subject based on the text descriptions, which has drawn increasing attention. Existing methods mainly resort to finetuning a pretrained generative model, where the identity-relevant information (e.g., the boy) and the identity-irrelevant information (e.g., the background or the pose of the boy) are entangled in the latent embedding space. However, the highly entangled latent embedding may lead to the failure of subject-driven text-to-image generation as follows: (i) the identity-irrelevant information hidden in the entangled embedding may dominate the generation process, resulting in the generated images heavily dependent on the irrelevant information while ignoring the given text descriptions; (ii) the identity-relevant information carried in the entangled embedding can not be appropriately preserved, resulting in identity change of the subject in the generated images. To tackle the problems, we propose DisenBooth, an identity-preserving disentangled tuning framework for subject-driven text-to-image generation. Specifically, DisenBooth finetunes the pretrained diffusion model in the denoising process. Different from previous works that utilize an entangled embedding to denoise each image, DisenBooth instead utilizes disentangled embeddings to respectively preserve the subject identity and capture the identity-irrelevant information. We further design the novel weak denoising and contrastive embedding auxiliary tuning objectives to achieve the disentanglement. Extensive experiments show that our proposed DisenBooth framework outperforms baseline models for subject-driven text-to-image generation with the identity-preserved embedding. Additionally, by combining the identity-preserved embedding and identity-irrelevant embedding, DisenBooth demonstrates more generation flexibility and controllability.
Hong Chen 0011, Yipeng Zhang 0003, Simin Wu, Xin Wang 0019, Xuguang Duan, Yuwei Zhou, Wenwu Zhu 0001
ICLR6
2024 CurBench: Curriculum Learning Benchmark
abstract
Curriculum learning is a training paradigm where machine learning models are trained in a meaningful order, inspired by the way humans learn curricula. Due to its capability to improve model generalization and convergence, curriculum learning has gained considerable attention and has been widely applied to various research domains. Nevertheless, as new curriculum learning methods continue to emerge, it remains an open issue to benchmark them fairly. Therefore, we develop CurBench, the first benchmark that supports systematic evaluations for curriculum learning. Specifically, it consists of 15 datasets spanning 3 research domains: computer vision, natural language processing, and graph machine learning, along with 3 settings: standard, noise, and imbalance. To facilitate a comprehensive comparison, we establish the evaluation from 2 dimensions: performance and complexity. CurBench also provides a unified toolkit that plugs automatic curricula into general machine learning processes, enabling the implementation of 15 core curriculum learning methods. On the basis of this benchmark, we conduct comparative experiments and make empirical analyses of existing methods. CurBench is open-source and publicly available at https://github.com/THUMNLab/CurBench.
Yuwei Zhou, Zirui Pan, Xin Wang 0019, Hong Chen 0011, Haoyang Li 0001, Yanwen Huang, Zhixiao Xiong, Fangzhou Xiong, Peiyang Xu, Shengnan Liu, Wenwu Zhu 0001
ICML1
2024 Simultaneous Super-resolution and Depth Estimation for Satellite Images Based on Diffusion Model
abstract
Satellite images provide an effective way to observe the earth surface on a large scale. 3D landscape models can provide critical structural information, such as forestry and crop growth. However, there has been very limited research to estimate the depth and the 3D models of the earth based on satellite images. LiDAR measurements on satellites are usually quite sparse. RGB images have higher resolution than LiDAR, but there has been little research on 3D surface measurements based on satellite RGB images. In comparison with in-situ sensing, satellite RGB images are usually low resolution. In this research, we explore the method that can enhance the satellite image resolution to generate super-resolution images and then conduct depth estimation and 3D reconstruction based on higher-resolution satellite images. Leveraging the strong generation capability of diffusion models, we developed a simultaneous diffusion model learning framework that can train diffusion models for both super-resolution images and depth estimation. With the super-resolution images and the corresponding depth maps, 3D surface reconstruction models with detailed landscape information can be generated. We evaluated the proposed methodology on multiple satellite datasets for both super-resolution and depth estimation tasks, which have demonstrated the effectiveness of our methodology.
Yuwei Zhou, Yangming Lee
IROS1
2024 Curriculum Learning for Multimedia in the Era of Large Language Models
abstract
This tutorial focuses on curriculum learning (CL), an important topic in machine learning, which gains an increasing amount of attention in the research community. CL is a learning paradigm that enables machines to learn from easy data to hard data, imitating the meaningful procedure of human learning with curricula. As an easy-to-use plug-in, CL has demonstrated its power in improving the generalization capacity and convergence rate of various models in a wide range of scenarios such as computer vision, natural language processing, reinforcement learning, etc. In particular, CL can also play an important role in multimedia applications. Therefore, it is essential to introduce CL to more scholars and researchers in the machine learning and multimedia community. However, there have been no tutorials on CL for multimedia so far, motivating the organization of this tutorial at ACM Multimedia 2024. To give a comprehensive tutorial on CL for multimedia, we plan to organize it from the following aspects: (1) theories, (2) approaches, (3) applications, (4) tools, and (5) future directions. First, we introduce the motivations, theories, and insights behind CL. Second, we advocate novel, high-quality approaches, as well as innovative solutions to the challenging problems in CL. Then we present the applications of CL in various scenarios, especially multimedia, followed by some relevant tools. In the end, we discuss open questions and future directions in the era of large language models. We believe this topic is at the core of the scope of ACM Multimedia and is attractive to the audience interested in machine learning and multimedia from both academia and industry.
Xin Wang 0019, Yuwei Zhou, Hong Chen 0011, Wenwu Zhu 0001
ACM Multimedia2
2024 DisenStudio: Customized Multi-Subject Text-to-Video Generation with Disentangled Spatial Control
abstract
Generating customized content in videos has received increasing attention recently. However, existing works primarily focus on customized text-to-video generation for single subject, suffering from subject-missing and attribute-binding problems when the video is expected to contain multiple subjects. Furthermore, existing models struggle to assign the desired actions to the corresponding subjects (action-binding problem), failing to achieve satisfactory multi-subject generation performance. To tackle the problems, in this paper, we propose DisenStudio, a novel framework that can generate text-guided videos for customized multiple subjects, given few images for each subject. Specifically, DisenStudio enhances a pretrained diffusion-based text-to-video model with our proposed spatial-disentangled cross-attention mechanism to associate each subject with the desired action. Then the model is customized for the multiple subjects with the proposed motion-preserved disentangled finetuning, which involves three tuning strategies: multi-subject co-occurrence tuning, masked single-subject tuning, and multi-subject motion-preserved tuning. The first two strategies guarantee the subject occurrence and preserve their visual attributes, and the third strategy helps the model maintain the temporal motion-generation ability when finetuning on static images. We conduct extensive experiments to demonstrate our proposed DisenStudio significantly outperforms existing methods in various metrics. Additionally, we show that DisenStudio can be used as a powerful tool for various controllable generation applications.
Hong Chen 0011, Xin Wang 0019, Yipeng Zhang 0003, Yuwei Zhou, Zeyang Zhang 0001, Siao Tang, Wenwu Zhu 0001
ACM Multimedia4
2024 DisenDreamer: Subject-Driven Text-to-Image Generation With Sample-Aware Disentangled Tuning
abstract
Subject-driven text-to-image generation aims to generate customized images of the given subject based on the text descriptions, which has drawn increasing attention recently. Existing methods mainly resort to finetuning a pretrained generative model, where the identity-relevant information (e.g., the boy) and the identity-irrelevant sample-specific information (e.g., the background or the pose of the boy) are entangled in the latent embedding space. However, the highly entangled latent embedding may lead to low subject identity fidelity and text prompt fidelity. To tackle the problems, we propose DisenDreamer, a sample-aware disentangled tuning framework for subject-driven text-to-image generation in this paper. Specifically, DisenDreamer finetunes the pretrained diffusion model in the denoising process. Different from previous works that utilize an entangled embedding to denoise, DisenDreamer instead utilizes a common text embedding to capture the identity-relevant information and a sample-specific visual embedding to capture the identity-irrelevant information. To disentangle the two embeddings, we further design the novel weak common denoising, weak sample-aware denoising, and the contrastive embedding auxiliary tuning objectives. Extensive experiments show that our proposed DisenDreamer framework outperforms baseline models for subject-driven text-to-image generation. Additionally, by combining the identity-relevant and the identity-irrelevant embedding, DisenDreamer demonstrates more generation flexibility and controllability.
Hong Chen 0011, Yipeng Zhang 0003, Xin Wang 0019, Xuguang Duan, Yuwei Zhou, Wenwu Zhu 0001
IEEE Trans. Circuits Syst. Video Technol.5
2023 Curriculum Co-disentangled Representation Learning across Multiple Environments for Social Recommendation
abstract
There exist complex patterns behind the decision-making processes of different individuals across different environments. For instance, in a social recommender system, various user behaviors are driven by highly entangled latent factors from two environments, i.e., consuming environment where users consume items and social environment where users connect with each other. Uncovering the disentanglement of these latent factors for users can benefit in enhanced explainability and controllability for recommendation. However, in literature there has been no work on social recommendation capable of disentangling user representations across consuming and social environments. To solve this problem, we study co-disentangled representation learning across different environments via proposing the curriculum co-disentangled representation learning (CurCoDis) model to disentangle the hidden factors for users across both consuming and social environments. To co-disentangle joint representations for user-item consumption and user-user social graph simultaneously, we partition the social graph into equal-size sub-graphs with minimum number of edges being cut, and design a curriculum weighing strategy for subgraph training through measuring the complexity of subgraphs via Descartes' rule of signs. We further develop the prototype-routing optimization mechanism, which achieves co-disentanglement of user representations across consuming and social environments. Extensive experiments for social recommendation demonstrate that our proposed CurCoDis model can significantly outperform state-of-the-art methods on several real-world datasets.
Xin Wang 0019, Zirui Pan, Yuwei Zhou, Hong Chen 0011, Chendi Ge, Wenwu Zhu 0001
ICML3
2023 Intra- and Inter-Modal Curriculum for Multimodal Learning
abstract
Multimodal learning has been widely studied and applied due to its improvement over previous unimodal tasks and its effectiveness on emerging multimodal challenges. However, it has been reported that modal encoders are under-optimized in multimodal learning in contrast to unimodal learning, especially when some modalities are dominant over others. Existing solutions to this problem suffer from two limitations: i) they merely focus on inter-modal balance, failing to consider the influence of intra-modal data on each modality; ii) their implementations heavily rely on unimodal performances or losses, thus being suboptimal for the tasks requiring modal interactions (e.g., visual question answering). To tackle these limitations, we propose I2MCL, a generic Intra- and Inter-Modal Curriculum Learning framework which simultaneously considers both data difficulty and modality balance for multimodal learning. In the intra-modal curriculum, we adopt a pretrained teacher model to obtain knowledge distillation loss as the difficulty measurer, which determines the data weights within the corresponding modality. In the inter-modal curriculum, we utilize a Pareto optimization strategy to measure and compare the gradients from distillation loss and task loss across modalities, capable of determining whether a modality should learn from the task or its teacher. Empirical experiments on various tasks including multimodal classification, visual question answering and visual entailment demonstrate that our proposed I2MCL is able to tackle the under-optimized modality problem and bring consistent improvement to multimodal learning.
Yuwei Zhou, Xin Wang 0019, Hong Chen 0011, Xuguang Duan, Wenwu Zhu 0001
ACM Multimedia1
2023 Global-Local GraphFormer: Towards Better Understanding of User Intentions in Sequential Recommendation
abstract
Transformer-based model has gained great success in the multimedia sequential recommendation task due to its strong ability to handle sequential data. However, existing Transformer-based models regard the items in the sequential data as a user-specific fully-connected graph (local graph) and only explicitly consider the temporal information in the local graph to capture the users’ intentions, ignoring the fact that the user-item bipartite graph (global graph) may carry important relation patterns to the sequential items. Additionally, it is still unclear whether (and how) the information hidden in the global graphs can help the Transformer-based models better understand the users’ sequential behavior according to the current literature. To investigate this important problem, we propose to utilize the global graph information to help the Transformer-based sequential recommendation, where the information from different modalities, i.e., user-item interactions in the global graph and the temporal patterns in the historical sequences, are taken into account jointly. In concrete, we propose two Global-Local (GL) GraphFormer models for utilizing both the global graph and local temporal information. One GL-GraphFormer is able to gift the Transformer-based model with both first- and second-order graph information through two specifically designed encodings. The other GL-GraphFormer transfers higher-order graph information into the local Transformer with pretrained Graph Neural Networks (GNNs). Extensive experiments on several real-world datasets demonstrate that i) our proposed GL-GraphFormers can bring substantial improvement over baseline methods, and ii) the benefits of different orders of global graph information vary with the dataset sparsity.
Hong Chen 0011, Bin Huang 0004, Xin Wang 0019, Yuwei Zhou, Wenwu Zhu 0001
MMAsia4
2023 Joint Data-Task Generation for Auxiliary Learning
abstract
Current auxiliary learning methods mainly adopt the methodology of reweighing losses for the manually collected auxiliary data and tasks. However, these methods heavily rely on domain knowledge during data collection, which may be hardly available in reality. Therefore, current methods will become less effective and even do harm to the primary task when unhelpful auxiliary data and tasks are employed. To tackle the problem, we propose a joint data-task generation framework for auxiliary learning (DTG-AuxL), which can bring benefits to the primary task by generating the new auxiliary data and task in a joint manner. The proposed DTG-AuxL framework contains a joint generator and a bi-level optimization strategy. Specifically, the joint generator contains a feature generator and a label generator, which are designed to be applicable and expressive for various auxiliary learning scenarios. The bi-level optimization strategy optimizes the joint generator and the task learning model, where the joint generator is effectively optimized in the upper level via the implicit gradient from the primary loss and the explicit gradient of our proposed instance regularization, while the task learning model is optimized in the lower level by the generated data and task. Extensive experiments show that our proposed DTG-AuxL framework consistently outperforms existing methods in various auxiliary learning scenarios, particularly when the manually collected auxiliary data and tasks are unhelpful.
Hong Chen 0011, Xin Wang 0019, Yuwei Zhou, Yijian Qin, Chaoyu Guan, Wenwu Zhu 0001
NeurIPS3
2023 Cross-domain Recommendation with Behavioral Importance Perception
abstract
Cross-domain recommendation (CDR) aims to leverage the source domain information to provide better recommendation for the target domain, which is widely adopted in recommender systems to alleviate the data sparsity and cold-start problems. However, existing CDR methods mostly focus on designing effective model architectures to transfer the source domain knowledge, ignoring the behavior-level effect during the loss optimization process, where behaviors regarding different aspects in the source domain may have different importance for the CDR model optimization. The ignorance of the behavior-level effect will cause the carefully designed model architectures ending up with sub-optimal parameters, which limits the recommendation performance. To tackle the problem, we propose a generic behavioral importance-aware optimization framework for cross-domain recommendation (BIAO). Specifically, we propose a behavioral perceptron which predicts the importance of each source behavior according to the corresponding item’s global impact and local user-specific impact. The joint optimization process of the CDR model and the behavioral perceptron is formulated as a bi-level optimization problem. In the lower optimization, only the CDR model is updated with weighted source behavior loss and the target domain loss, while in the upper optimization, the behavioral perceptron is updated with implicit gradient from a developing dataset obtained through the proposed reorder-and-reuse strategy. Extensive experiments show that our proposed optimization framework consistently improves the performance of different cross-domain recommendation models in 7 cross-domain scenarios, demonstrating that our method can serve as a generic and powerful tool for cross-domain recommendation1.
Hong Chen 0011, Xin Wang 0019, Ruobing Xie, Yuwei Zhou, Wenwu Zhu 0001
WWW4
2023 Degradation induced by holes in Si3N4/AlGaN/GaN MIS HEMTs under off-state stress with UV light
Qing Zhu 0013, Jiejie Zhu, Minhan Mi, Yuwei Zhou, Ziyue Zhao 0003, Xiaohua Ma 0001, Yue Hao 0001
Sci. China Inf. Sci.6
2023 Disentangled Representation Learning for Recommendation
abstract
There exist complex interactions among a large number of latent factors behind the decision making processes of different individuals, which drive the various user behavior patterns in recommender systems. These factors hidden in those diverse behaviors demonstrate highly entangled patterns, covering from high-level user intentions to low-level individual preferences. Uncovering the disentanglement of these latent factors can benefit in enhanced robustness, interpretability, and controllability during representation learning for recommendation. However, the large degree of entanglement within latent factors poses great challenges for learning representations that disentangle them, and remains largely unexplored in literature. In this paper, we present the SEMantic MACRo-mIcro Disentangled Variational Auto-Encoder (SEM-MacridVAE) model for learning disentangled representations from user behaviors, taking item semantic information into account. Our SEM-MacridVAE model achieves macro disentanglement by inferring the high-level concepts associated with user intentions (e.g., to buy a pair of shoes or a laptop) through a prototype routing mechanism, as well as capturing the individual preferences with respect to different concepts separately. The micro disentanglement is guaranteed through a micro-disentanglement regularizer stemming from an information-theoretic interpretation of VAEs, which forces each dimension of the representations to independently reflect an isolated low-level factor (e.g., the size or the color of a shirt). The semantic information including visual and categorical signals extracted from candidate items is utilized to further boost the recommendation performance of the proposed SEM-MacridVAE model. Empirical experiments demonstrate that our proposed approach is able to achieve significant improvement over the state-of-the-art baselines. We also show that the learned representations are interpretable and controllable, capable of potentially leading to a new paradigm for recommendation where users have fine-grained control over some target aspects of the recommendation candidates.
Xin Wang 0019, Hong Chen 0011, Yuwei Zhou, Wenwu Zhu 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Temporal MLP Network for PM 2.5 Estimation
abstract
Particulate Matter (PM) 2.5 is a critical factor to measure in the environment. However, accurate PM 2.5 measurement relies on special devices, which leads to limited spatial coverage. This paper focuses on the development of PM 2.5 estimation based on satellite images. We have developed a deep temporal multiple layer perception (MLP) neural network based on various satellite imaging inputs captured from different time periods together with meteorological inputs. In addition, the temporal MLP simultaneously learns both PM 2.5 and PM 10 which can improve the performance of each task. Experiments demonstrate that our method outperforms other machine learning algorithms used in PM 2.5 estimation.
Yuwei Zhou, John P. Kerekes
IGARSS1
2022 Curriculum-NAS: Curriculum Weight-Sharing Neural Architecture Search
abstract
Neural Architecture Search (NAS) is an effective way to automatically design neural architectures for various multimedia applications. Weight-sharing, as one of the most popular NAS strategies, has been widely adopted due to its search efficiency. Existing weight-sharing NAS methods overlook the influence of data distribution and treat each data sample equally. Contrastively, in this paper, we empirically discover that different data samples have different influences on architectures, e.g., some data samples are easy to fit by certain architectures but hard by others. Hence, there exist architectures with better performances on early data samples being more likely to be discovered in the whole NAS searching process, which leads to a suboptimal searching result. To tackle this problem, we propose Curriculum-NAS, a curriculum training framework on weight-sharing NAS, which dynamically changes the training data weights during the searching process. In particular, Curriculum-NAS utilizes the multiple subnets included in weight-sharing NAS to jointly assess data uncertainty, which serves as the difficulty criterion in a curriculum manner, so that the potentially optimal architectures can obtain higher probability of being fully trained and discovered. Extensive experiments on several image and text datasets demonstrate that our Curriculum-NAS can bring consistent improvement over existing weight-sharing NAS. The code is available online at https://github.com/zhouyw16/curriculum-nas.
Yuwei Zhou, Xin Wang 0019, Hong Chen 0011, Xuguang Duan, Chaoyu Guan, Wenwu Zhu 0001
ACM Multimedia1
2022 CurML: A Curriculum Machine Learning Library
abstract
Curriculum learning (CL) is a machine learning paradigm gradually learning from easy to hard, which is inspired by human curricula. As an easy-to-use and general training strategy, CL has been widely applied to various multimedia tasks covering images, texts, audios, videos, etc. The effectiveness of CL has recently facilitated an increasing number of new CL algorithms. However, there has been no open-source library for curriculum learning, making it hard to reproduce, evaluate and compare the numerous CL algorithms on fair benchmarks and settings. To ease and promote future research on CL, we develop CurML, the first Curriculum Machine L earning library to integrate existing CL algorithms into a unified framework. It is convenient to use and flexible to customize by calling the provided five APIs, which are designed for easily plugging into a general training process and conducting the data-oriented, model-oriented and loss-oriented curricula. Furthermore, we present empirical results obtained by CurML to demonstrate the advantages of our library. The code is available online at https://github.com/THUMNLab/CurML.
Yuwei Zhou, Hong Chen 0011, Zirui Pan, Chuanhao Yan, Fanqi Lin, Xin Wang 0019, Wenwu Zhu 0001
ACM Multimedia1
2022 Module-Aware Optimization for Auxiliary Learning
abstract
Auxiliary learning is a widely adopted practice in deep learning, which aims to improve the model performance on the primary task by exploiting the beneficial information in the auxiliary loss. Existing auxiliary learning methods only focus on balancing the auxiliary loss and the primary loss, ignoring the module-level auxiliary influence, i.e., an auxiliary loss will be beneficial for optimizing specific modules within the model but harmful to others, failing to make full use of auxiliary information. To tackle the problem, we propose a Module-Aware Optimization approach for Auxiliary Learning (MAOAL). The proposed approach considers the module-level influence through the learnable module-level auxiliary importance, i.e., the importance of each auxiliary loss to each module. Specifically, the proposed approach jointly optimizes the module-level auxiliary importance and the model parameters in a bi-level manner. In the lower optimization, the model parameters are optimized with the importance parameterized gradient, while in the upper optimization, the module-level auxiliary importance is updated with the implicit gradient from a small developing dataset. Extensive experiments show that our proposed MAOAL method consistently outperforms state-of-the-art baselines for different auxiliary losses on various datasets, demonstrating that our method can serve as a powerful generic tool for auxiliary learning.
Hong Chen 0011, Xin Wang 0019, Yue Liu 0025, Yuwei Zhou, Chaoyu Guan, Wenwu Zhu 0001
NeurIPS4
2022 Finding Feasible Systems for Subjective Constraints Using Recycled Observations
abstract
We consider the problem of finding a set of feasible or near-feasible systems among a finite number of simulated systems in the presence of stochastic constraints. When the constraints are subjective, a decision maker may want to test multiple threshold values for the constraints. Or the decision maker may simply want to determine how a set of feasible systems changes as constraints become more strict with the objective of pruning systems or finding the system with the best performance. When only the constraint thresholds change for the same set of underlying systems, it is natural to reuse observations collected from the feasibility check with a different threshold value. We present an indifference-zone procedure that recycles observations and provide an overall probability of correct decision for all threshold values. Our numerical experiments show that the proposed procedure performs well in reducing the required number of observations while providing a statistical guarantee on the probability of correct decision. Summary of Contribution: We consider the problem of determining the feasibility of a finite number of systems in the presence of subjective constraints on performance measures that can be estimated only by stochastic simulation. Specifically, our work focuses on the situation where the decision maker is willing to relax some constraint thresholds if necessary to achieve feasibility. Also, we discuss how our proposed procedures can help select the best system in the presence of multiple objectives. This is the first work that considers subjective constraints in the field of ranking and selection in simulation and it provides a practically useful decision-making tool with more flexibility on feasibility determination. History: Accepted by Bruno Tuffin, Area Editor for Simulation. Supplemental Material: The online supplement is available at https://doi.org/10.1287/ijoc.2022.1227 .
Yuwei Zhou, Sigrún Andradóttir, Seong-Hee Kim, Chuljin Park
INFORMS J. Comput.1
2021 PM2.5 Classification Through Convolutional Recurrent Neural Networks Applied to Modis AOD and TOA Reflectance Images
abstract
Traditional measurement of the air pollutant particulate matter known as PM2.5 is performed by ground monitoring stations. These measurements are sparse in spatial extent due to the limited number of monitoring stations and their uneven distribution. In this study we explored the use of satellite images together with a convolutional recurrent neural network to predict a category of PM2.5 concentration for continuous spatial samples across a region. Aerosol optical depth (AOD) and Top of Atmosphere (TOA) products from NASA's MODIS instrument were used to provide spatial and temporal data to the network trained to predict the PM2.5 category. Results demonstrate an over 70% classification accuracy for the PM2.5 category.
Yuwei Zhou, John P. Kerekes
IGARSS1
2020 CSABlock-based Cascade RCNN for Breast Mass Detection in Mammogram
abstract
Early screening and diagnosis of breast mass are essential for the prevention of breast cancer. There are some reasons to make mass detection be difficult and challenging. First, the resolution of mammography is very large and mass tissues are often subtle. Second, some mass overlap with the normal tissues which own similar texture. In this paper, we propose a novel attention module channel self-attention block (CSABlock), it can make better use of inter-layer features and strengthen the detection capability of the cascade R-CNN model. In order to further improve the detection quality, we also use a new domain-adaptive pre-training strategy. Experiments show that the proposed method achieves an average precision (AP) of 0.822 and average recall (AR) of 0.949, outperforms the state-of-the-art methods.
Qingfeng Wang 0004, Zhiqin Liu, Jun Huang 0005, Yuwei Zhou, Weiyun Xu
BIBM5