EDBT 2026 Demo / reviewers in the wild / expert
Yunhui Guo
dblp:165/3105
· DBLP profile ↗
42ranked-venue papers
19as first author
31since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 10 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 9 first-author · 18 since 2021Computer networks · 5 · 5 first-author · 4 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?abstractUnlike traditional visual segmentation, audio-visual segmentation (AVS) requires the model not only to identify and segment objects but also to determine whether they are sound sources. Recent AVS approaches have achieved impressive performance on standard benchmarks. Yet, an important question remains: Do these models genuinely integrate audio-visual cues to segment sounding objects? Our study reveals a fundamental bias in current methods: they tend to generate segmentation masks based predominantly on visual salience, irrespective of the audio context, resulting in unreliable predictions when sounds are absent or irrelevant. To address this challenge, we introduce AVSBench-Robust, a comprehensive benchmark incorporating diverse negative audio scenarios, including silence, noise, and off-screen sounds. We also propose a simple yet effective approach combining balanced training with negative samples and classifier-guided similarity learning. Our extensive experiments show that while state-of-the-art AVS methods consistently fail under negative audio conditions, our approach achieves remarkable improvements in both standard metrics and robustness measures, maintaining near-perfect false positive rates while preserving high-quality segmentation performance. Ziru Huang, Yunhui Guo, Yapeng Tian |
AAAI | 4 |
| 2026 | Effective Beamfocusing Design and Power Allocation for Near-Field Secure CommunicationsabstractThe beamfocusing property can be effectively utilized to enhance system performance in near-field communications. In this paper, the secure transmission for near-field extremely large-scale MIMO (XL-MIMO) communication systems is investigated, where the base station (BS) transmits private confidential information to legitimate users under the threat of a potential eavesdropper. To explore the unique characteristics of near-field physical layer security (PLS), a specific scenario involving a single legitimate user and a single eavesdropper is first considered. It is rigorously proved that 1) Artificial noise (AN) plays an essential role in near-field secure communications, facilitating the transformation of insecure systems into secure ones, and 2) allocating even a small portion of power for AN substantially improves security, compared with scenarios where AN is not employed. Additionally, for the general scenario with multiple legitimate users, a conventional scheme is initially introduced by employing semi-definite relaxation and successive convex approximation algorithms. Subsequently, an efficient low-complexity beamforming scheme incorporating AN is proposed to ensure secure transmission in the near-field region. Furthermore, the closed-form solution for optimal beamforming power allocation is derived. Extensive numerical results show that 1) the utilization of near-field beamfocusing exhibits significant performance gains in enhancing PLS in comparison with far-field communications, and 2) the proposed scheme achieves performance comparable to the high computational complexity conventional scheme, while significantly reducing computational complexity. Yunhui Guo, Yang Zhang 0013, Yaxin Ren, Minghao Shang, Lihua Pang, Yuanwei Liu |
IEEE Trans. Wirel. Commun. | 1 |
| 2025 | PALM: Pushing Adaptive Learning Rate Mechanisms for Continual Test-Time AdaptationabstractReal-world vision models in dynamic environments face rapid shifts in domain distributions, leading to decreased recognition performance. Using unlabeled test data, continuous test-time adaptation (CTTA) directly adjusts a pre-trained source discriminative model to these changing domains. A highly effective CTTA method involves applying layer-wise adaptive learning rates for selectively adapting pre-trained layers. However, it suffers from the poor estimation of domain shift and the inaccuracies arising from the pseudo-labels. This work aims to overcome these limitations by identifying layers for adaptation via quantifying model prediction uncertainty without relying on pseudo-labels. We utilize the magnitude of gradients as a metric, calculated by backpropagating the KL divergence between the softmax output and a uniform distribution, to select layers for further adaptation. Subsequently, for the parameters exclusively belonging to these selected layers, with the remaining ones frozen, we evaluate their sensitivity to approximate the domain shift and adjust their learning rates accordingly. We conduct extensive image classification experiments on CIFAR-10C, CIFAR-100C, and ImageNet-C, demonstrating the superior efficacy of our method compared to prior approaches. Sarthak Kumar Maharana, Baoming Zhang, Yunhui Guo |
AAAI | 3 |
| 2025 | H2ST: Hierarchical Two-Sample Tests for Continual Out-of-Distribution DetectionabstractTask Incremental Learning (TIL) is a specialized form of Continual Learning (CL) in which a model incrementally learns from non-stationary data streams. Existing TIL methodologies operate under the closed-world assumption, presuming that incoming data remains in-distribution (ID). However, in an open-world setting, incoming samples may originate from out-of-distribution (OOD) sources, with their task identities inherently unknown. Continually detecting OOD samples presents several challenges for current OOD detection methods: reliance on model outputs leads to excessive dependence on model performance, selecting suitable thresholds is difficult, hindering real-world deployment, and binary ID/OOD classification fails to provide task-level identification. To address these issues, we propose a novel continual OOD detection method called the Hierarchical Two-sample Tests (H2ST). H2ST eliminates the need for threshold selection through hypothesis testing and utilizes feature maps to better exploit model capabilities without excessive dependence on model performance. The proposed hierarchical architecture enables task-level detection with superior performance and lower overhead compared to non-hierarchical classifier two-sample tests. Extensive experiments and analysis validate the effectiveness of H2ST in open-world TIL scenarios and its superiority to the existing methods. Code is available at https://github.com/YuhangLiuu/H2ST. Yunhui Guo |
CVPR | 3 |
| 2025 | BATCLIP: Bimodal Online Test-Time Adaptation for CLIP
Sarthak Kumar Maharana, Baoming Zhang, Leonid Karlinsky, Rogério Feris, Yunhui Guo |
ICCV | 5 |
| 2025 | Adapting Pre-Trained Vision Models for Novel Instance Detection and SegmentationabstractNovel Instance Detection and Segmentation (NIDS) aims at detecting and segmenting novel object instances given a few examples of each instance. We propose a unified, simple, yet effective framework (NIDS-Net) comprising object proposal generation, embedding creation for both instance templates and proposal regions, and embedding matching for instance label assignment. Leveraging recent advancements in large vision methods, we utilize Grounding DINO and Segment Anything Model (SAM) to obtain object proposals with accurate bounding boxes and masks. Central to our approach is the generation of high-quality instance embeddings. We utilize foreground feature averages of patch embeddings from the DINOv2 ViT backbone, followed by refinement through a weight adapter mechanism that we introduce.We show experimentally that our weight adapter can adjust the embeddings locally within their feature space and effectively limit overfitting in the few-shot setting. Furthermore, the weight adapter optimizes weights to enhance the distinctiveness of instance embeddings during similarity computation. This methodology enables a straightforward matching strategy that results in significant performance gains. Our framework surpasses current state-of-the-art methods, demonstrating notable improvements in four detection datasets. In the segmentation tasks on seven core datasets of the BOP challenge, our method outperforms the leading published RGB methods and remains competitive with the best RGB-D method. We have also verified our method using real-world images from a Fetch robot and a RealSense camera.1 Yangxiao Lu, Jishnu Jaykumar, Yunhui Guo, Nicholas Ruozzi, Yu Xiang 0001 |
IROS | 3 |
| 2025 | <tt>AVROBUSTBENCH</tt>: Benchmarking the Robustness of Audio-Visual Recognition Models at Test-Time
Sarthak Kumar Maharana, Saksham Singh Kushwaha, Baoming Zhang, Adrian Rodriguez, Songtao Wei, Yapeng Tian, Yunhui Guo |
NeurIPS | 7 |
| 2025 | Modeling and Analysis of Spatial Correlation for Near-Field CommunicationsabstractThe near-field spatial correlation for multiple-input multiple-output (MIMO) communications in multi-path fading channels is analyzed. Based on the general non-uniform spherical wave (NUSW) model, an analytical integral-form expression for near-field spatial correlation is derived, which generalizes the conventional uniform plane wave (UPW)-based far-field spatial correlation. Furthermore, by considering the specific von Mises-Fisher distribution of scatterer locations, a simplified closed-form near-field spatial correlation expression is derived. It is rigorously proved that 1) in contrast to the far-field spatial correlation, the near-field spatial correlation no longer exhibits spatial stationary property, and 2) the NUSW-based near-field spatial correlation model depends on the power location spectrum, which encompasses both the angles and distances of the scatterers. Next, the developed near-field spatial correlation can be utilized to derive a closed-form expression for the effective degrees of freedom (EDoF). Additionally, a correlation-based stochastic channel model is constructed for MIMO communications, from which an optimal transmission strategy is devised. Subsequently, the power allocation is optimized to achieve the maximum ergodic spectral efficiency. Numerical results validate 1) the significance of near-field spatial correlation modeling for MIMO communications, 2) near-field MIMO exhibits a higher EDoF than the conventional far-field counterpart, and 3) the constructed correlation-based stochastic channel model facilitates the derivation of an optimal transmission strategy, thereby maximizing ergodic spectral efficiency. Yunhui Guo, Yang Zhang 0013, Zhaolin Wang 0001, Yuanwei Liu |
IEEE Trans. Commun. | 1 |
| 2025 | WonderHuman: Hallucinating Unseen Parts in Dynamic 3D Human ReconstructionabstractIn this paper, we present WonderHuman to reconstruct dynamic human avatars from a monocular video for high-fidelity novel view synthesis. Previous dynamic human avatar reconstruction methods typically require the input video to have full coverage of the observed human body. However, in daily practice, one typically has access to limited viewpoints, such as monocular front-view videos, making it a cumbersome task for previous methods to reconstruct the unseen parts of the human avatar. To tackle the issue, we present WonderHuman, which leverages 2D generative diffusion model priors to achieve high-quality, photorealistic reconstructions of dynamic human avatars from monocular videos, including accurate rendering of unseen body parts. Our approach introduces a Dual-Space Optimization technique, applying Score Distillation Sampling (SDS) in both canonical and observation spaces to ensure visual consistency and enhance realism in dynamic human reconstruction. Additionally, we present a View Selection strategy and Pose Feature Injection to enforce the consistency between SDS predictions and observed data, ensuring pose-dependent effects and higher fidelity in the reconstructed avatar. In the experiments, our method achieves SOTA performance in producing photorealistic renderings from the given monocular video, particularly for those challenging unseen parts. Zilong Wang 0013, Zhiyang Dou, Yuan Liu 0025, Cheng Lin 0001, Yunhui Guo, Xin Li 0003, Wenping Wang 0001, Xiaohu Guo |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2024 | Continual Learning in an Open and Dynamic WorldabstractBuilding autonomous agents that can process massive amounts of real-time sensor-captured data is essential for many real-world applications including autonomous vehicles, robotics and AI in medicine. As the agent often needs to explore in a dynamic environment, it is thus a desirable as well as challenging goal to enable the agent to learn over time without performance degradation. Continual learning aims to build a continual learner which can learn new concepts over the data stream while preserving previously learnt concepts. In the talk, I will survey three pieces of my recent research on continual learning (i) supervised continual learning, (ii) unsupervised continual learning, and (iii) multi-modal continual learning. In the first work, I will discuss a supervised continual learning algorithm called MEGA which dynamically balances the old tasks and the new task. In the second work, I will discuss unsupervised continual learning algorithms which learn representation continually without access to the labels. In the third work, I will elaborate an efficient continual learning algorithm that can learn multiple modalities continually without forgetting. Yunhui Guo |
AAAI | 1 |
| 2024 | Inconsistency-Based Data-Centric Active Open-Set AnnotationabstractActive learning, a method to reduce labeling effort for training deep neural networks, is often limited by the assumption that all unlabeled data belong to known classes. This closed-world assumption fails in practical scenarios with unknown classes in the data, leading to active open-set annotation challenges. Existing methods struggle with this uncertainty. We introduce NEAT, a novel, computationally efficient, data-centric active learning approach for open-set data. NEAT differentiates and labels known classes from a mix of known and unknown classes, using a clusterability criterion and a consistency mea- sure that detects inconsistencies between model predictions and feature distribution. In contrast to recent learning-centric solutions, NEAT shows superior performance in active open- set annotation, as our experiments confirm. Additional details on the further evaluation metrics, implementation, and archi- tecture of our method can be found in the public document at https://arxiv.org/pdf/2401.04923.pdf. Ruiyu Mao, Ouyang Xu, Yunhui Guo |
AAAI | 3 |
| 2024 | Unsupervised Feature Learning with Emergent Data-Driven PrototypicalityabstractGiven a set of images, our goal is to map each image to a point in a feature space such that, not only point proximity indicates visual similarity, but where it is located directly encodes how prototypical the image is according to the dataset. Our key insight is to perform unsupervised feature learning in hyperbolic instead of Euclidean space, where the distance between points still reflects image similarity, yet we gain additional capacity for representing prototypicality with the location of the point: The closer it is to the origin, the more prototypical it is. The latter property is simply emergent from optimizing the metric learning objective: The image similar to many training instances is best placed at the center of corresponding points in Euclidean space, but closer to the origin in hyperbolic space. We propose an unsupervised feature learning algorithm in Hyperbolic space with sphere pACKing. HACK first generates uniformly packed particles in the Poincaré ball of hyperbolic space and then assigns each image uniquely to a particle. With our feature mapper simply trained to spread out training instances in hyperbolic space, we observe that images move closer to the origin with congealing - a warping process that aligns all the images and makes them appear more common and similar to each other, validating our idea of unsupervised prototypicality discovery. We demonstrate that our data-driven prototypicality provides an easy and superior unsupervised instance selection to reduce sample complexity, increase model generalization with atypical instances and robustness with typical ones. Yunhui Guo, Youren Zhang, Yubei Chen, Stella X. Yu |
CVPR | 1 |
| 2024 | Segment Every Out-of-Distribution ObjectabstractSemantic segmentation models, while effective for indistribution categories, face challenges in real-world deployment due to encountering out-of-distribution (OoD) objects. Detecting these OoD objects is crucial for safety- critical applications. Existing methods rely on anomaly scores, but choosing a suitable threshold for generating masks presents difficulties and can lead to fragmentation and inaccuracy. This paper introduces a method to convert anomaly Score To segmentation Mask, called S2M, a simple and effective framework for OoD detection in semantic segmentation. Unlike assigning anomaly scores to pixels, S2M directly segments the entire OoD object. By transforming anomaly scores into prompts for a promptable segmentation model, S2M eliminates the need for thresh- old selection. Extensive experiments demonstrate that S2M outperforms the state-of-the-art by approximately 20% in IoU and 40% in mean F1 score, on average, across various benchmarks including Fishyscapes, Segment-Me-If- You-Can, and RoadAnomaly datasets. Code is available at https://github.com/WenjieZhao1/S2M. Yu Xiang 0001, Yunhui Guo |
CVPR | 5 |
| 2024 | Not Just Change the Labels, Learn the Features: Watermarking Deep Neural Networks with Multi-view Data
Sarthak Kumar Maharana, Yunhui Guo |
ECCV (65) | 3 |
| 2024 | Energy-Efficient Near-Field Wideband Beamforming with Circular ArraysabstractThe beamforming performance of the uniform circular array (UCA) in near-field wideband communication systems is investigated. Firstly, the unique beam squint effect in near-field wideband UCA systems is analyzed in both the distance and angular domains. It is demonstrated that the generated beams at different frequencies are focused at different locations, resulting in significant beamforming loss. To alleviate this unique beam squint effect and facilitate beamfocusing, an energy-efficient beamforming scheme based on the true-time delay (TTD) architecture is proposed. Specifically, the phase shifters (PSs) and the time delay of TTDs are designed based on the analytical formula for beamforming gain. Additionally, the minimum number of TTDs required to achieve a predetermined beamforming gain while minimizing energy consumption is quantified. Numerical results show that the proposed beamforming scheme effectively eliminates the near-field beam squint and outperforms the conventional schemes in terms of spectral efficiency and energy efficiency. Yunhui Guo, Yang Zhang 0013, Zhaolin Wang 0001, Yuanwei Liu, Zhiguo Ding 0001 |
GLOBECOM | 1 |
| 2024 | SkinCON: Towards Consensus for the Uncertainty of Skin Cancer Sub-typing Through Distribution Regularized Adaptive Predictive Sets (DRAPS)
Zhihang Ren, Xinrong Xie, Erik P. Duhaime, Kathy Fang, Tapabrata Chakraborti, Yunhui Guo, Stella X. Yu, David Whitney |
MICCAI (1) | 8 |
| 2024 | STONE: A Submodular Optimization Framework for Active 3D Object Detectionabstract3D object detection is fundamentally important for various emerging applications, including autonomous driving and robotics. A key requirement for training an accurate 3D object detector is the availability of a large amount of LiDAR-based point cloud data. Unfortunately, labeling point cloud data is extremely challenging, as accurate 3D bounding boxes and semantic labels are required for each potential object. This paper proposes a unified active 3D object detection framework, for greatly reducing the labeling cost of training 3D object detectors. Our framework is based on a novel formulation of submodular optimization, specifically tailored to the problem of active 3D object detection. In particular, we address two fundamental challenges associated with active 3D object detection: data imbalance and the need to cover the distribution of the data, including LiDAR-based point cloud data of varying difficulty levels. Extensive experiments demonstrate that our method achieves state-of-the-art performance with high computational efficiency compared to existing active learning methods. The code is available at [https://github.com/RuiyuM/STONE](https://github.com/RuiyuM/STONE) Ruiyu Mao, Sarthak Kumar Maharana, Rishabh Iyer 0001, Yunhui Guo |
NeurIPS | 4 |
| 2024 | Continual Audio-Visual Sound SeparationabstractIn this paper, we introduce a novel continual audio-visual sound separation task, aiming to continuously separate sound sources for new classes while preserving performance on previously learned classes, with the aid of visual guidance. This problem is crucial for practical visually guided auditory perception as it can significantly enhance the adaptability and robustness of audio-visual sound separation models, making them more applicable for real-world scenarios where encountering new sound sources is commonplace. The task is inherently challenging as our models must not only effectively utilize information from both modalities in current tasks but also preserve their cross-modal association in old tasks to mitigate catastrophic forgetting during audio-visual continual learning. To address these challenges, we propose a novel approach named ContAV-Sep ($\textbf{Cont}$inual $\textbf{A}$udio-$\textbf{V}$isual Sound $\textbf{Sep}$aration). ContAV-Sep presents a novel Cross-modal Similarity Distillation Constraint (CrossSDC) to uphold the cross-modal semantic similarity through incremental tasks and retain previously acquired knowledge of semantic similarity in old models, mitigating the risk of catastrophic forgetting. The CrossSDC can seamlessly integrate into the training process of different audio-visual sound separation frameworks. Experiments demonstrate that ContAV-Sep can effectively mitigate catastrophic forgetting and achieve significantly better performance compared to other continual learning baselines for audio-visual sound separation. Code is available at: https://github.com/weiguoPian/ContAV-Sep_NeurIPS2024. Weiguo Pian, Yiyang Nan, Shijian Deng, Shentong Mo, Yunhui Guo, Yapeng Tian |
NeurIPS | 5 |
| 2024 | Non-Uniform 3D Massive MIMO Arrays Topology Optimization for Near-Field CommunicationsabstractIn this paper, the design of non-uniform antenna topology for 3D massive MIMO arrays in near-field communications is investigated. Specifically, the near-field spherical wavefront radiation characteristics are considered to accurately model the variations of signal phase across array elements. Subsequently, the closed-form expressions of the per-user signal-to-interference noise ratio (SINR) and achievable sum rate for a multi-user MIMO system with maximum-ratio transmission (MRT) precoding are derived. The focus is on the maximization of the achievable sum rate by optimizing the non-uniform 3D antenna array topology. Since the optimization problem exhibits highly nonlinear and nonconvex characteristics, an enhanced particle swarm optimization (EPSO) algorithm is proposed to effectively solve it. Numerical results demonstrate the superiority of the proposed non-uniform 3D array topology in near-field communications for enhancing the achievable sum rate. Yunhui Guo, Yang Zhang 0013, Lihua Pang, Yuanwei Liu, Zhiguo Ding 0001 |
PIMRC | 1 |
| 2024 | VEATIC: Video-based Emotion and Affect Tracking in Context DatasetabstractHuman affect recognition has been a significant topic in psychophysics and computer vision. However, the currently published datasets have many limitations. For example, most datasets contain frames that contain only information about facial expressions. Due to the limitations of previous datasets, it is very hard to either understand the mechanisms for affect recognition of humans or generalize well on common cases for computer vision models trained on those datasets. In this work, we introduce a brand new large dataset, the Video-based Emotion and Affect Tracking in Context Dataset (VEATIC), that can conquer the limitations of the previous datasets. VEATIC has 124 video clips from Hollywood movies, documentaries, and home videos with continuous valence and arousal ratings of each frame via real-time annotation. Along with the dataset, we propose a new computer vision task to infer the affect of the selected character via both context and character information in each video frame. Additionally, we propose a simple model to benchmark this new computer vision task. We also compare the performance of the pretrained model using our dataset with other similar datasets. Experiments show the competing results of our pretrained model via VEATIC, indicating the generalizability of VEATIC. Our dataset is available at https://veatic.github.io. Zhihang Ren, Jefferson Ortega, Yunhui Guo, Stella X. Yu, David Whitney |
WACV | 5 |
| 2024 | Evolve: Enhancing Unsupervised Continual Learning with Multiple ExpertsabstractRecent years have seen significant progress in unsupervised continual learning methods. Despite their success in controlled settings, their practicality in real-world contexts remains uncertain. In this paper, we first empirically investigate existing self-supervised continual learning methods. We show that even with a replay buffer, existing methods cannot preserve the critical knowledge on videos with temporal-correlated input. Our insight is that the primary challenge of unsupervised continual learning stems from the unpredictable input and the absence of supervision as well as prior knowledge. Drawing inspiration from hybrid AI, we introduce Evolve, an innovative framework employing multiple pretrained models in the cloud, as experts, to bolster existing self-supervised learning methods on local clients. Evolve harnesses expert guidance through a novel expert aggregation loss, calculated and returned from the cloud. It also dynamically assigns weights to experts based on their confidence and tailored prior knowledge, thereby offering adaptive supervision for new streaming data. We extensively validate Evolve across several real-world data streams with temporal correlation. The results convincingly demonstrate that Evolve surpasses the best state-of-the-art unsupervised continual learning method by 6.1-53.7% in top-1 linear evaluation accuracy across various data streams, affirming the efficacy of diverse expert guidance. The codebase is at https://github.com/Orienfish/Evolve. Xiaofan Yu 0001, Tajana Rosing, Yunhui Guo |
WACV | 3 |
| 2024 | Wideband Beamforming for Near-Field Communications With Circular ArraysabstractThe three-dimensional (3D) beamforming property of the uniform circular array (UCA) in near-field wideband communication systems is comprehensively analyzed in both the distance and angular domains. It is rigorously demonstrated that the beam focal point only exists at a specific frequency in wideband UCA systems, resulting in significant beamforming loss. To facilitate near-field beamfocusing and alleviate the beam squint effect, the true-time delay (TTD)-based beamforming architecture is exploited. In particular, two wideband beamforming optimization approaches leveraging TTD units are proposed. 1)Analytical approach: In this approach, the phase shifters (PSs) and the time delay of TTD units are designed based on the analytical formula for beamforming gain. Following this design, the minimum number of TTD units required to achieve a predetermined beamforming gain is quantified. 2)Joint-optimization approach: In this method, the PSs and the TTD units are jointly optimized under practical maximum delay constraints to approximate the optimal unconstrained analog beamformer. Specifically, an efficient alternating optimization algorithm is proposed, where the PSs and the TTD units are alternately updated using either the closed-form solution or the low-complexity linear search approach. Extensive numerical results demonstrate that 1) the proposed beamforming schemes effectively mitigate the beam squint effect, and 2) the joint-optimization approach outperforms the analytical approach in terms of array gain and achievable spectral efficiency. Yunhui Guo, Yang Zhang 0013, Zhaolin Wang 0001, Yuanwei Liu |
IEEE Trans. Wirel. Commun. | 1 |
| 2023 | Audio-Visual Class-Incremental LearningabstractIn this paper, we introduce audio-visual class-incremental learning, a class-incremental learning scenario for audio-visual video recognition. We demonstrate that joint audio-visual modeling can improve class-incremental learning, but current methods fail to preserve semantic similarity between audio and visual features as incremental step grows. Furthermore, we observe that audio-visual correlations learned in previous tasks can be forgotten as incremental steps progress, leading to poor performance. To overcome these challenges, we propose AV-CIL, which incorporates Dual-Audio-Visual Similarity Constraint (D-AVSC) to maintain both instance-aware and class-aware semantic similarity between audio-visual modalities and Visual Attention Distillation (VAD) to retain previously learned audio-guided visual attentive ability. We create three audio-visual class-incremental datasets, AVE-Class-Incremental (AVE-CI), Kinetics-Sounds-Class-Incremental (K-S-CI), and VGGSound100-Class-Incremental (VS100-CI) based on the AVE, Kinetics-Sounds, and VGGSound datasets, respectively. Our experiments on AVE-CI, K-SCI, and VS100-CI demonstrate that AV-CIL significantly outperforms existing class-incremental learning methods in audio-visual class-incremental learning. Code and data are available at: https://github.com/weiguoPian/AV-CIL_ICCV2023. Weiguo Pian, Shentong Mo, Yunhui Guo, Yapeng Tian |
ICCV | 3 |
| 2023 | Antenna Topology Optimization for Massive MIMO Near-Field Wireless Communications with Line-of-Sight Deterministic ChannelsabstractIn this paper, we investigate the optimization of non-uniform planar array (NUPA) for massive multi-input multi-output (MIMO) near-field wireless communications with line-of-sight (LOS) channels. In particular, the NUPAs at both the transmitter and receiver are misaligned placed with certain placement angles. Our focus is on the maximization of the system channel capacity, by optimizing the antenna elements position in the transmit/receive NUPAs. With the near-field spherical wave channel modeling, and by taking into account the geometric structure relationship of the arrays, the mathematically expression of channel capacity with this NUPA deployment is derived. Then, the relationship between the eigenvalues and the misaligned angles is analyzed, which further reveals the effect of the misalignment angle on the channel capacity. We show that channel capacity under specific transmission distance is related to the offset angle but irrelevant to the placement angles. Numerical results demonstrate that our theoretical analysis are consistent with the simulation results and the superiority of our optimized NUPA topology in system capacity. Yunhui Guo, Yang Zhang 0013, Lihua Pang, Yijian Chen |
PIMRC | 1 |
| 2023 | AUC Maximization in Imbalanced Lifelong LearningabstractImbalanced data is ubiquitous in machine learning, such as medical or fine-grained image datasets. The existing continual learning methods employ various techniques such as balanced sampling to improve classification accuracy in this setting. However, classification accuracy is not a suitable metric for imbalanced data, and hence these methods may not obtain a good classifier as measured by other metrics (e.g., Area under the ROC Curve). In this paper, we propose a solution to enable efficient imbalanced continual learning by designing an algorithm to effectively maximize one widely used metric in an imbalanced data setting: Area Under the ROC Curve (AUC). We find that simply replacing accuracy with AUC will cause gradient interference problem due to the imbalanced data distribution. To address this issue, we propose a new algorithm, namely DIANA, which performs a novel synthesis of model DecouplIng ANd Alignment. In particular, the algorithm updates two models simultaneously: one focuses on learning the current knowledge while the other concentrates on reviewing previously-learned knowledge, and the two models gradually align during training. The results show that the proposed DIANA achieves state-of-the-art performance on all the imbalanced datasets compared with several competitive baselines. Yunhui Guo |
UAI | 3 |
| 2023 | Modeling Semantic Correlation and Hierarchy for Real-World Wildlife RecognitionabstractWe explore the challenges of human-in-the-loop frameworks to label wildlife recognition datasets with a neural network. In wildlife imagery, the main challenges for a model to assist human annotation are two-fold: (1) the training dataset is usually imbalanced, which makes the model's suggestion biased, and (2) there are complex taxonomies in the classes. We establish a simple and efficient baseline, including the debiasing loss function and the hyperbolic network architecture, to address these issues. Moreover, we propose leveraging the semantic correlation to train the model more effectively by adding a co-occurrence layer to our model during training. We demonstrate the efficacy of our method in both a real-world wildlife areal survey recognition dataset and the public image classification dataset, CIFAR100-LT, CIFAR10-LT, and iNaturalist. Dong-Jin Kim 0003, Zhongqi Miao, Yunhui Guo, Stella X. Yu |
IEEE Signal Process. Lett. | 3 |
| 2022 | CO-SNE: Dimensionality Reduction and Visualization for Hyperbolic DataabstractHyperbolic space can naturally embed hierarchies that often exist in real-world data and semantics. While high-dimensional hyperbolic embeddings lead to better representations, most hyperbolic models utilize low-dimensional embeddings, due to non-trivial optimization and visualization of high-dimensional hyperbolic data. We propose CO-SNE, which extends the Euclidean space visualization tool, t-SNE, to hyperbolic space. Like t-SNE, it converts distances between data points to joint probabilities and tries to minimize the Kullback-Leibler divergence between the joint probabilities of high-dimensional data$X$and low-dimensional embedding$Y$. However, unlike Euclidean space, hyperbolic space is inhomogeneous: A volume could contain a lot more points at a location far from the origin. CO-SNE thus uses hyperbolic normal distributions for$X$and hyperbolic Cauchy instead of t-SNE's Student's t-distribution for$Y$, and it additionally seeks to preserve$X$'s individual distances to the Origin in$Y$. We apply CO-SNE to naturally hyperbolic data and supervisedly learned hyperbolic features. Our results demonstrate that CO-SNE deflates high-dimensional hyperbolic data into a low-dimensional space without losing their hyperbolic characteristics, significantly outperforming popular visualization tools such as PCA, t-SNE, UMAP, and HoroPCA which is also designed for hyperbolic data. Yunhui Guo, Haoran Guo, Stella X. Yu |
CVPR | 1 |
| 2022 | Clipped Hyperbolic Classifiers Are Super-Hyperbolic ClassifiersabstractHyperbolic space can naturally embed hierarchies, unlike Euclidean space. Hyperbolic Neural Networks (HNNs) exploit such representational power by lifting Euclidean features into hyperbolic space for classification, outperforming Euclidean neural networks (ENNs) on datasets with known semantic hierarchies. However, HNNs underperform ENNs on standard benchmarks without clear hierarchies, greatly restricting HNNs' applicability in practice. Our key insight is that HNNs' poorer general classification performance results from vanishing gradients during backpropagation, caused by their hybrid architecture connecting Euclidean features to a hyperbolic classifier. We propose an effective solution by simply clipping the Euclidean feature magnitude while training HNNs. Our experiments demonstrate that clipped HNNs become super-hyperbolic classifiers: They are not only consistently better than HNNs which already outperform ENNs on hierarchical data, but also on-par with ENNs on MNIST, CIFAR10, CIFAR100 and ImageNet benchmarks, with better adversarial robustness and out-of-distribution detection. Yunhui Guo, Xudong Wang 0007, Yubei Chen, Stella X. Yu |
CVPR | 1 |
| 2022 | Unsupervised Hierarchical Semantic Segmentation with Multiview Cosegmentation and Clustering TransformersabstractUnsupervised semantic segmentation aims to discover groupings within and across images that capture object-and view-invariance of a category without external supervision. Grouping naturally has levels of granularity, creating ambiguity in unsupervised segmentation. Existing methods avoid this ambiguity and treat it as a factor outside modeling, whereas we embrace it and desire hierarchical grouping consistency for unsupervised segmentation. We approach unsupervised segmentation as a pixel-wise feature learning problem. Our idea is that a good representation shall reveal not just a particular level of grouping, but any level of grouping in a consistent and predictable manner. We enforce spatial consistency of grouping and bootstrap feature learning with co-segmentation among multiple views of the same image, and enforce semantic consistency across the grouping hierarchy with clustering transformers between coarse- and fine-grained features. We deliver the first data-driven unsupervised hierarchical semantic segmentation method called Hierarchical Segment Grouping (HSG). Capturing visual similarity and statistical co-occurrences, HSG also outperforms existing un-supervised segmentation methods by a large margin on five major object- and scene-centric benchmarks. Tsung-Wei Ke, Jyh-Jing Hwang, Yunhui Guo, Xudong Wang 0007, Stella X. Yu |
CVPR | 3 |
| 2021 | HyperRec: Efficient Recommender Systems with Hyperdimensional ComputingabstractRecommender systems are important tools for many commercial applications such as online shopping websites. There are several issues that make the recommendation task very challenging in practice. The first is that an efficient and compact representation is needed to represent users, items and relations. The second issue is that the online markets are changing dynamically, it is thus important that the recommendation algorithm is suitable for fast updates and hardware acceleration. In this paper, we propose a new hardware-friendly recommendation algorithm based on Hyperdimensional Computing, called HyperRec. Unlike existing solutions which leverages floating-point numbers for the data representation, in HyperRec, users and items are modeled with binary vectors in a high dimension. The binary representation enables to perform the reasoning process of the proposed algorithm only using Boolean operations, which is efficient on various computing platforms and suitable for hardware acceleration. In this work, we show how to utilize GPU and FPGA to accelerate the proposed HyperRec. When compared with the state-of-the-art methods for rating prediction, the CPU-based HyperRec implementation is 13.75x faster and consumes 87% less memory, while decreasing the mean squared error (MSE) for the prediction by as much as 31.84%. Our FPGA implementation is on average 67.0x faster and has 6.9x higher energy efficient as compared to CPU. Our GPU implementation further achieves on average 3.1x speedup as compared to FPGA, while providing only 1.2x lower energy efficiency. Yunhui Guo, Mohsen Imani, Jaeyoung Kang 0001, Sahand Salamat, Justin Morris, Baris Aksanli, Yeseong Kim, Tajana Rosing |
ASP-DAC | 1 |
| 2021 | MAT: Processing In-Memory Acceleration for Long-Sequence AttentionabstractAttention-based machine learning is used to model long-term dependencies in sequential data. Processing these models on long sequences can be prohibitively costly because of the large memory consumption. In this work, we propose MAT, a processing in-memory (PIM) framework, to accelerate long-sequence attention models. MAT adopts a memory-efficient processing flow for attention models to process sub-sequences in a pipeline with much smaller memory footprint. MAT utilizes a reuse-driven data layout and an optimal sample scheduling to optimize the performance of PIM attention. We evaluate the efficiency of MAT on two emerging long-sequence tasks including natural language processing and medical image processing. Our experiments show that MAT is $2.7 \times$ faster and $3.4 \times$ more energy efficient than the state-of-the-art PIM acceleration. As compared to TPU and GPU, MAT is $5.1 \times$ and $16.4 \times$ faster while consuming $27.5 \times$ and $41.0 \times$ less energy. Minxuan Zhou, Yunhui Guo, Bin Li 0064, Kevin W. Eliceiri, Tajana Rosing |
DAC | 2 |
| 2020 | AdaFilter: Adaptive Filter Fine-Tuning for Deep Transfer LearningabstractThere is an increasing number of pre-trained deep neural network models. However, it is still unclear how to effectively use these models for a new task. Transfer learning, which aims to transfer knowledge from source tasks to a target task, is an effective solution to this problem. Fine-tuning is a popular transfer learning technique for deep neural networks where a few rounds of training are applied to the parameters of a pre-trained model to adapt them to a new task. Despite its popularity, in this paper we show that fine-tuning suffers from several drawbacks. We propose an adaptive fine-tuning approach, called AdaFilter, which selects only a part of the convolutional filters in the pre-trained model to optimize on a per-example basis. We use a recurrent gated network to selectively fine-tune convolutional filters based on the activations of the previous layer. We experiment with 7 public image classification datasets and the results show that AdaFilter can reduce the average classification error of the standard fine-tuning by 2.54%. Yunhui Guo, Yandong Li, Liqiang Wang 0001, Tajana Rosing |
AAAI | 1 |
| 2020 | A Broader Study of Cross-Domain Few-Shot Learning
Yunhui Guo, Noel Codella, Leonid Karlinsky, James V. Codella, John R. Smith, Kate Saenko, Tajana Rosing, Rogério Feris |
ECCV (27) | 1 |
| 2020 | Efficient Distributed Training in Heterogeneous Mobile Networks with Active SamplingabstractMobile edge computing is an emerging research topic which aims at pushing the computation from the cloud to the edge devices. Most of the current machine learning (ML) algorithms, such as federated learning, are designed for homogeneous mobile networks, that is, all the devices collect the same type of data. In this paper, we address distributed training of ML algorithms in heterogeneous mobile networks where the features, rather than the samples, are distributed across multiple heterogeneous mobile devices. Training ML models in heterogeneous mobile networks incurs a large communication cost due to the necessity to deliver the local data to a central server. Inspired by active learning, which is traditionally used to reduce the labeling cost for training ML models, we propose an active sampling method to reduce the communication cost of learning in heterogeneous mobile networks. Instead of sending all the local data, the proposed active sampling method identifies and sends only informative data from each device to the central server. Extensive experiments on four real datasets, both with numerical simulation and on a networked mobile system, show that the proposed method can reduce the communication cost by up to 53% and energy consumption by up to 67% without accuracy degradation compared with the conventional approaches. Yunhui Guo, Xiaofan Yu 0001, Kamalika Chaudhuri, Tajana Rosing |
MSN | 1 |
| 2020 | Improved Schemes for Episodic Memory-based Lifelong LearningabstractCurrent deep neural networks can achieve remarkable performance on a single task. However, when the deep neural network is continually trained on a sequence of tasks, it seems to gradually forget the previous learned knowledge. This phenomenon is referred to as catastrophic forgetting and motivates the field called lifelong learning. Recently, episodic memory based approaches such as GEM and A-GEM have shown remarkable performance. In this paper, we provide the first unified view of episodic memory based approaches from an optimization's perspective. This view leads to two improved schemes for episodic memory based lifelong learning, called MEGA-\rom{1} and MEGA-\rom{2}. MEGA-\rom{1} and MEGA-\rom{2} modulate the balance between old tasks and the new task by integrating the current gradient with the gradient computed on the episodic memory. Notably, we show that GEM and A-GEM are degenerate cases of MEGA-\rom{1} and MEGA-\rom{2} which consistently put the same emphasis on the current task, regardless of how the loss changes over time. Our proposed schemes address this issue by using novel loss-balancing updating rules, which drastically improve the performance over GEM and A-GEM. Extensive experimental results show that the proposed schemes significantly advance the state-of-the-art on four commonly used lifelong learning benchmarks, reducing the error by up to 18%. Yunhui Guo, Tianbao Yang, Tajana Rosing |
NeurIPS | 1 |
| 2019 | Depthwise Convolution Is All You Need for Learning Multiple Visual DomainsabstractThere is a growing interest in designing models that can deal with images from different visual domains. If there exists a universal structure in different visual domains that can be captured via a common parameterization, then we can use a single model for all domains rather than one model per domain. A model aware of the relationships between different domains can also be trained to work on new domains with less resources. However, to identify the reusable structure in a model is not easy. In this paper, we propose a multi-domain learning architecture based on depthwise separable convolution. The proposed approach is based on the assumption that images from different domains share cross-channel correlations but have domain-specific spatial correlations. The proposed model is compact and has minimal overhead when being applied to new domains. Additionally, we introduce a gating mechanism to promote soft sharing between different domains. We evaluate our approach on Visual Decathlon Challenge, a benchmark for testing the ability of multi-domain models. The experiments show that our approach can achieve the highest score while only requiring 50% of the parameters compared with the state-of-the-art approaches. Yunhui Guo, Yandong Li, Liqiang Wang 0001, Tajana Rosing |
AAAI | 1 |
| 2019 | SpotTune: Transfer Learning Through Adaptive Fine-TuningabstractTransfer learning, which allows a source task to affect the inductive bias of the target task, is widely used in computer vision. The typical way of conducting transfer learning with deep neural networks is to fine-tune a model pretrained on the source task using data from the target task. In this paper, we propose an adaptive fine-tuning approach, called SpotTune, which finds the optimal fine-tuning strategy per instance for the target data. In SpotTune, given an image from the target task, a policy network is used to make routing decisions on whether to pass the image through the fine-tuned layers or the pre-trained layers. We conduct extensive experiments to demonstrate the effectiveness of the proposed approach. Our method outperforms the traditional fine-tuning approach on 12 out of 14 standard datasets. We also compare SpotTune with other state-of-the-art fine-tuning strategies, showing superior performance. On the Visual Decathlon datasets, our method achieves the highest score across the board without bells and whistles. Yunhui Guo, Humphrey Shi, Abhishek Kumar 0001, Kristen Grauman, Tajana Rosing, Rogério Feris |
CVPR | 1 |
| 2017 | Understanding Users' Budgets for Recommendation with Hierarchical Poisson FactorizationabstractPeople consume and rate products in online shopping websites. The historical purchases of customers reflect their personal consumption habits and indicate their future shopping behaviors. Traditional preference-based recommender systems try to provide recommendations by analyzing users' feedback such as ratings and clicks. But unfortunately, most of the existing recommendation algorithms ignore the budget of the users. So they cannot avoid recommending users with products that will exceed their budgets. And they also cannot understand how the users will assign their budgets to different products. In this paper, we develop a generative model named collaborative budget-aware Poisson factorization (CBPF) to connect users' ratings and budgets. The CBPF model is intuitive and highly interpretable. We compare the proposed model with several state-of-the-art budget-unaware recommendation methods on several real-world datasets. The results show the advantage of uncovering users' budgets for recommendation. Yunhui Guo, Congfu Xu, Hanzhang Song, Xin Wang 0060 |
IJCAI | 1 |
| 2016 | Constrained Preference Embedding for Item Recommendation
Xin Wang 0060, Congfu Xu, Yunhui Guo, Hui Qian 0001 |
IJCAI | 3 |
| 2016 | LBMF: Log-Bilinear Matrix Factorization for Recommender Systems
Yunhui Guo, Xin Wang 0060, Congfu Xu |
PAKDD (1) | 1 |
| 2016 | Collaborative Expert Recommendation for Community-Based Question Answering
Congfu Xu, Xin Wang 0060, Yunhui Guo |
ECML/PKDD (1) | 3 |
| 2015 | Recommendation Algorithms for Optimizing Hit Rate, User Satisfaction and Website Revenue
Xin Wang 0060, Yunhui Guo, Congfu Xu |
IJCAI | 2 |