EDBT 2026 Demo / reviewers in the wild / expert
Guang Dai
dblp:51/4674
· DBLP profile ↗
72ranked-venue papers
13as first author
46since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 54 · 7 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 33 · 9 first-author · 21 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Privacy Leaks by Adversaries: Adversarial Iterations for Membership Inference AttackabstractMembership inference attack (MIA) has become one of the most widely used and effective methods for evaluating the privacy risks of machine learning models. This attack aims to determine whether a specific sample is part of the model's training set by analyzing the model's output. While traditional membership inference attacks focus on leveraging the model’s posterior output, such as confidence on the target sample, we propose IMIA, a novel attack strategy that utilizes the process of generating adversarial samples to infer membership. We propose to infer the member properties of the target sample using the number of iterations required to generate its adversarial sample. We conduct experiments across multiple models and datasets, and our results demonstrate that the number of iterations for generating an adversarial sample is a reliable feature for membership inference, achieving strong performance both in black-box and white-box attack scenarios. This work provides a new perspective for evaluating model privacy and highlights the potential of adversarial example-based features for privacy leakage assessment. Zhishen Sun, Haishan Ye, Luo Luo, Xiangyu Chang, Guang Dai |
AAAI | 6 |
| 2026 | HALO: Hardness-aware bilevel-inspired contrastive graph clustering
Kuang Zhou, Haishan Ye, Guang Dai, Ivor W. Tsang |
Int. J. Approx. Reason. | 4 |
| 2026 | Warm-start or cold-start? A comparison of generalizability in gradient-based hyperparameter tuning
Yubo Zhou, Chengli Tan, Haishan Ye, Quanziang Wang, Junmin Liu, Deyu Meng, Ivor W. Tsang, Guang Dai |
Neural Networks | 9 |
| 2026 | MA-FSAR: Multimodal Adaptation of CLIP for few-shot action recognition
Jiazheng Xing, Jian Zhao 0006, Chao Xu 0023, Mengmeng Wang 0005, Guang Dai, Yong Liu 0007, Jingdong Wang 0001, Xuelong Li 0001 |
Pattern Recognit. | 5 |
| 2026 | Corrigendum to "MA-FSAR: Multimodal Adaptation of CLIP for few-shot action recognition" [Pattern Recognition 169 (2026) 111902]
Jiazheng Xing, Jian Zhao 0006, Chao Xu 0023, Mengmeng Wang 0005, Guang Dai, Yong Liu 0007, Jingdong Wang 0001, Xuelong Li 0001 |
Pattern Recognit. | 5 |
| 2025 | SpotActor: Training-Free Layout-Controlled Consistent Image GenerationabstractText-to-image diffusion models significantly enhance the efficiency of artistic creation with high-fidelity image generation. However, in typical application scenarios like comic book production, they can neither place each subject into its expected spot nor maintain the consistent appearance of each subject across images. For these issues, we pioneer a novel task, Layout-to-Consistent-Image (L2CI) generation, which produces consistent and compositional images in accordance with the given layout conditions and text prompts. To accomplish this challenging task, we present a new formalization of dual energy guidance with optimization in a dual semantic-latent space and thus propose a training-free pipeline, SpotActor, which features a layout-conditioned optimizing stage and a consistent sampling stage. In the optimizing stage, we innovate a nuanced layout energy function to mimic the attention activations with a sigmoid-like objective. While in the sampling stage, we design Regional Interconnection Self-Attention (RISA) and Semantic Fusion Cross-Attention (SFCA) mechanisms that allow mutual interactions across images. To evaluate the performance, we present ActorBench, a specified benchmark with hundreds of reasonable prompt-box pairs stemming from object detection datasets. Comprehensive experiments are conducted to demonstrate the effectiveness of our method. The results prove that SpotActor fulfills the expectations of this task and showcases the potential for practical applications with superior layout alignment, subject consistency, prompt conformity and background diversity. Jiahao Wang 0004, Caixia Yan, Weizhan Zhang, Haonan Lin, Mengmeng Wang 0005, Guang Dai, Tieliang Gong, Hao Sun 0015, Jingdong Wang 0001 |
AAAI | 6 |
| 2025 | On the Risk of Evidence Pollution for Malicious Social Text Detection in the Era of LLMsabstractEvidence-enhanced detectors present remarkable abilities in identifying malicious social text.However, the rise of large language models (LLMs) brings potential risks of evidence pollution to confuse detectors.This paper explores potential manipulation scenarios including basic pollution, and rephrasing or generating evidence by LLMs.To mitigate the negative impact, we propose three defense strategies from the data and model sides, including machine-generated text detection, a mixture of experts, and parameter updating.Extensive experiments on four malicious social text detection tasks with ten datasets illustrate that evidence pollution significantly compromises detectors, where the generating strategy causes up to a 14.4% performance drop.Meanwhile, the defense strategies could mitigate evidence pollution, but they faced limitations for practical employment.Further analysis illustrates that polluted evidence (i) is of high quality, evaluated by metrics and humans; (ii) would compromise the model calibration, increasing expected calibration error up to 21.6%; and (iii) could be integrated to amplify the negative impact, especially for encoder-based LMs, where the accuracy drops by 21.8%. Herun Wan, Minnan Luo, Zhixiong Su, Guang Dai, Xiang Zhao 0002 |
ACL (1) | 4 |
| 2025 | IMOL: Incomplete-Modality-Tolerant Learning for Multi-Domain Fake News Video DetectionabstractZhi Zeng, Jiaying Wu, Minnan Luo, Herun Wan, Xiangzheng Kong, Zihan Ma, Guang Dai, Qinghua Zheng. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zhi Zeng 0001, Minnan Luo, Herun Wan, Xiangzheng Kong, Zihan Ma 0001, Guang Dai |
ACL (1) | 7 |
| 2025 | Low-Biased General Annotated Dataset GenerationabstractPre-training backbone networks on a general annotated dataset (e.g., ImageNet) that comprises numerous manually collected images with category annotations has proven to be indispensable for enhancing the generalization capacity of downstream visual tasks. However, those manually collected images often exhibit bias, which is non-transferable across either categories or domains, thus causing the model’s generalization capacity degeneration. To mitigate this problem, we present a low-biased general annotated dataset generation framework (lbGen). Instead of expensive manual collection, we aim at directly generating low-biased images with category annotations. To achieve this goal, we propose to leverage the advantage of a multimodal foundation model (e.g., CLIP), in terms of aligning images in a low-biased semantic space defined by language. Specifically, we develop a bi-level semantic alignment loss, which not only forces all generated images to be consistent with the semantic distribution of all categories belonging to the target dataset in an adversarial learning manner, but also requires each generated image to match the semantic description of its category name. In addition, we further cast an existing image quality scoring model into a quality assurance loss to preserve the quality of the generated image. By leveraging these two loss functions, we can obtain a low-biased image generation model by simply fine-tuning a pre-trained diffusion model using only all category names in the target dataset as input. Experimental results confirm that, compared with the manually labeled dataset or other synthetic datasets, the utilization of our generated low-biased dataset leads to stable generalization capacity enhancement of different backbone networks across various tasks, especially in tasks where the manually labeled samples are scarce. Code is available at: https://github.com/vvvvvjdy/lbGen Dengyang Jiang, Haoyu Wang 0016, Lei Zhang 0054, Wei Wei 0008, Guang Dai, Yanning Zhang 0001 |
CVPR | 5 |
| 2025 | Action Detail Matters: Refining Video Recognition with Local Action QueriesabstractVideo action recognition involves interpreting both global context and specific details to accurately identify actions. While previous models are effective at capturing spatiotemporal features, they often lack a focused representation of key action details. To address this, we introduce FocusVideo, a framework designed for refining video action recognition through integrated global and local feature learning. Inspired by human visual cognition theory, our approach balances the focus on both broad contextual changes and action-specific details, minimizing the influence of irrelevant background noise. We first employ learnable action queries to selectively emphasize action-relevant regions without requiring region-specific labels. Next, these queries are learned by a local action streaming branch that enables progressive query propagation. Moreover, we introduce a parameter-free feature interaction mechanism for effective multi-scale interaction between global and local features with minimal additional overhead. Extensive experiments demonstrate that FocusVideo achieves state-of-the-art performance across multiple action recognition datasets, validating its effectiveness and robustness in handling action-relevant details. Mengmeng Wang 0005, Zeyi Huang, Xiangjie Kong 0001, Guojiang Shen, Guang Dai, Jingdong Wang 0001, Yong Liu 0007 |
CVPR | 5 |
| 2025 | How Do Social Bots Participate in Misinformation Spread? A Comprehensive Dataset and AnalysisabstractSocial media platforms provide an ideal environment to spread misinformation, where social bots can accelerate the spread.This paper explores the interplay between social bots and misinformation on the Sina Weibo platform.We construct a large-scale dataset that includes annotations for both misinformation and social bots.From the misinformation perspective, the dataset is multimodal, containing 11,393 pieces of misinformation and 16,416 pieces of verified information.From the social bot perspective, this dataset contains 65,749 social bots and 345,886 genuine accounts, annotated using a weakly supervised annotator.Extensive experiments demonstrate the comprehensiveness of the dataset, the clear distinction between misinformation and real information, and the high quality of social bot annotations.Further analysis illustrates that: (i) social bots are deeply involved in information spread; (ii) misinformation with the same topics has similar content, providing the basis of echo chambers, and social bots would amplify this phenomenon; and (iii) social bots generate similar content aiming to manipulate public opinions. Herun Wan, Minnan Luo, Zihan Ma 0001, Guang Dai, Xiang Zhao 0002 |
EMNLP | 4 |
| 2025 | Density-aware and Depth-aware Visual Representation for Zero-Shot Object CountingabstractPrevious methods often utilize CLIP semantic classifiers with class names for zero-shot object counting. However, they ignore crucial density and depth knowledge for counting tasks. Thus, we propose a density-aware and depth-aware prompt counting model, which captures density information via learning density-aware prompts based on density-aware contrastive loss and incorporates depth guidance with predefined depth-aware prompts. To facilitate the training process, we design two strategies for standard counting loss and the contrastive loss, where the former prioritizes larger and sparser objects initially, gradually focusing on smaller and denser objects, and the latter adopts coarse-to-fine density learning. Besides, we construct a dataset named LVIS-372 with more real-world scenarios and balanced instance distribution compared to existing ones. Finally, the experimental results demonstrate the effectiveness of our proposed method. Feng Tian 0002, Ni Zhang 0001, Nian Liu 0002, Haonan Miao, Guang Dai, Mengmeng Wang 0005 |
ICASSP | 6 |
| 2025 | Exploring Triple Knowledge Cues for Zero-Shot Human-Object Interaction DetectionabstractCurrent zero-shot human-object interaction detection methods often follow a two-phase pipeline, which uses a pre-trained detector to detect instances and then adopts CLIP to perform interaction prediction. During the second phase, they either obtain pairwise representations by directly performing RoI-Align on CLIP features or designing additional queries and decoders to fuse CLIP features. However, CLIP visual features often lack fine-grained information, thus being detrimental to capturing complex HOI interactions. Besides, extra decoders might increase computation costs. Thus, we propose a triple knowledge cues exploration model without extra decoders to explore various knowledge guidance for improving CLIP representations. First, we incorporate position distribution and semantic priors to delineate a layout from the predicted boxes and inject semantics by using the CLIP text embeddings. Next, we explore object priors by leveraging predefined class names and the text encoder to obtain saliency maps for humans and objects. Then, we design three types of holistic tokens to capture diverse attribute cues for human, object, and interaction, respectively. The above cues are finally integrated into a vanilla two-stage CLIP-based baseline. The experimental results on HICO-DET demonstrate the effectiveness of our proposed model. Ni Zhang 0001, Qidong Liu 0002, Guang Dai, Yan Chen 0031, Feng Tian 0002 |
ICASSP | 5 |
| 2025 | ProAdvPrompter: A Two-Stage Journey to Effective Adversarial Prompting for LLMsabstractAs large language models (LLMs) are increasingly being integrated into various real-world applications, the identification of their vulnerabilities to jailbreaking attacks becomes an essential component of ensuring the safety and reliability of LLMs.
Previous studies have developed LLM assistants, known as the adversarial prompter, to automatically generate suffixes that manipulate target LLMs into generating harmful and undesirable outputs.
However, these approaches often suffer from low performance or generate semantically meaningless prompts, which can be easily identified by perplexity-based defenses.
In this paper, we introduce a novel two-stage method, $\texttt{ProAdvPrompter}$, that significantly improves the performance of adversarial prompters.
In $\texttt{ProAdvPrompter}$, the first stage (Exploration) utilizes the loss information to guide the adversarial prompter in generating suffixes that are more likely to elicit harmful responses.
Then the second stage (Exploitation) iteratively fine-tunes the prompter using high-quality generated adversarial suffixes to further boost performance.
Additionally, we incorporate the prompt template to aid in the Exploration stage and propose a filtering mechanism to accelerate the training process in the Exploitation stage.
We evaluate $\texttt{ProAdvPrompter}$ against the well-aligned LLMs (i.e., Llama2-Chat-7B and Llama3-chat-8B), achieving attack success rates of 99.68% and 97.12% respectively after 10 trials on the AdvBench dataset, thereby enhancing performance by $\sim 2$ times compared to previous works.
Moreover, $\texttt{ProAdvPrompter}$ reduces training time by 20% on Llama3-Instruct-8B, generates more generalized adversarial suffixes, and demonstrates resilience against the perplexity defense.
An ablation study further evaluates the effects of key components in $\texttt{ProAdvPrompter}$ (the prompt template and the filtering mechanism). Hao Di, Haishan Ye, Yinghui Huang 0001, Xiangyu Chang, Guang Dai, Ivor W. Tsang |
ICLR | 6 |
| 2025 | Manifold Constraint Reduces Exposure Bias in Accelerated Diffusion SamplingabstractDiffusion models have demonstrated significant potential for generating high-quality images, audio, and videos. However, their iterative inference process entails substantial computational costs, limiting practical applications. Recently, researchers have introduced accelerated sampling methods that enable diffusion models to generate samples with far fewer timesteps than those used during training. Nonetheless, as the number of sampling steps decreases, the prediction errors significantly degrade the quality of generated outputs. Additionally, the exposure bias in diffusion models further amplifies these errors. To address these challenges, we leverage a manifold hypothesis to explore the exposure bias problem in depth. Based on this geometric perspective, we propose a manifold constraint that effectively reduces exposure bias during accelerated sampling of diffusion models. Notably, our method involves no additional training and requires only minimal hyperparameter tuning. Extensive experiments demonstrate the effectiveness of our approach, achieving a FID score of 15.60 with 10-step SDXL on MS-COCO, surpassing the baseline by a reduction of 2.57 in FID. Yuzhe Yao, Jun Chen 0023, Zeyi Huang, Haonan Lin, Mengmeng Wang 0005, Guang Dai, Jingdong Wang 0001 |
ICLR | 6 |
| 2025 | Second-Order Fine-Tuning without Pain for LLMs: A Hessian Informed Zeroth-Order OptimizerabstractFine-tuning large language models (LLMs) is necessary for specific downstream tasks, but classic first-order optimizer entails prohibitive GPU memory because of the back propagation. Recent works such as MeZO have turned to zeroth-order optimizers for fine-tuning, which reduce substantial memory by using two forward passes. However, heterogeneous curvatures across different parameter dimensions in LLMs often cause model convergence instability or even failure. In this work, we propose HiZOO, a diagonal Hessian informed Zeroth-Order Optimizer , which is the first work to leverage the diagonal Hessian to enhance ZOO for fine-tuning LLMs. We provide theoretical proof for HiZOO and visualize the optimization trajectories on test functions to illustrate how it improves convergence in handling heterogeneous curvatures. Extensive experiments on various models (RoBERTa, OPT, Phi-2 and LLama3, with 350M$\sim$66B parameters) indicate that HiZOO significantly reduces training steps and enhances model accuracy, while keeping the memory advantage of ZOO. For example, on SST2 task HiZOO achieves $8\times$ speedup and better accuracy over MeZO across different models. We also propose HiZOO-L, which reduces the Hessian memory cost to 10\% of the MeZO, while maintaining almost same performance. Compared with ZO-Adam, HiZOO-L achieves a 4.3\% improvement, just using 50\% of the GPU memory. Code is available at https://anonymous.4open.science/r/HiZOO-27F8. Yanjun Zhao 0001, Sizhe Dang, Haishan Ye, Guang Dai, Yi Qian 0004, Ivor W. Tsang |
ICLR | 4 |
| 2025 | DynaMind: Reasoning over Abstract Video Dynamics for Embodied Decision-MakingabstractIntegrating natural language instructions and visual perception with decision-making is a critical challenge for embodied agents. Existing methods often struggle to balance the conciseness of language commands with the richness of video content. To bridge the gap between modalities, we propose extracting key spatiotemporal patterns from video that capture visual saliency and temporal evolution, referred to as dynamic representation. Building on this, we introduce DynaMind, a framework that enhances decision-making through dynamic reasoning. Specifically, we design an adaptive FrameScorer to evaluate video frames based on semantic consistency and visual saliency, assigning each frame an importance score. These scores are used to filter redundant video content and synthesize compact dynamic representations. Leveraging these representations, we predict critical future dynamics and apply a dynamic-guided policy to generate coherent and context-aware actions. Extensive results demonstrate that DynaMind significantly outperforms the baselines across several simulation benchmarks and real-world scenarios. Ziru Wang, Mengmeng Wang 0005, Jade Dai, Teli Ma, Guo-Jun Qi, Yong Liu 0007, Guang Dai, Jingdong Wang 0001 |
ICML | 7 |
| 2025 | Instructing Text-to-Image Diffusion Models via Classifier-Guided Semantic OptimizationabstractText-to-image diffusion models have emerged as powerful tools for high-quality image generation and editing. Many existing approaches rely on text prompts as editing guidance. However, these methods are constrained by the need for manual prompt crafting, which can be time-consuming, introduce irrelevant details, and significantly limit editing performance. In this work, we propose optimizing semantic embeddings guided by attribute classifiers to steer text-to-image models toward desired edits, without relying on text prompts or requiring any training or fine-tuning of the diffusion model. We utilize classifiers to learn precise semantic embeddings at the dataset level. The learned embeddings are theoretically justified as the optimal representation of attribute semantics, enabling disentangled and accurate edits. Experiments further demonstrate that our method achieves high levels of disentanglement and strong generalization across different domains of data. Code is available at https://github.com/Chang-yuanyuan/CASO. Yuanyuan Chang, Yinghua Yao, Mengmeng Wang 0005, Ivor W. Tsang, Guang Dai |
IJCAI | 6 |
| 2025 | VidEvo: Evolving Video Editing through Exhaustive Temporal ModelingabstractText-guided video editing (TGVE) has become a recent hotspot due to its entertainment value and practical applications. To reduce overhead, existing methods primarily extend from text-to-image diffusion models and typically involve reconstruction and editing phases. However, challenges persist, particularly in enhancing temporal consistency of a video while adhering to textual alignment requirements. A crucial factor leading to the aforementioned issue is the inadequate and implicit tuning of the attention module within existing methods, which is specifically designed to capture temporal information. In light of this, we introduce VidEvo, a novel one-shot video editing method that leverages explicit cues derived from the original video to enhance temporal modeling. By integrating null-video embedding (NVE) and window-frame attention (WFA) components, VidEvo facilitates the smooth and coherent generation of videos from global and local perspectives simultaneously. To be specific, NVE learns a set of multi-scale temporal embeddings within the visual space during the reconstruction phase. These embeddings are subsequently directly injected into the attention module of the editing phase, explicitly augmenting the temporal consistency of the entire video. On the other hand, WFA enhances local temporal modeling by dynamically optimizing attention mechanisms between adjacent frames, which improves temporal coherence with reduced computational costs. Experimental evaluations show that VidEvo enhances frame-to-frame temporal consistency. Ablation studies confirm NVE and WFA’s effectiveness and their plug-and-play capability with other methods. Sizhe Dang, Huan Liu 0001, Mengmeng Wang 0005, Guang Dai, Jingdong Wang 0001 |
IJCAI | 5 |
| 2025 | TriCLIP-3D: A Unified Parameter-Efficient Framework for Tri-Modal 3D Visual Grounding based on CLIPabstract3D visual grounding allows an embodied agent to understand visual information in real-world 3D environments based on human instructions, which is crucial for embodied intelligence. Existing 3D visual grounding methods typically rely on separate encoders for different modalities (e.g., RGB images, text, and 3D point clouds), resulting in large and complex models that are inefficient to train. While some approaches use pre-trained 2D multi-modal models like CLIP for 3D tasks, they still struggle with aligning point cloud data to 2D encoders. As a result, these methods continue to depend on 3D encoders for feature extraction, further increasing model complexity and training inefficiency. In this paper, we propose a unified 2D pre-trained multi-modal network to process all three modalities (RGB images, text, and point clouds), significantly simplifying the architecture. By leveraging a 2D CLIP bi-modal model with adapter-based fine-tuning, this framework effectively adapts to the tri-modal setting, improving both adaptability and performance across modalities. Our Geometric-Aware 2D-3D Feature Recovery and Fusion (GARF) module is designed to fuse geometric multi-scale features from point clouds and images. We then integrate textual features for final modality fusion and introduce a multi-modal decoder to facilitate deep cross-modal understanding. Together, our method achieves unified feature extraction and fusion across the three modalities, enabling an end-to-end 3D visual grounding model. Compared to the baseline, our method reduces the number of trainable parameters by approximately 58%, while achieving a 6.52% improvement in the 3D detection task and a 6.25% improvement in the 3D visual grounding task. Zanyi Wang, Zeyi Huang, Guang Dai, Jingdong Wang 0001, Mengmeng Wang 0005 |
ACM Multimedia | 4 |
| 2025 | Understand, Refine and Summarize: Multi-View Knowledge Progressive Enhancement Learning for Fake News Video DetectionabstractAs short videos become a dominant medium for news dissemination, fake news videos pose increasing threats to public trust and information integrity. Existing methods primarily focus on learning multimodal representations to predict binary veracity labels, yet they overlook the use of external evidence, which is important for identifying more sophisticated fake news that subtly exploits psychological cues and cognitive biases. Moreover, these approaches do not provide fine-grained attribution labels, which are essential for interpretable misinformation governance. To address these limitations, we introduce EvidSV, the first comprehensive benchmark supporting evidence- and attribution-aware fake news video detection. Drawing inspiration from the human cognitive process of interpreting news-related content, we propose MUKE, a multi-view knowledge progressive enhancement learning framework. By jointly analyzing both the news content and supporting evidence, MUKE (1) facilitates the understanding of news semantics to (2) progressively refine shared domain knowledge, and (3) adaptively summarizes multi-view knowledge to assess news veracity. Extensive experiments demonstrate that MUKE consistently outperforms existing methods in both fake news detection and attribution, and generalizes effectively to previously unseen domains. Our code is available at https://github.com/zzeng1998/EvidSV. Zhi Zeng 0001, Minnan Luo, Xiangzheng Kong, Zihan Ma 0001, Guang Dai |
ACM Multimedia | 6 |
| 2025 | MonoLift: Learning 3D Manipulation Policies from Monocular RGB via DistillationabstractAlthough learning 3D manipulation policies from monocular RGB images is lightweight and deployment-friendly, the lack of structural information often leads to inaccurate action estimation. While explicit 3D inputs can mitigate this issue, they typically require additional sensors and introduce data acquisition overhead. An intuitive alternative is to incorporate a pre-trained depth estimator; however, this often incurs substantial inference-time cost. To address this, we propose MonoLift, a tri-level knowledge distillation framework that transfers spatial, temporal, and action-level knowledge from a depth-guided teacher to a monocular RGB student. By jointly distilling geometry-aware features, temporal dynamics, and policy behaviors during training, MonoLift enables the student model to perform 3D-aware reasoning and precise control at deployment using only monocular RGB input.
Extensive experiments on both simulated and real-world manipulation tasks show that MonoLift not only outperforms existing monocular approaches but even surpasses several methods that rely on explicit 3D input, offering a resource-efficient and effective solution for vision-based robotic control. The video demonstration is available on our project page: https://robotasy.github.io/MonoLift/. Ziru Wang, Mengmeng Wang 0005, Guang Dai, Yongliu Long, Jingdong Wang 0001 |
NeurIPS | 3 |
| 2025 | VisualNetO&M: A Digital Twin-Based Collaborative Visualization System for Power System Communication Network Operation and MaintenanceabstractThe operation and maintenance (O&M) of the communication network supporting a power system are essential for ensuring grid reliability. This paper presents VisualNetO&M, a collaborative visualization system integrated with a digital process twin of the communication network to enhance O&M efficiency. It provides visualizations for key tasks and facilitates collaboration among operators, technicians, and managers. We validated its effectiveness in Xi’an City, China, where it reduced the O&M workflow completion time from 16 hours to just 1 hour. This improvement resulted in a significant economic benefit of nearly 2/3 million USD over 10 months, highlighting the value of VisualNetO&M. Le Liu 0008, Chuhua Yang, Guang Dai, Kaifeng Bai, Siming Chen 0001, Peng Wang 0015 |
VINCI | 4 |
| 2025 | TryOn-Adapter: Efficient Fine-Grained Clothing Identity Adaptation for High-Fidelity Virtual Try-On
Jiazheng Xing, Chao Xu 0023, Yijie Qian, Yang Liu 0356, Guang Dai, Baigui Sun, Yong Liu 0007, Jingdong Wang 0001 |
Int. J. Comput. Vis. | 5 |
| 2025 | PSDiff: Diffusion Model for Person Search With Iterative and Collaborative RefinementabstractDominant Person Search methods aim to localize and recognize query persons in a unified network, which jointly optimizes the two sub-tasks of pedestrian detection and Re-Identification (ReID). Despite significant progress, current methods face two primary challenges: 1) the pedestrian candidates learned within detectors are suboptimal for the ReID task. 2) the potential for collaboration between two sub-tasks is overlooked. To address these issues, we present a novel Person Search framework based on the Diffusion model, PSDiff. PSDiff formulates the person search as a dual denoising process from noisy boxes and ReID embeddings to ground truths. Distinct from the conventional Detection-to-ReID approach, our denoising paradigm discards prior pedestrian candidates generated by detectors, thereby avoiding the local optimum problem of the ReID task. Following the new paradigm, we further design a new Collaborative Denoising Layer (CDL) to optimize detection and ReID sub-tasks in an iterative and collaborative way, which makes two sub-tasks mutually beneficial. Extensive experiments on the standard benchmarks show that PSDiff achieves state-of-the-art performance with fewer parameters and elastic computing overhead. Chengyou Jia, Minnan Luo, Zhuohang Dang, Guang Dai, Xiaojun Chang, Jingdong Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Disentangled Noisy Correspondence LearningabstractCross-modal retrieval is crucial in understanding latent correspondences across modalities. However, existing methods implicitly assume well-matched training data, which is impractical as real-world data inevitably involves imperfect alignments, i.e., noisy correspondences. Although some works explore similarity-based strategies to address such noise, they suffer from sub-optimal similarity predictions influenced by modality-exclusive information (MEI), e.g., background noise in images and abstract definitions in texts. This issue arises as MEI is not shared across modalities, thus aligning it in training can markedly mislead similarity predictions. Moreover, although intuitive, directly applying previous cross-modal disentanglement methods suffers from limited noise tolerance and disentanglement efficacy. Inspired by the robustness of information bottlenecks against noise, we introduce DisNCL, a novel information-theoretic framework for feature Disentanglement in Noisy Correspondence Learning, to adaptively balance the extraction of modality-invariant information (MII) and MEI with certifiable optimal cross-modal disentanglement efficacy. DisNCL then enhances similarity predictions in modality-invariant subspace, thereby greatly boosting similarity-based alleviation strategy for noisy correspondences. Furthermore, DisNCL introduces soft matching targets to model noisy many-to-many relationships inherent in multi-modal inputs for noise-robust and accurate cross-modal alignment. Extensive experiments confirm DisNCL's efficacy by 2% average recall improvement. Mutual information estimation and visualization results show that DisNCL learns meaningful MII/MEI subspaces, validating our theoretical analyses. Zhuohang Dang, Minnan Luo, Jihong Wang 0003, Chengyou Jia, Haochen Han, Herun Wan, Guang Dai, Xiaojun Chang, Jingdong Wang 0001 |
IEEE Trans. Image Process. | 7 |
| 2024 | Noisy Correspondence Learning with Self-Reinforcing Errors MitigationabstractCross-modal retrieval relies on well-matched large-scale datasets that are laborious in practice. Recently, to alleviate expensive data collection, co-occurring pairs from the Internet are automatically harvested for training. However, it inevitably includes mismatched pairs, i.e., noisy correspondences, undermining supervision reliability and degrading performance. Current methods leverage deep neural networks' memorization effect to address noisy correspondences, which overconfidently focus on similarity-guided training with hard negatives and suffer from self-reinforcing errors. In light of above, we introduce a novel noisy correspondence learning framework, namely Self-Reinforcing Errors Mitigation (SREM). Specifically, by viewing sample matching as classification tasks within the batch, we generate classification logits for the given sample. Instead of a single similarity score, we refine sample filtration through energy uncertainty and estimate model's sensitivity of selected clean samples using swapped classification entropy, in view of the overall prediction distribution. Additionally, we propose cross-modal biased complementary learning to leverage negative matches overlooked in hard-negative training, further improving model optimization stability and curbing self-reinforcing errors. Extensive experiments on challenging benchmarks affirm the efficacy and efficiency of SREM. Zhuohang Dang, Minnan Luo, Chengyou Jia, Guang Dai, Xiaojun Chang, Jingdong Wang 0001 |
AAAI | 4 |
| 2024 | SSMG: Spatial-Semantic Map Guided Diffusion Model for Free-Form Layout-to-Image GenerationabstractDespite significant progress in Text-to-Image (T2I) generative models, even lengthy and complex text descriptions still struggle to convey detailed controls. In contrast, Layout-to-Image (L2I) generation, aiming to generate realistic and complex scene images from user-specified layouts, has risen to prominence. However, existing methods transform layout information into tokens or RGB images for conditional control in the generative process, leading to insufficient spatial and semantic controllability of individual instances. To address these limitations, we propose a novel Spatial-Semantic Map Guided (SSMG) diffusion model that adopts the feature map, derived from the layout, as guidance. Owing to rich spatial and semantic information encapsulated in well-designed feature maps, SSMG achieves superior generation quality with sufficient spatial and semantic controllability compared to previous works. Additionally, we propose the Relation-Sensitive Attention (RSA) and Location-Sensitive Attention (LSA) mechanisms. The former aims to model the relationships among multiple objects within scenes while the latter is designed to heighten the model's sensitivity to the spatial information embedded in the guidance. Extensive experiments demonstrate that SSMG achieves highly promising results, setting a new state-of-the-art across a range of metrics encompassing fidelity, diversity, and controllability. Chengyou Jia, Minnan Luo, Zhuohang Dang, Guang Dai, Xiaojun Chang, Mengmeng Wang 0005, Jingdong Wang 0001 |
AAAI | 4 |
| 2024 | A Multimodal, Multi-Task Adapting Framework for Video Action RecognitionabstractRecently, the rise of large-scale vision-language pretrained models like CLIP, coupled with the technology of Parameter-Efficient FineTuning (PEFT), has captured substantial attraction in video action recognition. Nevertheless, prevailing approaches tend to prioritize strong supervised performance at the expense of compromising the models' generalization capabilities during transfer. In this paper, we introduce a novel Multimodal, Multi-task CLIP adapting framework named M2-CLIP to address these challenges, preserving both high supervised performance and robust transferability. Firstly, to enhance the individual modality architectures, we introduce multimodal adapters to both the visual and text branches. Specifically, we design a novel visual TED-Adapter, that performs global Temporal Enhancement and local temporal Difference modeling to improve the temporal representation capabilities of the visual encoder. Moreover, we adopt text encoder adapters to strengthen the learning of semantic label information. Secondly, we design a multi-task decoder with a rich set of supervisory signals, including the original contrastive learning head, a cross-modal classification head, a cross-modal masked language modeling head, and a visual classification head. This multi-task decoder adeptly satisfies the need for strong supervised performance within a multimodal framework. Experimental results validate the efficacy of our approach, demonstrating exceptional performance in supervised learning while maintaining strong generalization in zero-shot scenarios. Mengmeng Wang 0005, Jiazheng Xing, Boyuan Jiang, Jun Chen 0023, Jianbiao Mei, Xingxing Zuo 0001, Guang Dai, Jingdong Wang 0001, Yong Liu 0007 |
AAAI | 7 |
| 2024 | Stress Diffuser: A Biofeedback Agent for Stress Management in Children During Homework with Parent InvolvementabstractParent involvement in children’s homework has emerged as an important component of their education. However, this involvement can generate tension between parents and children, particularly when coupled with a heavy workload, potentially exacerbating stress levels in children. Some children desire to express and share their stress with their parents. However, limited emotional regulation abilities and a lack of stress awareness often hinder effective stress management and communication by children. To address this, we designed a stress biofeedback system named Stress Diffuser, comprising a wearable biosensor with interactive embodied devices, that display the child’s stress during homework sessions. The system intends to manage children’s stress by enhancing awareness among both the child and the parent, and influencing their educational strategies and the way of interaction. We conducted experiments with eight families, including primary school children aged eight to eleven and their parents. The results of our quantitative and qualitative analysis indicated that all participants enhanced their stress awareness while using Stress Diffuser during homework. Most parents adjusted their supervisory strategies and interaction styles in response to their children’s stress data. The study sheds light on children’s needs for a designed agent to communicate their stress and regulate their emotions. Furthermore, the study explores the use of interactive biofeedback technology in homework settings with parent involvement. It reveals a favorable influence on facilitating stress management in children and on enhancing the quality of supervision and interaction between children and parents. Jing Li 0133, Pinhao Wang, Emilia I. Barakova, Jun Hu 0001, Guang Dai |
IDC | 5 |
| 2024 | Learning to Rematch Mismatched Pairs for Robust Cross-Modal RetrievalabstractCollecting well-matched multimedia datasets is crucial for training cross-modal retrieval models. However, in real-world scenarios, massive multimodal data are harvested from the Internet, which inevitably contains Partially Mis-matched Pairs (PMPs). Undoubtedly, such semantical irrelevant data will remarkably harm the cross-modal retrieval performance. Previous efforts tend to mitigate this problem by estimating a soft correspondence to down-weight the contribution of PMPs. In this paper, we aim to address this challenge from a new perspective: the potential semantic similarity among unpaired samples makes it possible to excavate useful knowledge from mismatched pairs. To achieve this, we propose L2RM, a general framework based on Optimal Transport (OT) that learns to rematch mismatched pairs. In detail, L2RM aims to generate refined alignments by seeking a minimal-cost transport plan across different modalities. To formalize the rematching idea in OT, first, we propose a self-supervised cost function that automatically learns from explicit similarity-cost mapping relation. Second, we present to model a partial OT problem while restricting the transport among false positives to further boost refined alignments. Extensive experiments on three benchmarks demonstrate our L2RM significantly improves the robustness against PMPs for existing models. The code is available at https://github.com/hhc1997/L2RM. Haochen Han, Guang Dai, Minnan Luo, Jingdong Wang 0001 |
CVPR | 3 |
| 2024 | Timestep-Aware Correction for Quantized Diffusion Models
Yuzhe Yao, Feng Tian 0002, Jun Chen 0023, Haonan Lin, Guang Dai, Yong Liu 0007, Jingdong Wang 0001 |
ECCV (66) | 5 |
| 2024 | Decentralized Riemannian Conjugate Gradient Method on the Stiefel ManifoldabstractThe conjugate gradient method is a crucial first-order optimization method that generally converges faster than the steepest descent method, and its computational cost is much lower than that of second-order methods. However, while various types of conjugate gradient methods have been studied in Euclidean spaces and on Riemannian manifolds, there is little study for those in distributed scenarios. This paper proposes a decentralized Riemannian conjugate gradient descent (DRCGD) method that aims at minimizing a global function over the Stiefel manifold. The optimization problem is distributed among a network of agents, where each agent is associated with a local function, and the communication between agents occurs over an undirected connected graph. Since the Stiefel manifold is a non-convex set, a global function is represented as a finite sum of possibly non-convex (but smooth) local functions. The proposed method is free from expensive Riemannian geometric operations such as retractions, exponential maps, and vector transports, thereby reducing the computational complexity required by each agent. To the best of our knowledge, DRCGD is the first decentralized Riemannian conjugate gradient algorithm to achieve global convergence over the Stiefel manifold. Jun Chen 0023, Haishan Ye, Mengmeng Wang 0005, Tianxin Huang, Guang Dai, Ivor W. Tsang, Yong Liu 0007 |
ICLR | 5 |
| 2024 | Double Stochasticity Gazes Faster: Snap-Shot Decentralized Stochastic Gradient Tracking MethodsabstractIn decentralized optimization, $m$ agents form a network and only communicate with their neighbors, which gives advantages in data ownership, privacy, and scalability. At the same time, decentralized stochastic gradient descent ($\texttt{SGD}$) methods, as popular decentralized algorithms for training large-scale machine learning models, have shown their superiority over centralized counterparts. Distributed stochastic gradient tracking $\texttt{DSGT}$ has been recognized as the popular and state-of-the-art decentralized $\texttt{SGD}$ method due to its proper theoretical guarantees. However, the theoretical analysis of $\texttt{DSGT}$ shows that its iteration complexity is $\tilde{\mathcal{O}} \left(\frac{\bar{\sigma}^2}{m\mu \varepsilon} + \frac{\sqrt{L}\bar{\sigma}}{\mu(1 - \lambda_2(W))^{1/2} C_W \sqrt{\varepsilon} }\right)$, where the doubly stochastic matrix $W$ represents the network topology and $ C_W $ is a parameter that depends on $W$. Thus, it indicates that the convergence property of $\texttt{DSGT}$ is heavily affected by the topology of the communication network. To overcome the weakness of $\texttt{DSGT}$, we resort to the snap-shot gradient tracking skill and propose two novel algorithms, snap-shot $\texttt{DSGT}$ ($\texttt{SS-DSGT}$) and accelerated snap-shot $\texttt{DSGT}$ ($\texttt{ASS-DSGT}$). We further justify that $\texttt{SS-DSGT}$ exhibits a lower iteration complexity compared to $\texttt{DSGT}$ in the general communication network topology. Additionally, $\texttt{ASS-DSGT}$ matches $\texttt{DSGT}$'s iteration complexity $\mathcal{O}\left( \frac{\bar{\sigma}^2}{m\mu \varepsilon} + \frac{\sqrt{L}\bar{\sigma}}{\mu (1 - \lambda_2(W))^{1/2}\sqrt{\varepsilon}} \right)$ under the same conditions as $\texttt{DSGT}$. Numerical experiments validate $\texttt{SS-DSGT}$'s superior performance performance in the general communication network topology and exhibit better practical performance of $\texttt{ASS-DSGT}$ on the specified $W$ compared to $\texttt{DSGT}$. Hao Di, Haishan Ye, Xiangyu Chang, Guang Dai, Ivor W. Tsang |
ICML | 4 |
| 2024 | Double Variance Reduction: A Smoothing Trick for Composite Optimization Problems without First-Order GradientabstractVariance reduction techniques are designed to decrease the sampling variance, thereby accelerating convergence rates of first-order (FO) and zeroth-order (ZO) optimization methods. However, in composite optimization problems, ZO methods encounter an additional variance called the coordinate-wise variance, which stems from the random gradient estimation. To reduce this variance, prior works require estimating all partial derivatives, essentially approximating FO information. This approach demands $\mathcal{O}(d)$ function evaluations ($d$ is the dimension size), which incurs substantial computational costs and is prohibitive in high-dimensional scenarios. This paper proposes the Zeroth-order Proximal Double Variance Reduction ($\texttt{ZPDVR}$) method, which utilizes the averaging trick to reduce both sampling and coordinate-wise variances. Compared to prior methods, $\texttt{ZPDVR}$ relies solely on random gradient estimates, calls the stochastic zeroth-order oracle (SZO) in expectation $\mathcal{O}(1)$ times per iteration, and achieves the optimal $\mathcal{O}(d(n + \kappa)\log (\frac{1}{\epsilon}))$ SZO query complexity in the strongly convex and smooth setting, where $\kappa$ represents the condition number and $\epsilon$ is the desired accuracy. Empirical results validate $\texttt{ZPDVR}$’s linear convergence and demonstrate its superior performance over other related methods. Hao Di, Haishan Ye, Yueling Zhang, Xiangyu Chang, Guang Dai, Ivor W. Tsang |
ICML | 5 |
| 2024 | Can Gaussian Sketching Converge Faster on a Preconditioned Landscape?abstractThis paper focuses on the large-scale optimization which is very popular in the big data era. The gradient sketching is an important technique in the large-scale optimization. Specifically, the random coordinate descent algorithm is a kind of gradient sketching method with the random sampling matrix as the sketching matrix. In this paper, we propose a novel gradient sketching called GSGD (Gaussian Sketched Gradient Descent). Compared with the classical gradient sketching methods such as the random coordinate descent and SEGA (Hanzely et al., 2018), our GSGD does not require the importance sampling but can achieve a fast convergence rate matching the ones of these methods with importance sampling. Furthermore, if the objective function has a non-smooth regularization term, our GSGD can also exploit the implicit structure information of the regularization term to achieve a fast convergence rate. Finally, our experimental results substantiate the effectiveness and efficiency of our algorithm. Haishan Ye, Guang Dai, Ivor W. Tsang |
ICML | 3 |
| 2024 | Generating Action-conditioned Prompts for Open-vocabulary Video Action Recognition
Chengyou Jia, Minnan Luo, Xiaojun Chang, Zhuohang Dang, Mingfei Han 0002, Mengmeng Wang 0005, Guang Dai, Sizhe Dang, Jingdong Wang 0001 |
ACM Multimedia | 7 |
| 2024 | Schedule Your Edit: A Simple yet Effective Diffusion Noise Schedule for Image EditingabstractText-guided diffusion models have significantly advanced image editing, enabling high-quality and diverse modifications driven by text prompts. However, effective editing requires inverting the source image into a latent space, a process often hindered by prediction errors inherent in DDIM inversion.
These errors accumulate during the diffusion process, resulting in inferior content preservation and edit fidelity, especially with conditional inputs.
We address these challenges by investigating the primary contributors to error accumulation in DDIM inversion and identify the singularity problem in traditional noise schedules as a key issue.
To resolve this, we introduce the *Logistic Schedule*, a novel noise schedule designed to eliminate singularities, improve inversion stability, and provide a better noise space for image editing. This schedule reduces noise prediction errors, enabling more faithful editing that preserves the original content of the source image. Our approach requires no additional retraining and is compatible with various existing editing methods.
Experiments across eight editing tasks demonstrate the Logistic Schedule's superior performance in content preservation and edit fidelity compared to traditional noise schedules, highlighting its adaptability and effectiveness.
The project page is available at https://lonelvino.github.io/SYE/. Haonan Lin, Yan Chen 0031, Jiahao Wang 0004, Wenbin An, Mengmeng Wang 0005, Feng Tian 0002, Yong Liu 0007, Guang Dai, Jingdong Wang 0001, Qianying Wang 0002 |
NeurIPS | 8 |
| 2024 | Flipped Classroom: Aligning Teacher Attention with Student in Generalized Category DiscoveryabstractRecent advancements have shown promise in applying traditional Semi-Supervised Learning strategies to the task of Generalized Category Discovery (GCD). Typically, this involves a teacher-student framework in which the teacher imparts knowledge to the student to classify categories, even in the absence of explicit labels. Nevertheless, GCD presents unique challenges, particularly the absence of priors for new classes, which can lead to the teacher's misguidance and unsynchronized learning with the student, culminating in suboptimal outcomes. In our work, we delve into why traditional teacher-student designs falter in generalized category discovery as compared to their success in closed-world semi-supervised learning. We identify inconsistent pattern learning as the crux of this issue and introduce FlipClass—a method that dynamically updates the teacher to align with the student's attention, instead of maintaining a static teacher reference. Our teacher-attention-update strategy refines the teacher's focus based on student feedback, promoting consistent pattern recognition and synchronized learning across old and new classes. Extensive experiments on a spectrum of benchmarks affirm that FlipClass significantly surpasses contemporary GCD methods, establishing new standards for the field. Haonan Lin, Wenbin An, Jiahao Wang 0004, Yan Chen 0031, Feng Tian 0002, Mengmeng Wang 0005, Qianying Wang 0002, Guang Dai, Jingdong Wang 0001 |
NeurIPS | 8 |
| 2024 | OneActor: Consistent Subject Generation via Cluster-Conditioned GuidanceabstractText-to-image diffusion models benefit artists with high-quality image generation. Yet their stochastic nature hinders artists from creating consistent images of the same subject. Existing methods try to tackle this challenge and generate consistent content in various ways. However, they either depend on external restricted data or require expensive tuning of the diffusion model. For this issue, we propose a novel one-shot tuning paradigm, termed OneActor. It efficiently performs consistent subject generation solely driven by prompts via a learned semantic guidance to bypass the laborious backbone tuning. We lead the way to formalize the objective of consistent subject generation from a clustering perspective, and thus design a cluster-conditioned model. To mitigate the overfitting challenge shared by one-shot tuning pipelines, we augment the tuning with auxiliary samples and devise two inference strategies: semantic interpolation and cluster guidance. These techniques are later verified to significantly improve the generation quality. Comprehensive experiments show that our method outperforms a variety of baselines with satisfactory subject consistency, superior prompt conformity as well as high image quality. Our method is capable of multi-subject generation and compatible with popular diffusion extensions. Besides, we achieve a $4\times$ faster tuning speed than tuning-based baselines and, if desired, avoid increasing the inference time. Furthermore, our method can be naturally utilized to pre-train a consistent subject generation network from scratch, which will implement this research task into more practical applications. (Project page: https://johnneywang.github.io/OneActor-webpage/) Jiahao Wang 0004, Caixia Yan, Haonan Lin, Weizhan Zhang, Mengmeng Wang 0005, Tieliang Gong, Guang Dai, Hao Sun 0015 |
NeurIPS | 7 |
| 2024 | Learning Discretized Neural Networks under Ricci FlowabstractIn this paper, we study Discretized Neural Networks (DNNs) composed of low-precision weights and activations, which suffer from either infinite or zero gradients due to the non-differentiable discrete function during training. Most training-based DNNs in such scenarios employ the standard Straight-Through Estimator (STE) to approximate the gradient w.r.t. discrete values. However, the use of STE introduces the problem of gradient mismatch, arising from perturbations in the approximated gradient. To address this problem, this paper reveals that this mismatch can be interpreted as a metric perturbation in a Riemannian manifold, viewed through the lens of duality theory. Building on information geometry, we construct the Linearly Nearly Euclidean (LNE) manifold for DNNs, providing a background for addressing perturbations. By introducing a partial differential equation on metrics, i.e., the Ricci flow, we establish the dynamical stability and convergence of the LNE metric with the $L^2$-norm perturbation. In contrast to previous perturbation theories with convergence rates in fractional powers, the metric perturbation under the Ricci flow exhibits exponential decay in the LNE manifold. Experimental results across various datasets demonstrate that our method achieves superior and more stable performance for DNNs compared to other representative training-based methods. Jun Chen 0023, Hanwen Chen, Mengmeng Wang 0005, Guang Dai, Ivor W. Tsang, Yong Liu 0007 |
J. Mach. Learn. Res. | 4 |
| 2024 | Disentangled Representation Learning With Transmitted Information BottleneckabstractEncoding only the task-related information from the raw data, i.e., disentangled representation learning, can greatly contribute to the robustness and generalizability of models. Although significant advances have been made by regularizing the information in representations with information theory, two major challenges remain: 1) the representation compression inevitably leads to performance drop; 2) the disentanglement constraints on representations are in complicated optimization. To these issues, we introduce Bayesian networks with transmitted information to formulate the interaction among input and representations during disentanglement. Building upon this framework, we propose DisTIB (Transmitted Information Bottleneck for Disentangled representation learning), a novel objective that navigates the balance between information compression and preservation. We employ variational inference to derive a tractable estimation for DisTIB. This estimation can be simply optimized via standard gradient descent with a reparameterization trick. Moreover, we theoretically prove that DisTIB can achieve optimal disentanglement, underscoring its superior efficacy. To solidify our claims, we conduct extensive experiments on various downstream tasks to demonstrate the appealing efficacy of DisTIB and validate our theoretical analyses. Zhuohang Dang, Minnan Luo, Chengyou Jia, Guang Dai, Jihong Wang 0003, Xiaojun Chang, Jingdong Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Disentangled Generation With Information Bottleneck for Enhanced Few-Shot LearningabstractFew-shot learning (FSL) poses a significant challenge in classifying unseen classes with limited samples, primarily stemming from the scarcity of data. Although numerous generative approaches have been investigated for FSL, their generation process often results in entangled outputs, exacerbating the distribution shift inherent in FSL. Consequently, this considerably hampers the overall quality of the generated samples. Addressing this concern, we present a pioneering framework called DisGenIB, which leverages an Information Bottleneck (IB) approach for Disentangled Generation. Our framework ensures both discrimination and diversity in the generated samples, simultaneously. Specifically, we introduce a groundbreaking Information Theoretic objective that unifies disentangled representation learning and sample generation within a novel framework. In contrast to previous IB-based methods that struggle to leverage priors, our proposed DisGenIB effectively incorporates priors as invariant domain knowledge of sub-features, thereby enhancing disentanglement. This innovative approach enables us to exploit priors to their full potential and facilitates the overall disentanglement process. Moreover, we establish the theoretical foundation that reveals certain prior generative and disentanglement methods as special instances of our DisGenIB, underscoring the versatility of our proposed framework. To solidify our claims, we conduct comprehensive experiments on demanding FSL benchmarks, affirming the remarkable efficacy and superiority of DisGenIB. Furthermore, the validity of our theoretical analyses is substantiated by the experimental results. Our code is available at https://github.com/eric-hang/DisGenIB. Zhuohang Dang, Minnan Luo, Jihong Wang 0003, Chengyou Jia, Caixia Yan, Guang Dai, Xiaojun Chang |
IEEE Trans. Image Process. | 6 |
| 2023 | Boosting Few-shot Action Recognition with Graph-guided Hybrid MatchingabstractClass prototype construction and matching are core aspects of few-shot action recognition. Previous methods mainly focus on designing spatiotemporal relation modeling modules or complex temporal alignment algorithms. Despite the promising results, they ignored the value of class prototype construction and matching, leading to unsatisfactory performance in recognizing similar categories in every task. In this paper, we propose GgHM, a new framework with Graph-guided Hybrid Matching. Concretely, we learn task-oriented features by the guidance of a graph neural network during class prototype construction, optimizing the intra- and inter-class feature correlation explicitly. Next, we design a hybrid matching strategy, combining frame-level and tuple-level matching to classify videos with multivariate styles. We additionally propose a learnable dense temporal modeling module to enhance the video feature temporal representation to build a more solid foundation for the matching process. GgHM shows consistent improvements over other challenging baselines on several few-shot datasets, demonstrating the effectiveness of our method. The code will be publicly available at https://github.com/jiazheng-xing/GgHM. Jiazheng Xing, Mengmeng Wang 0005, Yudi Ruan, Bofan Chen, Yaowei Guo, Boyu Mu, Guang Dai, Jingdong Wang 0001, Yong Liu 0007 |
ICCV | 7 |
| 2023 | SUBP: Soft Uniform Block Pruning for 1×N Sparse CNNs Multithreading Acceleration
Jingyang Xiang, Siqi Li 0009, Jun Chen 0023, Guang Dai, Shipeng Bai, Yukai Ma, Yong Liu 0007 |
NeurIPS | 4 |
| 2022 | Synchronization of complex-valued stochastic coupled systems with hybrid impulses via discrete-time state observations control
Guang Dai, Zhen Guan, Yan Liu 0081 |
Neural Comput. Appl. | 1 |
| 2017 | A Balanced Assignment Mechanism for Online Taxi RecommendationabstractMajority of taxi recommender systems mainly focused on satisfaction of passengers without considering fairness in assignment of taxi drivers. In this paper we propose a balanced assignment mechanism for online taxi recommendation (BAMOTR). BAMOTR provides a mechanism for fair assignment of drivers at some locations with specific routes to pick up passengers and ensures a short waiting time for passengers. Fair assignment is intended to minimize the differences in income among the taxi drivers. Analysis shows out that fair assignment of drivers and shortening the time the passenger wait before pick up is a trade-off problem. In this paper, we set a regulatory factor that can adjust the trade-off between fair assignment of drivers and shortening of waiting time of passengers. We also propose an efficient range refinement algorithm to solve online taxi recommendation problem in BAMOTR. It is theoretically and experimentally proved that range refinement algorithm ensures the same recommendation result as brute-force algorithm, however it greatly reduces the time overhead. We validate the performances of BAMOTR with extensive evaluations. Experimental results show that BAMOTR achieve better recommendation fairness than compared approaches and guarantee a short waiting time for passengers to be picked up. Guang Dai, Stephen Manko Wambura, Heli Sun |
MDM | 1 |
| 2014 | Multicategory large margin classification methods: Hinge losses vs. coherence functions
Zhihua Zhang 0004, Cheng Chen 0015, Guang Dai, Wu-Jun Li, Dit-Yan Yeung |
Artif. Intell. | 3 |
| 2013 | An iterative SVM approach to feature selection and classification in high-dimensional datasets
Dehua Liu, Hui Qian 0001, Guang Dai, Zhihua Zhang 0008 |
Pattern Recognit. | 3 |
| 2012 | Coherence functions with applications in large-margin classification methods
Zhihua Zhang 0004, Dehua Liu, Guang Dai, Michael I. Jordan |
J. Mach. Learn. Res. | 3 |
| 2011 | A non-convex relaxation approach to sparse dictionary learningabstractDictionary learning is a challenging theme in computer vision. The basic goal is to learn a sparse representation from an overcomplete basis set. Most existing approaches employ a convex relaxation scheme to tackle this challenge due to the strong ability of convexity in computation and theoretical analysis. In this paper we propose a non-convex online approach for dictionary learning. To achieve the sparseness, our approach treats a so-called minimax concave (MC) penalty as a nonconvex relaxation of the ℓ0penalty. This treatment expects to obtain a more robust and sparse representation than existing convex approaches. In addition, we employ an online algorithm to adaptively learn the dictionary, which makes the non-convex formulation computationally feasible. Experimental results on the sparseness comparison and the applications in image denoising and image inpainting demonstrate that our approach is more effective and flexible. Jianping Shi, Guang Dai, Jingdong Wang 0001 |
CVPR | 3 |
| 2011 | Bayesian Generalized Kernel Mixed Models
Zhihua Zhang 0004, Guang Dai, Michael I. Jordan |
J. Mach. Learn. Res. | 2 |
| 2010 | Sparse Unsupervised Dimensionality Reduction Algorithms
Wenjun Dou, Guang Dai, Congfu Xu |
ECML/PKDD (1) | 2 |
| 2010 | Regularized Discriminant Analysis, Ridge Regression and Beyond
Zhihua Zhang 0004, Guang Dai, Congfu Xu, Michael I. Jordan |
J. Mach. Learn. Res. | 2 |
| 2010 | A regularization framework for multiclass classification: A deterministic annealing approach
Zhihua Zhang 0004, Gang Wang 0004, Dit-Yan Yeung, Guang Dai, Frederick H. Lochovsky |
Pattern Recognit. | 4 |
| 2009 | Optimal Scoring for Unsupervised LearningabstractWe are often interested in casting classification and clustering problems in a regression framework, because it is feasible to achieve some statistical properties in this framework by imposing some penalty criteria. In this paper we illustrate optimal scoring, which was originally proposed for performing Fisher linear discriminant analysis by regression, in the application of unsupervised learning. In particular, we devise a novel clustering algorithm that we call optimal discriminant clustering (ODC). We associate our algorithm with the existing unsupervised learning algorithms such as spectral clustering, discriminative clustering and sparse principal component analysis. Thus, our work shows that optimal scoring provides a new approach to the implementation of unsupervised learning. This approach facilitates the development of new unsupervised learning algorithms. Guang Dai |
NIPS | 2 |
| 2009 | A Flexible and Efficient Algorithm for Regularized Fisher Discriminant Analysis
Zhihua Zhang 0004, Guang Dai, Michael I. Jordan |
ECML/PKDD (2) | 2 |
| 2008 | A Scalable Kernel-Based Semisupervised Metric Learning Algorithm with Out-of-Sample Generalization AbilityabstractIn recent years, metric learning in the semisupervised setting has aroused a lot of research interest. One type of semisupervised metric learning utilizes supervisory information in the form of pairwise similarity or dissimilarity constraints. However, most methods proposed so far are either limited to linear metric learning or unable to scale well with the data set size. In this letter, we propose a nonlinear metric learning method based on the kernel approach. By applying low-rank approximation to the kernel matrix, our method can handle significantly larger data sets. Moreover, our low-rank approximation scheme can naturally lead to out-of-sample generalization. Experiments performed on both artificial and real-world data show very promising results. Dit-Yan Yeung, Hong Chang 0001, Guang Dai |
Neural Comput. | 3 |
| 2007 | Kernel selection forl semi-supervised kernel machinesabstractExisting semi-supervised learning methods are mostly based on either the cluster assumption or the manifold assumption. In this paper, we propose an integrated regularization framework for semi-supervised kernel machines by incorporating both the cluster assumption and the manifold assumption. Moreover, it supports kernel learning in the form of kernel selection. The optimization problem involves joint optimization over all the labeled and unlabeled data points, a convex set of basic kernels, and a discrete space of unknown labels for the unlabeled data. When the manifold assumption is incorporated, graph Laplacian kernels are used as the basic kernels for learning an optimal convex combination of graph Laplacian kernels. Comparison with related methods on the USPS data set shows very promising results. Guang Dai, Dit-Yan Yeung |
ICML | 1 |
| 2007 | Boosting Kernel Discriminant Analysis and Its Application on Tissue Classification of Gene Expression Data
Guang Dai, Dit-Yan Yeung |
IJCAI | 1 |
| 2007 | A Scalable Kernel-Based Algorithm for Semi-Supervised Metric Learning
Dit-Yan Yeung, Hong Chang 0001, Guang Dai |
IJCAI | 3 |
| 2007 | Face recognition using a kernel fractional-step discriminant analysis algorithm
Guang Dai, Dit-Yan Yeung, Yuntao Qian |
Pattern Recognit. | 1 |
| 2007 | Learning the kernel matrix by maximizing a KFD-based class separability criterion
Dit-Yan Yeung, Hong Chang 0001, Guang Dai |
Pattern Recognit. | 3 |
| 2006 | Tensor Embedding Methods
Guang Dai, Dit-Yan Yeung |
AAAI | 1 |
| 2006 | Extending Kernel Fisher Discriminant Analysis with the Weighted Pairwise Chernoff Criterion
Guang Dai, Dit-Yan Yeung, Hong Chang 0001 |
ECCV (4) | 1 |
| 2006 | Local Discriminant Embedding with Tensor RepresentationabstractWe present a subspace learning method, called local discriminant embedding with tensor representation (LDET), that addresses simultaneously the generalization and data representation problems in subspace learning. LDET learns multiple interrelated subspaces for obtaining a lower-dimensional embedding by incorporating both class label information and neighborhood information. By encoding each object as a second- or higher-order tensor, LDET can capture higher-order structures in the data without requiring a large sample size. Extensive empirical studies have been performed to compare LDET with a second- or third-order tensor representation and the original LDE on their face recognition performance. Not only does LDET have a lower computational complexity than LDE, but LDET is also superior to LDE in terms of its recognition accuracy. Jian Xia, Dit-Yan Yeung, Guang Dai |
ICIP | 3 |
| 2005 | Nonlinear dimensionality reduction for classification using kernel weighted subspace methodabstractWe study the use of kernel subspace methods that learn low-dimensional subspace representations for classification tasks. In particular, we propose a new method called kernel weighted nonlinear discriminant analysis (KWNDA) which possesses several appealing properties. First, like all kernel methods, it handles nonlinearity in a disciplined manner that is also computationally attractive. Second, by introducing weighting functions into the discriminant criterion, it outperforms existing kernel discriminant analysis methods in terms of the classification accuracy. Moreover, it also effectively deals with the small sample size problem. We empirically compare different subspace methods with respect to their classification performance of facial images based on the simple nearest neighbor rule. Experimental results show that KWNDA substantially outperforms competing linear as well as nonlinear subspace methods. Guang Dai, Dit-Yan Yeung |
ICIP (2) | 1 |
| 2004 | Face Recognition Using Novel LDA-Based Algorithms
Guang Dai, Yuntao Qian |
ECAI | 1 |
| 2004 | Modified kernel-based nonlinear feature extraction [face recognition example]abstractFeature extraction techniques are widely used in many applications to pre-process data in order to reduce the complexity of subsequent processes. A group of kernel-based Fisher discriminant analysis (KFDA) algorithms has attracted much attention due to their high performance. In this paper, the inherent limitations of those KFDA algorithms have been discussed and a novel algorithm is proposed to effectively overcome those limitations. Experimental results on face recognition suggest that this proposed algorithm is superior to the existing methods in terms of correct classification rate. Guang Dai, Yuntao Qian, Sen Jia 0001 |
ICASSP (5) | 1 |
| 2004 | A robust feature extraction framework for face recognitionabstractThe kernel fractional-stop nonlinear discriminant analysis (KF-NDA) method not only extends the fractional-step linear discriminant analysis (F-LDA) method to a nonlinear version, but also further improves the generalization ability of traditional kernel nonlinear discriminant analysis (K-NDA). On the other hand, the Gabor transformed face images exhibit strong characteristics of spatial locality, scale and orientation selectivity, similar to those displayed by Gabor wavelets. Such characteristics produce salient local features that are most suitable for face recognition (FR). Hence, the augmented Gabor feature vector (AGFV) derived from a set of downsampled Gabor wavelet representations of face images is robust to the various of face images and simultaneously exhibits the more discriminatory information. Based on the AGFV and the KF-NDA, a robust feature extraction framework, i.e., the Gabor KF-NDA (GKF-NDA), is proposed for FR. In this framework, the KF-NDA method is directly applied to extract the robust nonlinear feature from the AGFV. Experimental results tested on the popular databases show that the GKF-NDA is more effective than oilier existing FR approaches. Guang Dai, Yasutoshi Otani |
ICIP | 1 |
| 2004 | Kernel generalized nonlinear discriminant analysis algorithm for pattern recognition
Guang Dai, Yuntao Qian |
ICIP | 1 |
| 2004 | A Gabor direct fractional-step LDA algorithm for face recognitionabstractRecently, a direct fractional-step linear discriminant analysis (DF-LDA) algorithm was proposed and successfully applied to face recognition (FR). However, the classification performance of DFT-LDA is degraded by the limitations of the direct linear discriminant analysis (D-LDA) used in DF-LDA. We describe a novel DF-LDA to solve this problem, and based on this novel DF-LDA, a novel Gabor DF-LDA (GDF-LDA), which directly applies the novel DF-LDA to the high-dimensional augmented Gabor feature vectors (AGFV) derived from the Gabor wavelet representation of face images, has been proposed for FR. The GDF-LDA not only is robust to facial variations, but also overcomes the limitations of the previous DF-LDA. The comparative results on the ORL database show that the GDF-LDA is more effective than existing FR methods. Guang Dai, Yuntao Qian |
ICME | 1 |