VLDB 2026 Research / reviewers in the wild / expert
Hong-Han Shuai
dblp:86/10294
· DBLP profile ↗
137ranked-venue papers
11as first author
95since 2021 · last 2026
0000-0003-2216-077XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 80 · 1 first-author · 65 since 2021Artificial intelligence and machine learning · 57 · 3 first-author · 40 since 2021Databases, data management, data science and information retrieval · 34 · 8 first-author · 10 since 2021Computer networks · 10 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FreeCond: Free Lunch in the Input Conditions of Text-Guided InpaintingabstractText-to-image inpainting models often exhibit an unpredictable balance among image coherence and prompt adherence. This rigidity limits their adaptability across diverse scenarios, including coarse masks, non-object, and interaction prompts. Recognizing this instability as an indicator of learned generation diversity, we aim to control model behavior for given objective. We propose Empirical Feature Intervention (EFI), a metric-agnostic framework that precomputes how feature interventions influence evaluation metrics—such as CLIP, Human Preference Score (HPS), and Image Reward (IR). Building on EFI, we introduce FreeCond, a free-of-cost framework that applies two simple input interventions (Image Frequency and Mask Value Modulation), these interventions can be further optimized via Surrogate Intervention Optimization (SIO) based on a surrogate model regressed with precomputed EFI data. FreeCond enables real-time, user-interactive control of pretrained models without retraining or architectural modifications. Also, to benchmark performance on challenging settings, we present FCIBench. Experiments on EditBench, BrushBench, and FCIBench demonstrate that FreeCond substantially improves CLIP, HPS, and IR metrics by up to 22%, 8%, and 54%, respectively. Teng-Fang Hsiao, Bo-Kai Ruan, Sung-Lin Tsai, Yi-Lun Wu, Hong-Han Shuai |
WACV | 5 |
| 2026 | DNA: Dual-branch Network with Adaptation for Open-Set Online Handwriting GenerationabstractOnline handwriting generation (OHG) enhances handwriting recognition models by synthesizing diverse, human-like samples. However, existing OHG methods struggle to generate unseen characters, particularly in glyph-based languages like Chinese, limiting their real-world applicability. In this paper, we introduce our method for OHG, where the writer’s style and the characters generated during testing are unseen during training. To tackle this challenge, we propose a Dual-branch Network with Adaptation (DNA), which comprises an adaptive style branch and an adaptive content branch. The style branch learns stroke attributes such as writing direction, spacing, placement, and flow to generate realistic handwriting. Meanwhile, the content branch is designed to generalize effectively to unseen characters by decomposing character content into structural information and texture details, extracted via local and global encoders, respectively. Extensive experiments demonstrate that our DNA model is well-suited for the unseen OHG setting, achieving state-of-the-art performance. Tsai-Ling Huang, Nhat-Tuong Do-Tran, Ngoc-Hoang-Lam Le, Hong-Han Shuai |
WACV | 4 |
| 2026 | Learnable Query-Enhanced Pose TransformationabstractPose-Guided Person Image Synthesis (PGPIS) aims to transfer a person from a source image to a target pose (e.g., skeleton) while preserving their original appearance. Although existing methods can produce high-quality results at first glance, they often suffer from noticeable distortions in fine details. We identify the root cause of these issues as the heavy reliance on pre-trained encoders for extracting visual features from the source image. To address this, we propose a novel Query Enhancement Network composed of two key components: the Query-based Feature Fusion Transformer (QFFT) and Pose-Masked Attention (PMA). The QFFT uses learnable queries to fuses multi-scale features from high to low resolution extracted by the backbone encoder, thereby significantly enhancing the realism of texture details in the generated images. To better capture the relationship between pose information and visual features from the source image, we introduce PMA that uses the pose skeleton as a mask to guide the attention mechanism to focus on the pose regions. Our method produces high-quality, visually coherent results and outperforms existing approaches on standard evaluation metrics, including FID, SSIM, and LPIPS, demonstrating its effectiveness on the DeepFashion dataset. Yi-Zhen Wang, Hong-Han Shuai |
WACV | 2 |
| 2026 | HyperRAG: Reasoning N-ary Facts over Hypergraphs for Retrieval Augmented GenerationabstractGraph-based Retrieval-Augmented Generation (RAG) typically operates on binary Knowledge Graphs (KGs). However, decomposing complex facts into binary triples often leads to semantic fragmentation and longer reasoning paths, increasing the risk of retrieval drift and computational overhead. In contrast, n-ary hypergraphs preserve high-order relational integrity, enabling shallower and more semantically cohesive inference. To exploit this topology, we propose HyperRAG, a framework tailored for n-ary hypergraphs featuring two complementary retrieval paradigms: (i) HyperRetriever learns structural-semantic reasoning over n-ary facts to construct query-conditioned relational chains. It enables accurate factual tracking, adaptive high-order traversal, and interpretable multi-hop reasoning under context constraints. (ii) HyperMemory leverages the LLM's parametric memory to guide beam search, dynamically scoring n-ary facts and entities for query-aware path expansion. Extensive evaluations on WikiTopics (11 closed-domain datasets) and three open-domain QA benchmarks (HotpotQA, MuSiQue, and 2WikiMultiHopQA) validate HyperRAG's effectiveness. HyperRetriever achieves the highest answer accuracy overall, with average gains of 2.95% in MRR and 1.23% in Hits@10 over the strongest baseline. Qualitative analysis further shows that HyperRetriever bridges reasoning gaps through adaptive and interpretable n-ary chain construction, benefiting both open and closed-domain QA. Our codes are publicly available at https://github.com/Vincent-Lien/HyperRAG.git. Wen-Sheng Lien, Yu-Kai Chan, Hao-Lung Hsiao, Bo-Kai Ruan, Meng-Fen Chiang, Chien-An Chen, Yi-Ren Yeh, Hong-Han Shuai |
WWW | 8 |
| 2025 | Future Sight and Tough Fights: Revolutionizing Sequential Recommendation with FENRecabstractSequential recommendation (SR) systems predict user preferences by analyzing time-ordered interaction sequences. A common challenge for SR is data sparsity, as users typically interact with only a limited number of items. While contrastive learning has been employed in previous approaches to address the challenges, these methods often adopt binary labels, missing finer patterns and overlooking detailed information in subsequent behaviors of users. Additionally, they rely on random sampling to select negatives in contrastive learning, which may not yield sufficiently hard negatives during later training stages. In this paper, we propose Future data utilization with Enduring Negatives for contrastive learning in sequential Recommendation (FENRec). Our approach aims to leverage future data with time-dependent soft labels and generate enduring hard negatives from existing data, thereby enhancing the effectiveness in tackling data sparsity. Experiment results demonstrate our state-of-the-art performance across four benchmark datasets, with an average improvement of 6.16% across all metrics. Yu-Hsuan Huang 0002, Ling Lo, Hong-Han Shuai, Wen-Huang Cheng |
AAAI | 4 |
| 2025 | Training-and-Prompt-Free General Painterly Harmonization via Zero-Shot Disentenglement on Style and Content ReferencesabstractPainterly image harmonization aims at seamlessly blending disparate visual elements within a single image. However, previous approaches often struggle due to limitations in training data or reliance on additional prompts, leading to inharmonious and content-disrupted output. To surmount these hurdles, we design a Training-and-prompt-Free General Painterly Harmonization method (TF-GPH). TF-GPH incorporates a novel “Similarity Disentangle Mask”, which disentangles the foreground content and background image by redirecting their attention to corresponding reference images, enhancing the attention mechanism for multi-image inputs. Additionally, we propose a “Similarity Reweighting” mechanism to balance harmonization between stylization and content preservation. This mechanism minimizes content disruption by prioritizing the content-similar features within the given background style reference. Finally, we address the deficiencies in existing benchmarks by proposing novel range-based evaluation metrics and a new benchmark to better reflect real-world applications. Extensive experiments demonstrate the efficacy of our method across benchmarks. Teng-Fang Hsiao, Bo-Kai Ruan, Hong-Han Shuai |
AAAI | 3 |
| 2025 | Memory-Augmented Re-Completion for 3D Semantic Scene CompletionabstractSemantic Scene Completion (SSC) aims to reconstruct a 3D voxel representation occupied by semantic classes based on ordinary inputs such as 2D RGB images, depth maps, or point clouds. Given the cost-effective and promising applications in autonomous driving, camera-based SSC has attracted considerable attention to developing various approaches. However, current methods mainly focus on precise 2D-to-3D projection while overlooking the challenge of completing invisible regions, leading to numerous false negatives and suboptimal SSC performance. To address this issue, we propose a novel architecture, Memory-augmented Re-completion (MARE), designed to enhance completion capability. Our MARE model encapsulates regional relationships by incorporating a memory bank that stores vital region-tokens while two protocols concerning diversity and age are adopted to optimize the bank adversarially. Additionally, we introduce a Re-completion pipeline incorporated with an Information Spreading module to progressively complete the invisible regions while bridging the scale gap between region-level and voxel-level information. Extensive experiments conducted on the SSCBench-KITTI-360 and SemanticKITTI datasets validate the effectiveness of our approach. Yu-Wen Tseng, Sheng-Ping Yang, Jhih-Ciang Wu, I-Bin Liao, Yung-Hui Li, Hong-Han Shuai, Wen-Huang Cheng |
AAAI | 6 |
| 2025 | From Diffusion to Decision: A Diffusion-ReRanking in Scene Text DetectionabstractDiffusion models have recently shown great potential in object detection and instance segmentation, yet their application to scene text detection, with its unique challenges such as instance variability and subjective human annotations, remains unexplored. In this paper, we propose DRR (Diffusion ReRanking), a method that adapts diffusion-based instance segmentation for scene text detection. Traditional instance segmentation often relies on classification scores for ranking, potentially overlooking the accuracy of bounding boxes and mask quality. DRR addresses this by incorporating two networks: a diffusion network, trained with a combination of projection loss and pairwise loss in the mask branch to produce more precise and tightly-bound segmentations, and a reranking network, which refines the results by evaluating bounding box accuracy and mask quality. Extensive experiments demonstrate the effectiveness of DRR, achieving a precision of 86.7%, recall of 81.7%, and an F-measure of 84.1% on CTW1500, highlighting DRR’s potential to advance scene text detection. Jia-Ying Yong, Hong-Han Shuai, Yung-Hui Li, Wen-Huang Cheng |
AVSS | 2 |
| 2025 | MonoTAKD: Teaching Assistant Knowledge Distillation for Monocular 3D Object DetectionabstractMonocular 3D object detection (Mono3D) holds noteworthy promise for autonomous driving applications owing to the cost-effectiveness and rich visual context of monocular camera sensors. However, depth ambiguity poses a significant challenge, as it requires extracting precise 3D scene geometry from a single image, resulting in suboptimal performance when transferring knowledge from a LiDARbased teacher model to a camera-based student model. To facilitate effective distillation, we introduce Monocular Teaching Assistant Knowledge Distillation (MonoTAKD), which proposes a camera-based teaching assistant (TA) model to transfer robust 3D visual knowledge to the student model, leveraging the smaller feature representation gap. Additionally, we define 3D spatial cues as residual features that capture the differences between the teacher and the TA models. We then leverage these cues to improve the student model's 3D perception capabilities. Experimental results show that our MonoTAKD achieves state-of-the-art performance on the KITTI3D dataset. Furthermore, we evaluate the performance on nuScenes and KITTI raw datasets to demonstrate the generalization of our model to multi-view 3D and unsupervised data settings. Our code is available at https://github.com/hoiliu-0801/MonoTAKD. Hou-I Liu, Christine Wu, Jen-Hao Cheng, Wenhao Chai, Shian-Yun Wang, Gaowen Liu, Hugo Latapie, Jhih-Ciang Wu, Jenq-Neng Hwang, Hong-Han Shuai, Wen-Huang Cheng |
CVPR | 10 |
| 2025 | TF-TI2I: Training-Free Text-And-Image-To-Image Generation via Multi-Modal Implicit-Context Learning in Text-To-Image ModelsabstractText-and-Image-To-Image (TI2I), an extension of Text-To-Image (T2I), integrates image inputs with textual instructions to enhance image generation. Existing methods often partially utilize image inputs, focusing on specific elements like objects or styles, or they experience a decline in generation quality with complex, multi-image instructions. To overcome these challenges, we introduce Training-Free Text-and-Image-to-Image (TF-TI2I), which adapts cutting-edge T2I models such as SD3 without the need for additional training. Our method capitalizes on the MM-DiT architecture, in which we point out that textual tokens can implicitly learn visual information from vision tokens. We enhance this interaction by extracting a condensed visual representation from reference images, facilitating selective information sharing through Reference Contextual Masking -- this technique confines the usage of contextual tokens to instruction-relevant visual information. Additionally, our Winner-Takes-All module mitigates distribution shifts by prioritizing the most pertinent references for each vision token. Addressing the gap in TI2I evaluation, we also introduce the FG-TI2I Bench, a comprehensive benchmark tailored for TI2I and compatible with existing T2I methods. Our approach shows robust performance across various benchmarks, confirming its effectiveness in handling complex image-generation tasks. Teng-Fang Hsiao, Bo-Kai Ruan, Yi-Lun Wu, Tzu-Ling Lin, Hong-Han Shuai |
ICCV | 5 |
| 2025 | When Anchors Meet Cold Diffusion: A Multi-Stage Approach to Lane Detection
Bo-Lun Huang, Zi-Xiang Ni, Feng-Kai Huang, Hong-Han Shuai, Wen-Huang Cheng |
ICCV | 4 |
| 2025 | Perspective-Aware Teaching: Adapting Knowledge for Heterogeneous Distillation
Jhe-Hao Lin, Chan-Feng Hsu, Hong-Han Shuai, Wen-Huang Cheng |
ICCV | 5 |
| 2025 | FPW: Frequency-Domain Pixel-by-Pixel Watermarking Against Unauthorized Images Used on Training Generative ModelabstractThe proliferation of diffusion-based generative models has significantly advanced the field of synthetic image generation. However, the lack of transparency in training datasets, often not open-sourced, raises substantial concerns about copyright infringement, potentially involving the unauthorized use of proprietary images. To address this issue, we introduce Forensic Provenance Watermarking (FPW), a novel watermarking approach designed specifically for diffusion models. Specifically, FPW embeds a detectable pattern into the training data that, when used without permission, manifests in the generated images. This unique pattern serves as forensic evidence of data misuse, enabling data owners to legally challenge entities that misuse their copyrighted material. Our method simplifies watermark designs and utilizes spatial relationships within the images, employing a surrogate generative model to ensure robustness across various conditions. Experimental results on multiple datasets manifest that FPW maintains high detection rates even at low poisoning ratios, making it an effective tool for copyright protection in the era of advanced generative models. The code is available on the project page. Jan Chi-Yuan, Hong-Han Shuai |
ICIP | 2 |
| 2025 | IterDiff: Training-Free Iterative Face Editing Via Efficient Clip-Guided Memory BankabstractThe rise of generative models has transformed image generation and editing, enabling high-quality, user-guided outputs. Iterative face editing, essential for applications like virtual makeup and entertainment, allows users to refine images progressively. However, this process often leads to artifact accumulation, semantic inconsistency, and quality degradation over multiple edits. Existing methods, while effective in single-step modifications, struggle with sequential edits. To robustly maintain fidelity and consistency in iterative face editing across multiple sessions, we propose IterDiff, a training-free framework leveraging diffusion models with a novel Training-Free Feature Preservation (TF2P) approach to tackle these challenges by storing and retrieving key-value (KV) pairs from self-attention layers. Additionally, we further improve its efficiency and feasibility by Efficient CLIP-guided Memory Bank (ECMB). Experiments on the proposed benchmark show that IterDiff excels in prompt alignment, content consistency, and image quality, providing a robust solution for iterative facial attribute editing. Code, dataset and supplementary materials are available at https://github.com/david20571015/IterDiff. Chun-Yao Chiu, Feng-Kai Huang, Teng-Fang Hsiao, Hong-Han Shuai, Wen-Huang Cheng |
ICIP | 4 |
| 2025 | Unraveling Vanishing Point And Calibrating Tiny Objects For Semantic Scene CompletionabstractSemantic Scene Completion (SSC) aims to jointly predict semantic categories and 3D occupancy of a scene from coarse inputs, which is crucial for providing reliable perception in autonomous driving. In this paper, we enhance existing SSC models by unveiling the vanishing point region, specifically addressing challenges posed by tiny objects and voxels distant from the monocular camera. At the core of our method, we propose the Vanishing Point Aggregator (VPA) to prior-itize features in high-density central areas. The proposed VPA seamlessly integrates the Vanishing Point Query (VPQ) with the vanilla instance query via a cross-attention fusion mechanism to refine feature representation. To evaluate the effectiveness of our method, we conduct comprehensive experiments on two standard SSC benchmarks and demonstrate that our method achieves SOTA performance. Our approach significantly improves the performance across various semantic classes, including a notable gain of 0.37 mIoU on SemanticKITTI and 0.5 mIoU on SSCBench-KITTI-360 for tiny objects. Ablation studies further validate the efficacy of our innovative query fusion strategy, showcasing its capability in long-range predictions for SSC tasks. Sheng-Ping Yang, Yu-Wen Tseng, Yung-Chieh Yang, I-Bin Liao, Chi-En Huang, Shen-Hsuan Liu, Yung-Hui Li, Jhih-Ciang Wu, Hong-Han Shuai, Wen-Huang Cheng |
ICIP | 9 |
| 2025 | Foreground Focus: Enhancing Coherence and Fidelity in Camouflaged Image GenerationabstractCamouflaged image generation is emerging as a solution to data scarcity in camouflaged vision perception, offering a cost-effective alternative to data collection and labeling. Recently, the state-of-the-art approach successfully generates camouflaged images using only foreground objects. However, it faces two critical weaknesses: 1) the background knowledge does not integrate effectively with foreground features, resulting in a lack of foreground-background coherence (e.g., color discrepancy); 2) the generation process does not prioritize the fidelity of foreground objects, which leads to distortion, particularly for small objects. To address these issues, we propose a Foreground-Aware Camouflaged Image Generation (FACIG) model. Specifically, we introduce a Foreground-Aware Feature Integration Module (FAFIM) to strengthen the integration between foreground features and background knowledge. In addition, a Foreground-Aware Denoising Loss is designed to enhance foreground reconstruction supervision. Experiments on various datasets show our method outperforms previous methods in overall camouflaged image quality and foreground fidelity. Pei-Chi Chen, Chan-Feng Hsu, Hung-Jen Chen 0001, Hong-Han Shuai, Wen-Huang Cheng |
ICME | 6 |
| 2025 | Flowing Crowd to Count Flows: A Self-Supervised Framework for Video Individual CountingabstractVideo Individual Counting (VIC), which seeks to count unique individuals across video sequences without duplication, has broader applications than traditional Video Crowd Counting (VCC), including urban planning, event management, and safety monitoring. However, although current VIC approaches have demonstrated strong capabilities, their reliance on identity-level or group-level annotations necessitates substantial labeling effort and expense. To reduce the high costs of manual annotation, we introduce VIC-SSL, a novel self-supervised learning approach that utilizes unlabeled data along with the innovative feature-level augmentation technique called Foreground-driven ShiftMix (F-ShiftMix). By blending and shifting in the feature space rather than the image space, F-ShiftMix generates realistic crowd motion without explicit annotations, while preserving global semantic coherence. Furthermore, VIC-SSL integrates the Cost-guided Flow Prompt (CFP) and the Distinction-aware Cross-Attention (DCA) to enhance flow-aware localization and inter-frame correspondence learning. Our extensive experiments across three datasets, including SenseCrowd, CroHD, and CARLA, demonstrate that VIC-SSL substantially outperforms existing methods, achieving state-of-the-art results with significantly reduced data requirements. These results showcase VIC-SSL's potential to dramatically lower annotation costs and improve the deployment feasibility of VIC systems in complex scenarios. The project website is available at https://leohuang0511.github.io/vic-ssl. Feng-Kai Huang, Bo-Lun Huang, Li-Wu Tsao, Jhih-Ciang Wu, Hong-Han Shuai, Wen-Huang Cheng |
ACM Multimedia | 5 |
| 2025 | OinkTrack: An Ultra-Long-Term Dataset for Multi-Object Tracking and Re-Identification of Group-Housed PigsabstractLong-term multi-animal tracking in densely group-housed agricultural settings is critical for automated behavior monitoring and early anomaly detection in precision livestock farming. However, it poses significant challenges due to persistent occlusions from feeders and water dispensers, high inter-individual appearance similarity, and drastic visual changes across day and night cycles. Existing multi-object tracking datasets rarely capture the combined difficulty of these real-world conditions. To address this, we introduce OinkTrack, a large-scale benchmark for continuous multi-pig tracking in commercial farm environments. The dataset comprises over five hours of annotated video across sixteen sequences, covering day, night, night-to-day, and day-to-night transitions. Each sequence ranges from one minute to one hour, featuring an average of thirty-six pigs per frame. In total, OinkTrack provides 573,700 bounding boxes linked to 574 consistent pig identities. It enables detailed behavior analysis under varying lighting and crowding conditions. We describe the data collection and annotation process, present statistical insights into tracking difficulty, and benchmark 11 state-of-the-art tracking methods. OinkTrack provides a robust foundation for developing long-term tracking models and supports downstream applications such as individual activity profiling and early detection of abnormal behavior in real-world, high-density animal populations. The complete dataset and supplementary materials are publicly accessible at https://leohuang0511.github.io/oinktrack-page. Feng-Kai Huang, Hong-Wei Xu, Chu-Chuan Lee, Hong-Yi Tu, Hong-Han Shuai, Wen-Huang Cheng |
ACM Multimedia | 5 |
| 2025 | Color Me Correctly: Bridging Perceptual Color Spaces and Text Embeddings for Improved Diffusion GenerationabstractAccurate color alignment in text-to-image (T2I) generation is critical for applications such as fashion, product visualization, and interior design, yet current diffusion models struggle with nuanced and compound color terms (e.g., Tiffany blue, baby pink), often producing images that are misaligned with human intent. Existing approaches rely on cross-attention manipulation, reference images, or fine-tuning but fail to systematically resolve ambiguous color descriptions. To precisely render colors under prompt ambiguity, we propose a training-free framework that enhances color fidelity by leveraging a large language model (LLM) to disambiguate color-related prompts and guiding color blending operations directly in the text embedding space. Our method first employs a large language model (LLM) to resolve ambiguous color terms in the text prompt, and then refines the text embeddings based on the spatial relationships of the resulting color terms in the CIELab color space. Unlike prior methods, our approach improves color accuracy without requiring additional training or external reference images. Experimental results demonstrate that our framework improves color alignment without compromising image quality, bridging the gap between text semantics and visual generation. All supplementary materials are available at https://Sung-Lin.github.io/TintBench/. Sung-Lin Tsai, Bo-Lun Huang, Yu-Ting Shen, Cheng-Yu Yeo, Chiang Tseng, Bo-Kai Ruan, Wen-Sheng Lien, Hong-Han Shuai |
ACM Multimedia | 8 |
| 2025 | RecipeGen: A Step-Aligned Multimodal Benchmark for Real-World Recipe GenerationabstractCreating recipe images is a key challenge in food computing, with applications in culinary education and multimodal recipe assistants. However, existing datasets lack fine-grained alignment between recipe goals, step-wise instructions, and visual content. We present RecipeGen, the first large-scale, real-world benchmark for recipe-based Text-to-Image (T2I), Image-to-Video (I2V), and Text-to-Video (T2V) generation. RecipeGen contains 26,435 recipes, 196,724 images, and 4,491 videos, covering diverse ingredients, cooking procedures, styles, and dish types. We further propose domain-specific evaluation metrics to assess ingredient fidelity and interaction modeling, benchmark representative T2I, I2V, and T2V models, and provide insights for future recipe generation models. Project page is available at https://wenbin08.github.io/RecipeGen. Ruoxuan Zhang, Jidong Gao, Bin Wen 0001, Chenming Zhang, Hong-Han Shuai, Wen-Huang Cheng |
ACM Multimedia | 6 |
| 2025 | CookAnything: A Framework for Flexible and Consistent Multi-Step Recipe Image GenerationabstractCooking is a sequential and visually grounded activity, where each step such as chopping, mixing, or frying carries both procedural logic and visual semantics. While recent diffusion models have shown strong capabilities in text-to-image generation, they struggle to handle structured multi-step scenarios like recipe illustration. Additionally, current recipe illustration methods are unable to adjust to the natural variability in recipe length, generating a fixed number of images regardless of the actual instructions structure. To address these limitations, we present CookAnything, a flexible and consistent diffusion-based framework that generates coherent, semantically distinct image sequences from textual cooking instructions of arbitrary length. The framework introduces three key components: (1) Step-wise Regional Control (SRC), which aligns textual steps with corresponding image regions within a single denoising process; (2) Flexible RoPE, a step-aware positional encoding mechanism that enhances both temporal coherence and spatial diversity; and (3) Cross-Step Consistency Control (CSCC), which maintains fine-grained ingredient consistency across steps. Experimental results on recipe illustration benchmarks show that CookAnything performs better than existing methods in training-based and training-free settings. The proposed framework supports scalable, high-quality visual synthesis of complex multi-step instructions and holds significant potential for broad applications in instructional media, and procedural content creation. More details are at https://github.com/zhangdaxia22/CookAnything. Ruoxuan Zhang, Bin Wen 0001, Songhan Zuo, Jian-Yu Jiang-Lin, Hong-Han Shuai, Wen-Huang Cheng |
ACM Multimedia | 7 |
| 2025 | Ranking-based Preference Optimization for Diffusion Models from Implicit User FeedbackabstractDirect preference optimization (DPO) methods have shown strong potential in aligning text-to-image diffusion models with human preferences by training on paired comparisons. These methods improve training stability by avoiding the REINFORCE algorithm but still struggle with challenges such as accurately estimating image probabilities due to the non-linear nature of the sigmoid function and the limited diversity of offline datasets. In this paper, we introduce Diffusion Denoising Ranking Optimization (Diffusion-DRO), a new preference learning framework grounded in inverse reinforcement learning. Diffusion-DRO removes the dependency on a reward model by casting preference learning as a ranking problem, thereby simplifying the training objective into a denoising formulation and overcoming the non-linear estimation issues found in prior methods. Moreover, Diffusion-DRO uniquely integrates offline expert demonstrations with online policy-generated negative samples, enabling it to effectively capture human preferences while addressing the limitations of offline data. Comprehensive experiments show that Diffusion-DRO delivers improved generation quality across a range of challenging and unseen prompts, outperforming state-of-the-art baselines in both both quantitative metrics and user studies. Our source code and pre-trained models are available at https://github.com/basiclab/DiffusionDRO. Yi-Lun Wu, Bo-Kai Ruan, Chiang Tseng, Hong-Han Shuai |
NeurIPS | 4 |
| 2025 | RSMAE: Radiometric Resolution and Scale-Aware Masked Autoencoder for SAR Ship RecognitionabstractSynthetic Aperture Radar (SAR) has emerged as an indispensable tool for maritime surveillance, providing reliable all-weather, day-and-night imaging capabilities. However, automated ship recognition in SAR imagery presents significant challenges, particularly due to variations in patch sizes, with small-size SAR image patches posing the greatest difficulties. To address this challenge, we propose Radiometric Resolution and Scale-aware Masked Autoencoder (RSMAE), a novel framework designed for SAR ship recognition. Our method incorporates three key innovations: (1) a scale-aware augmentation that adapts to images of varying image sizes for masked image modeling, enabling the model to learn multi-scale features and reconstruct fine-grained details lost during upscaling; (2) a foreground-background balanced masking strategy that independently handles the ship region and its surrounding region, ensuring the ship area is neither over-masked nor under-masked; and (3) a radiometric resolution-aware (RadRe-aware) reweighting mechanism that leverages SAR-specific radiometric characteristics to enhance reconstruction of challenging samples. Experimental results on the OpenSARShip dataset demonstrate that the proposed RSMAE consistently outperforms state-of-the-art methods by at least 2.07% in terms of recognition accuracy. These findings highlight the robustness and efficiency of the proposed RSMAE, making it a compelling solution for SAR ship recognition tasks. Wei-Lun Tseng, Yi-Lun Wu, Ming-Chun Lee, Hong-Han Shuai |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2025 | Swapped logit distillation via bi-level teacher alignment
Stephen Ekaputra Limantoro, Jhe-Hao Lin, Chih-Yu Wang 0006, Yi-Lung Tsai, Hong-Han Shuai, Wen-Huang Cheng |
Multim. Syst. | 5 |
| 2025 | Compressing Deep Neural Networks with Goal-Specific Pruning and Self-DistillationabstractNeural network (NN) compression aims at reducing the model size and receives much research attention. Nevertheless, we observe that when compressing convolutional neural networks (CNNs), previous approaches may not well measure the impact of filters to loss, resulting in a significant performance degradation after compression. On the other hand, for compressing the fully connected neural networks (FCNNs), we observe that converting the weight matrix to the block diagonal structure would result in better compression. Therefore, for compressing CNNs, we propose a new pipeline in this article, named Retraining-Aware Pruning (RAP) , with a new self-distillation approach, named High-Level Activation-Guided Attention-Preserving Self-Distillation (HAP) and a novel filter pruning strategy, named Normalized Gradients and Geometric Median (NGGM) to effectively improve the accuracy and reduce the model size. Further, for reducing the model size of FCNNs, we formulate a new research problem, i.e., Compression with Difference-Minimized Block Diagonal Structure (COMIS) , and propose a new algorithm, Memory-Efficient and Structure-Aware Compression (MESA) to effectively prune the weights into a block diagonal structure to significantly boost the compression rate. Extensive experiments on different models show that our approaches significantly outperform the state-of-the-art baselines in terms of compression rate, accuracy, and inference speed-up. Fa-You Chen, Yun-Jui Hsu, Chia-Hsun Lu, Hong-Han Shuai, Lo-Yao Yeh |
ACM Trans. Knowl. Discov. Data | 4 |
| 2024 | Revealing Hidden Context in Camouflage Instance Segmentation
Thanh Hai Phung, Hong-Han Shuai |
ACCV (8) | 2 |
| 2024 | Capture Concept Through Comparison: Vision-and-Language Representation Learning with Intrinsic Information Mining
Yun-Zhu Song, Yi-Syuan Chen, Tzu-Ling Lin, Bei Liu 0001, Jianlong Fu, Hong-Han Shuai |
ACCV (3) | 6 |
| 2024 | Distraction is All You Need: Memory-Efficient Image Immunization against Diffusion-Based Image EditingabstractRecent text-to-image (T2I) diffusion models have revolutionized image editing by empowering users to control out-comes using natural language. However, the ease of image manipulation has raised ethical concerns, with the poten-tial for malicious use in generating deceptive or harmful content. To address the concerns, we propose an image im-munization approach named semantic attack to protect our images from being manipulated by malicious agents using diffusion models. Our approach focuses on disrupting the semantic understanding of T2I diffusion models regarding specific content. By attacking the cross-attention mecha-nism that encodes image features with text messages during editing, we distract the model's attention regarding the con-tent of our concern. Our semantic attack renders the model uncertain about the areas to edit, resulting in poorly edited images and contradicting the malicious editing attempts. In addition, by shifting the attack target towards intermediate attention maps from the final generated image, our approach substantially diminishes computational burden and alleviates GPU memory constraints in comparison to pre-vious methods. Moreover, we introduce timestep universal gradient updating to create timestep-agnostic perturbations effective across different input noise levels. By treating the full diffusion process as discrete denoising timesteps during the attack, we achieve equivalent or even superior immu-nization efficacy with nearly half the memory consumption of the previous method. Our contributions include a prac-tical and effective approach to safeguard images against malicious editing, and the proposed method offers robust immunization against various image inpainting and editing approaches, showcasing its potential for real-world appli-cations. Ling Lo, Cheng Yu Yeo, Hong-Han Shuai, Wen-Huang Cheng |
CVPR | 3 |
| 2024 | EmoVIT: Revolutionizing Emotion Insights with Visual Instruction TuningabstractVisual Instruction Tuning represents a novel learning paradigm involving the fine-tuning of pre-trained language models using task-specific instructions. This paradigm shows promising zero-shot results in various natural language processing tasks but is still unexplored in vision emotion understanding. In this work, we focus on enhancing the model's proficiency in understanding and adhering to instructions related to emotional contexts. Initially, we identify key visual clues critical to visual emotion recognition. Subsequently, we introduce a novel GPT-assisted pipeline for generating emotion visual instruction data, effectively addressing the scarcity of annotated instruction data in this domain. Expanding on the groundwork established by InstructBLIP, our proposed EmoVIT architecture incorporates emotion-specific instruction data, leveraging the powerful capabilities of Large Language Models to enhance performance. Through extensive experiments, our model showcases its proficiency in emotion classification, adeptness in affective reasoning, and competence in comprehending humor. The comparative analysis provides a robust benchmark for Emotion Visual Instruction Tuning in the era of LLMs, providing valuable insights and opening avenues for future exploration in this domain. Our code is available at https://github.com/aimmemotion/EmoVIT. Chu-Jun Peng, Yu-Wen Tseng, Hung-Jen Chen 0001, Chan-Feng Hsu, Hong-Han Shuai, Wen-Huang Cheng |
CVPR | 6 |
| 2024 | DQ-DETR: DETR with Dynamic Query for Tiny Object Detection
Yi-Xin Huang, Hou-I Liu, Hong-Han Shuai, Wen-Huang Cheng |
ECCV (76) | 3 |
| 2024 | DetailSemNet: Elevating Signature Verification Through Detail-Semantic Integration
Meng-Cheng Shih, Tsai-Ling Huang, Yu-Heng Shih, Hong-Han Shuai, Hsuan-Tung Liu, Yi-Ren Yeh |
ECCV (25) | 4 |
| 2024 | TrajPrompt: Aligning Color Trajectory with Vision-Language Representations
Li-Wu Tsao, Hao-Tang Tsui, Yu-Rou Tuan, Pei-Chi Chen, Kuan-Lin Wang, Jhih-Ciang Wu, Hong-Han Shuai, Wen-Huang Cheng |
ECCV (41) | 7 |
| 2024 | The Fabrication of Reality and Fantasy: Scene Generation with LLM-Assisted Prompt Interpretation
Chan-Feng Hsu, Jhe-Hao Lin, Terence Lin, Yi-Ning Huang, Hong-Han Shuai, Wen-Huang Cheng |
ECCV (22) | 7 |
| 2024 | Toward Low Artifact Virtual Try-On Via Pre-Warping Partitioned Clothing AlignmentabstractMost image-based try-on methods adopt a warping model to deform the in-shop clothes directly, but they often encounter distortion or corrupt results when dealing with complex body poses or testing on wild data. To address the challenge, we propose a pre-warping partitioned clothing alignment method toward artifact-free virtual try-on in wild data. Specifically, we use perspective transformation to warp different parts of the in-shop clothes and then adjust the results using the Warping-Parsing Condition Generator module, simultaneously generating human parsing. This approach simplifies the learning objectives of the clothing warping module, eliminating the need for significant displacements or rotations when dealing with complex poses. Experimental results demonstrate that our approach is more stable and reduces artifacts compared to state-of-the-art methods. Wei-Chian Liang, Chieh-Yun Chen, Hong-Han Shuai |
ICIP | 3 |
| 2024 | Hierarchically Aggregated Identification Transformer Network for Camouflaged Object DetectionabstractCamouflaged object detection (COD) targets the segmentation of objects hidden in intricate environments, a task complicated by the pronounced similarities between objects and their surroundings. The diverse appearances of camouflaged objects, such as different view angles, partial visibilities, and ambiguous forms, further exacerbate this challenge. To address these issues, we introduce the Hierarchically Aggregated Identification Transformer Network (HAIT-Net). HAITNet harnesses local and global features to refine object localization by employing multi-scale transformer features unified through the Feature Cascaded Fusion Module (FCFM). To tackle ambiguity from indistinct textures, we present the Graph-based Low-level Feature Enhancement Module (GLFEM) and Graph-based Feature Aggregation Module (GFAM). GLFEM enhances texture representation in ambiguous areas, while GFAM reduces false positives and refines prediction maps by discerning contextual relationships. Experimental results on three widely used datasets demonstrate that the proposed HAITNet outperforms the state-of-the-art approaches. Our code is available at https://github.com/underlmao/HAITNet. Thanh Hai Phung, Hung-Jen Chen 0001, Hong-Han Shuai |
ICME | 3 |
| 2024 | Learning Efficient Interaction Anchor for HOI DetectionabstractHuman-object interaction (HOI) detection seeks complicated relationships between humans and objects, yet struggles persist in correctly associating multiple objects with a single human in complex interaction scenarios. In this paper, we tackle such an issue by introducing a novel interaction anchor that employs flexible strategies across different decoder layers with Barlow constraint and Interactivity-Instance Fusion. The proposed modules are both additive and easily implementable in existing approaches, offering computational efficiency within transformer-based models to compact cross-interactivity. Extensive experiments validate the effectiveness of our method, demonstrating comparable performance on HICO-DET and V-COCO for HOI detection. Lirong Xue, Kang-Yang Huang, Rong Chao, Jhih-Ciang Wu, Hong-Han Shuai, Yung-Hui Li, Wen-Huang Cheng |
ICME | 5 |
| 2024 | ReCorD: Reasoning and Correcting Diffusion for HOI GenerationabstractDiffusion models revolutionize image generation by leveraging natural language to guide the creation of multimedia content. Despite significant advancements in such generative models, challenges persist in depicting detailed human-object interactions, especially regarding pose and object placement accuracy. We introduce a training-free method named Reasoning and Correcting Diffusion (ReCorD) to address these challenges. Our model couples Latent Diffusion Models with Visual Language Models to refine the generation process, ensuring precise depictions of HOIs. We propose an interaction-aware reasoning module to improve the interpretation of the interaction, along with an interaction correcting module to refine the output image for more precise HOI generation delicately. Through a meticulous process of pose selection and object positioning, ReCorD achieves superior fidelity in generated images while efficiently reducing computational requirements. We conduct comprehensive experiments on three benchmarks to demonstrate the significant progress in solving text-to-image generation tasks, showcasing ReCorD's ability to render complex interactions accurately by outperforming existing methods in HOI classification score, as well as FID and Verb CLIP-Score. Project website is available at https://alberthkyhky.github.io/ReCorD/ . Jian-Yu Jiang-Lin, Kang-Yang Huang, Ling Lo, Yi-Ning Huang, Terence Lin, Jhih-Ciang Wu, Hong-Han Shuai, Wen-Huang Cheng |
ACM Multimedia | 7 |
| 2024 | A Cat Is A Cat (Not A Dog!): Unraveling Information Mix-ups in Text-to-Image Encoders through Causal Analysis and Embedding OptimizationabstractThis paper analyzes the impact of causal manner in the text encoder of text-to-image (T2I) diffusion models, which can lead to information bias and loss. Previous works have focused on addressing the issues through the denoising process. However, there is no research discussing how text embedding contributes to T2I models, especially when generating more than one object. In this paper, we share a comprehensive analysis of text embedding: i) how text embedding contributes to the generated images and ii) why information gets lost and biases towards the first-mentioned object. Accordingly, we propose a simple but effective text embedding balance optimization method, which is training-free, with an improvement of 125.42\% on information balance in stable diffusion. Furthermore, we propose a new automatic evaluation metric that quantifies information loss more accurately than existing methods, achieving 81\% concordance with human assessments. This metric effectively measures the presence and accuracy of objects, addressing the limitations of current distribution scores like CLIP's text-image similarities. Chieh-Yun Chen, Chiang Tseng, Li-Wu Tsao, Hong-Han Shuai |
NeurIPS | 4 |
| 2024 | Personalized EDM Subject Generation via Co-factored User-Subject Embedding
Yu-Hsiu Chen, Zhi Rui Tam, Hong-Han Shuai |
PAKDD (2) | 3 |
| 2024 | Natural Light Can Also be Dangerous: Traffic Sign Misinterpretation Under Adversarial Natural Light AttacksabstractCommon illumination sources like sunlight or artificial light may introduce hidden vulnerabilities to AI systems. Our paper delves into these potential threats, offering a novel approach to simulate varying light conditions, including sunlight, headlights, and flashlight illuminations. Moreover, unlike typical physical adversarial attacks requiring conspicuous alterations, our method utilizes a model-agnostic black-box attack integrated with the Zeroth Order Optimization (ZOO) algorithm to identify deceptive patterns in a physically-applicable space. Consequently, attackers can recreate these simulated conditions, deceiving machine learning models with seemingly natural light. Empirical results demonstrate the efficacy of our method, misleading models trained on the GTSRB and LISA datasets under natural-like physical environments with an attack success rate exceeding 70% across all digital datasets, and remaining effective against all evaluated real-world traffic signs. Importantly, after adversarial training using samples generated from our approach, models showcase enhanced robustness, underscoring the dual value of our work in both identifying and mitigating potential threats.1 Teng-Fang Hsiao, Bo-Lun Huang, Zi-Xiang Ni, Yan-Ting Lin, Hong-Han Shuai, Yung-Hui Li, Wen-Huang Cheng |
WACV | 5 |
| 2024 | Semantic Fusion Augmentation and Semantic Boundary Detection: A Novel Approach to Multi-Target Video Moment RetrievalabstractGiven an untrimmed video and a natural language query, video moment retrieval (VMR) aims to retrieve video moments described by the query. However, most existing VMR methods assume a one-to-one mapping between the input query and the target video moment (single-target VMR), disregarding the possibility that a video may contain multiple target moments that match the query description (multi-target VMR). Previous methods tackle multi-target VMR by incorporating false negative moments with the original target moment for multi-target training. However, existing methods cannot properly work when no false negative moments exist in the video, or when the identified false negative moments are noisy but are still being utilized as pseudo-labels. In this paper, we propose to tackle multi-target VMR by Semantic Fusion Augmentation and Semantic Boundary Detection (SFABD). Specifically, we use feature-level augmentation to generate augmented target moments, along with an intra-video contrastive loss to ensure feature consistency. Meanwhile, we perform semantic boundary detection to adaptively remove all false negatives from the negative set of contrastive loss to avoid semantic confusion. Extensive experiments conducted on Charades-STA, ActivityNet Captions, and QVHighlights show that our method achieves state-of-the-art performance on multi-target metrics and single-target metrics. The source code is available at https://github.com/basiclab/SFABD. Yi-Lun Wu, Hong-Han Shuai |
WACV | 3 |
| 2024 | Arbitrary-Resolution and Arbitrary-Scale Face Super-Resolution with Implicit Representation NetworksabstractFace super-resolution (FSR) is a critical technique for enhancing low-resolution facial images and has significant implications for face-related tasks. However, existing FSR methods are limited by fixed up-sampling scales and sensitivity to input size variations. To address these limitations, this paper introduces an Arbitrary-Resolution and Arbitrary-Scale FSR method with implicit representation networks (ARASFSR), featuring three novel designs. First, ARASFSR employs 2D deep features, local relative coordinates, and up-sampling scale ratios to predict RGB values for each target pixel, allowing super-resolution at any up-sampling scale. Second, a local frequency estimation module captures high-frequency facial texture information to reduce the spectral bias effect. Lastly, a global coordinate modulation module guides FSR to leverage prior facial structure knowledge and achieve resolution adaptation effectively. Quantitative and qualitative evaluations demonstrate the robustness of ARASFSR over existing state-of-the-art methods while super-resolving facial images across various input sizes and up-sampling scales. Yi-Ting Tsai, Yu Wei Chen, Hong-Han Shuai |
WACV | 3 |
| 2024 | Density-Based Flow Mask Integration via Deformable Convolution for Video People Flux EstimationabstractCrowd counting is currently applied in many areas, such as transportation hubs and streets. However, most of the research still focuses on counting the number of people in a single image, and there is little research on solving the problem of calculating the number of non-repeated people in a video segment. Currently, multiple object tracking is mainly relied upon for video counting, but this method is not suitable for situations where the crowd density is too high. Therefore, we propose a Flow Mask Integration Deformable Convolution network (FMDC) combined with Inter-Frame Head Contrastive Learning (IFHC) to predict the situation of people entering and exiting the screen in a density-based manner. We verify that our proposed method is highly effective in densely populated situations and diverse scenes, and the experimental results show that our proposed method surpasses existing methods. Chang-Lin Wan, Feng-Kai Huang, Hong-Han Shuai |
WACV | 3 |
| 2024 | Predicting and Exploring Abandonment Signals in a Banking Task-Oriented Chatbot ServiceabstractIn this study, we developed predictive models to address the problem of chatbot abandonment, a problem that can result in losing business opportunities. Specifically, we target on a conversation log dataset of a banking chatbot involving 1,373 users and hand-crafted features. By leveraging a pre-trained BERT model on the textural features and the hand-crafted features, the model achieved an F1-score of 0.89 in predicting discontinued conversation and 0.80 in predicting abandonment. Our findings indicate that textual features help capture more abandonment, while hand-crafted features improve detection precision. Our analysis with SHAP and LIME revealed that user typing, the chatbot expressing inability of recognizing intent, and the chatbot asking what users want to do during an ongoing conversation are top signals of user abandoning the chatbot. These findings suggest that chatbot designers should consider providing pre-set options or constraints for user inputs and presenting possible intents to the user and avoid expressing inability, incompetence, or ignoring the users’ current attempt. Chieh Hsu, Hsin-Chien Tung, Hong-Han Shuai, Yung-Ju Chang |
Int. J. Hum. Comput. Interact. | 3 |
| 2024 | CA-FER: Mitigating Spurious Correlation With Counterfactual Attention in Facial Expression RecognitionabstractAlthough facial expression recognition based on deep learning has become a major trend, existing methods have been found to prefer learning spurious statistical correlations and non-robust features during training. This degenerates the model's generalizability in practical situations. One of the research fields mitigating such misperception of correlations as causality is causal reasoning. In this paper, we propose a learnable counterfactual attention mechanism, CA-FER, that uses causal reasoning to simultaneously optimize feature discrimination and diversity to mitigate spurious correlations in expression datasets. To the best of our knowledge, this is the first work to study the spurious correlations in facial expression recognition from a counterfactual attention perspective. Extensive experiments on a synthetic dataset and four public datasets demonstrate that our method outperforms previous methods, which shows the effectiveness and generalizability of our learnable counterfactual attention mechanism for the expression recognition task. Pin-Jui Huang, Hung-Cheng Huang, Hong-Han Shuai, Hao-Wen Cheng |
IEEE Trans. Affect. Comput. | 4 |
| 2024 | Modeling Uncertainty for Low-Resolution Facial Expression RecognitionabstractRecently, facial expression recognition techniques have made significant progress on high-resolution web images. However, in real-world applications, the obtained images are often with low resolution since they are mostly captured in a wide range of public spaces. As a result, the ambiguity of the expression labels hinders recognition performance due to not only subjective emotion annotations but also ambiguous images. Existing approaches tend to perform poorly when the resolution of face images decreases. In this work, we aim to model the aleatoric uncertainty induced by low-image-resolution and label ambiguity for robust facial expression recognition. We propose probabilistic data uncertainty learning to capture the ambiguity induced by poor image resolution. Additionally, we introduce the emotion wheel to learn the label-uncertainty-aware embedding. Moreover, we exploit the ambiguous nature of neutrality and propose a neutral expression constraint to learn more robust features for facial expression recognition. To the best of our knowledge, this is the first work utilizing the intrinsic nature of neutrality as a regularization to benefit model training. Extensive experimental results show the effectiveness and robustness of our approach. Under low-resolution conditions, our proposed method outperforms the state-of-the-art approaches by 3.02% and 3.16% in terms of accuracy on RAF-DB and FERPlus, respectively. Ling Lo, Bo-Kai Ruan, Hong-Han Shuai, Hao-Wen Cheng |
IEEE Trans. Affect. Comput. | 3 |
| 2024 | A DeNoising FPN With Transformer R-CNN for Tiny Object DetectionabstractDespite notable advancements in the field of computer vision, the precise detection of tiny objects continues to pose a significant challenge, largely owing to the minuscule pixel representation allocated to these objects in imagery data. This challenge resonates profoundly in the domain of geoscience and remote sensing, where high-fidelity detection of tiny objects can facilitate a myriad of applications ranging from urban planning to environmental monitoring. In this paper, we propose a new framework, namely, DeNoising FPN with Trans R-CNN (DNTR), to improve the performance of tiny object detection. DNTR consists of an easy plug-in design, DeNoising FPN (DN-FPN), and an effective Transformer-based detector, Trans R-CNN. Specifically, feature fusion in the feature pyramid network is important for detecting multiscale objects. However, noisy features may be produced during the fusion process since there is no regularization between the features of different scales. Therefore, we introduce a DN-FPN module that utilizes contrastive learning to suppress noise in each level’s features in the top-down path of FPN. Second, based on the two-stage framework, we replace the obsolete R-CNN detector with a novel Trans R-CNN detector to focus on the representation of tiny objects with self-attention. Experimental results manifest that our DNTR outperforms the baselines by at least 17.4% in terms of APvton the AI-TOD dataset and 9.6% in terms of AP on the VisDrone dataset, respectively. Our code will be available at https://github.com/hoiliu-0801/DNTR. Hou-I Liu, Yu-Wen Tseng, Kai-Cheng Chang, Pin-Jyun Wang, Hong-Han Shuai, Wen-Huang Cheng |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Temporal Difference-Aware Graph Convolutional Reinforcement Learning for Multi-Intersection Traffic Signal ControlabstractTraffic light control plays a crucial role in intelligent transportation systems. This paper introduces Temporal Difference-Aware Graph Convolutional Reinforcement Learning (TeDA-GCRL), a decentralized RL-based method for efficient multi-intersection traffic signal control. Specifically, we put forward a new graph architecture using each lane as a node for considering intersection relations. Additionally, we propose two new rewards by considering temporal information, namely Temporal-Aware Pressure on Incoming Lanes (TAPIL) and Temporal-Aware Action Consistency (TAAC), which enhance learning efficiency and time-interval sensitivity. Experimental results on five datasets show the superiority of TeDA-GCRL over state-of-the-art methods by at least 9.5% in average travel time. Wei-Yu Lin, Yun-Zhu Song, Bo-Kai Ruan, Hong-Han Shuai, Li-Chun Wang 0001, Yung-Hui Li |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Language-guided Residual Graph Attention Network and Data Augmentation for Visual GroundingabstractVisual grounding is an essential task in understanding the semantic relationship between the given text description and the target object in an image. Due to the innate complexity of language and the rich semantic context of the image, it is still a challenging problem to infer the underlying relationship and to perform reasoning between the objects in an image and the given expression. Although existing visual grounding methods have achieved promising progress, cross-modal mapping across different domains for the task is still not well handled, especially when the expressions are complex and long. To address the issue, we propose a language-guided residual graph attention network for visual grounding (LRGAT-VG), which enables us to apply deeper graph convolution layers with the assistance of residual connections between them. This allows us to better handle long and complex expressions than other graph-based methods. Furthermore, we perform a Language-guided Data Augmentation (LGDA), which is based on copy-paste operations on pairs of source and target images to increase the diversity of training data while maintaining the relationship between the objects in the image and the expression. With extensive experiments on three visual grounding benchmarks, including RefCOCO, RefCOCO+, and RefCOCOg, LRGAT-VG with LGDA achieves competitive performance with other state-of-the-art graph network-based referring expression approaches and demonstrates its effectiveness. Jia Wang 0020, Hong-Han Shuai, Yung-Hui Li, Wen-Huang Cheng |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | CMAF: Cross-Modal Augmentation via Fusion for Underwater Acoustic Image RecognitionabstractUnderwater image recognition is crucial for underwater detection applications. Fish classification has been one of the emerging research areas in recent years. Existing image classification models usually classify data collected from terrestrial environments. However, existing image classification models trained with terrestrial data are unsuitable for underwater images, as identifying underwater data is challenging due to their incomplete and noisy features. To address this, we propose a cross-modal augmentation via fusion ( CMAF ) framework for acoustic-based fish image classification. Our approach involves separating the process into two branches: visual modality and sonar signal modality, where the latter provides a complementary character feature. We augment the visual modality, design an attention-based fusion module, and adopt a masking-based training strategy with a mask-based focal loss to improve the learning of local features and address the class imbalance problem. Our proposed method outperforms the state-of-the-art methods. Our source code is available at https://github.com/WilkinsYang/CMAF . Shih-Wei Yang, Li-Hsiang Shen, Hong-Han Shuai, Kai-Ten Feng |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | Zero-Shot Face-Based Voice Conversion: Bottleneck-Free Speech Disentanglement in the Real-World ScenarioabstractOften a face has a voice. Appearance sometimes has a strong relationship with one's voice. In this work, we study how a face can be converted to a voice, which is a face-based voice conversion. Since there is no clean dataset that contains face and speech, voice conversion faces difficult learning and low-quality problems caused by background noise or echo. Too much redundant information for face-to-voice also causes synthesis of a general style of speech. Furthermore, previous work tried to disentangle speech with bottleneck adjustment. However, it is hard to decide on the size of the bottleneck. Therefore, we propose a bottleneck-free strategy for speech disentanglement. To avoid synthesizing the general style of speech, we utilize framewise facial embedding. It applied adversarial learning with a multi-scale discriminator for the model to achieve better quality. In addition, the self-attention module is added to focus on content-related features for in-the-wild data. Quantitative experiments show that our method outperforms previous work. Shao-En Weng, Hong-Han Shuai, Wen-Huang Cheng |
AAAI | 2 |
| 2023 | Beyond Detection: A Defend-and-Summarize Strategy for Robust and Interpretable Rumor Analysis on Social MediaabstractAs the impact of social media gradually escalates, people are more likely to be exposed to indistinguishable fake news.Therefore, numerous studies have attempted to detect rumors on social media by analyzing the textual content and propagation paths.However, fewer works on rumor detection tasks consider the malicious attacks commonly observed at response level.Moreover, existing detection models have poor interpretability.To address these issues, we propose a novel framework named Defend-And-Summarize (DAS) based on the concept that responses sharing similar opinions should exhibit similar features.Specifically, DAS filters out the attack responses and summarizes the responsive posts of each conversation thread in both extractive and abstractive ways to provide multi-perspective prediction explanations.Furthermore, we enhance our detection architecture with the transformer and Bi-directional Graph Convolutional Networks.Experiments on three public datasets, i.e., RumorEval2019, Twitter15, and Twitter16, demonstrate that our DAS defends against malicious attacks and provides prediction explanations, and the proposed detection model achieves state-of-the-art. 1 Yi-Ting Chang, Yun-Zhu Song, Yi-Syuan Chen, Hong-Han Shuai |
EMNLP | 4 |
| 2023 | Size Does Matter: Size-aware Virtual Try-on via Clothing-oriented Transformation Try-on NetworkabstractVirtual try-on tasks aim at synthesizing realistic try-on results by trying target clothes on humans. Most previous works relied on the Thin Plate Spline or appearance flows to warp clothes to fit human body shapes. However, both approaches cannot handle complex warping, leading to over distortion or misalignment. Furthermore, there is a critical unaddressed challenge of adjusting clothing sizes for try-on. To tackle these issues, we propose a Clothing-Oriented Transformation Try-On Network (COTTON). COTTON leverages clothing structure with landmarks and segmentation to design a novel landmark-guided transformation for precisely deforming clothes, allowing for size adjustment during try-on. Additionally, to properly remove the clothing region from the human image without losing significant human characteristics, we propose a clothing elimination policy based on both transformed clothes and human segmentation. This method enables users to try on clothes tucked-in or untucked while retaining more human characteristics. Both qualitative and quantitative results show that COTTON outperforms the state-of-the-art high-resolution virtual try-on approaches. All the code is available at https://github.com/cotton6/COTTON-size-does-matter. Chieh-Yun Chen, Yi-Chung Chen, Hong-Han Shuai, Wen-Huang Cheng |
ICCV | 3 |
| 2023 | SINC: Self-Supervised In-Context Learning for Vision-Language TasksabstractLarge Pre-trained Transformers exhibit an intriguing capacity for in-context learning. Without gradient updates, these models can rapidly construct new predictors from demonstrations presented in the inputs. Recent works promote this ability in the vision-language domain by incorporating visual information into large language models that can already make in-context predictions. However, these methods could inherit issues in the language domain, such as template sensitivity and hallucination. Also, the scale of these language models raises a significant demand for computations, making learning and operating these models resource-intensive. To this end, we raise a question: "How can we enable in-context learning without relying on the intrinsic in-context ability of large language models?". To answer it, we propose a succinct and general framework, Self-supervised IN-Context learning (SINC), that introduces a meta-model to learn on self-supervised prompts consisting of tailored demonstrations. The learned models can be transferred to downstream tasks for making in-context predictions on-the-fly. Extensive experiments show that SINC outperforms gradient-based methods in various vision-language tasks under few-shot settings. Furthermore, the designs of SINC help us investigate the benefits of in-context learning across different tasks, and the analysis further reveals the essential components for the emergence of in-context learning in the vision-language domain. Yi-Syuan Chen, Yun-Zhu Song, Cheng Yu Yeo, Bei Liu 0001, Jianlong Fu, Hong-Han Shuai |
ICCV | 6 |
| 2023 | Most Important Person-guided Dual-branch Cross-Patch Attention for Group Affect RecognitionabstractGroup affect refers to the subjective emotion that is evoked by an external stimulus in a group, which is an important factor that shapes group behavior and outcomes. Recognizing group affect involves identifying important individuals and salient objects among a crowd that can evoke emotions. However, most existing methods lack attention to affective meaning in group dynamics and fail to account for the contextual relevance of faces and objects in group-level images. In this work, we propose a solution by incorporating the psychological concept of the Most Important Person (MIP), which represents the most noteworthy face in a crowd and has affective semantic meaning. We present the Dual-branch Cross-Patch Attention Transformer (DCAT) which uses global image and MIP together as inputs. Specifically, we first learn the informative facial regions produced by the MIP and the global context separately. Then, the Cross-Patch Attention module is proposed to fuse the features of MIP and global context together to complement each other. Our proposed method outperforms state-of-the-art methods on GAF 3.0, GroupEmoW, and HECO datasets. Moreover, we demonstrate the potential for broader applications by showing that our proposed model can be transferred to another group affect task, group cohesion, and achieve comparable results. Ming-Xian Lee, Tzu-Jui Chen, Hung-Jen Chen 0001, Hou-I Liu, Hong-Han Shuai, Wen-Huang Cheng |
ICCV | 6 |
| 2023 | Shilling Black-box Review-based Recommender Systems through Fake Review GenerationabstractReview-Based Recommender Systems (RBRS) have attracted increasing research interest due to their ability to alleviate well-known cold-start problems. RBRS utilizes reviews to construct the user and items representations. However, in this paper, we argue that such a reliance on reviews may instead expose systems to the risk of being shilled. To explore this possibility, in this paper, we propose the first generation-based model for shilling attacks against RBRSs. Specifically, we learn a fake review generator through reinforcement learning, which maliciously promotes items by forcing prediction shifts after adding generated reviews to the system. By introducing the auxiliary rewards to increase text fluency and diversity with the aid of pre-trained language models and aspect predictors, the generated reviews can be effective for shilling with high fidelity. Experimental results demonstrate that the proposed framework can successfully attack three different kinds of RBRSs on the Amazon corpus with three domains and Yelp corpus. Furthermore, human studies also show that the generated reviews are fluent and informative. Finally, equipped with Attack Review Generators (ARGs), RBRSs with adversarial training are much more robust to malicious reviews. Hung-Yun Chiang, Yi-Syuan Chen, Yun-Zhu Song, Hong-Han Shuai, Jason S. Chang |
KDD | 4 |
| 2023 | A Source You Prefer, or Majority? Investigating User Responses to Conflicting Opinions in Multi-Platform Restaurant-Review ListsabstractCustomers learn about restaurants in various ways, and integrating this disparate information could give them access to a greater diversity of perspectives. Conflicting opinions between restaurant-review platforms are inevitable. However, such conflicts’ influences on users’ perceptions remain unclear, especially when the opinion of a user’s preferred platform conflicts with the majority of others. This study’s experiment with a sample of 304 users found that, when such situations occurred, the preferred platform’s influence differed depending on whether the user was shown a sequence of whole-platform aggregations vs. a sequence of individual reviews drawn from multiple platforms. That is, the participants accepted the majority view most of the time, but when looking at aggregated lists, if their preferred platform expressed a minority positive opinion based on a high quantity of reviews, that minority opinion could prevail over the majority one. Between-platform conflicts were also found to have a greater impact on user reactions than within-platform ones did. Tzu-Hao Lin, Yen-Yun Liu, Hong-Han Shuai, Fang-Hsin Hsu, Yung-Ju Chang |
Int. J. Hum. Comput. Interact. | 3 |
| 2023 | Seeing the unseen: Wifi-based 2D human pose estimation via an evolving attentive spatial-Frequency network
Yi-Chung Chen, Zhi-Kai Huang, Lu Pang 0008, Jian-Yu Jiang-Lin, Chia-Han Kuo, Hong-Han Shuai, Wen-Huang Cheng |
Pattern Recognit. Lett. | 6 |
| 2023 | Decoupling-Cooperative Framework for Referring Expression ComprehensionabstractReferring Expression Comprehension (REC) aims to locate a specific object within an image by interpreting a referring expression articulated in natural language. This task comprises two essential branches: understanding and localizing. The former entails processing cognitive information from multimodal data, while the latter involves realizing the predictions in the perceptive visual space. Although various advanced approaches have been developed for each of these branches separately, existing REC approaches are unable to effectively leverage them due to the specific designs of architectures or objectives for REC, which bind understanding and localizing inseparably. To overcome this challenge, we propose the Decoupling-Cooperative Framework (DCF). The decoupling scheme in DCF enables us to utilize up-to-date methods for understanding and localizing with minimal constraints. Meanwhile, the proposed cooperative modules enable better integration of the strengths from both branches to achieve further enhancements. Extensive experiments demonstrate that DCF achieves state-of-the-art performance across four benchmarks, thus highlighting the generalizability of DCF. Yun-Zhu Song, Yi-Syuan Chen, Hong-Han Shuai |
IEEE Signal Process. Lett. | 3 |
| 2023 | General then Personal: Decoupling and Pre-training for Personalized Headline GenerationabstractAbstract Personalized Headline Generation aims to generate unique headlines tailored to users’ browsing history. In this task, understanding user preferences from click history and incorporating them into headline generation pose challenges. Existing approaches typically rely on predefined styles as control codes, but personal style lacks explicit definition or enumeration, making it difficult to leverage traditional techniques. To tackle these challenges, we propose General Then Personal (GTP), a novel framework comprising user modeling, headline generation, and customization. We train the framework using tailored designs that emphasize two central ideas: (a) task decoupling and (b) model pre-training. With the decoupling mechanism separating the task into generation and customization, two mechanisms, i.e., information self-boosting and mask user modeling, are further introduced to facilitate the training and text control. Additionally, we introduce a new evaluation metric to address existing limitations. Extensive experiments conducted on the PENS dataset, considering both zero-shot and few-shot scenarios, demonstrate that GTP outperforms state-of-the-art methods. Furthermore, ablation studies and analysis emphasize the significance of decoupling and pre-training. Finally, the human evaluation validates the effectiveness of our approaches.1 Yun-Zhu Song, Yi-Syuan Chen, Lu Wang 0008, Hong-Han Shuai |
Trans. Assoc. Comput. Linguistics | 4 |
| 2023 | An Overview of Facial Micro-Expression Analysis: Data, Methodology and ChallengeabstractFacial micro-expressions indicate brief and subtle facial movements that appear during emotional communication. In comparison to macro-expressions, micro-expressions are more challenging to be analyzed due to the short span of time and the fine-grained changes. In recent years, micro-expression recognition (MER) has drawn much attention because it can benefit a wide range of applications, e.g. police interrogation, clinical diagnosis, depression analysis, and business negotiation. In this survey, we offer a fresh overview to discuss new research directions and challenges these days for MER tasks. For example, we review MER approaches from three novel aspects: macro-to-micro adaptation, recognition based on key apex frames, and recognition based on facial action units. Moreover, to mitigate the problem of limited and biased ME data, synthetic data generation is surveyed for the diversity enrichment of micro-expression data. Since micro-expression spotting can boost micro-expression analysis, the state-of-the-art spotting works are also introduced in this paper. At last, we discuss the challenges in MER research and provide potential solutions as well as possible directions for further investigation. Ling Lo, Hong-Han Shuai, Wen-Huang Cheng |
IEEE Trans. Affect. Comput. | 3 |
| 2023 | SPEC: Summary Preference Decomposition for Low-Resource Abstractive SummarizationabstractNeural abstractive summarization has been widely studied and achieved great success with large-scale corpora. However, the considerable cost of annotating data motivates the need for learning strategies under low-resource settings. In this paper, we investigate the problems of learning summarizers with only few examples and propose corresponding methods for improvements. First, typical transfer learning methods are prone to be affected by data properties and learning objectives in the pretext tasks. Therefore, based on pretrained language models, we further present a meta learning framework to transfer few-shot learning processes from source corpora to the target corpus. Second, previous methods learn from training examples without decomposing thecontentandpreference. The generated summaries could therefore be constrained by the preference bias in the training set, especially under low-resource settings. As such, we propose decomposing the contents and preferences during learning through the parameter modulation, which enables control over preferences during inference. Third, given a target application, specifying required preferences could be non-trivial because the preferences may be difficult to derive through observations. Therefore, we propose a novel decoding method to automatically estimate suitable preferences and generate corresponding summary candidates from the few training examples. Extensive experiments demonstrate that our methods achieve state-of-the-art performance on six diverse corpora with 30.11%/33.95%/27.51% and 26.74%/31.14%/24.48% average improvements on ROUGE-1/2/L under 10- and 100-example settings. Yi-Syuan Chen, Yun-Zhu Song, Hong-Han Shuai |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Referring Expression Comprehension Via Enhanced Cross-modal Graph Attention NetworksabstractReferring expression comprehension aims to localize a specific object in an image according to a given language description. It is still challenging to comprehend and mitigate the gap between various types of information in the visual and textual domains. Generally, it needs to extract the salient features from a given expression and match the features of expression to an image. One challenge in referring expression comprehension is the number of region proposals generated by object detection methods is far more than the number of entities in the corresponding language description. Remarkably, the candidate regions without described by the expression will bring a severe impact on referring expression comprehension. To tackle this problem, we first propose a novel Enhanced Cross-modal Graph Attention Networks (ECMGANs) that boosts the matching between the expression and the entity position of an image. Then, an effective strategy named Graph Node Erase (GNE) is proposed to assist ECMGANs in eliminating the effect of irrelevant objects on the target object. Experiments on three public referring expression comprehension datasets show unambiguously that our ECMGANs framework achieves better performance than other state-of-the-art methods. Moreover, GNE is able to obtain higher accuracies of visual-expression matching effectively. Jia Wang 0020, Jingcheng Ke, Hong-Han Shuai, Yung-Hui Li, Wen-Huang Cheng |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | ShuttleNet: Position-Aware Fusion of Rally Progress and Player Styles for Stroke Forecasting in BadmintonabstractThe increasing demand for analyzing the insights in sports has stimulated a line of productive studies from a variety of perspectives, e.g., health state monitoring, outcome prediction. In this paper, we focus on objectively judging what and where to return strokes, which is still unexplored in turn-based sports. By formulating stroke forecasting as a sequence prediction task, existing works can tackle the problem but fail to model information based on the characteristics of badminton. To address these limitations, we propose a novel Position-aware Fusion of Rally Progress and Player Styles framework (ShuttleNet) that incorporates rally progress and information of the players by two modified encoder-decoder extractors. Moreover, we design a fusion network to integrate rally contexts and contexts of the players by conditioning on information dependency and different positions. Extensive experiments on the badminton dataset demonstrate that ShuttleNet significantly outperforms the state-of-the-art methods and also empirically validates the feasibility of each component in ShuttleNet. On top of that, we provide an analysis scenario for the stroke forecasting problem. Wei-Yao Wang, Hong-Han Shuai, Kai-Shiang Chang, Wen-Chih Peng |
AAAI | 2 |
| 2022 | Social-SSL: Self-supervised Cross-Sequence Representation Learning Based on Transformers for Multi-agent Trajectory Prediction
Li-Wu Tsao, Yan-Kai Wang, Hao-Siang Lin, Hong-Han Shuai, Lai-Kuan Wong, Wen-Huang Cheng |
ECCV (22) | 4 |
| 2022 | Residual Graph Attention Network and Expression-Respect Data Augmentation Aided Visual GroundingabstractVisual grounding aims to localize a target object in an image based on a given text description. Due to the innate complexity of language, it is still a challenging problem to perform reasoning of complex expressions and to infer the underlying relationship between the expression and the object in an image. To address these issues, we propose a residual graph attention network for visual grounding. The proposed approach first builds an expression-guided relation graph and then performs multi-step reasoning followed by matching the target object. It allows performing better visual grounding with complex expressions by using deeper layers than other graph network approaches. Moreover, to increase the diversity of training data, we perform an expression-respect data augmentation based on copy-paste operations to pairs of source and target images. The proposed approach achieves better performance with extensive experiments than other state-of-the-art graph network-based approaches and demonstrates its effectiveness. Jia Wang 0020, Hung-Yi Wu, Jun-Cheng Chen, Hong-Han Shuai, Wen-Huang Cheng |
ICIP | 4 |
| 2022 | Finding the Achilles Heel: Progressive Identification Network for Camouflaged Object DetectionabstractCamouflaged object detection (COD) aims to segment objects assimilating into their surroundings. The key challenge for COD is that there are existing high intrinsic similarities between the target object and the background. To solve this challenging problem, we propose the Cascaded Decamouflage Module to progressively improve the prediction map, where each decamouflage module is composed of the region enhancement block and the reverse attention mining block to accurately detect the camouflaged object and obtain complete target objects. In addition, we introduce the classification-based label reweighting to produce the gated label maps as the supervision for assisting the network to capture the most conspicuous region of a camouflaged object and obtain the target object entirely. Extensive experiments on three challenging datasets demonstrate that the proposed model outperforms state-of-the-art methods under different evaluation metrics. Mu-Chun Chou, Hung-Jen Chen 0001, Hong-Han Shuai |
ICME | 3 |
| 2022 | The Hierarchical Ensemble Model for Network Intrusion Detection in the Real-world DatasetabstractNetwork intrusion detection is an indispensable defense in the critical era fulling of cyberattacks. However, it faces a severe class imbalanced issue, and most of the researches are conducted on simulated data. Therefore, this work introduces a hierarchical ensemble architecture with machine learning approaches. It is trained on the latest and real-world dataset to solve the above problems. The experiments show that we outperform state-of-the-art methods on real network traffic data. Shao-En Weng, Chu-Jun Peng, Yin-Chi Li, Hong-Han Shuai, Wen-Huang Cheng |
ISCAS | 5 |
| 2022 | Efficiency-reinforced Learning with Auxiliary Depth Reconstruction for Autonomous Navigation of Mobile DevicesabstractIn this paper, we take Unmanned Aerial Vehicles (UAVs) as the mobile devices to study the problem of autonomous navigation since UAVs have been adopted as intelligent vehicles for executing complex tasks such as bridge structure examination, crowd estimation, target searching, and package delivery. As Deep Reinforcement Learning (DRL) has achieved great success in many control tasks, it is envisaged to exploit DRL for autonomous navigation. Nevertheless, as the navigation path becomes distant, searching in a large number of states and action spaces becomes very challenging to DRL. In this paper, we provide a novel reinforcement learning framework to facilitate the autonomous navigation in complicated environments by jointly considering the temporal abstractions and policy efficiency to dynamically select the frequency of the action decisions with the efficiency regularization. Moreover, to bootstrap the learning procedure, we further add an auxiliary task of depth map reconstruction to accelerate the learning process. Experimental results on 3D UAV simulator and DeepMind Lab environments manifest that the proposed framework improves the state-of-the-art methods in terms of success rates in different environments. Cheng-Chun Li, Hong-Han Shuai, Li-Chun Wang 0004 |
MDM | 2 |
| 2022 | Towards Understanding Cross Resolution Feature Matching for Surveillance Face RecognitionabstractCross-resolution face recognition (CRFR) in an open-set setting is a practical application for surveillance scenarios where low-resolution (LR) probe faces captured via surveillance cameras require being matched to a watchlist of high-resolution (HR) galleries. Although CRFR is to be of practical use, it sees a performance drop of more than 10% compared to that of high-resolution face recognition protocols. The challenges of CRFR are multifold, including the domain gap induced by the HR and LR images, the pose/texture variations, etc. To this end, this work systematically discusses possible issues and their solutions that affect the accuracy of CRFR. First, we explore the effect of resolution changes and conclude that resolution matching is the key for CRFR. Even simply downscaling the HR faces to match the LR ones brings a performance gain. Next, to further boost the accuracy of matching cross-resolution faces, we found that a well-designed super-resolution network, which can (a) represent the images continuously, is (b) suitable for real-world degradation kernel, (c) adaptive to different input resolutions, and (d) guided by an identity-preserved loss, is necessary to upsample the LR faces with discriminative enhancement. Here, the proposed identity-preserved loss plays the role of reconciling the objective discrepancy of super-resolution between human perception and machine recognition. Finally, we emphasize that removing the pose variations is an essential step before matching faces for recognition in the super-resolved feature space. Our method is evaluated on benchmark datasets, including SCface, cross-resolution LFW, and QMUL-Tinyface. The results show that the proposed method outperforms the SOTA methods by a clear margin and narrows the performance gap compared to the high-resolution face recognition protocol. Chiawei Kuo, Yi-Ting Tsai, Hong-Han Shuai, Yi-Ren Yeh |
ACM Multimedia | 3 |
| 2022 | Mimicking the Annotation Process for Recognizing the Micro ExpressionsabstractMicro-expression recognition (MER) has recently become a popular research topic due to its wide applications, e.g., movie rating and recognizing the neurological disorder. By virtue of deep learning techniques, the performance of MER has been significantly improved and reached unprecedented results. This paper proposes a novel architecture to mimic how the expressions are annotated. Specifically, during the annotation process in several datasets, the AU labels are first obtained with FACS, and the expression labels are then decided based on the combinations of the AU labels. Meanwhile, these AU labels describe either the eyes or mouth movements (mutually-exclusive). Following this idea, we design a dual-branch structure with a new augmentation method to separately capture the eyes and mouth features and teach the model what the general expressions should be. Moreover, to adaptively fuse the area features for different expressions, we propose Area Weighted Module to assign different weights to each region. Additionally, we set up an auxiliary task to align the AU similarity scores to help our model capture facial patterns further with AU labels. The proposed approach outperforms other state-of-the-art methods in terms of accuracy on the CASME II and SAMM datasets. Moreover, we provide a new visualization approach to show the relationship between the facial regions and AU features. Bo-Kai Ruan, Ling Lo, Hong-Han Shuai, Wen-Huang Cheng |
ACM Multimedia | 3 |
| 2022 | Improving Multi-Document Summarization through Referenced Flexible Extraction with Credit-AwarenessabstractA notable challenge in Multi-Document Summarization (MDS) is the extremely-long length of the input.In this paper, we present an extract-then-abstract Transformer framework to overcome the problem.Specifically, we leverage pre-trained language models to construct a hierarchical extractor for salient sentence selection across documents and an abstractor for rewriting the selected contents as summaries.However, learning such a framework is challenging since the optimal contents for the abstractor are generally unknown.Previous works typically create pseudo extraction oracle to enable the supervised learning for both the extractor and the abstractor.Nevertheless, we argue that the performance of such methods could be restricted due to the insufficient information for prediction and inconsistent objectives between training and testing.To this end, we propose a loss weighting mechanism that makes the model aware of the unequal importance for the sentences not in the pseudo extraction oracle, and leverage the fine-tuned abstractor to generate summary references as auxiliary signals for learning the extractor.Moreover, we propose a reinforcement learning method that can efficiently apply to the extractor for harmonizing the optimization between training and testing.Experiment results show that our framework substantially outperforms strong baselines with comparable model sizes and achieves the best results on the Multi-News, Multi-XScience, and WikiCatSum corpora. 1 Yun-Zhu Song, Yi-Syuan Chen, Hong-Han Shuai |
NAACL-HLT | 3 |
| 2022 | Improving Entity Disambiguation Using Knowledge Graph Regularization
Zhi Rui Tam, Yi-Lun Wu, Hong-Han Shuai |
PAKDD (1) | 3 |
| 2022 | On Extracting Socially Tenuous Groups for Online Social Networks With $k$k-TrianglesabstractExisting research on finding social groups mostly focuses on dense subgraphs in social networks. However, finding socially tenuous groups also has many important applications. In this paper, we introduce the notion of k-triangles to measure the tenuity of a group. We then formulate a new research problem, Minimum k-Triangle Disconnected Group with No-Pair Constraint (MkTG), to find a socially tenuous group from the online social network. We prove that MkTG is NP-hard and inapproximable within any ratio. Two algorithms, namely TERA and TERA-ADV, are designed for solving MkTG effectively and efficiently. Further, we examine the MkTG problem on tree-based social networks, due to their structural resemblance with corporate social networks built upon the supervision relation. Accordingly, we devise an efficient algorithm, namely Tenuity Maximization for Trees (TMT), to obtain the optimal solution in polynomial time. In addition, we study a more general version of MkTG, named Generalized Minimum k-Triangle Disconnected Group without No-Pair Constraint (MkTG-G). We formulate MkTG-G, analyze its inapproximability, and propose a randomized approximation algorithm, named Randomized Ranking with Limited Neighborhood Participation (RLNP). Experimental results on real datasets manifest that the proposed algorithms outperform the baselines in terms of both efficiency and solution quality. Hong-Han Shuai, De-Nian Yang, Guang-Siang Lee, Liang-Hao Huang, Wang-Chien Lee, Ming-Syan Chen |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2022 | Learning to Solve Task-Optimized Group Search for Social Internet of ThingsabstractWith the maturity and popularity of Internet of Things (IoT), the notion of Social Internet of Things (SIoT) has been proposed to support novel applications and networking services for the IoT in more effective and efficient ways. Although there are many works for SIoT, they focus on designing the architectures and protocols for SIoT under the specific schemes. How to efficiently utilize the collaboration capability of SIoT to complete complex tasks remains unexplored. Therefore, we propose a new problem family, namely,Task-Optimized SIoT Selection (TOSS), to find the best group of IoT objects for a given set of tasks in the task pool. TOSS aims to select the target SIoT group such that the target SIoT group is able to easily communicate with each other while maximizing the accuracy of performing the given tasks. We propose two problem formulations, namedBounded Communication-loss TOSS (BC-TOSS)andRobustness Guaranteed TOSS (RG-TOSS), for different scenarios and prove that they are both NP-hard and inapproximable. We propose a polynomial-time algorithm with a performance guarantee for BC-TOSS, and an efficient polynomial-time algorithm to obtain good solutions for RG-TOSS. Moreover, as RG-TOSS is NP-hard and inapproximable within any factor, we further proposeStructure-Aware Reinforcement Learning (SARL)to leverage the Graph Convolutional Networks (GCN) and Deep Reinforcement Learning (DRL) to effectively solve RG-TOSS. Further, since we use graph models to simulate the problem instance for DRL, which is different from the real ones, we proposeStructure-Aware Meta Reinforcement Learning (SAMRL)for fast adapting to new domains. Experimental results on multiple real datasets indicate that our proposed algorithms outperform the other deterministic and learning-based baseline approaches. Chen-Hsu Yang, Hong-Han Shuai, Ming-Syan Chen |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2022 | Facial Chirality: From Visual Self-Reflection to Robust Facial Feature LearningabstractAs a fundamental vision task, facial expression recognition has made substantial progress recently. However, the recognition performance often degrades significantly in real-world scenarios due to the lack of robust facial features. In this paper, we propose an effective facial feature learning method that takes the advantage of facial chirality to discover the discriminative features for facial expression recognition. Most previous studies implicitly assume that human faces are symmetric. However, our work reveals that the facial asymmetric effect can be a crucial clue. Given a face image and its reflection without additional labels, we decouple the emotion-invariant facial features from the input image pair to better capture the emotion-related facial features. Moreover, as our model aligns emotion-related features of the image pair to enhance the recognition performance, the value of precise facial landmark alignment as a pre-processing step is reconsidered in this paper. Experiments demonstrate that the learned emotion-related features outperform the state of the art methods on several facial expression recognition benchmarks as well as real-world occlusion datasets, which manifests the effectiveness and robustness of the proposed model. Ling Lo, Hong-Han Shuai, Wen-Huang Cheng |
IEEE Trans. Multim. | 3 |
| 2022 | Spatiotemporal Dilated Convolution With Uncertain Matching for Video-Based Crowd EstimationabstractIn this paper, we propose a novel SpatioTemporal convolutional Dense Network (STDNet) to address the video-based crowd counting problem, which contains the decomposition of 3D convolution and the 3D spatiotemporal dilated dense convolution to alleviate the rapid growth of the model size caused by the Conv3D layer. Moreover, since the dilated convolution extracts the multiscale features, we combine the dilated convolution with the channel attention block to enhance the feature representations. Due to the error that occurs from the difficulty of labeling crowds, especially for videos, imprecise or standard-inconsistent labels may lead to poor convergence for the model. To address this issue, we further propose a new patch-wise regression loss (PRL) to improve the original pixel-wise loss. Experimental results on three video-based benchmarks, i.e., the UCSD, Mall and WorldExpo’10 datasets, show that STDNet outperforms both image- and video-based state-of-the-art methods. The source codes are released athttps://github.com/STDNet/STDNet. Yu-Jen Ma, Hong-Han Shuai, Wen-Huang Cheng |
IEEE Trans. Multim. | 2 |
| 2022 | Template-Free Try-On Image Synthesis via Semantic-Guided OptimizationabstractThe virtual try-on task is so attractive that it has drawn considerable attention in the field of computer vision. However, presenting the 3-D physical characteristic (e.g., pleat and shadow) based on a 2-D image is very challenging. Although there have been several previous studies on 2-D-based virtual try-on work, most: 1) required user-specified target poses that are not user-friendly and may not be the best for the target clothing and 2) failed to address some problematic cases, including facial details, clothing wrinkles, and body occlusions. To address these two challenges, in this article, we propose an innovative template-free try-on image synthesis (TF-TIS) network. The TF-TIS first synthesizes the target pose according to the user-specified in-shop clothing. Afterward, given an in-shop clothing image, a user image, and a synthesized pose, we propose a novel model for synthesizing a human try-on image with the target clothing in the best fitting pose. The qualitative and quantitative experiments both indicate that the proposed TF-TIS outperforms the state-of-the-art methods, especially for difficult cases. Chien-Lung Chou, Chieh-Yun Chen, Chia-Wei Hsieh, Hong-Han Shuai, Jiaying Liu 0001, Wen-Huang Cheng |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Improving Crowd Density Estimation by Fusing Aerial Images and Radio SignalsabstractA recent line of research focuses on crowd density estimation from RGB images for a variety of applications, for example, surveillance and traffic flow control. The performance drops dramatically for low-quality images, such as occlusion, or poor light conditions. However, people are equipped with various wireless devices, allowing the received signals to be easily collected at the base station. As such, another line of research utilizes received signals for crowd counting. Nevertheless, received signals offer only information regarding the number of people, while an accurate density map cannot be derived. As unmanned aerial vehicles (UAVs) are now treated as flying base stations and equipped with cameras, we make the first attempt to leverage both RGB images and received signals for crowd density estimation on UAVs. Specifically, we propose a novel network to effectively fuse the RGB images and received signal strength (RSS) information. Moreover, we design a new loss function that considers the uncertainty from RSS and makes the prediction consistent with the received signals. Experimental results show that the proposed method successfully helps break the limit of traditional crowd density estimation methods and achieves state-of-the-art performance. The proposed dataset is released as a public download for future research. Kai-Wei Yang, Yen-Yun Huang, Jen-Wei Huang, Ya-Rou Hsu, Chang-Lin Wan, Hong-Han Shuai, Li-Chun Wang 0001, Wen-Huang Cheng |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2022 | Mask or Non-Mask? Robust Face Mask Detector via Triplet-Consistency Representation LearningabstractIn the absence of vaccines or medicines to stop COVID-19, one of the effective methods to slow the spread of the coronavirus and reduce the overloading of healthcare is to wear a face mask. Nevertheless, to mandate the use of face masks or coverings in public areas, additional human resources are required, which is tedious and attention-intensive. To automate the monitoring process, one of the promising solutions is to leverage existing object detection models to detect the faces with or without masks. As such, security officers do not have to stare at the monitoring devices or crowds, and only have to deal with the alerts triggered by the detection of faces without masks. Existing object detection models usually focus on designing the CNN-based network architectures for extracting discriminative features. However, the size of training datasets of face mask detection is small, while the difference between faces with and without masks is subtle. Therefore, in this article, we propose a face mask detection framework that uses the context attention module to enable the effective attention of the feed-forward convolution neural network by adapting their attention maps’ feature refinement. Moreover, we further propose an anchor-free detector with Triplet-Consistency Representation Learning by integrating the consistency loss and the triplet loss to deal with the small-scale training data and the similarity between masks and occlusions. Extensive experimental results show that our method outperforms the other state-of-the-art methods. The source code is released as a public download to improve public health at https://github.com/wei-1006/MaskFaceDetection . Chun-Wei Yang, Thanh Hai Phung, Hong-Han Shuai, Wen-Huang Cheng |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2021 | Meta-Transfer Learning for Low-Resource Abstractive SummarizationabstractNeural abstractive summarization has been studied in many pieces of literature and achieves great success with the aid of large corpora. However, when encountering novel tasks, one may not always benefit from transfer learning due to the domain shifting problem, and overfitting could happen without adequate labeled examples. Furthermore, the annotations of abstractive summarization are costly, which often demand domain knowledge to ensure the ground-truth quality. Thus, there are growing appeals for Low-Resource Abstractive Summarization, which aims to leverage past experience to improve the performance with limited labeled examples of target corpus. In this paper, we propose to utilize two knowledge-rich sources to tackle this problem, which are large pre-trained models and diverse existing corpora. The former can provide the primary ability to tackle summarization tasks; the latter can help discover common syntactic or semantic information to improve the generalization ability. We conduct extensive experiments on various summarization corpora with different writing styles and forms. The results demonstrate that our approach achieves the state-of-the-art on 6 corpora in low-resource scenarios, with only 0.7% of trainable parameters compared to previous work. Yi-Syuan Chen, Hong-Han Shuai |
AAAI | 2 |
| 2021 | Explainable Health State Prediction for Social IoTs through Multi-Channel AttentionabstractThe core technology of Industry 4.0 is to enable the intelligence of manufacturing. One of the important tasks is anomaly detection. Although existing anomaly detection methods have achieved high accuracy, the basis of judgments cannot provide explainability, which greatly reduces the possibility for improving the model or facilitating human-machine cooperation. Therefore, in this paper, the goal is to provide the explainability for machine fault detection for social IoTs and realize the health monitoring and prognosis of the bearings simultaneously. Specifically, vibration signals from multiple sensors are transformed into spectrograms by short-time Fourier transform. Afterward, the features of frequency-domain data are extracted by the Squeeze-and-Excitation block and self-attention mechanism to assess the degradation of whole system. As such, when the process enters the early degradation, the source of components that causes the abnormality can be identified through the attention weight distribution. Experimental results show that the proposed approach achieves high accuracy in run-to-failure tests. Moreover, the proposed approach shows a better ability to explain the predicted results than the state-of-the-art bearing detection methods. Yu-Li Chan, Hong-Han Shuai |
GLOBECOM | 2 |
| 2021 | FashionMirror: Co-attention Feature-remapping Virtual Try-on with Sequential Template PosesabstractVirtual try-on tasks have drawn increased attention. Prior arts focus on tackling this task via warping clothes and fusing the information at the pixel level with the help of semantic segmentation. However, conducting semantic segmentation is time-consuming and easily causes error accumulation over time. Besides, warping the information at the pixel level instead of the feature level limits the performance (e.g., unable to generate different views) and is unstable since it directly demonstrates the results even with a misalignment. In contrast, fusing information at the feature level can be further refined by the convolution to obtain the final results. Based on these assumptions, we propose a co-attention feature-remapping framework, namely FashionMirror, that generates the try-on results according to the driven-pose sequence in two stages. In the first stage, we consider the source human image and the target try-on clothes to predict the removed mask and the try-on clothing mask, which replaces the pre-processed semantic segmentation and reduces the inference time. In the second stage, we first remove the clothes on the source human via the removed mask and warp the clothing features conditioning on the try-on clothing mask to fit the next frame human. Meanwhile, we predict the optical flows from the consecutive 2D poses and warp the source human to the next frame at the feature level. Then, we enhance the clothing features and source human features in every frame to generate realistic try-on results with spatiotemporal smoothness. Both qualitative and quantitative results show that FashionMirror outperforms the state-of-the-art virtual try-on approaches. Chieh-Yun Chen, Ling Lo, Pin-Jui Huang, Hong-Han Shuai, Wen-Huang Cheng |
ICCV | 4 |
| 2021 | Gradient Normalization for Generative Adversarial NetworksabstractIn this paper, we propose a novel normalization method called gradient normalization (GN) to tackle the training instability of Generative Adversarial Networks (GANs) caused by the sharp gradient space. Unlike existing work such as gradient penalty and spectral normalization, the proposed GN only imposes a hard 1-Lipschitz constraint on the discriminator function, which increases the capacity of the discriminator. Moreover, the proposed gradient normalization can be applied to different GAN architectures with little modification. Extensive experiments on four datasets show that GANs trained with gradient normalization outperform existing methods in terms of both Frechet Inception Distance and Inception Score. Yi-Lun Wu, Hong-Han Shuai, Zhi Rui Tam, Hong-Yu Chiu |
ICCV | 2 |
| 2021 | Attack as the Best Defense: Nullifying Image-to-image Translation GANs via Limit-aware Adversarial AttackabstractDue to the great success of image-to-image (Img2Img) translation GANs, many applications with ethics issues arise, e.g., DeepFake and DeepNude, presenting a challenging problem to prevent the misuse of these techniques. In this work, we tackle the problem by a new adversarial attack scheme, namely the Nullifying Attack, which cancels the image translation process and proposes a corresponding framework, the Limit-Aware Self-Guiding Gradient Sliding Attack (LaS-GSA) under a black-box setting. In other words, by processing the image with the proposed LaS-GSA before publishing, any image translation functions can be nullified, which prevents the images from malicious manipulations. First, we introduce the limit-aware RGF and the gradient sliding mechanism to estimate the gradient that adheres to the adversarial limit, i.e., the pixel value limitations of the adversarial example. We theoretically prove that our model is able to avoid the error caused by the projection in both the direction and the length. Then, an effective self-guiding prior is extracted solely from the threat model and the target image to efficiently leverage the prior information and guide the gradient estimation process. Extensive experiments demonstrate that LaS-GSA requires fewer queries to nullify the image translation process with higher success rates than 4 state-of-the-art methods. Chin-Yuan Yeh, Hsi-Wen Chen, Hong-Han Shuai, De-Nian Yang, Ming-Syan Chen |
ICCV | 3 |
| 2021 | Structure-Aware Parameter-Free Group Query via Heterogeneous Information Network TransformerabstractOwing to a wide range of important applications, such as team formation, dense subgraph discovery, and activity attendee suggestions on online social networks, Group Query attracts a lot of attention from the research community. However, most existing works are constrained by a unified social tightness k (e.g., for k-core, or k-plex), without considering the diverse preferences of social cohesiveness in individuals. In this paper, we introduce a new group query, namely Parameter-free Group Query (PGQ), and propose a learning-based model, called PGQN, to find a group that accommodates personalized requirements on social contexts and activity topics. First, PGQN extracts node features by a GNN-based method on Heterogeneous Activity Information Network (HAIN). Then, we transform the PGQ into a graph-to-set (Graph2Set) problem to learn the diverse user preference on topics and members, and find new attendees to the group. Experimental results manifest that our proposed model outperforms nine state-of-the-art methods by at least 51% in terms of F1-score on three public datasets. Hsi-Wen Chen, Hong-Han Shuai, De-Nian Yang, Wang-Chien Lee, Chuan Shi 0001, Philip S. Yu, Ming-Syan Chen |
ICDE | 2 |
| 2021 | Facial Chirality: Using Self-Face Reflection to Learn Discriminative Features for Facial Expression RecognitionabstractAs a fundamental vision task, facial expression recognition has made substantial progress recently. However, the recognition performance often degrades largely in real-world scenarios due to the lack of robust facial features. In this paper, we propose a simple but effective facial feature learning method that takes the advantage of facial chirality to discover the discriminative features for facial expression recognition. Most previous studies implicitly assume that human faces are symmetric. However, our work reveals that the facial asymmetric effect can be a crucial clue. Given a face image and its reflection without additional labels, we decouple the reflection-invariant facial features from the input image pair and then demonstrate that the new features with a standard and lightweight learning model (e.g. ResNet-18) are sufficiently robust to outperform the state-of-the-art methods (e.g. SCN in CVPR 2020 and ESRs in AAAI 2020). Our experiments also show the potential of the new features for other facial vision tasks such as expression image retrieval. Ling Lo, Hong-Han Shuai, Wen-Huang Cheng |
ICME | 3 |
| 2021 | Heterogeneous Federated Learning Through Multi-Branch NetworkabstractRecently, federated learning has gained increasing attention for privacy-preserving computation since the learning paradigm allows to train models without the need for exchanging the data across different institutions distributively. However, heterogeneity of computational capabilities of edge devices is seldom discussed and analyzed in the current literature for heterogeneous federated learning. To address this issue, we propose a novel heterogeneous federated learning framework based on multi-branch deep neural network models which enable the selection of a proper sub-branch model for the client devices according to their computational capabilities. Meanwhile, we also present an aggregation method for model training, MFedAvg, that performs branch-wise averaging-based aggregation. With extensive experiments on MNIST, FashionMNIST, MedMNIST, and CIFAR-10, it demonstrates that our proposed approaches can achieve satisfactory performance with guaranteed convergence and effectively utilize all the available resources for training across different devices with lower communication cost than its homogeneous counterpart. Ching-Hao Wang, Kang-Yang Huang, Jun-Cheng Chen, Hong-Han Shuai, Wen-Huang Cheng |
ICME | 4 |
| 2021 | Self-Attentive Recommendation for Multi-Source Review PackageabstractWith the diversified sources satisfying users' needs, many online service platforms collect information from multiple sources in order to provide a set of useful information to the users. However, existing recommendation systems are mostly designed for single-source data, and thus fail to recommend multi-source review packages since the interplay between the reviews of different sources is not properly modeled. In fact, modeling the interplay between different sources is challenging because 1) two reviews may conflict with each other, 2) different users have different preferences on review sources, and 3) users' preferences to each source may change under different scenarios. To address these challenges, we propose Self-Attentive Recommendation for multi-source review Package (SARP), for predicting how useful the user feels to the package, while simultaneously reflecting how much the user is affected by each review. Specifically, SARP jointly considers the relationships of every user, purpose, and review source to learn better latent representations. A self-attention module is further used for integrating source representations and the review ratings, following a multi-layer perceptron (MLP) for the prediction tasks. Experimental results on the self-constructed dataset and public dataset demonstrate that the proposed model outperforms the state-of-the-art approaches. Yu-Hsiu Chen, Hong-Han Shuai, Yung-Ju Chang |
IJCNN | 3 |
| 2021 | Re-Attention Is All You Need: Memory-Efficient Scene Text Detection via Re-Attention on Uncertain RegionsabstractScene text detection plays an important role on vision-based robot navigation to many potential landmarks such as nameplates, information signs, floor button in the elevators. Recently, scene text detection with segmentation-based methods has been receiving more and more attention. The segmentation results can be used to efficiently predict scene text of various shapes, such as irregular text in most scene text images. However, two kinds of texts remain unsolved: 1) tiny and 2) blurry instances. Moreover, the annotations for tiny/blurry texts are usually ignored during training, while tiny/blurry texts can still offer visual auxiliaries for robots to understand the world. Therefore, in this paper, we propose a new approach to effectively detect both clear and blurry texts. Specifically, we propose a re-attention module without increasing the learnable parameters, which first predicts the region of texts as the candidate region and leverages the same network to detect the candidate region again for reducing the required memory. Moreover, to avoid the errors from the first detection propagating to the re-attended area, we propose a new fusion module that learns to integrate the results of the re-attended regions and the first prediction. Experimental results manifest that the proposed method outperforms state-of-the-art methods on four challenging datasets. Hsiang-Chun Chang, Hung-Jen Chen 0001, Yu-Chia Shen, Hong-Han Shuai, Wen-Huang Cheng |
IROS | 4 |
| 2021 | Reproducibility Companion Paper: Knowledge Enhanced Neural Fashion Trend ForecastingabstractThis companion paper supports the replication of the fashion trend forecasting experiments with the KERN (Knowledge Enhanced Recurrent Network) method that we presented in the ICMR 2020. We provide an artifact that allows the replication of the experiments using a Python implementation. The artifact is easy to deploy with simple installation, training and evaluation. We reproduce the experiments conducted in the original paper and obtain similar performance as previously reported. The replication results of the experiments support the main claims in the original paper. Yunshan Ma 0002, Yujuan Ding, Xun Yang 0001, Lizi Liao, Wai Keung Wong, Tat-Seng Chua, Jinyoung Moon, Hong-Han Shuai |
ICMR | 8 |
| 2021 | Face-based Voice Conversion: Learning the Voice behind a FaceabstractZero-shot voice conversion (VC) trained by non-parallel data has gained a lot of attention in recent years. Previous methods usually extract speaker embeddings from audios and use them for converting the voices into different voice styles. Since there is a strong relationship between human faces and voices, a promising approach would be to synthesize various voice characteristics from face representation. Therefore, we introduce a novel idea of generating different voice styles from different human face photos, which can facilitate new applications, e.g., personalized voice assistants. However, the audio-visual relationship is implicit. Moreover, the existing VCs are trained on laboratory-collected datasets without speaker photos, while the datasets with both photos and audios are in-the-wild datasets. Directly replacing the target audio with the target photo and training on the in-the-wild dataset leads to noisy results. To address these issues, we propose a novel many-to-many voice conversion network, namely Face-based Voice Conversion (FaceVC), with a 3-stage training strategy. Quantitative and qualitative experiments on the LRS3-Ted dataset show that the proposed FaceVC successfully performs voice conversion according to the target face photos. Audio samples can be found on the demo website at https://facevc.github.io/. Hsiao-Han Lu, Shao-En Weng, Ya-Fan Yen, Hong-Han Shuai, Wen-Huang Cheng |
ACM Multimedia | 4 |
| 2021 | DensER: Density-imbalance-Eased Representation for LiDAR-based Whole Scene UpsamplingabstractWith the development of depth sensors, 3D point cloud upsampling that generates a high-resolution point cloud given a sparse input becomes emergent. However, many previous works focused on single 3D object reconstruction and refinement. Although a few recent works began to discuss 3D structure refine-ment for a more complex scene, they do not target LiDAR-based point clouds, which have density imbalance issues from near to far. This paper proposed DensER, a Density-imbalance-Eased regional Representation. Notably, to learn robust representations and model local geometry under imbalance point density, we designed density-aware multiple receptive fields to extract the regional features. Moreover, founded on the patch reoccurrence property of a nature scene, we proposed a density-aided attentive module to enrich the extracted features of point-sparse areas by referring to other non-local regions. Finally, by coupling with novel manifold-based upsamplers, DensER shows the ability to super-resolve LiDAR-based whole-scene point clouds. The exper-imental results show DensER outperforms related works both in qualitative and quantitative evaluation. We also demonstrate that the enhanced point clouds can improve downstream tasks such as 3D object detection and depth completion. Tso-Yuan Chen, Ching-Chun Hsiao, Wen-Huang Cheng, Hong-Han Shuai, Peter Chen |
VCIP | 4 |
| 2021 | ROSNet: Robust one-stage network for CT lesion detectionabstractAutomatic lesion detection from computed tomography (CT) scans is an important task in medical diagnosis. However, three frequent properties of medical data make CT lesion detection a challenging task: (1) Scale variance: Large scale variation is across lesion instances. Especially, it is extremely difficult to detect small lesions; (2) Imbalanced data: The data distributions are highly imbalanced, where few classes account for the majority of data; (3) Prediction stability: Based on our observations, an input lesion image with slightly pixel shift or translation can lead to drastic output mispredictions and this is not allowed for medical applications. To address these challenges, this paper proposes a Robust One-Stage Network (ROSNet) for robust CT lesion detection. Specifically, a novel nested structure of neural networks is developed to generate a series of feature pyramids for detecting CT lesions in various scales, an effective data sensitive class-balanced loss as well as a shift-invariant downsampling strategy are also introduced to improve the detection performance. Experiments are conducted on a large-scale and diverse dataset, DeepLesion, showing that ROSNet outperforms the best performance in MICCAI 2019 by 3.95% (2-class detection task) and 25.41% (8-class detection task) in terms of mean average precision (mAP). Kuan-Yu Lung, Chi-Rung Chang, Shao-En Weng, Hao-Siang Lin, Hong-Han Shuai, Wen-Huang Cheng |
Pattern Recognit. Lett. | 5 |
| 2021 | Introduction to the Special Issue on Explainable AI on Multimedia ComputingabstractNo abstract available. Wen-Huang Cheng, Jiaying Liu 0001, Nicu Sebe, Junsong Yuan 0001, Hong-Han Shuai |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2020 | TemPEST: Soft Template-Based Personalized EDM Subject Generation through Collaborative Summarization
Yu-Hsiu Chen, Hong-Han Shuai, Wen-Chih Peng |
AAAI | 3 |
| 2020 | Attractive or Faithful? Popularity-Reinforced Learning for Inspired Headline GenerationabstractWith the rapid proliferation of online media sources and published news, headlines have become increasingly important for attracting readers to news articles, since users may be overwhelmed with the massive information. In this paper, we generate inspired headlines that preserve the nature of news articles and catch the eye of the reader simultaneously. The task of inspired headline generation can be viewed as a specific form of Headline Generation (HG) task, with the emphasis on creating an attractive headline from a given news article. To generate inspired headlines, we propose a novel framework called POpularity-Reinforced Learning for inspired Headline Generation (PORL-HG). PORL-HG exploits the extractive-abstractive architecture with 1) Popular Topic Attention (PTA) for guiding the extractor to select the attractive sentence from the article and 2) a popularity predictor for guiding the abstractor to rewrite the attractive sentence. Moreover, since the sentence selection of the extractor is not differentiable, techniques of reinforcement learning (RL) are utilized to bridge the gap with rewards obtained from a popularity score predictor. Through quantitative and qualitative experiments, we show that the proposed PORL-HG significantly outperforms the state-of-the-art headline generation models in terms of attractiveness evaluated by both human (71.03%) and the predictor (at least 27.60%), while the faithfulness of PORL-HG is also comparable to the state-of-the-art generation model. Yun-Zhu Song, Hong-Han Shuai, Sung-Lin Yeh, Yi-Lun Wu, Lun-Wei Ku, Wen-Chih Peng |
AAAI | 2 |
| 2020 | Live Multi-Streaming and Donation Recommendations via Coupled Donation-Response Tensor FactorizationabstractIn contrast to traditional online videos, live multi-streaming supports real-time social interactions between multiple streamers and viewers, such as donations. However, donation and multi-streaming channel recommendations are challenging due to complicated streamer and viewer relations, asymmetric communications, and the tradeoff between personal interests and group interactions. In this paper, we introduce Multi-Stream Party (MSP) and formulate a new multi-streaming recommendation problem, called Donation and MSP Recommendation (DAMRec). We propose Multi-stream Party Recommender System (MARS) to extract latent features via socio-temporal coupled donation-response tensor factorization for donation and MSP recommendations. Experimental results on Twitch and Douyu manifest that MARS significantly outperforms existing recommenders by at least 38.8% in terms of hit ratio and mean average precision. Hsu-Chao Lai, Jui-Yi Tsai, Hong-Han Shuai, Jiun-Long Huang, Wang-Chien Lee, De-Nian Yang |
CIKM | 3 |
| 2020 | On Minimizing Diagonal Block-Wise Differences for Neural Network CompressionabstractDeep neural networks have achieved great success on a wide spectrum of applications. However, neural network (NN) mod- els often include a massive number of weights and consume much memory. To reduce the NN model size, we observe that the struc- ture of the weight matrix can be further re-organized for a better compression, i.e., converting the weight matrix to the block diag- onal structure. Therefore, in this paper, we formulate a new re- search problem to consider the structural factor of the weight ma- trix, named Compression with Difference-Minimized Block Diagonal Structure (COMIS), and propose a new algorithm, Memory-Efficient and Structure-Aware Compression (MESA), which effectively prunes the weights into a block diagonal structure to significantly boost the compression rate. Extensive experiments on different models show that MESA achieves 135× to 392× compression rates for different models, which are 1.8 to 3.03 times the compression rates of the state-of-the-art approaches. In addition, our approach provides an in- ference speed-up from 2.6× to 5.1×, a speed-up up to 44% to the state-of-the-art approaches. Yun-Jui Hsu, Yi-Ting Chang, Hong-Han Shuai, Wei-Lun Tseng, Chen-Hsu Yang |
ECAI | 4 |
| 2020 | Character-Preserving Coherent Story Visualization
Yun-Zhu Song, Zhi Rui Tam, Hung-Jen Chen 0001, Huiao-Han Lu, Hong-Han Shuai |
ECCV (17) | 5 |
| 2020 | Trajectory Prediction in Heterogeneous Environment via Attended Ecology EmbeddingabstractTrajectory prediction is a highly desirable feature for safe navigation or autonomous vehicle in complex traffic. In this paper, we consider the practical environment of predicting trajectory in the heterogeneous traffic ecology. The proposed method has various applications in trajectory prediction problems and also in applied fields beyond tracking. One challenge stands out of the trajectory prediction-heterogeneous environment. Particularly, many factors should be considered in the environments, i.e., multiple types of road-agents, social interactions and terrains. The information is complicated and large that may result in inaccurate trajectory prediction. We propose two social and visual enforced attention modules to circumvent the problem and a variant of an Info-GAN structure to predict the trajectory with multi-modal behaviors. Experimental results show that the proposed method significantly outperforms state-of-the-art methods in both heterogeneous and homogeneous real environments. Wei-Cheng Lai, Zi-Xiang Xia, Hao-Siang Lin, Lien-Feng Hsu, Hong-Han Shuai, I-Hong Jhuo, Wen-Huang Cheng |
ACM Multimedia | 5 |
| 2020 | Domain-Adaptive Object Detection via Uncertainty-Aware Distribution AlignmentabstractDomain adaptation aims to transfer knowledge from the source data with annotations to scarcely-labeled data in the target domain, which has attracted a lot of attention in recent years and facilitated many multimedia applications. Recent approaches have shown the effectiveness of using adversarial learning to reduce the distribution discrepancy between the source and target images by aligning distribution between source and target images at both image and instance levels. However, this remains challenging since two domains may have distinct background scenes and different objects. Moreover, complex combinations of objects and a variety of image styles deteriorate the unsupervised cross-domain distribution alignment. To address these challenges, in this paper, we design an end-to-end approach for unsupervised domain adaptation of object detector. Specifically, we propose a Multi-level Entropy Attention Alignment (MEAA) method that consists of two main components: (1) Local Uncertainty Attentional Alignment (LUAA) module to accelerate the model better perceiving structure-invariant objects of interest by utilizing information theory to measure the uncertainty of each local region via the entropy of the pixel-wise domain classifier and (2) Multi-level Uncertainty-Aware Context Alignment (MUCA) module to enrich domain-invariant information of relevant objects based on the entropy of multi-level domain classifiers. The proposed MEAA is evaluated in four domain-shift object detection scenarios. Experiment results demonstrate state-of-the-art performance on three challenging scenarios and competitive performance on one benchmark dataset. Dang-Khoa Nguyen, Wei-Lun Tseng, Hong-Han Shuai |
ACM Multimedia | 3 |
| 2020 | S2SiamFC: Self-supervised Fully Convolutional Siamese Network for Visual TrackingabstractTo exploit rich information from unlabeled data, in this work, we propose a novel self-supervised framework for visual tracking which can easily adapt the state-of-the-art supervised Siamese-based trackers into unsupervised ones by utilizing the fact that an image and any cropped region of it can form a natural pair for self-training. Besides common geometric transformation-based data augmentation and hard negative mining, we also propose adversarial masking which helps the tracker to learn other context information by adaptively blacking out salient regions of the target. The proposed approach can be trained offline using images only without any requirement of manual annotations and temporal information from multiple consecutive frames. Thus, it can be used with any kind of unlabeled data, including images and video frames. For evaluation, we take SiamFC as the base tracker and name the proposed self-supervised method as S2SiamFC. Extensive experiments and ablation studies on the challenging VOT2016 and VOT2018 datasets are provided to demonstrate the effectiveness of the proposed method which not only achieves comparable performance to its supervised counterpart and other unsupervised methods requiring multiple frames. Chon-Hou Sio, Yu-Jen Ma, Hong-Han Shuai, Jun-Cheng Chen, Wen-Huang Cheng |
ACM Multimedia | 3 |
| 2020 | AU-assisted Graph Attention Convolutional Network for Micro-Expression RecognitionabstractMicro-expressions (MEs) are important clues for reflecting the real feelings of humans, and micro-expression recognition (MER) can thus be applied in various real-world applications. However, it is difficult to perceive and interpret MEs correctly. With the advance of deep learning technologies, the accuracy of micro-expression recognition is improved but still limited by the lack of large-scale datasets. In this paper, we propose a novel micro-expression recognition approach by combining Action Units (AUs) and emotion category labels. Specifically, based on facial muscle movements, we model different AUs based on relational information and integrate the AUs recognition task with MER. Besides, to overcome the shortcomings of limited and imbalanced training samples, we propose a data augmentation method that can generate nearly indistinguishable image sequences with AU intensity of real-world micro-expression images, which effectively improve the performance and are compatible with other micro-expression recognition methods. Experimental results on three mainstream micro-expression datasets, i.e., CASME II, SAMM, and SMIC, manifest that our approach outperforms other state-of-the-art methods on both single database and cross-database micro-expression recognition. Ling Lo, Hong-Han Shuai, Wen-Huang Cheng |
ACM Multimedia | 3 |
| 2020 | Quality-Aware Streaming Network Embedding with Memory Refreshing
Hsi-Wen Chen, Hong-Han Shuai, Sheng-De Wang, De-Nian Yang |
PAKDD (1) | 2 |
| 2020 | Optimizing Item and Subgroup Configurations for Social-Aware VR ShoppingabstractShopping in VR malls has been regarded as a paradigm shift for E-commerce, but most of the conventional VR shopping platforms are designed for a single user. In this paper, we envisage a scenario of VR group shopping, which brings major advantages over conventional group shopping in brick-and-mortar stores and Web shopping: 1) configure flexible display of items and partitioning of subgroups to address individual interests in the group, and 2) support social interactions in the subgroups to boost sales. Accordingly, we formulate the Social-aware VR Group-Item Configuration (SVGIC) problem to configure a set of displayed items for flexibly partitioned subgroups of users in VR group shopping. We prove SVGIC is APX-hard and also NP-hard to approximate within [EQUATION]. We design a 4-approximation algorithm based on the idea of Co-display Subgroup Formation (CSF) to configure proper items for display to different subgroups of friends. Experimental results on real VR datasets and a user study with hTC VIVE manifest that our algorithms outperform baseline approaches by at least 30.1% of solution quality. Shao-Heng Ko, Hsu-Chao Lai, Hong-Han Shuai, Wang-Chien Lee, Philip S. Yu, De-Nian Yang |
Proc. VLDB Endow. | 3 |
| 2019 | Social-Aware VR Configuration Recommendation via Multi-Feedback Coupled Tensor FactorizationabstractRecent technological advent in virtual reality (VR) has attracted a lot of attention to the VR shopping, which thus far is designed for a single user. In this paper, we envision the scenario of VR group shopping, where VR supports: 1) flexible display of items to address diverse personal preferences, and 2) convenient view switching between personal and group views to foster social interactions. We formulate the Multiview-Enabled Configuration Recommendation (MECR) problem to rank a set of displayed items for a VR shopping user. We design the Multiview-Enabled Configuration Ranking System (MEIRS) that first extracts discriminative features based on Marketing theories and then introduces a new coupled tensor factorization model to learn the representation of users, Multi-View Display (MVD) configurations, and multiple feedback with content features. Experimental results manifest that the proposed approach outperforms personalized recommendations and group recommendations by at least 30.8% in large-scale datasets and 63.3% in the user study in terms of hit ratio and mean average precision. Hsu-Chao Lai, Hong-Han Shuai, De-Nian Yang, Jiun-Long Huang, Wang-Chien Lee, Philip S. Yu |
CIKM | 2 |
| 2019 | BeautyGlow: On-Demand Makeup Transfer Framework With Reversible Generative NetworkabstractAs makeup has been widely-adopted for beautification, finding suitable makeup by virtual makeup applications becomes popular. Therefore, a recent line of studies proposes to transfer the makeup from a given reference makeup image to the source non-makeup one. However, it is still challenging due to the massive number of makeup combinations. To facilitate on-demand makeup transfer, in this work, we propose BeautyGlow that decompose the latent vectors of face images derived from the Glow model into makeup and non-makeup latent vectors. Since there is no paired dataset, we formulate a new loss function to guide the decomposition. Afterward, the non-makeup latent vector of a source image and makeup latent vector of a reference image and are effectively combined and revert back to the image domain to derive the results. Experimental results show that the transfer quality of BeautyGlow is comparable to the state-of-the-art methods, while the unique ability to manipulate latent vectors allows BeautyGlow to realize on-demand makeup transfer. Hung-Jen Chen 0001, Ka-Ming Hui, Szu-Yu Wang, Li-Wu Tsao, Hong-Han Shuai, Wen-Huang Cheng |
CVPR | 5 |
| 2019 | Optimizing Social-Topic Engagement on Social Network and Knowledge GraphabstractExisting research on social networks manifests two crucial criteria to improve activity engagement of users: (1) user interests in the activity topics and (2) opportunities of making new friends with some acquaintances. However, current online platforms still involve massive manual selection for activity attendees and contents without proper recommendations. In this paper, therefore, we formulate a new activity organization problem, named Social Knowledge Group Query (SKGQ), to recommend attendees and topic-related contents simultaneously. We prove that SKGQ is NP-hard and design an approximation algorithm, named Social cOntent Knowledge Exploration (SOKE), to jointly choose the activity attendees and topic-related contents according to social-oriented and topic- oriented strategies. Simulation results manifest that the solution acquired by SOKE is close to the optimal solution and outperforms various baselines. Ya-Wen Teng, Yishuo Shi, Jui-Yi Tsai, Hong-Han Shuai, Chih-Hua Tai, De-Nian Yang |
GLOBECOM | 4 |
| 2019 | Fit-me: Image-Based Virtual Try-on With Arbitrary PosesabstractThe image-based virtual try-on system has raised research attention recently, but it still requires to upload an image of a user with the target pose. We present a novel learning model, Fit-Me network, to seamlessly fit in-shop clothing into a person image and simultaneously transform the pose of the person image to another given one. The proposed Fit-Me network helps users not only save the time used to change clothes physically but also provide comprehensive information about how suitable the clothes are. By facilitating the arbitrary pose transformation, we can generate consecutive poses to help users get more information for deciding whether to buy the clothes or not from different aspects. Chia-Wei Hsieh, Chieh-Yun Chen, Chien-Lung Chou, Hong-Han Shuai, Wen-Huang Cheng |
ICIP | 4 |
| 2019 | Dressing for Attention: Outfit Based Fashion Popularity PredictionabstractAnalysis of fashion trends is crucial. However, existing predictive algorithms of fashion popularity are restricted to be feasible on the coarse style level but not a finer item level. That is, they are only predictive in the future popularity of a given type of fashion styles (e.g., Rocker), but cannot be precisely down to a particular outfit look chosen by individuals. This paper thus proposes the first solution directly aimed at predicting the fine-grained fashion popularity of an outfit look by taking social media as the learning source. Particularly, a deep temporal sequence learning framework is developed and the proposed framework is evaluated on a real dataset of 380,000 street fashion images collected from the fashion website lookbook.nu. The experimental results show that our proposed framework outperforms the state-of-the-art approaches, with a relative increase of 11.51% to 27.62% (MSE metric) and 7.02% to 32.61% (CSE metric) in the prediction accuracy. Ling Lo, Chia-Lin Liu, Rong-An Lin, Bo Wu 0018, Hong-Han Shuai, Wen-Huang Cheng |
ICIP | 5 |
| 2019 | FashionOn: Semantic-guided Image-based Virtual Try-on with Detailed Human and Clothing InformationabstractThe image-based virtual try-on system has attracted a lot of research attention. The virtual try-on task is challenging since synthesizing try-on images involves the estimation of 3D transformation from 2D images, which is an ill-posed problem. Therefore, most of the previous virtual try-on systems cannot solve difficult cases, e.g., body occlusions, wrinkles of clothes, and details of the hair. Moreover, the existing systems require the users to upload the image for the target pose, which is not user-friendly. In this paper, we aim to resolve the above challenges by proposing a novel FashionOn network to synthesize user images fitting different clothes in arbitrary poses to provide comprehensive information about how suitable the clothes are. Specifically, given a user image, an in-shop clothing image, and a target pose (can be arbitrarily manipulated by joint points), FashionOn learns to synthesize the try-on images by three important stages: pose-guided parsing translation, segmentation region coloring, and salient region refinement. Extensive experiments demonstrate that FashionOn maintains the details of clothing information (e.g., logo, pleat, lace), as well as resolves the body occlusion problem, and thus achieves the state-of-the-art virtual try-on performance both qualitatively and quantitatively. Chia-Wei Hsieh, Chieh-Yun Chen, Chien-Lung Chou, Hong-Han Shuai, Jiaying Liu 0001, Wen-Huang Cheng |
ACM Multimedia | 4 |
| 2019 | Stop Hiding Behind Windshield: A Windshield Image Enhancer Based on a Two-way Generative Adversarial NetworkabstractWindshield images captured by surveillance cameras are usually difficult to be seen through due to severe image degradation such as reflection, motion blur, low light, haze, and noise. Such image degradation hinders the capability of identifying and tracking people. In this paper, we aim to address this challenging windshield images enhancement task by presenting a novel deep learning model based on a two-way generative adversarial network, called Two-way Individual Normalization Perceptual Adversarial Network, TWIN-PAN. TWIN-PAN is an unpaired learning network which does not require pairs of degraded and corresponding ground truth images for training. Also, unlike existing image restoration algorithms which only address one specific type of degradation at once, TWIN-PAN can restore the image from various types of degradation. To restore the content inside the extremely degraded windshield and ensure the semantic consistency of the image, we introduce cyclic perceptual loss to the network and combine it with cycle-consistency loss. Moreover, to generate better restoration images, we introduce individual instance normalization layers for the generators, which can help our generators better adapt to their own input distributions. Furthermore, we collect a large high-quality windshield image dataset (WIE-Dataset) to train our network and to validate the robustness of our method in restoring degraded windshield images. Experimental results on human detection, vehicle ReID and user study manifest that the proposed method is effective for windshield image restoration. Chi-Rung Chang, Kuan-Yu Lung, Yi-Chung Chen, Zhi-Kai Huang, Hong-Han Shuai, Wen-Huang Cheng |
MMAsia | 5 |
| 2019 | Multiple Fisheye Camera Tracking via Real-Time Feature ClusteringabstractRecently, Multi-Target Multi-Camera Tracking (MTMC) makes a breakthrough due to the release of DukeMTMC and show the feasibility of related applications. However, most of the existing MTMC methods focus on the batch methods which attempt to find the global optimal solution from the entire image sequence and thus are not suitable for the real-time applications, e.g., customer tracking in unmanned stores. In this paper, we propose a low-cost online tracking algorithm, namely, Deep Multi-Fisheye-Camera Tracking (DeepMFCT) to identify the customers and locate the corresponding positions from multiple overlapping fisheye cameras. Based on any single camera tracking algorithm (e.g., Deep SORT), our proposed algorithm establishes the correlation between different single camera tracks. Owing to the lack of well-annotated multiple overlapping fisheye cameras dataset, the main challenge of this issue is to efficiently overcome the domain gap problem between normal cameras and fisheye cameras based on existed deep learning based model. To address this challenge, we integrate a single camera tracking algorithm with cross camera clustering including location information that achieves great performance on the unmanned store dataset and Hall dataset. Experimental results show that the proposed algorithm improves the baselines by at least 7% in terms of MOTA on the Hall dataset. Chon-Hou Sio, Hong-Han Shuai, Wen-Huang Cheng |
MMAsia | 2 |
| 2019 | RNN-Assisted Network Coding for Secure Heterogeneous Internet of Things With Unreliable StorageabstractWith the rapid growth of Internet of Things (IoT), integrating a variety of IoT can result in novel applications. However, IoT devices are often deployed in an open environment where IoT are inclined to be malfunctioned. Although data reliability can be achieved by data recovery with conventional replication, the communication between IoT is susceptible to eavesdropping. Therefore, in this paper, we study the eavesdropping prevention of data repair in IoT environments based on network coding. We theoretically derive the relation between security level and storage in heterogeneous IoT systems. To further reduce the repair bandwidth, we exploit recurrent neural network for the storage failure prediction. Under the condition when failure probability and workloads of storage devices are considered, two allocation algorithms are proposed to avoid data repair. Finally, we show the relation between storage cost and reliability with different numbers of IoT devices. Experimental results manifest that the proposed allocation algorithms can outperform the baseline case by 18.4% in terms of the security level. Chen-Hung Liao, Hong-Han Shuai, Li-Chun Wang 0001 |
IEEE Internet Things J. | 2 |
| 2018 | Eavesdropping prevention for heterogeneous Internet of Things systemsabstractWith the rapid growth of Internet of Things (IoT), the idea of integrating a variety of IoTs has been proposed to support a variety of novel applications. However, IoT devices are often deployed in an open environment in which IoTs are inclined to be malfunctioned. Although data reliability can be achieved by data recovery with conventional replication, the communication between IoTs is susceptible to eavesdropping. Since most of the current work focuses on designing the architectures for IoTs under the specific scenarios, the eavesdropping prevention for heterogeneous IoT systems remains unexplored. Therefore, in this paper, we consider using the network-coding-based distributed storage systems for security and show that repair bandwidth can be reduced by increasing storage per node. Moreover, we theoretically derive the relation between repair bandwidth and storage in heterogeneous IoT systems. Finally, we show the relation between storage cost and reliability with regard to different amounts of IoT devices. Chen-Hung Liao, Hong-Han Shuai, Li-Chun Wang 0001 |
CCNC | 2 |
| 2018 | Newsfeed Filtering and Dissemination for Behavioral Therapy on Social Network AddictionsabstractWhile the popularity of online social network (OSN) apps continues to grow, little attention has been drawn to the increasing cases of Social Network Addictions (SNAs). In this paper, we argue that by mining OSN data in support of online intervention treatment, data scientists may assist mental healthcare professionals to alleviate the symptoms of users with SNA in early stages. Our idea, based on behavioral therapy, is to incrementally substitute highly addictive newsfeeds with safer, less addictive, and more supportive newsfeeds. To realize this idea, we propose a novel framework, called Newsfeed Substituting and Supporting System (N3S), for newsfeed filtering and dissemination in support of SNA interventions. New research challenges arise in 1) measuring the addictive degree of a newsfeed to an SNA patient, and 2) properly substituting addictive newsfeeds with safe ones based on psychological theories. To address these issues, we first propose the Additive Degree Model (ADM) to measure the addictive degrees of newsfeeds to different users. We then formulate a new optimization problem aiming to maximize the efficacy of behavioral therapy without sacrificing user preferences. Accordingly, we design a randomized algorithm with a theoretical bound. A user study with 716 Facebook users and 11 mental healthcare professionals around the world manifests that the addictive scores can be reduced by more than 30%. Moreover, experiments show that the correlation between the SNA scores and the addictive degrees quantified by the proposed model is much greater than that of state-of-the-art preference based models. Hong-Han Shuai, Yen-Chieh Lien, De-Nian Yang, Yi-Feng Lan, Wang-Chien Lee, Philip S. Yu |
CIKM | 1 |
| 2018 | On Accelerating Multi-Layered Heterogeneous Network Embedding LearningabstractWith the rapid growth of online social networks and IoT networks, mining valuable knowledge from the graph data become important. Meanwhile, as machine learning algorithms show their powers in prediction, different machine learning algorithms are proposed for different applications, e.g., personal recommendation, price prediction, communication anomaly detection. However, it is challenging to extract network features from graph data as the inputs for machine learning algorithms. One of the promising approaches is to use graph embedding approach, which extracts the valuable information of networks from each node into low dimensional vectors. However, the graph embedding approaches on a large-scale network require tremendous training time. Therefore, in this paper, we propose NOde Differentiation for Graph Embedding (NODGE) to prioritize the nodes, while high priority nodes are allocated with more resources to train their representations. We also theoretically analyze the proposed NODGE. Experimental results show that the proposed method reduces the training time of state-of-the-art method by at least 30.7%. Hong-Han Shuai, Cheng-Ming Tsai, Yun-Jui Hsu, Ta-Che Hsiao |
GLOBECOM | 1 |
| 2018 | Highly Parallel Sequential Pattern Mining on a Heterogeneous PlatformabstractSequential pattern mining can be applied to various fields such as disease prediction and stock analysis. Many algorithms have been proposed for sequential pattern mining, together with acceleration methods. In this paper, we show that a heterogeneous platform with CPU and GPU is more suitable for sequential pattern mining than traditional CPU-based approaches since the support counting process is inherently succinct and repetitive. Therefore, we propose the PArallel SequenTial pAttern mining algorithm, referred to as PASTA, to accelerate sequential pattern mining by combining the merits of CPU and GPU computing. Explicitly, PASTA adopts the vertical bitmap representation of database to exploits the GPU parallelism. In addition, a pipeline strategy is proposed to ensure that both CPU and GPU on the heterogeneous platform operate concurrently to fully utilize the computing power of the platform. Furthermore, we develop a swapping scheme to mitigate the limited memory problem of the GPU hardware without decreasing the performance. Finally, comprehensive experiments are conducted to analyze PASTA with different baselines. The experiments show that PASTA outperforms the state-of-the-art algorithms by orders of magnitude on both real and synthetic datasets. Yu-Heng Hsieh, Chun-Chieh Chen, Hong-Han Shuai, Ming-Syan Chen |
ICDM | 3 |
| 2018 | Maximizing Social Influence on Target Users
Yu Ting Wen, Wen-Chih Peng, Hong-Han Shuai |
PAKDD (3) | 3 |
| 2018 | Customer Purchase Behavior Prediction from Payment DatasetsabstractWith the advances in the development of mobile payments, a huge amount of payment data are collected by banks. User payment data offer a good dataset to depict customer behavior patterns. A comprehensive understanding of customers' purchase behavior is crucial to developing good marketing strategies, which may trigger much greater purchase amounts. For example, by exploring customer behavior patterns, given a target store, a set of potential customers is able to be identified. Yu Ting Wen, Pei-Wen Yeh, Tzu-Hao Tsai, Wen-Chih Peng, Hong-Han Shuai |
WSDM | 5 |
| 2018 | QMSampler: Joint Sampling of Multiple Networks with Quality GuaranteeabstractBecause Online Social Networks (OSNs) have become increasingly important in the last decade, they have motivated a great deal of research on Social Network Analysis (SNA). Currently, SNA algorithms are evaluated on real datasets obtained from large-scale OSNs, which are usually sampled by Breadth-First-Search (BFS), Random Walk (RW), or some variations of the latter. However, none of the released datasets provides any statistical guarantees on the difference between the sampled datasets and the ground truth. Moreover, all existing sampling algorithms only focus on sampling a single OSN, but each OSN is actually a sampling of a complete social network. Hence, even if the whole dataset from a single OSN is sampled, the results may still be skewed and may not fully reflect the properties of the complete social network. To address the above issues, we have made the first attempt to explore the joint sampling of multiple OSNs and propose an approach called Quality-guaranteed Multi-network Sampler (QMSampler) that can jointly sample multiple OSNs. QMSampler provides a statistical guarantee on the difference between the sampled real dataset and the ground truth (the perfect integration of all OSNs). Our experimental results demonstrate that the proposed approach generates a much smaller bias than any existing method. QMSampler has also been released as a free download. Hong-Han Shuai, De-Nian Yang, Philip S. Yu, Ming-Syan Chen |
IEEE Trans. Big Data | 1 |
| 2018 | A Comprehensive Study on Social Network Mental Disorders Detection via Online Social Media MiningabstractThe explosive growth in popularity of social networking leads to the problematic usage. An increasing number of social network mental disorders (SNMDs), such as Cyber-Relationship Addiction, Information Overload, and Net Compulsion, have been recently noted. Symptoms of these mental disorders are usually observed passively today, resulting in delayed clinical intervention. In this paper, we argue that mining online social behavior provides an opportunity to actively identify SNMDs at an early stage. It is challenging to detect SNMDs because the mental status cannot be directly observed from online social activity logs. Our approach, new and innovative to the practice of SNMD detection, does not rely on self-revealing of those mental factors via questionnaires in Psychology. Instead, we propose a machine learning framework, namely, Social Network Mental Disorder Detection (SNMDD), that exploits features extracted from social network data to accurately identify potential cases of SNMDs. We also exploit multi-source learning in SNMDD and propose a new SNMD-based Tensor Model (STM) to improve the accuracy. To increase the scalability of STM, we further improve the efficiency with performance guarantee. Our framework is evaluated via a user study with 3,126 online social network users. We conduct a feature analysis, and also apply SNMDD on large-scale datasets and analyze the characteristics of the three SNMD types. The results manifest that SNMDD is promising for identifying online social network users with potential SNMDs. Hong-Han Shuai, De-Nian Yang, Yi-Feng Lan, Wang-Chien Lee, Philip S. Yu, Ming-Syan Chen |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2017 | Task-Optimized Group Search for Social Internet of Things
Hong-Han Shuai, Kuo-Feng Hsu, Ming-Syan Chen |
EDBT | 2 |
| 2017 | On Finding Socially Tenuous Groups for Online Social NetworksabstractExisting research on finding social groups mostly focuses on dense subgraphs in social networks. However, finding socially tenuous groups also has many important applications. In this paper, we introduce the notion of k-triangles to measure the tenuity of a group. We then formulate a new research problem, Minimum k-Triangle Disconnected Group (MkTG), to find a socially tenuous group from online social networks. We prove that MkTG is NP-Hard and inapproximable within any ratio in arbitrary graphs but polynomial-time tractable in threshold graphs. Two algorithms, namely TERA and TERA-ADV, are designed to exploit graph-theoretical approaches for solving MkTG on general graphs effectively and efficiently. Experimental results on seven real datasets manifest that the proposed algorithms outperform existing approaches in both efficiency and solution quality. Liang-Hao Huang, De-Nian Yang, Hong-Han Shuai, Wang-Chien Lee, Ming-Syan Chen |
KDD | 4 |
| 2017 | Distributed and scalable sequential pattern mining through stream processing
Chun-Chieh Chen, Hong-Han Shuai, Ming-Syan Chen |
Knowl. Inf. Syst. | 2 |
| 2016 | When Social Influence Meets Item InferenceabstractResearch issues and data mining techniques for product recommendation and viral marketing have been widely studied. Existing works on seed selection in social networks do not take into account the effect of product recommendations in e-commerce stores. In this paper, we investigate the seed selection problem for viral marketing that considers both effects of social influence and item inference (for product recommendation). We develop a new model, Social Item Graph (SIG), that captures both effects in the form of hyperedges. Accordingly, we formulate a seed selection problem, called Social Item Maximization Problem (SIMP), and prove the hardness of SIMP. We design an efficient algorithm with performance guarantee, called Hyperedge-Aware Greedy (HAG), for SIMP and develop a new index structure, called SIG-index, to accelerate the computation of diffusion process in HAG. Moreover, to construct realistic SIG models for SIMP, we develop a statistical inference based framework to learn the weights of hyperedges from data. Finally, we perform a comprehensive evaluation on our proposals with various baselines. Experimental result validates our ideas and demonstrates the effectiveness and efficiency of the proposed model and algorithms over baselines. Hui-Ju Hung, Hong-Han Shuai, De-Nian Yang, Liang-Hao Huang, Wang-Chien Lee, Jian Pei 0001, Ming-Syan Chen |
KDD | 2 |
| 2016 | Mining Online Social Data for Detecting Social Network Mental DisordersabstractAn increasing number of social network mental disorders (SNMDs), such as Cyber-Relationship Addiction, Information Overload, and Net Compulsion, have been recently noted. Symptoms of these mental disorders are usually observed passively today, resulting in delayed clinical intervention. In this paper, we argue that mining online social behavior provides an opportunity to actively identify SNMDs at an early stage. It is challenging to detect SNMDs because the mental factors considered in standard diagnostic criteria (questionnaire) cannot be observed from online social activity logs. Our approach, new and innovative to the practice of SNMD detection, does not rely on self-revealing of those mental factors via questionnaires. Instead, we propose a machine learning framework, namely, Social Network Mental Disorder Detection (SNMDD), that exploits features extracted from social network data to accurately identify potential cases of SNMDs. We also exploit multi-source learning in SNMDD and propose a new SNMDbased Tensor Model (STM) to improve the performance. Our framework is evaluated via a user study with 3126 online social network users. We conduct a feature analysis, and also apply SNMDD on large-scale datasets and analyze the characteristics of the three SNMD types. The results show that SNMDD is promising for identifying online social network users with potential SNMDs. Hong-Han Shuai, De-Nian Yang, Yi-Feng Lan, Wang-Chien Lee, Philip S. Yu, Ming-Syan Chen |
WWW | 1 |
| 2016 | A Comprehensive Study on Willingness Maximization for Social Activity Planning with Quality GuaranteeabstractStudies show that a person is willing to join a social group activity if the activity is interesting, and if some close friends also join the activity as companions. The literature has demonstrated that the interests of a person and the social tightness among friends can be effectively derived and mined from social networking websites. However, even with the above two kinds of information widely available, social group activities still need to be coordinated manually, and the process is tedious and time-consuming for users, especially for a large social group activity, due to complications of social connectivity and the diversity of possible interests among friends. To address the above important need, this paper proposes to automatically select and recommend potential attendees of a social group activity, which could be very useful for social networking websites as a value-added service. We first formulate a new problem, named Willingness mAximization for Social grOup (WASO). This paper points out that the solution obtained by a greedy algorithm is likely to be trapped in a local optimal solution. Thus, we design a new randomized algorithm to effectively and efficiently solve the problem. Given the available computational budgets, the proposed algorithm is able to optimally allocate the resources and find a solution with an approximation ratio. We implement the proposed algorithm in Facebook, and the user study demonstrates that social groups obtained by the proposed algorithm significantly outperform the solutions manually configured by users. Hong-Han Shuai, De-Nian Yang, Philip S. Yu, Ming-Syan Chen |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2015 | Revenue maximization for telecommunications company with social viral marketingabstractViral marketing, a marketing strategy that leverages the influence power in intimate relationship, has become more prevalent due to the popularity of online social networking services in recent years. Consumers are more likely to make a purchase based on social media referrals. Since marketing through social media and traditional channels may target on different audiences, how to maximize the revenue of a telecommunications company by employing different advertising ways and selecting initial users for advertisements is a critical problem. Therefore, in this paper, we formulate a new research problem, namely Cost-Aware Multi-wAy Influence maXimization (CAMAIX) to address the need mentioned above. We design a 1/2-approximation algorithm with various pruning and budget allocation strategies to solve CAMAIX efficiently. We conduct extensive experiments on a large-scale real dataset from a telecommunications company. The results show that our proposed algorithm outperforms the baseline algorithms in both solution quality and efficiency. Hong-Han Shuai, Hsiang-Chun Hsu, De-Nian Yang, Chung-Kuang Chou, Jihg-Hong Lin, Ming-Syan Chen |
IEEE BigData | 1 |
| 2015 | Forming Online Support Groups for Internet and Behavior Related AddictionsabstractWhile online social networks have become a part of many people's daily lives, Internet and social network addictions (ISNAs) have been noted recently. With increased patients in addictive Internet use, clinicians often form support groups to help patients. This has become a trend because groups organized around therapeutic goals can effectively enrich members with insight and guidance while holding everyone accountable along the way. With the emergence of online social network services, there is a trend to form support groups online with the aid of mental health professionals. Nevertheless, it becomes impractical for a psychiatrist to manually select the group members because she faces an enormous number of candidates, while the selection criteria are also complicated since they span both the social and symptom dimensions. To effectively address the need of mental healthcare professionals, this paper makes the first attempt to study a new problem, namely Member Selection for Online Support Group (MSSG). The problem aims to maximize the similarity of the symptoms of all selected members, while ensuring that any two members are unacquainted to each other. We prove that MSSG is NP-Hard and inapproximable within any ratio, and design a 3-approximation algorithm with a guaranteed error bound. We evaluate MSSG via a user study with 11 mental health professionals, and the results manifest that MSSG can effectively find support group members satisfying the member selection criteria. Experimental results on large-scale real datasets also demonstrate that our proposed algorithm outperforms other baselines in terms of solution quality and efficiency. Hong-Han Shuai, De-Nian Yang, Yi-Feng Lan, Wang-Chien Lee, Philip S. Yu, Ming-Syan Chen |
CIKM | 2 |
| 2015 | Scale-Adaptive Group Optimization for Social Activity Planning
Hong-Han Shuai, De-Nian Yang, Philip S. Yu, Ming-Syan Chen |
PAKDD (1) | 1 |
| 2014 | Identifying Your Customers in Social NetworksabstractPersonal social networks are considered as one of the most influential sources in shaping a customer's attitudes and behaviors. However, the interactions with friends or colleagues in social networks of individual customers are barely observable in most e-commerce companies. In this paper, we study the problem of customer identification in social networks, i.e., connecting customer accounts at e-commerce sites to the corresponding user accounts in online social networks such as Twitter. Identifying customers in social networks is a crucial prerequisite for many potential marketing applications. These applications, for example, include personalized product recommendation based on social correlations, discovering community of customers, and maximizing product adoption and profits over social networks. Chun-Ta Lu, Hong-Han Shuai, Philip S. Yu |
CIKM | 2 |
| 2014 | Low-Density Cut Based Tree Decomposition for Large-Scale SVM ProblemsabstractThe current trend of growth of information reveals that it is inevitable that large-scale learning problems become the norm. In this paper, we propose and analyze a novel Low-density Cut based tree Decomposition method for large-scale SVM problems, called LCD-SVM. The basic idea here is divide and conquer: use a decision tree to decompose the data space and train SVMs on the decomposed regions. Specifically, we demonstrate the application of low density separation principle to devise a splitting criterion for rapidly generating a high-quality tree, thus maximizing the benefits of SVMs training. Extensive experiments on 14 real-world datasets show that our approach can provide a significant improvement in training time over state-of-the-art methods while keeps comparable test accuracy with other methods, especially for very large-scale datasets. Lifang He 0001, Hong-Han Shuai, Xiangnan Kong, Xiaowei Yang 0003, Philip S. Yu |
ICDM | 2 |
| 2013 | On Pattern Preserving Graph GenerationabstractReal datasets always play an essential role in graph mining and analysis. However, nowadays most available real datasets only support millions of nodes. Therefore, the literature on Big Data analysis utilizes statistical graph generators to generate a massive graph (e.g., billions of nodes) for evaluating the scalability of an algorithm. Nevertheless, current popular statistical graph generators are properly designed to preserve only the statistical metrics, such as the degree distribution, diameter, and clustering coefficient of the original social graphs. Recently, the importance of frequent graph patterns has been recognized in the various works on graph mining, but unfortunately this crucial criterion has not been noticed in the existing graph generators. To address this important need, we make the first attempt to design a Pattern Preserving Graph Generation (PPGG) algorithm to generate a graph including all frequent patterns and three most popular statistical parameters: degree distribution, clustering coefficient, and average vertex degree. The experimental results show that PPGG, which we have released as a free download, is efficient and able to generate a billion-node graph in approximately 10 minutes, much faster than the existing graph generators. Hong-Han Shuai, De-Nian Yang, Philip S. Yu, Ming-Syan Chen |
ICDM | 1 |
| 2013 | Willingness Optimization for Social Group ActivityabstractStudies show that a person is willing to join a social group activity if the activity is interesting, and if some close friends also join the activity as companions. The literature has demonstrated that the interests of a person and the social tightness among friends can be effectively derived and mined from social networking websites. However, even with the above two kinds of information widely available, social group activities still need to be coordinated manually, and the process is tedious and time-consuming for users, especially for a large social group activity, due to complications of social connectivity and the diversity of possible interests among friends. To address the above important need, this paper proposes to automatically select and recommend potential attendees of a social group activity, which could be very useful for social networking websites as a value-added service. We first formulate a new problem, named Willingness mAximization for Social grOup (WASO). This paper points out that the solution obtained by a greedy algorithm is likely to be trapped in a local optimal solution. Thus, we design a new randomized algorithm to effectively and efficiently solve the problem. Given the available computational budgets, the proposed algorithm is able to optimally allocate the resources and find a solution with an approximation ratio. We implement the proposed algorithm in Facebook, and the user study demonstrates that social groups obtained by the proposed algorithm significantly outperform the solutions manually configured by users. Hong-Han Shuai, De-Nian Yang, Philip S. Yu, Ming-Syan Chen |
Proc. VLDB Endow. | 1 |
| 2011 | MobiUP: An Upsampling-Based System Architecture for High-Quality Video Streaming on Mobile DevicesabstractNowadays, mobile video streaming enables people to access digital content, such as online TV shows, music videos, sports reports, and news programs, anytime, anywhere. However, current streaming services in mobile networks are subject to the available wireless bandwidth shared among many users and can only provide videos with limited resolutions. Moreover, on recently developed high-resolution mobile devices, such as iPhone, Google Nexus One, Nokia N97, and SonyEricsson X10, the resolution of video streaming is much lower than the devices can actually support. As a result, existing video upsampling schemes usually introduce visual artifacts. In response to the above problem, we bridge the resolution gap between streaming videos and client screens, and propose a novel upsampling-based system architecture, called MobiUP, to enable high-quality video streaming onto mobile devices. To avoid modifying existing codecs for video streaming, MobiUP upsamples videos with decoded frames and appends a limited amount of metadata to the streaming videos for facilitating high-quality and real-time conversion from low resolution to high fullscreen resolution on the client side. In other words, the proposed upsampling architecture complements current systems. Therefore, MobiUP is generic and flexible, and it can be implemented easily on mobile devices for practical use. The implementation results demonstrate that, although the appended metadata is less than 8% of the total transmitted data, it improves the quality of the upsampled video significantly. Meanwhile, the computation time of MobiUP Client is close to that of bilinear upsampling algorithms implemented on mobile devices. Hong-Han Shuai, De-Nian Yang, Wen-Huang Cheng, Ming-Syan Chen |
IEEE Trans. Multim. | 1 |