EDBT 2026 Demo / reviewers in the wild / expert
Yixuan Sun
dblp:221/1711 · also Yi-Xuan Sun
· DBLP profile ↗
30ranked-venue papers
9as first author
28since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 8 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 11 since 2021Systems, architecture and hardware · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Tracing the Heart's Pathways: ECG Representation Learning from a Cardiac Conduction PerspectiveabstractThe multi-lead electrocardiogram (ECG) stands as a cornerstone of cardiac diagnosis. Recent strides in electrocardiogram self-supervised learning (eSSL) have brightened prospects for enhancing representation learning without relying on high-quality annotations. Yet earlier eSSL methods suffer a key limitation: they focus on consistent patterns across leads and beats, overlooking the inherent differences in heartbeats rooted in cardiac conduction processes, while subtle but significant variations carry unique physiological signatures. Moreover, representation learning for ECG analysis should align with ECG diagnostic guidelines, which progress from individual heartbeats to single leads and ultimately to lead combinations. This sequential logic, however, is often neglected when applying pre-trained models to downstream tasks. To address these gaps, we propose CLEAR-HUG, a two-stage framework designed to capture subtle variations in cardiac conduction across leads while adhering to ECG diagnostic guidelines. In the first stage, we introduce an eSSL model termed Conduction-LEAd Reconstructor (CLEAR), which captures both specific variations and general commonalities across heartbeats. Treating each heartbeat as a distinct entity, CLEAR employs a simple yet effective sparse attention mechanism to reconstruct signals without interference from other heartbeats. In the second stage, we implement a Hierarchical lead-Unified Group head (HUG) for disease diagnosis, mirroring clinical workflow. Experimental results across six tasks show a 6.84% improvement, validating the effectiveness of CLEAR-HUG. This highlights its ability to enhance representations of cardiac conduction and align patterns with expert diagnostic guidelines. Tan Pan, Yixuan Sun, Chen Jiang 0006, Qiong Gao, Xingmeng Zhang, Zhenqi Yang, Limei Han, Yixiu Liang, Kaiyu Guo |
AAAI | 2 |
| 2026 | Proximity Alert: Ipelets for Neighborhood Graphs and Clustering (Media Exposition)abstractNeighborhood graphs and clustering algorithms are fundamental structures in both computational geometry and data analysis. Visualizing them can help build insight into their behavior and properties. The Ipe extensible drawing editor, developed by Otfried Cheong, is a widely used software system for generating figures. One particular aspect of Ipe is the ability to add Ipelets, which extend its functionality. Here we showcase a set of Ipelets designed to help visualize neighborhood graphs and clustering algorithms. These include: ε-neighbor graphs, furthest-neighbor graphs, Gabriel graphs, k-nearest neighbor graphs, k-th-nearest neighbor graphs, k-mutual neighbor graphs, k-th-mutual neighbor graphs, asymmetric k-nearest neighbor graphs, asymmetric k-th-nearest neighbor graphs, relative-neighbor graphs, sphere-of-influence graphs, Urquhart graphs, Yao graphs, and clustering algorithms including complete-linkage, DBSCAN, HDBSCAN, k-means, k-means++, k-medoids, mean shift, and single-linkage. Our Ipelets are all programmed in Lua and are freely available. Gitan Balogh, June Cagan, Bea Fatima, Auguste H. Gezalyan, Danesh Sivakumar, Arushi Srinivasan, Yixuan Sun, Vahe Zaprosyan, David M. Mount |
SoCG | 7 |
| 2026 | Poster: Hybrid Frequency Crossover for Multi-Scale Field Reconstruction in Wireless Sensor Networks
Guoqing Lu, Yixuan Sun, Yiwen Jiang, Dongxu Xia |
SECON | 2 |
| 2026 | A hybrid intelligence framework with uncertainty quantification for reliable surface roughness prediction in micro-milling
Zongbao He, Yixuan Sun, Shuaiqi Yuan, Fengze Qin, Chi Fai Cheung, Huajun Cao, Chunjin Wang |
Expert Syst. Appl. | 2 |
| 2025 | QA-MDT: Quality-aware Masked Diffusion Transformer for Enhanced Music GenerationabstractText-to-music (TTM) generation, which converts textual descriptions into audio, opens up innovative avenues for multimedia creation. Achieving high quality and diversity in this process demands extensive, high-quality data, which are often scarce in available datasets. Most open-source datasets frequently suffer from issues like low-quality waveforms and low text-audio consistency, hindering the advancement of music generation models. To address these challenges, we propose a novel quality-aware training paradigm for generating high-quality, high-musicality music from large-scale, quality-imbalanced datasets. Additionally, by leveraging unique properties in the latent space of musical signals, we adapt and implement a masked diffusion transformer (MDT) model for the TTM task, showcasing its capacity for quality control and enhanced musicality. Furthermore, we introduce a three-stage caption refinement approach to address low-quality captions' issue. Experiments show state-of-the-art (SOTA) performance on benchmark datasets including MusicCaps and the Song-Describer Dataset with both objective and subjective metrics. Demo audio samples are available at https://qa-mdt.github.io/, code and pretrained checkpoints are open-sourced at https://github.com/ivcylc/OpenMusic. Ruoyu Wang 0029, Jun Du 0002, Yixuan Sun, Zilu Guo, Zhengrong Zhang, Jianqing Gao |
IJCAI | 5 |
| 2025 | DSparse: A Distributed Training Method for Edge Clusters Based on Sparse Update
Xiaohui Peng 0002, Yixuan Sun, Zhenghui Zhang, Yifan Wang 0005 |
J. Comput. Sci. Technol. | 2 |
| 2024 | Pixel-Level Semantic Correspondence Through Layout-Aware Representation Learning and Multi-Scale Matching IntegrationabstractEstablishing precise semantic correspondence across object instances in different images is a fundamental and challenging task in computer vision. In this task, difficulty arises often due to three challenges: confusing regions with similar appearance, inconsistent object scale, and indistinguishable nearby pixels. Recognizing these challenges, our paper proposes a novel semantic matching pipeline named LPMFlow toward extracting fine-grained semantics and geometry layouts for building pixel-level semantic correspondences. LPMFlow consists of three modules, each addressing one of the aforementioned challenges. The layout-aware representation learning module uniformly encodes source and target tokens to distinguish pixels or regions with similar appearances but different geometry semantics. The progressive feature superresolution module outputs four sets of 4D correlation tensors to generate accurate semantic flow between objects in different scales. Finally, the matching flow integration and refinement module is exploited to fuse matching flow in different scales to give the final flow predictions. The whole pipeline can be trained end-to-end, with a balance of computational cost and correspondence details. Extensive experiments based on benchmarks such as SPair-71K, PF-PASCAL, and PF-WILLOW have proved that the proposed method can well tackle the three challenges and outperform the previous methods, es-pecially in more stringent settings. Code is available at https://github.com/YXSUNMADMAX/LPMFlow. Yixuan Sun, Zhangyue Yin, Haibo Wang 0006, Yan Wang 0068, Xipeng Qiu, Weifeng Ge |
CVPR | 1 |
| 2024 | The Merit of River Network Topology for Neural Flood ForecastingabstractClimate change exacerbates riverine floods, which occur with higher frequency and intensity than ever. The much-needed forecasting systems typically rely on accurate river discharge predictions. To this end, the SOTA data-driven approaches treat forecasting at spatially distributed gauge stations as isolated problems, even within the same river network. However, incorporating the known topology of the river network into the prediction model has the potential to leverage the adjacency relationship between gauges. Thus, we model river discharge for a network of gauging stations with GNNs and compare the forecasting performance achieved by different adjacency definitions. Our results show that the model fails to benefit from the river network topology information, both on the entire network and small subgraphs. The learned edge weights correlate with neither of the static definitions and exhibit no regular pattern. Furthermore, the GNNs struggle to predict sudden, narrow discharge spikes. Our work hints at a more general underlying phenomenon of neural prediction not always benefitting from graphical structure and may inspire a systematic study of the conditions under which this happens. Nikolas Kirschstein, Yixuan Sun |
ICML | 2 |
| 2024 | Self-cognitive Denoising in the Presence of Multiple Noisy Label SourcesabstractThe strong performance of neural networks typically hinges on the availability of extensive labeled data, yet acquiring ground-truth labels is often challenging. Instead, noisy supervisions from multiple sources, e.g., by multiple well-designed rules, are more convenient to collect. In this paper, we focus on the realistic problem of learning from multiple noisy label sources, and argue that prior studies have overlooked the crucial self-cognition ability of neural networks, i.e., the inherent capability of autonomously distinguishing noise during training. We theoretically analyze this ability of neural networks when meeting multiple noisy label sources, which reveals that neural networks possess the capability to recognize both instance-wise noise within each single noisy label source and annotator-wise quality among multiple noisy label sources. Inspired by the theoretical analyses, we introduce an approach named Self-cognitive Denoising for Multiple noisy label sources (SDM), which exploits the self-cognition ability of neural networks to denoise during training. Furthermore, we build a selective distillation module following the theoretical insights to optimize computational efficiency. The experiments on various datasets demonstrate the superiority of our method. Yixuan Sun, Ya-Lin Zhang 0001, Jun Zhou 0011 |
ICML | 1 |
| 2024 | Depth-Aware Multi-Modal Fusion for Generalized Zero-Shot LearningabstractRealizing Generalized Zero-Shot Learning (GZSL) based on large models is emerging as a prevailing trend. However, most existing methods merely regard large models as black boxes, solely leveraging the features output by the final layer while disregarding potential performance enhancements from other layers. Indeed, numerous researchers have visually depicted variations in the features learned across different layers of neural networks. Motivated by this observation, we propose a Vision Transformer (ViT)-based GZSL method named Depth-Aware Multi-Modal ViT (DAM2ViT), which exploits multi-level features of ViT. DAM2ViT incorporates a multi-modal interaction block to align semantic information of categories across multiple layers, thereby augmenting the model's capacity to learn associations between visual and semantic spaces. Extensive experiments conducted on three benchmark datasets (i.e., CUB, SUN, AWA2) have showcased that DAM2ViT achieves competitive results compared to state-of-the-art methods. Weipeng Cao, Xuyang Yao, Zhiwu Xu 0001, Yinghui Pan, Yixuan Sun, Dachuan Li, Bohua Qiu, Muheng Wei |
INDIN | 5 |
| 2024 | Weakly Supervised Gaussian Contrastive Grounding with Large Multimodal Models for Video Question AnsweringabstractVideo Question Answering (VideoQA) aims to answer natural language questions based on the information observed in videos. Despite the recent success of Large Multimodal Models (LMMs) in image-language understanding, they deal with VideoQA insufficiently, by simply taking uniformly sampled frames as visual inputs, which ignores question-relevant visual clues. Moreover, there are no human annotations for question-critical timestamps in existing VideoQA datasets. In light of this, we propose a weakly supervised framework to enforce the LMMs to reason out the answers with question-critical moments as visual inputs. Specifically, we fuse the question and answer pairs as event descriptions to find multiple keyframes as target moments and pseudo-labels. With these pseudo-labeled keyframes as additionally weak supervision, we devise a lightweight Gaussian-based Contrastive Grounding (GCG) module. GCG learns multiple Gaussian masks to characterize the temporal structure of the video, and sample question-critical frames as positive moments to be the visual inputs of LMMs. Extensive experiments on several benchmarks verify the effectiveness of our framework, and we achieve substantial improvements compared to previous state-of-the-art methods. Haibo Wang 0006, Chenghang Lai, Yixuan Sun, Weifeng Ge |
ACM Multimedia | 3 |
| 2024 | Collaborative Refining for Learning from Inaccurate LabelsabstractThis paper considers the problem of learning from multiple sets of inaccurate labels, which can be easily obtained from low-cost annotators, such as rule-based annotators. Previous works typically concentrate on aggregating information from all the annotators, overlooking the significance of data refinement. This paper presents a collaborative refining approach for learning from inaccurate labels. To refine the data, we introduce the annotator agreement as an instrument, which refers to whether multiple annotators agree or disagree on the labels for a given sample. For samples where some annotators disagree, a comparative strategy is proposed to filter noise. Through theoretical analysis, the connections among multiple sets of labels, the respective models trained on them, and the true labels are uncovered to identify relatively reliable labels. For samples where all annotators agree, an aggregating strategy is designed to mitigate potential noise. Guided by theoretical bounds on loss values, a sample selection criterion is introduced and modified to be more robust against potentially problematic values. Through these two methods, all the samples are refined during training, and these refined samples are used to train a lightweight model simultaneously. Extensive experiments are conducted on benchmark and real-world datasets to demonstrate the superiority of our methods. Yixuan Sun, Ya-Lin Zhang 0001, Libang Zhang, Jun Zhou 0011, Guo Ye, Huimei He |
NeurIPS | 2 |
| 2024 | Enabling 13K-Atom Excited-State GW Calculations via Low-Rank Approximations and HPC on the New Sunway SupercomputerabstractGW approximation is a powerful approach to accurately describe the excited-state of semiconductors. However, GW incurs high computational cost $\mathcal{O}\left(N^{4}\right)$ and large memory usage $\mathcal{O}\left(N^{3}\right)$, limiting its applications to thousands of (2,742) atoms even on leadership supercomputers. Herein we present a massively parallel implementation of accurate and efficient cubic-scaling plane-wave GW calculations by using low-rank approximations and high-performance computing on leadership supercomputers. By using a series of low rank approximations, we can reduce the expensive GW calculations to the cubic-scaling computational cost $\mathcal{O}\left(N^{3}\right)$ and quadratic memory usage $\mathcal{O}\left(N^{2}\right)$. With the help of parallel and communication optimization, the plane-wave GW calculations gain an overall speedup of over 70x and efficiently scale up to 13,824 atoms within a few minutes using 449,280 cores on new Sunway supercomputer. This accomplishment paves the way for excited-state quantum mechanical material simulations at mesoscopic scale (10K atoms) and for the design of next-generation semiconductor devices. Wentiao Wu, Zhengbang Zhou, Qingcai Jiang, Junwei Feng, Xinming Qin, Huanhuan Ma, Zhenwei Cao, Junshi Chen 0003, Xinyong Meng, Bingkun Hou, Yuanfan Xiong, Linhao Wang, Yixuan Sun, Hong An, Jinlong Yang 0003, Wei Hu 0006 |
SC | 14 |
| 2024 | A safe reinforcement learning algorithm for supervisory control of power plants
Yixuan Sun, Sami Khairy, Richard B. Vilim, Akshay J. Dave |
Knowl. Based Syst. | 1 |
| 2023 | Treatment Effect Estimation across DomainsabstractTreatment effect estimation is essential in the causal inference literature, which has attracted increasing attention in recent years. Most previous methods assume that the training and test data are drawn from the same distribution, which may not hold in practice since the effect estimators may need to be deployed across domains. Meanwhile, in real-world applications, little or no targeted treatments may be conducted in the new domain. Therefore, we focus on a more realistic scenario in this paper, where treatments and outcomes can be observed in the source domain, but the target domain only contains some unlabeled data, i.e., only features are available. In this scenario, thedistribution shift exists not only in the source data due to the selection bias between the control and treated groups, but also between the source and target data. We propose a novel direct learning framework along with the distribution adaptation and reliable scoring modules. In the distribution adaptation module, we design three specialized density ratio estimators to aid the issue of complex distribution shifts. Even so, we may face the challenge of unreliable pseudo-effects in this framework. To address that, we also design the uncertainty-based reliable scoring module as a vital support, which makes the method more reliable. The experiments are conducted on synthetic data and benchmark datasets, which demonstrate the superiority of our method. Yixuan Sun, Ya-Lin Zhang 0001, Wei Wang 0028, Jun Zhou 0011 |
CIKM | 1 |
| 2023 | MISC210K: A Large-Scale Dataset for Multi-Instance Semantic CorrespondenceabstractSemantic correspondence have built up a new way for object recognition. However current single-object matching schema can be hard for discovering commonalities for a category and far from the real-world recognition tasks. To fill this gap, we design the multi-instance semantic correspondence task which aims at constructing the correspondence between multiple objects in an image pair. To support this task, we build a multi-instance semantic correspondence (MISC) dataset from COCO Detection 2017 task called MISC210K. We construct our dataset as three steps: (1) category selection and data cleaning; (2) keypoint design based on 3D models and object description rules; (3) human-machine collaborative annotation. Following these steps, we select 34 classes of objects with 4,812 challenging images annotated via a well designed semi-automatic workflow, and finally acquire 218,179 image pairs with instance masks and instance-level keypoint pairs annotated. We design a dual-path collaborative learning pipeline to train instance-level co-segmentation task and fine-grained level correspondence task together. Benchmark evaluation and further ablation results with detailed analysis are provided with three future directions proposed. Our project is available on https://github.com/YXSUNMADMAX/MISC210K. Yixuan Sun, Haijing Guo, Yuzhou Zhao, Runmin Wu, Yizhou Yu, Weifeng Ge |
CVPR | 1 |
| 2023 | Correspondence Transformers with Asymmetric Feature Learning and Matching Flow Super-ResolutionabstractThis paper solves the problem of learning dense visual correspondences between different object instances of the same category with only sparse annotations. We decompose this pixel-level semantic matching problem into two easier ones: (i) First, local feature descriptors of source and target images need to be mapped into shared semantic spaces to get coarse matching flows. (ii) Second, matching flows in low resolution should be refined to generate accurate point-to-point matching results. We propose asymmetric feature learning and matching flow super-resolution based on vision transformers to solve the above problems. The asymmetric feature learning module exploits a biased cross-attention mechanism to encode token features of source images with their target counterparts. Then matching flow in low resolutions is enhanced by a super-resolution network to get accurate correspondences. Our pipeline is built upon vision transformers and can be trained in an end-to-end manner. Extensive experimental results on several popular benchmarks, such as PF-PASCAL, PF-WILLOW, and SPair-71 K, demonstrate that the proposed method can catch subtle semantic differences in pixels efficiently. Code is available on https://github.com/YXSUNMADMAX/ACTR. Yixuan Sun, Dongyang Zhao, Zhangyue Yin, Tao Gui, Weifeng Ge |
CVPR | 1 |
| 2023 | Is One Epoch All You Need For Multi-Fidelity Hyperparameter Optimization?abstractHyperparameter optimization (HPO) is crucial for fine-tuning machine learning models, but it can be computationally expensive.To reduce costs, Multi-fidelity HPO (MF-HPO) leverages intermediate accuracy levels in the learning process and discards low-performing models early on.We conducted a comparison of various representative MF-HPO methods against a simple baseline on classical benchmark data.The baseline involved discarding all models except the Top-K after training for only one epoch, followed by further training to select the best model.Surprisingly, this baseline achieved similar results to its counterparts, while requiring an order of magnitude less computation.Upon analyzing the learning curves of the benchmark data, we observed a few dominant learning curves, which explained the success of our baseline.This suggests that researchers should (1) always use the suggested baseline in benchmarks and (2) broaden the diversity of MF-HPO benchmarks to include more complex cases. Romain Egele, Isabelle Guyon, Yixuan Sun, Prasanna Balaprakash |
ESANN | 3 |
| 2023 | Weakly Supervised Learning of Semantic Correspondence through Cascaded Online Correspondence RefinementabstractIn this paper, we develop a weakly supervised learning algorithm to learn robust semantic correspondences from large-scale datasets with only image-level labels. Following the spirit of multiple instance learning (MIL), we decompose the weakly supervised correspondence learning problem into three stages: image-level matching, region-level matching, and pixel-level matching. We propose a novel cascaded online correspondence refinement algorithm to integrate MIL and the correspondence filtering and refinement procedure into a single deep network and train this network end-to-end with only image-level supervision, i.e., without point-to-point matching information. During the correspondence learning process, pixel-to-pixel matching pairs inferred from weak supervision are propagated, filtered, and enhanced through masked correspondence voting and calibration. Besides, we design a correspondence consistency check algorithm to select images with discriminative key points to generate pseudo-labels for classical matching algorithms. Finally, we filter out about 110,000 images from the ImageNet ILSVRC training set to formulate a new dataset, called SC-ImageNet. Experiments on several popular benchmarks indicate that pre-training on SC-ImageNet can improve the performance of state-ofthe-art algorithms efficiently. Our project is available on https://github.com/21210240056/SC-ImageNet. Yixuan Sun, Chenghang Lai, Qing Xu 0017, Xuli Shen, Weifeng Ge |
ICCV | 2 |
| 2023 | Hierarchical Visual Categories Modeling: A Joint Representation Learning and Density Estimation Framework for Out-of-Distribution DetectionabstractDetecting out-of-distribution inputs for visual recognition models has become critical in safe deep learning. This paper proposes a novel hierarchical visual category modeling scheme to separate out-of-distribution data from in-distribution data through joint representation learning and statistical modeling. We learn a mixture of Gaussian models for each in-distribution category. There are many Gaussian mixture models to model different visual categories. With these Gaussian models, we design an in-distribution score function by aggregating multiple Mahalanobis-based metrics. We don’t use any auxiliary outlier data as training samples, which may hurt the generalization ability of out-of-distribution detection algorithms. We split the ImageNet-1k dataset into ten folds randomly. We use one fold as the in-distribution dataset and the others as out-of-distribution datasets to evaluate the proposed method. We also conduct experiments on seven popular benchmarks, including CIFAR, iNaturalist, SUN, Places, Textures, ImageNet-O, and OpenImage-O. Extensive experiments indicate that the proposed method outperforms state-of-the-art algorithms clearly. Meanwhile, we find that our visual representation has a competitive performance when compared with features learned by classical methods. These results demonstrate that the proposed method hasn’t weakened the discriminative ability of visual recognition models and keeps high efficiency in detecting out-of-distribution samples. Xinyu Zhou 0006, Pinxue Guo, Yixuan Sun, Weifeng Ge |
ICCV | 4 |
| 2023 | Collision-free Coverage Path Planning for the Variable-speed Curvature-constrained RobotabstractDubins coverage has been extensively researched to address the coverage path planning (CPP) problem of a known environment for the curvature-constrained robot. However, its fixed-speed assumption prevents the robot from accelerating to reduce the time and limits its flexibility to avoid obstacles. Therefore, this paper presents a collision-free CPP approach (CFC) for the obstacle-constrained environment, which enhances time efficiency by constructing the variable-speed Dubins paths and ensures robot safety by building a risk potential surface for representing the possibility of collision. Furthermore, CFC models the CPP problem as an asymmetric traveling salesman problem (ATSP) and utilizes a graph pruning strategy to reduce the computational cost. Comparison tests with other Dubins coverage methods demonstrate that CFC provides shorter coverage times and better runtimes than the other Dubins coverage methods while preventing collision risk between the robot and obstacles. Physical experiments in a laboratory setting demonstrate the applicability of CFC to the physical robot. Lin Li 0075, Dian-xi Shi, Songchang Jin, Yixuan Sun, Xing Zhou 0004, Shaowu Yang, Hengzhu Liu |
ICRA | 4 |
| 2023 | A Framework for Detecting Frauds from Extremely Few LabelsabstractIn this paper, we present a framework to deal with the fraud detection task with extremely few labeled frauds. We involve human intelligence in the loop in a labor-saving manner and introduce several ingenious designs to the model construction process. Namely, a rule mining module is introduced, and the learned rules will be refined with expert knowledge. The refined rules will be used to relabel the unlabeled samples and get the potential frauds. We further present a model to learn with the reliable frauds, the potential frauds, and the rest normal samples. Note that the label noise problem, class imbalance problem, and confirmation bias problem are all addressed with specific strategies when building the model. Experimental results are reported to demonstrate the effectiveness of the framework. Ya-Lin Zhang 0001, Yixuan Sun, Meng Li 0068, Yeyu Zhao, Wei Wang 0028, Jun Zhou 0011, Jinghua Feng |
WSDM | 2 |
| 2023 | Path guided motion synthesis for Drosophila larvaeabstractThe deformability and high degree of freedom of mollusks bring challenges in mathematical modeling and synthesis of motions. Traditional analytical and statistical models are limited by either rigid skeleton assumptions or model capacity, and have difficulty in generating realistic and multi-pattern mollusk motions. In this work, we present a large-scale dynamic pose dataset of Drosophila larvae and propose a motion synthesis model named Path2Pose to generate a pose sequence given the initial poses and the subsequent guiding path. The Path2Pose model is further used to synthesize long pose sequences of various motion patterns through a recursive generation method. Evaluation analysis results demonstrate that our novel model synthesizes highly realistic mollusk motions and achieves state-of-the-art performance. Our work proves high performance of deep neural networks for mollusk motion synthesis and the feasibility of long pose sequence synthesis based on the customized body shape and guiding path. Yixuan Sun, Ziao Liu, Zhefeng Gong, Nenggan Zheng |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2022 | Exploiting Mixed Unlabeled Data for Detecting Samples of Seen and Unseen Out-of-Distribution ClassesabstractOut-of-Distribution (OOD) detection is essential in real-world applications, which has attracted increasing attention in recent years. However, most existing OOD detection methods require many labeled In-Distribution (ID) data, causing a heavy labeling cost. In this paper, we focus on the more realistic scenario, where limited labeled data and abundant unlabeled data are available, and these unlabeled data are mixed with ID and OOD samples. We propose the Adaptive In-Out-aware Learning (AIOL) method, in which we employ the appropriate temperature to adaptively select potential ID and OOD samples from the mixed unlabeled data and consider the entropy over them for OOD detection. Moreover, since the test data in realistic applications may contain OOD samples whose classes are not in the mixed unlabeled data (we call them unseen OOD classes), data augmentation techniques are brought into the method to further improve the performance. The experiments are conducted on various benchmark datasets, which demonstrate the superiority of our method. Yixuan Sun, Wei Wang 0028 |
AAAI | 1 |
| 2022 | FERV39k: A Large-Scale Multi-Scene Dataset for Facial Expression Recognition in VideosabstractCurrent benchmarks for facial expression recognition (FER) mainly focus on static images, while there are limited datasets for FER in videos. It is still ambiguous to evaluate whether performances of existing methods remain satisfactory in real-world application-oriented scenes. For example, the “Happy” expression with high intensity in Talk-Show is more discriminating than the same expression with low intensity in Official-Event. To fill this gap, we build a large-scale multi-scene dataset, coined as FERV39k. We analyze the important ingredients of constructing such a novel dataset in three aspects: (1) multi-scene hierarchy and expression class, (2) generation of candidate video clips, (3) trusted manual labelling process. Based on these guidelines, we select 4 scenarios subdivided into 22 scenes, annotate 86k samples automatically obtained from 4k videos based on the well-designed workflow, and finally build 38,935 video clips labeled with 7 classic expressions. Experiment benchmarks on four kinds of baseline frame-works were also provided and further analysis on their performance across different scenes and some challenges for future research were given. Besides, we systematically investigate key components of DFER by ablation studies. The baseline framework and our project are available on https://github.com/wangyanckxx/FERV39k. Yan Wang 0068, Yixuan Sun, Zhongying Liu, Shuyong Gao, Wei Zhang 0016, Weifeng Ge |
CVPR | 2 |
| 2022 | Hebo System: Trusted Copyright Authorization in Computer NetworksabstractThe DRM (Digital Rights Management) systems protect owners’ copyrights by controlling consumers’ access to digital works. However, they fail to provide authorization evidence if customers use digital works on other platforms outside the DRM systems. There is no such evidence that can trustworthily be disseminated in computer networks, which results in many copyright lawsuits. To address the problem, we propose the Copyrights Authorization Model (CAM) and the Hebo system to ensure the consensus on copyright authorization in computer networks. The CAM proves that participants agree on copyright authorizations if they are traceable, integrated, and non-repudiated. Based on the CAM, we design the ledger of trusted authorization forest that keeps the three properties and independent zones that maintain the ledger. The Hebo system is composed of these zones. It can provide authorization evidence for consumers to avoid copyright disputes. Besides, it has the advantages of a beneficial locality. The TPS (transactions per second) can increase with the number of independent zones, nodes can flexibly choose ledgers according to their capability, the system reduces redundant storage, and zones allow asynchrony. Finally, we evaluate the system on five different platforms, and the average time costs of operations are less than 60 milliseconds. The TPS of a single node depends on the configuration of the hardware, which is an average of 138, 19, 24, 21, and 98 in Server, MacBook, J-Nano, Pi4B, and J-TX2, respectively. Yixuan Sun, Xiaohui Peng 0002 |
ICPADS | 3 |
| 2022 | DPCNet: Dual Path Multi-Excitation Collaborative Network for Facial Expression Representation Learning in VideosabstractCurrent works of facial expression learning in video consume significant computational resources to learn spatial channel feature representations and temporal relationships. To mitigate this issue, we propose a Dual Path multi-excitation Collaborative Network (DPCNet) to learn the critical information for facial expression representation from fewer keyframes in videos. Specifically, the DPCNet learns the important regions and keyframes from a tuple of four view-grouped frames by multi-excitation modules and produces dual-path representations of one video with consistency under two regularization strategies. A spatial-frame excitation module and a channel-temporal aggregation module are introduced consecutively to learn spatial-frame representation and generate complementary channel-temporal aggregation, respectively. Moreover, we design a multi-frame regularization loss to enforce the representation of multiple frames in the dual view to be semantically coherent. To obtain consistent prediction probabilities from the dual path, we further propose a dual path regularization loss, aiming to minimize the divergence between the distributions of two-path embeddings. Extensive experiments and ablation studies show that the DPCNet can significantly improve the performance of video-based FER and achieve state-of-the-art results on the large-scale DFEW dataset. Yan Wang 0068, Yixuan Sun, Wei Song 0007, Shuyong Gao, Zhaoyu Chen 0001, Weifeng Ge |
ACM Multimedia | 2 |
| 2022 | A survey of human-in-the-loop for machine learning
Xingjiao Wu, Luwei Xiao, Yixuan Sun, Junhang Zhang, Tianlong Ma, Liang He 0001 |
Future Gener. Comput. Syst. | 3 |
| 2020 | Feasibility Analysis and Suitable Antenna Directions of IGNSS-R Altimetry Measurement for Avoiding the Intersatellite InterferenceabstractThe interferometric Global Navigation Satellite System Reflectometry (iGNSS-R) exploits the full spectrum of GNSS signal to improve the ranging performance for sea surface height applications. However, simultaneous reception of multiple satellite signals introduces the intersatellite interference degrading the altimetry accuracy, which occurs in both ground-based and spaceborne scenarios. By simulating situations of ground-based stations in Beijing and Hainan, the feasibility of ground-based iGNSS-R altimetry measurement without the intersatellite interference is verified, for more than half of the total observation areas are available and up to 34.73% and 74.64% of a day is suitable in Beijing and Hainan. Some antenna directions are recommended in the northern (southern) hemisphere: for long-time measurement, the lower (upper) parabola formed by GLONASS and the area of specular reflection points of GEO satellites are preferred; For continuous measurement, GEO satellites are recommended; For all-day measurement, the range of elevation and azimuth is 30°-60° and ±300(150°-210°). Yixuan Sun, Dongkai Yang, Junming Xia, Cong Yin |
IGARSS | 1 |
| 2018 | Local Feature Sufficiency Exploration for Predicting Security-Constrained Generation Dispatch in Multi-area Power SystemsabstractDeriving generation dispatch is essential for efficient and secure operation of electric power systems. This is usually achieved by solving a security-constrained optimal power flow (SCOPF) problem, which is by nature non-convex, usually nonlinear and thus computationally intensive. The state-of-the-art optimization approaches are not able to solve this problem for large-scale power systems within the power system operation time window (usually 5 minutes). In this work, we developed supervised learning approaches to determine security-constrained generation dispatch within a much shorter time window. More importantly, the physical constraint of only accessing to local measurements and other information in most utilities' real-time operation can not be ignored for the predictive models. The feasibility and accuracy of utilizing only local features (measurements and grid information in one area) to predict optimal local generation dispatch (dispatch of all generators in the corresponding area) in multi-area power systems has been explored. The results showed optimal local generation dispatch can be predicted with local features with high accuracy, which is comparable to the results obtained with global features. Yixuan Sun, Xiaoyuan Fan, Qiuhua Huang, Xinya Li 0002, Renke Huang, Tianzhixi Yin, Guang Lin 0001 |
ICMLA | 1 |