EDBT 2026 Demo / reviewers in the wild / expert
Junlin Yang
dblp:226/5634
· DBLP profile ↗
12ranked-venue papers
6as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 4 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Computer networks · 3 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VibraGait: Multi-User Gait Recognition based on Footstep-Induced Floor Vibrations via mmWave
Junlin Yang, Jiadi Yu, Linghe Kong, Yanmin Zhu 0006, Daqiang Zhang 0001, Yi-Chao Chen 0001 |
INFOCOM | 1 |
| 2026 | Zero-Shot Face Authentication in Multi-User Scenarios Using mmWave SignalsabstractFace authentication has become an increasingly attractive and prevalent technology in human-computer interaction. Recent works have explored radio frequency (RF) signals for illumination-robust and privacy-preserving face authentication. However, these approaches require all users to undergo compulsory prior registration and cannot effectively handle unregistered users, which restricts their practical applicability in real-world scenarios. In this paper, we present ammWave-basedmulti-user face authentication system,m3FacePass, which performs zero-shot face authentication in complex multi-user scenarios using a commercial off-the-shelf (COTS) mmWave radar. First, m3FacePass collects mmWave signals and reconstructs spatial mappings by generating 3D heatmaps. Based on the 3D heatmaps,m3FacePassdetects and separates faces in multi-user scenarios through template matching. Then, unique facial features are extracted and transformed to a hypersphere manifold, which reserves feature space for unregistered users. By introducing dummy classifiers in the authentication model,m3FacePassoptimizes decision boundaries between registered and unregistered users. Furthermore,m3FacePassenhances its generalization ability for zero-shot authentication by incorporating pseudo unknowns during model training. Extensive experiments in real-world environments demonstrate thatm3FacePassachieves an authentication accuracy of 96.6% for registered users and 91.1% for unregistered users in multi-user scenarios. Junlin Yang, Jiadi Yu, Hao Kong 0004, Yanmin Zhu 0006 |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | $m^{3}FacePass$: mmWave-Based Multi-User Face Authentication with Zero-Shot LearningabstractFace authentication has become an increasingly attractive and prevalent technology in human-computer interaction. Recent works have explored radio frequency (RF) signals for illumination-robust and privacy-preserving face authentication. However, these approaches require all users to undergo compulsory prior registration and cannot effectively handle unregistered users, which restricts their practical applicability in real-world scenarios. In this paper, we present a mmWave-based multiuser face authentication system,$m^{3}$FacePass, which performs zero-shot face authentication in complex multi-user scenarios using a commercial off-the-shelf (COTS) mm Wave radar. First,$m^{3}$FacePass collects mmWave signals and reconstructs spatial mappings by generating 3D heatmaps. Based on the 3D heatmaps,$m^{3}$FacePass detects and separates faces in multi-user scenarios through template matching. Then, unique facial features are extracted and transformed to a hypersphere manifold, which reserves feature space for unregistered users. By introducing dummy classifiers in the authentication model,$\boldsymbol{m}^{\mathbf{3}}$FacePass optimizes decision boundaries between registered and unregistered users. Furthermore,$\boldsymbol{m}^{\mathbf{3}}$FacePass enhances its generalization ability for zero-shot authentication by incorporating pseudo unknowns during model training. Extensive experiments in realworld environments demonstrate that$\boldsymbol{m}^{3}$FacePass achieves an authentication accuracy of 96.6% for registered users and 91.1% for unregistered users in multi-user scenarios. Junlin Yang, Jiadi Yu, Hao Kong 0004, Yanmin Zhu 0006 |
ICPADS | 1 |
| 2025 | OpenCUA: Open Foundations for Computer-Use AgentsabstractVision-language models have demonstrated impressive capabilities as computer-use agents (CUAs) capable of automating diverse computer tasks. As their commercial potential grows, critical details of the most capable CUA systems remain closed. As these agents will increasingly mediate digital interactions and execute consequential decisions on our behalf, the research community needs access to open CUA frameworks to study their capabilities, limitations, and risks. To bridge this gap, we propose OpenCUA, a comprehensive open-source framework for scaling CUA data and foundation models. Our framework consists of: (1) an annotation infrastructure that seamlessly captures human computer-use demonstrations; (2) AgentNet, the first large-scale computer-use task dataset spanning 3 operating systems and 200+ applications and websites; (3) a scalable pipeline that transforms demonstrations into state–action pairs with reflective long Chain-of-Thought reasoning that sustain robust performance gains as data scales. Our end-to-end agent models demonstrate strong performance across CUA benchmarks. In particular, OpenCUA-72B achieves an average success rate of 45.0% on OSWorld‑Verified, establishing a new state-of-the-art (SOTA) among open-source models. Further analysis confirms that our approach generalizes well across domains and benefits significantly from increased test-time computation. We release our annotation tool, datasets, code, and models to build open foundations for further CUA research. Xinyuan Wang 0010, Dunjie Lu, Junlin Yang, Tianbao Xie, Xiaole Guo, Yiheng Xu, Chen Henry Wu, Zhennan Shen, Zhuokai Li, Xiaochuan Li 0003, Junda Chen, Boyuan Zheng 0001, Peihang Li, Fangyu Lei, Ruisheng Cao, Yeqiao Fu, Dongchan Shin, Martin Shin, Jiarui Hu 0006, Jixuan Chen, Yuxiao Ye, Yipu Wang, Diyi Yang, Victor Zhong, Y. Charles, Tao Yu 0009 |
NeurIPS | 4 |
| 2025 | Scaling Computer-Use Grounding via User Interface Decomposition and SynthesisabstractGraphical user interface (GUI) grounding, the ability to map natural language instructions to specific actions on graphical user interfaces, remains a critical bottleneck in computer use agent development. Current benchmarks oversimplify grounding tasks as short referring expressions, failing to capture the complexity of real-world interactions that require software commonsense, layout understanding, and fine-grained manipulation capabilities. To address these limitations, we introduce OSWorld-G, a comprehensive benchmark comprising 564 finely annotated samples across diverse task types including text matching, element recognition, layout understanding, and precise manipulation. Additionally, we synthesize and release the largest computer use grounding dataset Jedi, which contains 4 million examples through multi-perspective decoupling of tasks. Our multi-scale models trained on Jedi demonstrate its effectiveness by outperforming existing approaches on ScreenSpot-v2, ScreenSpot-Pro, and our OSWorld-G. Furthermore, we demonstrate that improved grounding with Jedi directly enhances agentic capabilities of general foundation models on complex computer tasks with state-of-the-art performance, improving from 23% to 51% on OSWorld. Through detailed ablation studies, we identify key factors contributing to grounding performance and verify that combining specialized data for different interface elements enables compositional generalization to novel interfaces. All benchmark, data, checkpoints, and code are open-sourced and available at https://osworld-grounding.github.io. Tianbao Xie, Xiaochuan Li 0003, Junlin Yang, Haoyuan Wu, Jixuan Chen, Wenjing Hu, Xinyuan Wang 0010, Yiheng Xu, Doyen Sahoo, Tao Yu 0009, Caiming Xiong |
NeurIPS | 4 |
| 2025 | Toward Open-World-Aware User Authentication Based on Human Bodies Using mmWave SignalsabstractUser authentication is evolving with expanded applications and innovative techniques. New authentication approaches utilize RF signals to sense specific human characteristics, offering a contactless and nonintrusive solution. However, these RF signal-based methods struggle with challenges in open-world scenarios, i.e., dynamic environments, daily behaviors with unrestricted postures, and identification of unauthorized users with security threats. In this paper, we present an open-world user authentication system, OpenAuth, which leverages a commercial off-the-shelf (COTS) mmWave radar to sense unrestricted human postures and behaviors for identifying individuals. First, OpenAuth utilizes a MUSIC-based neural network imaging model to eliminate environmental clutter and generate environment-independent human silhouette images. Then, the human silhouette images are normalized to consistent topological structures of human postures, ensuring robustness against unrestricted human postures. Next, fine-grained body features are extracted from these environment-independent and posture-independent human silhouette images using a metric learning model. To eliminate potential security threats that arise from unauthorized users, OpenAuth synthesizes data placeholders for enhancing unauthorized user identification. Finally, a k-NN-based authentication model is constructed to authenticate users' identities. Experiments in real environments show that the proposed OpenAuth achieves an average authentication accuracy of 93.4% and false acceptance rate (FAR) of 1.8% in open-world scenarios. Junlin Yang, Jiadi Yu, Linghe Kong, Yanmin Zhu 0006, Hongning Dai |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | mmHand: 3D Hand Pose Estimation Leveraging mmWave SignalsabstractHand pose estimation is a key support for a variety of interactive applications including user interface control, sign language understanding, virtual reality modeling, etc. Existing approaches mainly exploit wearable devices such as gloves or bracelets to estimate hand poses, which may introduce high deploying costs and intrusive user experience. Others rely on vision technologies whereas they could face complicated illuminations and privacy leakage. In this paper, we present a millimeter wave (mmWave) signal-based 3D hand pose estimation system, mmHand, which utilizes a mmWave radar to generate 3D hand skeletons and reconstruct 3D hand meshes. mmHand first leverages mmWave signals to sense a hand and pre-process the signals. Then, mmHand extracts spatial and temporal features using a designed attention-based hourglass network (mmSpaceNet) and Long Short-Term Memory (LSTM), respectively. Based on the extracted features, mmHand further regresses hand joints in 3D space to generate 3D hand skeletons. Finally, 3D hand meshes that continuously describe hand poses with detailed surfaces are reconstructed through a hand Model with Articulated and Non-rigid defOrmations (MANO). Extensive experiments demonstrate that mmHand can accurately generate 3D hand skeletons with 18.3mm mean per joint position error and 95.1 % of correct key points, which indicates the effectiveness of mmHand on hand pose estimation. Hao Kong 0004, Haoxin Lyu, Jiadi Yu, Linghe Kong, Junlin Yang, Yanzhi Ren, Hongbo Liu 0002, Yingying Chen 0001 |
ICDCS | 5 |
| 2024 | OpenAuth: Human Body-Based User Authentication Using mmWave Signals in Open-World ScenariosabstractUser authentication is evolving with expanded application scenarios and innovative techniques. New authentication approaches utilize RF signals to sense specific human behaviors and characteristics, such as faces, specific gestures, etc., offering a contactless and nonintrusive solution. However, these RF signal-based methods struggle with challenges in open-world scenarios, i.e., dynamic environments, daily behaviors with unrestricted postures, and identification of unauthorized users with security threats. In this paper, we present an open-world user authentication system, OpenAuth, which leverages a commercial off-the-shelf (COTS) mmWave radar to sense unrestricted human postures and behaviors for identifying individuals. First, OpenAuth utilizes a MUSIC-based neural network imaging model to eliminate environmental clutter and generates environment-independent human silhouette images. Then, the human silhouette images are normalized to consistent topological structures of human postures, ensuring robustness against unrestricted human postures. Based on the environment-independent and posture-independent human silhouette images, OpenAuth further extracts fine-grained body features through a metric learning model for user authentication. To eliminate potential security threats that arise from frequent accesses by unauthorized users, OpenAuth synthesizes data placeholders for enhancing the applicability of unauthorized user identification. Finally, a k-NN-based authentication model is constructed based on the extracted body features to authenticate users' identities. Experiments in real environments show that the proposed OpenAuth achieves an average authentication accuracy of 93.4 % and false acceptance rate (FAR) of 1.8% in open-world scenarios. Junlin Yang, Jiadi Yu, Linghe Kong, Yanmin Zhu 0006, Hongning Dai |
ICDCS | 1 |
| 2021 | Semantic Segmentation With Generative Models: Semi-Supervised Learning and Strong Out-of-Domain GeneralizationabstractTraining deep networks with limited labeled data while achieving a strong generalization ability is key in the quest to reduce human annotation efforts. This is the goal of semi-supervised learning, which exploits more widely available unlabeled data to complement small labeled data sets. In this paper, we propose a novel framework for discriminative pixel-level tasks using a generative model of both images and labels. Concretely, we learn a generative adversarial network that captures the joint image-label distribution and is trained efficiently using a large set of un-labeled images supplemented with only few labeled ones. We build our architecture on top of StyleGAN2 [45], augmented with a label synthesis branch. Image labeling at test time is achieved by first embedding the target image into the joint latent space via an encoder network and test-time optimization, and then generating the label from the inferred embedding. We evaluate our approach in two important domains: medical image segmentation and part-based face segmentation. We demonstrate strong in-domain performance compared to several baselines, and are the first to showcase extreme out-of-domain generalization, such as transferring from CT to MRI in medical imaging, and photographs of real faces to paintings, sculptures, and even cartoons and animal faces. Project Page: https://nv-tlabs.github.io/semanticGAN/ Daiqing Li, Junlin Yang, Karsten Kreis, Antonio Torralba 0001, Sanja Fidler |
CVPR | 2 |
| 2021 | Robust and Generalizable Visual Representation Learning via Random Convolutions
Zhenlin Xu, Deyi Liu, Junlin Yang, Colin Raffel, Marc Niethammer |
ICLR | 3 |
| 2020 | Layer Embedding Analysis in Convolutional Neural Networks for Improved Probability Calibration and ClassificationabstractIn this project, our goal is to develop a method for interpreting how a neural network makes layer-by-layer embedded decisions when trained for a classification task, and also to use this insight for improving the model performance. To do this, we first approximate the distribution of the image representations in these embeddings using random forest models, the output of which, termed embedding outputs, are used for measuring how the network classifies each sample. Next, we design a pipeline to use this layer embedding output to calibrate the original model output for improved probability calibration and classification. We apply this two-steps method in a fully convolutional neural network trained for a liver tissue classification task on our institutional dataset that contains 20 3D multi-parameter MR images for patients with hepatocellular carcinoma, as well as on a public dataset with 131 3D CT images. The results show that our method is not only able to provide visualizations that are easy to interpret, but that the embedded decision-based information is also useful for improving model performance in terms of probability calibration and classification, achieving the best performance compared to other baseline methods. Moreover, this method is computationally efficient, easy to implement, and robust to hyper-parameters. Fan Zhang 0009, Nicha C. Dvornek, Junlin Yang, Julius Chapiro, James S. Duncan |
IEEE Trans. Medical Imaging | 3 |
| 2019 | Unsupervised Domain Adaptation via Disentangled Representations: Application to Cross-Modality Liver Segmentation
Junlin Yang, Nicha C. Dvornek, Fan Zhang 0009, Julius Chapiro, Ming De Lin, James S. Duncan |
MICCAI (2) | 1 |