Skip to main navigation Skip to search Skip to main content

Semi-SwinUNeTR: Towards 3D Swin Vision Transformer-Based UNet for Medical Image Segmentation with Limited Annotations

  • Yinbing Tian
  • , Ziyang Wang
  • , Li Guo*
  • *Corresponding author for this work

Research output: Contribution to journalArticlepeer-review

2 Downloads (Pure)

Abstract

Accurate brain tumor segmentation from magnetic resonance imaging (MRI) is essential for computer-assisted diagnosis, treatment planning, and disease monitoring. However, brain tumors usually exhibit irregular, heterogeneous, and multi-scale spatial patterns with complex and ambiguous boundaries. At the same time, the performance of deep segmentation models is often constrained by the limited availability of voxel-level annotations, which are expensive and time-consuming to obtain. To address these challenges, this paper proposes Semi-SwinUNeTR, a semi-supervised framework for 3D brain tumor segmentation with limited annotated data. The proposed method adopts SwinUNeTR as the segmentation backbone, enabling hierarchical volumetric representation learning through shifted-window self-attention while preserving the encoder-decoder structure required for dense prediction. On top of this backbone, we introduce a dual-consistency semi-supervised learning strategy, consisting of mean teacher-based model consistency and interpolation consistency-based data consistency. In addition, voxel-wise consistency weights are used to redistribute semi-supervised supervision toward structurally complex and boundary-irregular tumor regions without changing the SwinUNeTR backbone. Experiments on the BraTS 2019 benchmark demonstrate that the proposed framework achieves strong performance across different annotation ratios. The original Semi-SwinUNeTR achieves Dice scores of 84.93%, 86.25%, 87.05%, and 87.83% under the 10%, 20%, 40%, and 80% labeled-data settings, respectively. With the weighted consistency extension, the Dice scores are further improved to 85.64%, 87.94%, and 88.59% under the 10%, 20%, and 80% labeled-data settings, respectively, while the corresponding HD 95 values are reduced to 8.9826, 8.1854, and 7.4533. These results indicate that combining a SwinUNeTR backbone with complementary model consistency, data consistency, and voxel-wise consistency weighting is an effective strategy for semi-supervised volumetric medical image segmentation under limited annotation.

Original languageEnglish
Article number695
Number of pages25
JournalBioengineering
Volume13
Issue number6
Early online date17 Jun 2026
DOIs
Publication statusPublished - 17 Jun 2026

Bibliographical note

Copyright © 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.

Data Access Statement

The BraTS 2019 dataset used in this study is available from the official BraTS challenge organizers subject to their data access policy.

Funding

This work was supported in part by Key Research and Development Program of Hainan Province under Grant ZDYF2025 (LALH) 002, and in part by the Beijing Natural Science Foundation under Grant L232039.

Funder number
ZDYF2025 (LALH) 002
L232039

    Keywords

    • SwinUNeTR
    • auxiliary weighting
    • biomedical image analysis
    • interpolation consistency
    • semi-supervised learning
    • vision transformer
    • volumetric medical image segmentation

    Fingerprint

    Dive into the research topics of 'Semi-SwinUNeTR: Towards 3D Swin Vision Transformer-Based UNet for Medical Image Segmentation with Limited Annotations'. Together they form a unique fingerprint.

    Cite this