Abstract
Accurate anatomical landmark detection is important for orthodontic analysis, surgical planning, and morphometric measurement, but fully supervised methods usually require large expert-annotated datasets. This work studies a one-shot setting, where only a single annotated template image is used for training. We propose a foundation-model-based landmark detection framework using a frozen DINO Vision Transformer (ViT) backbone. The proposed framework integrates three complementary components: a Multi-Layer Multi-Facet (MLMF) module that adaptively fuses key and value features from multiple ViT layers through global source-wise reweighting; a Mamba-Based Long-Range Context Aggregation (MLCA) module that injects global anatomical context into fused patch descriptors with linear complexity; and a Topology-Constrained Graph Refinement (TCGR) module that refines the predicted landmark configuration using anatomical graph constraints. Experiments on the Cephalometric dataset and the Hand X-ray dataset demonstrate that the proposed method achieves strong performance. Overall, the results show that jointly exploiting multi-source foundation-model representations, efficient long-range context aggregation, and topology-aware refinement improves annotation-efficient anatomical landmark detection.
| Original language | English |
|---|---|
| Article number | 2414 |
| Number of pages | 19 |
| Journal | Electronics |
| Volume | 15 |
| Issue number | 11 |
| Early online date | 2 Jun 2026 |
| DOIs | |
| Publication status | Published - 2 Jun 2026 |
Bibliographical note
Copyright © 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.Data Access Statement
The ISBI 2015 Cephalometric Landmark Challenge dataset and the Hand X-ray landmark dataset used in this study are publicly available from their original providers, subject to the corresponding data access and usage policies.Funding
This work was supported by the Key Research and Development Program of Hainan Province (No. ZDYF2025(LALH)002) and the Beijing Natural Science Foundation (No. L232039).
Keywords
- one-shot landmark detection
- foundation model
- multi-layer feature fusion
- Mamba
- state space model
- graph attention network
Fingerprint
Dive into the research topics of 'Foundation Model-Based One-Shot Anatomical Landmark Detection with Mamba and Graph Refinement'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver