Skip to main navigation Skip to search Skip to main content

Query, Align, and Distill: Navigation-Aware Cross-Modal Interaction for Efficient Vision-and-Language Navigation

Research output: Contribution to journalArticlepeer-review

6 Downloads (Pure)

Abstract

Recent large-scale Vision-and-Language Navigation (VLN) models deliver strong accuracy but remain costly to deployment due to heavy parameters and computation. We tackle efficient VLN in two steps. First, we build a high-performing teacher that makes navigation evidence selection explicit and compressible. Concretely, the teacher introduces a small set of learnable query slots to extract global and local action-sufficient navigable evidence from panoramic observations via a Navigable Query Generator, and progressively grounds these evidence tokens to the instruction with an Instruction–Query Aligner for policy prediction. Second, leveraging this explicit query bottleneck as a distillation interface, we train a compact student by transferring both where to attend and what to do: we distill the teacher’s global/local navigable queries with a navigation-aware token-adaptive objective, and further match action distributions during fine-tuning. Experiments on standard VLN benchmarks demonstrate that our student nearly matches the teacher’s navigation performance while reducing the number of parameters by 93.65% compared to the teacher.
Original languageEnglish
Number of pages14
JournalIEEE Transactions on Circuits and Systems For Video Technology
DOIs
Publication statusPublished - 25 Aug 2026

Bibliographical note

This is an accepted manuscript of an article published in: Z. Chen, Y. Ge, Z. Wang, P. Cao and L. Yang, "Query, Align, and Distill: Navigation-Aware Cross-Modal Interaction for Efficient Vision-and-Language Navigation," in IEEE Transactions on Circuits and Systems for Video Technology, doi: 10.1109/TCSVT.2026.3727199. For the purposes of open access the author/s has/ve applied a Creative Commons Attribution (CC BY) licence to any Author Accepted Manuscript (AAM) version arising from this submission.

Fingerprint

Dive into the research topics of 'Query, Align, and Distill: Navigation-Aware Cross-Modal Interaction for Efficient Vision-and-Language Navigation'. Together they form a unique fingerprint.

Cite this