Multimodal Foundation Models Integrated with Extended Reality (XR) for Adaptive Human–AI Collaboration in Intelligent Environments

Amit Verma, Priya Verma, Sandeep Kumar

Abstract


ABSTRACT
The convergence of Extended Reality (XR)—encompassing Virtual,
Augmented, and Mixed Realities—with Multimodal Foundation Models
(MFMs) represents a seminal paradigm shift in physical-cyber interaction.
Intelligent environments, ranging from industrial smart factories and robotic
surgical suites to collaborative design studios, require context-aware, real
time, and adaptive human-AI collaboration. Traditional interactive systems in
XR rely on rigid command frameworks or single-modal inputs, failing under
dynamic, high-cognitive-load conditions. This review paper presents a
synthesis of state-of-the-art developments, architectural methodologies, and
performance trade-offs in integrating MFMs into XR ecosystems. We propose
a
taxonomy of multimodal fusion strategies operating across spatial,
temporal, and semantic dimensions, evaluating how cross-attention
mechanisms reduce latency and cognitive load. Through comparative analysis
of edge-cloud deployment trade-offs, spatial grounding pipelines, and closed
loop adaptive interaction frameworks, we demonstrate that MFM-XR
integration elevates intent recognition accuracy to over 94% while
maintaining sub-20 ms end-to-end latency in edge-accelerated environments.
Furthermore, this study identifies critical research gaps concerning spatial
hallucination, real-time compute overhead, multi-user concurrency, and
sensory privacy, formulating strategic trajectories to guide spatial computing.

KEYWORDS: Multimodal Foundation Models (MFMs), Extended Reality
(XR), Human-AI Collaboration, Intelligent Environments, Vision-Language
Action Models, Embodied Intelligence, Spatial Computing.

Full Text:

PDF 84-95