ClickAIXR: On-Device Multimodal Vision-Language Interaction with Real-World Objects in Extended Reality

ClickAIXR interaction flow for selecting a physical object and querying an on-device vision-language model Figure from the author-uploaded arXiv version.

At a glance

  • Brings vision-language interaction to an XR headset, without sending the visual query to the cloud.
  • Pairs an on-device VLM with controller-based 3D cropping to resolve which real-world object a user means.
  • Targets private, responsive question answering about selected physical objects.
Publication
arXiv preprint (arXiv)

Motivation

Vision-language assistants can answer questions about the physical world, but cloud inference introduces privacy, connectivity, and latency concerns. In XR, they must also know precisely which nearby object the user intends to ask about.

Method

ClickAIXR combines local vision-language inference with a controller-driven object-selection interface on an XR headset. Users place and adjust a 3D crop around a physical object using depth, width, and height controls; the selected visual context is then passed to an on-device VLM with the spoken or typed question.

Evaluation

The paper evaluates the interaction on real-world objects, considering the accuracy of the selected context as well as the practical responsiveness of the on-device workflow. The system is positioned against cloud-based and gaze-based alternatives that leave privacy or referential ambiguity unresolved.

Limitations

The approach depends on the capabilities and compute budget of the headset-resident model, and careful selection remains necessary when objects are occluded, small, or spatially close together.

Dominik Engel
Dominik Engel
Deep Learning Researcher