Figure from the author-uploaded arXiv version.Vision-language assistants can answer questions about the physical world, but cloud inference introduces privacy, connectivity, and latency concerns. In XR, they must also know precisely which nearby object the user intends to ask about.
ClickAIXR combines local vision-language inference with a controller-driven object-selection interface on an XR headset. Users place and adjust a 3D crop around a physical object using depth, width, and height controls; the selected visual context is then passed to an on-device VLM with the spoken or typed question.
The paper evaluates the interaction on real-world objects, considering the accuracy of the selected context as well as the practical responsiveness of the on-device workflow. The system is positioned against cloud-based and gaze-based alternatives that leave privacy or referential ambiguity unresolved.
The approach depends on the capabilities and compute budget of the headset-resident model, and careful selection remains necessary when objects are occluded, small, or spatially close together.