Toggle menu
Toggle preferences menu
Toggle personal menu
Not logged in
Your IP address will be publicly visible if you make any edits.
NVIDIA XR AI
Information
Type Software development kit
Industry Enterprise extended reality, artificial intelligence
Developer NVIDIA
Written In Python, JavaScript, C++, Swift, Kotlin
Operating System Linux (server; Ubuntu 22.04 or 24.04 recommended)
License Apache License 2.0
Supported Devices Web browsers, Android, iOS and visionOS, native C++ clients; AR glasses and XR headsets
Release Date June 16, 2026 (public beta)
Website https://developer.nvidia.com/xr/xr-ai


NVIDIA XR AI is an open-source developer library from NVIDIA for building multimodal AI agents that run alongside AR glasses and other extended reality (XR) devices. The agents receive live camera, microphone and device data from the wearer, pass it to GPU-hosted models and tools, and answer with speech or data in the same session.[1][2] NVIDIA released it as a public beta on June 16, 2026, during the Augmented World Expo (AWE) 2026, and publishes the code on GitHub under the Apache 2.0 license.[3][4][5]

The repository describes XR AI as "an open-source foundation for multimodal, real-time conversational AI in the NVIDIA CloudXR ecosystem." It is aimed at hands-busy enterprise work such as factory maintenance, laboratory procedures and surgical assistance, and it can optionally pair the agent with NVIDIA CloudXR remote rendering when an application needs 3D content in the headset.[5][2][1] As of September 2026 the software is still in beta; the latest release is v0.2.0-beta.2 from September 10, 2026, and NVIDIA warns that its APIs and behavior may change.[6][5]

Reviewed 29 September 2026. Checked names, dates, versions, model and requirements tables, partner uses, VITURE Helix details, LabOS preprint and press quotes against the NVIDIA blogs, GitHub releases, versioned docs, VITURE release, arXiv PDF and cited press. About review dates.

History

A pre-release titled "[PID] Alpha 0.1 Release (Private)" was published in NVIDIA's GitHub repository on May 6, 2026, before the first public version.[7] The public beta, v0.1.0-beta, went out on June 16, 2026. That day NVIDIA posted an announcement on its corporate blog by David Chu and a technical walkthrough by Greg Barbone, who leads XR product and partner management at NVIDIA.[3][1][2] VITURE's launch release for its Helix glasses said NVIDIA would "officially unveil its XR AI solution" in David Chu's keynote at AWE 2026 in Long Beach, California, with Helix shown as the first AI glasses built on the platform.[4]

VITURE says Helix builds on an earlier collaboration between VITURE, NVIDIA, the Le Cong Lab at Stanford University and the Mengdi Wang Lab at Princeton University.[4] A preprint called "LabOS: The AI-XR Co-Scientist That Sees and Works With Humans", first posted in October 2025, describes a system in which researchers wear XR glasses that stream what they see and hear to a local GPU server (or the cloud) for real-time agent inference; the glasses then display the agent's feedback. Its authors come from Stanford University School of Medicine, Princeton University and other institutions, and six of them, including Greg Barbone, list NVIDIA as their affiliation. The paper's acknowledgements say the AR glasses hardware was procured from Viture.[8] The preprint does not use the name XR AI. NVIDIA's June 2026 announcement says Rana, an AutoBio company, is "introducing its LabOS system on NVIDIA XR AI" at the Cong Lab at Stanford and the Wang Lab at Princeton.[1]

Version 0.2.0-beta.2 (September 10, 2026) revised the agent architecture around typed AgentRuntime and VoiceAgent components and retired the in-tree MCP servers and the public Pipecat pipeline APIs. It also added lab-instrument-monitoring and tea-making samples, moved text-to-speech from Piper to Pocket-TTS, and made breaking package, API and configuration changes, including renaming the media hub to DeviceIOHub.[6][9]

Version Date Notes
[PID] Alpha 0.1 (private) May 6, 2026 Private pre-release tag[7]
v0.1.0-beta June 16, 2026 First public release: XR Media Hub, agent SDK, GPU-backed VLM, LLM, STT and TTS services, six MCP servers, two end-to-end samples, CloudXR integration layer[3]
v0.2.0-beta.2 September 10, 2026 Revised agent architecture, two new samples, Pocket-TTS, versioned documentation, breaking changes including the DeviceIOHub rename[6]

Architecture

NVIDIA's announcement groups the platform into four capabilities: ingesting video, audio, depth, pose and sensor data from AR and XR devices; connecting agents to tools such as NVIDIA Metropolis, its video search and summarization (VSS) blueprint and NeMo Retriever for retrieval-augmented generation; supporting models including Nemotron reasoning models and Cosmos Reason; and agent orchestration through NeMo Agent Toolkit, with DGX Spark, DGX Station and RTX PRO systems running inference in the cloud, data center or at the edge.[1]

In the code, a client device joins a session and publishes audio, video and data to the DeviceIOHub (called the XR Media Hub in the first beta). The hub tags each stream with the participant's identity and fans the events out to agent workers. Video frames stay in shared memory, and a worker fetches pixels only when a task needs an image. Replies are routed back only to the participant who asked. LiveKit carries the client connection, but workers talk to the hub over msgpack and ZeroMQ and never address LiveKit directly. A typical voice agent chains speech recognition, a language or vision-language model and text-to-speech, and can call tools in between.[10][2] Because several clients can join one hub and several agents can watch the same streams, NVIDIA presents the design as usable for multi-user and multi-agent sessions.[2]

At launch, enterprise data and XR functions were exposed to agents through the Model Context Protocol (MCP). The first beta shipped MCP servers for visual question answering, video queries, scene rendering, OpenXR spatial information, vector math and transcripts.[2][3] The xr-render-demo sample launches the hub, a CloudXR runtime, model services and an agent that creates and moves objects in the user's space by voice; CloudXR then streams the rendered scene from GPU servers to the device. CloudXR needs a separate license and is not bundled with XR AI.[2][3]

Bundled AI services

The reference stack runs NVIDIA open models locally by default, most of them served through vLLM. Developers can instead point the language and vision models at hosted or self-hosted NVIDIA NIM endpoints or other OpenAI-compatible APIs; the v0.2 documentation notes that speech recognition and text-to-speech cannot use hosted NIM speech services and stay on local or self-hosted servers.[11][2]

Role v0.1.0-beta model v0.2.0-beta.2 model
Vision-language model Cosmos-Reason1-7B[12] Cosmos3 Nano Reasoner[11]
Speech-to-text parakeet-tdt-0.6b-v3[12] parakeet-tdt-0.6b-v3[11]
Text-to-speech magpie_tts_multilingual_357m; Piper voices[12] magpie_tts_multilingual_357m; kyutai/pocket-tts[11]
Fast language model Llama-3.1-Nemotron-Nano-8B-v1[12] Llama-3.1-Nemotron-Nano-8B-v1[11]
Tool-calling language model NVIDIA-Nemotron-3-Nano-30B-A3B[12] NVIDIA-Nemotron-3-Nano-30B-A3B[11]
Multimodal (text and video) model Nemotron-3-Nano-Omni-30B-A3B-Reasoning[12] Nemotron-3-Nano-Omni-30B-A3B-Reasoning[11]
Embeddings None listed[12] llama-nemotron-embed-1b-v2[11]

The XR render demo uses a small model for quick acknowledgments and status updates while a larger model handles reasoning and tool calls in the background.[2]

Requirements

The server side runs on Linux, with Ubuntu 22.04 or 24.04 recommended, Python 3.11 or 3.12, NVIDIA driver 580 or later, Docker 24 or later and the NVIDIA Container Toolkit. WSL2 on Windows is not officially supported. The bundled GPU profiles target two 48 GB NVIDIA Ada GPUs, one RTX PRO 6000 Blackwell workstation GPU, or a DGX Spark; the full local model stack needs about 55 GB of GPU-visible memory. Developers who use cloud-hosted models still need an NVIDIA GPU with NVENC and NVDEC on the hub machine.[13]

Supported clients

The repository includes sample clients for each platform. They share a common StreamKit design in which one backend layer is the only code that imports a LiveKit SDK.[14]

Client Transport Build
Web LiveKit from CDN None[14]
Web-XR Local LiveKit and CloudXR bundles Build script[14]
Android LiveKit Android Android Studio or Gradle[14]
iOS and visionOS LiveKit Swift and CloudXRKit Xcode[14]
Native C++ LiveKit C++ CMake[14]

The first beta listed Android XR passthrough camera and immersive session APIs as not yet wired up.[3] NVIDIA's announcement names smart glasses from Meta, Rokid and Viture as compatible with the LabOS system built on XR AI.[1] In August 2026 the XR site The Ghost Howls reported running the XR AI client in a browser on the RayNeo X3 Pro glasses.[15]

Uses

NVIDIA's launch post listed these early users and demonstrations.[1]

Organization Field Described use
Siemens Manufacturing Research into helping factory engineers wearing glasses find maintenance information, troubleshoot programmable logic controller issues, verify work and record shop-floor activity, using DGX Spark[1]
Rana (AutoBio) Life sciences LabOS, hands-free guidance for stem cell therapy and gene-editing experiments at Stanford and Princeton labs[1]
VITURE Wearables Integrated XR AI into a wearable interface for hands-free workplace guidance[1]
Surreality Lab, University of Pittsburgh Medical Center Healthcare Context-aware assistance for surgical teams, running on DGX Station[1]
Innoactive Automotive design Capturing information from design reviews, showrooms and digital twins, powered by DGX Spark[1]
Atlantic Studios Immersive media Voice-guided exploration of a scan of the Titanic wreck[1]

VITURE introduced its Helix AI safety glasses at AWE 2026 as "the first AI safety eyewear platform built on NVIDIA's XR AI solution." In the arrangement VITURE describes, NVIDIA provides the AI infrastructure and VITURE supplies the hardware and edge software. VITURE said Helix would start shipping in Q1 2027 at prices from US$600.[4]

Reception

DEVELOP3D (July 2026) and Auganix (August 2026) reported on the public beta, summarizing NVIDIA's announcement and partner examples.[16][17] In a hands-on note, The Ghost Howls author Skarredghost described a backend that runs language models locally on a workstation so that company data does not leave it. He had trouble running it under WSL on Windows because it is made for Linux, but once running it "worked really well" and answered questions about what he was looking at through the RayNeo X3 Pro "out of the box."[15]

See also

References

  1. ↑ 1.00 1.01 1.02 1.03 1.04 1.05 1.06 1.07 1.08 1.09 1.10 1.11 1.12 David Chu (2026-06-16). "Hands Free, AIs Forward: NVIDIA XR AI Brings Agents to AR Glasses". NVIDIA Blog. NVIDIA. https://blogs.nvidia.com/blog/nvidia-xr-ai/. Retrieved 2026-09-29.
  2. ↑ 2.0 2.1 2.2 2.3 2.4 2.5 2.6 2.7 2.8 Greg Barbone (2026-06-16). "Building AI Agents for AR Glasses and XR Devices with NVIDIA XR AI". NVIDIA Technical Blog. NVIDIA. https://developer.nvidia.com/blog/building-ai-agents-for-ar-glasses-and-xr-devices-with-nvidia-xr-ai/. Retrieved 2026-09-29.
  3. ↑ 3.0 3.1 3.2 3.3 3.4 3.5 "XR AI v0.1.0 Beta Release". GitHub. NVIDIA. 2026-06-16. https://github.com/NVIDIA/xr-ai/releases/tag/v0.1.0-beta. Retrieved 2026-09-29.
  4. ↑ 4.0 4.1 4.2 4.3 "VITURE Unveils Helix, the First AI Safety Glasses Built on NVIDIA's XR AI Solution, at AWE 2026". PR Newswire. VITURE. 2026-06-16. https://www.prnewswire.com/news-releases/viture-unveils-helix-the-first-ai-safety-glasses-built-on-nvidias-xr-ai-solution-at-awe-2026-302802005.html. Retrieved 2026-09-29.
  5. ↑ 5.0 5.1 5.2 "NVIDIA/xr-ai: XR AI". GitHub. NVIDIA. https://github.com/NVIDIA/xr-ai. Retrieved 2026-09-29.
  6. ↑ 6.0 6.1 6.2 "XR AI v0.2.0 Beta". GitHub. NVIDIA. 2026-09-10. https://github.com/NVIDIA/xr-ai/releases/tag/v0.2.0-beta.2. Retrieved 2026-09-29.
  7. ↑ 7.0 7.1 "Releases - NVIDIA/xr-ai". GitHub. NVIDIA. https://github.com/NVIDIA/xr-ai/releases. Retrieved 2026-09-29.
  8. ↑ Le Cong, David Smerkous, Xiaotong Wang, Di Yin, Zaixi Zhang, et al., Mengdi Wang (2025-10-16). "LabOS: The AI-XR Co-Scientist That Sees and Works With Humans". arXiv preprint 2510.14861. https://arxiv.org/abs/2510.14861. Retrieved 2026-09-29.
  9. ↑ "Release migration". XR AI documentation. NVIDIA. https://nvidia.github.io/xr-ai/v0.2.0-beta.2/reference/migrations.html. Retrieved 2026-09-29.
  10. ↑ "Architecture". XR AI documentation. NVIDIA. https://nvidia.github.io/xr-ai/v0.2.0-beta.2/overview/architecture.html. Retrieved 2026-09-29.
  11. ↑ 11.0 11.1 11.2 11.3 11.4 11.5 11.6 11.7 "AI inference servers". XR AI documentation. NVIDIA. https://nvidia.github.io/xr-ai/v0.2.0-beta.2/components/ai-services.html. Retrieved 2026-09-29.
  12. ↑ 12.0 12.1 12.2 12.3 12.4 12.5 12.6 "AI inference servers". XR AI documentation (v0.1.0-beta). NVIDIA. https://nvidia.github.io/xr-ai/v0.1.0-beta/components/ai-services.html. Retrieved 2026-09-29.
  13. ↑ "Requirements". XR AI documentation. NVIDIA. https://nvidia.github.io/xr-ai/v0.2.0-beta.2/getting_started/requirements.html. Retrieved 2026-09-29.
  14. ↑ 14.0 14.1 14.2 14.3 14.4 14.5 "Connecting clients". XR AI documentation. NVIDIA. https://nvidia.github.io/xr-ai/v0.2.0-beta.2/getting_started/clients.html. Retrieved 2026-09-29.
  15. ↑ 15.0 15.1 Skarredghost (2026-08-10). "The XR Week Peek (2026.10.08): New rumors on Steam Frame, new VR mods, and more!". The Ghost Howls. https://skarredghost.com/2026/08/10/steam-frame-vr-mods/. Retrieved 2026-09-29.
  16. ↑ "NVIDIA XR AI developer library launched". DEVELOP3D. 2026-07-10. https://develop3d.com/ai/nvidia-xr-ai-developer-library-launched/. Retrieved 2026-09-29.
  17. ↑ Sam Sprigg (2026-08-04). "NVIDIA XR AI Enters Public Beta for AR Glasses and XR Devices". Auganix. https://www.auganix.org/xr-news-nvidia-xr-ai-public-beta/. Retrieved 2026-09-29.