NVIDIA XR AI
More actions
| NVIDIA XR AI | |
|---|---|
| Information | |
| Type | Software development kit |
| Industry | Enterprise extended reality, artificial intelligence |
| Developer | NVIDIA |
| Written In | Python, JavaScript, C++, Swift, Kotlin |
| Operating System | Linux (server; Ubuntu 22.04 or 24.04 recommended) |
| License | Apache License 2.0 |
| Supported Devices | Web browsers, Android, iOS and visionOS, native C++ clients; AR glasses and XR headsets |
| Release Date | June 16, 2026 (public beta) |
| Website | https://developer.nvidia.com/xr/xr-ai |
NVIDIA XR AI is an open-source developer library from NVIDIA for building multimodal AI agents that run alongside AR glasses and other extended reality (XR) devices. The agents receive live camera, microphone and device data from the wearer, pass it to GPU-hosted models and tools, and answer with speech or data in the same session.[1][2] NVIDIA released it as a public beta on June 16, 2026, during the Augmented World Expo (AWE) 2026, and publishes the code on GitHub under the Apache 2.0 license.[3][4][5]
The repository describes XR AI as "an open-source foundation for multimodal, real-time conversational AI in the NVIDIA CloudXR ecosystem." It is aimed at hands-busy enterprise work such as factory maintenance, laboratory procedures and surgical assistance, and it can optionally pair the agent with NVIDIA CloudXR remote rendering when an application needs 3D content in the headset.[5][2][1] As of September 2026 the software is still in beta; the latest release is v0.2.0-beta.2 from September 10, 2026, and NVIDIA warns that its APIs and behavior may change.[6][5]
History
A pre-release titled "[PID] Alpha 0.1 Release (Private)" was published in NVIDIA's GitHub repository on May 6, 2026, before the first public version.[7] The public beta, v0.1.0-beta, went out on June 16, 2026. That day NVIDIA posted an announcement on its corporate blog by David Chu and a technical walkthrough by Greg Barbone, who leads XR product and partner management at NVIDIA.[3][1][2] VITURE's launch release for its Helix glasses said NVIDIA would "officially unveil its XR AI solution" in David Chu's keynote at AWE 2026 in Long Beach, California, with Helix shown as the first AI glasses built on the platform.[4]
VITURE says Helix builds on an earlier collaboration between VITURE, NVIDIA, the Le Cong Lab at Stanford University and the Mengdi Wang Lab at Princeton University.[4] A preprint called "LabOS: The AI-XR Co-Scientist That Sees and Works With Humans", first posted in October 2025, describes a system in which researchers wear XR glasses that stream what they see and hear to a local GPU server (or the cloud) for real-time agent inference; the glasses then display the agent's feedback. Its authors come from Stanford University School of Medicine, Princeton University and other institutions, and six of them, including Greg Barbone, list NVIDIA as their affiliation. The paper's acknowledgements say the AR glasses hardware was procured from Viture.[8] The preprint does not use the name XR AI. NVIDIA's June 2026 announcement says Rana, an AutoBio company, is "introducing its LabOS system on NVIDIA XR AI" at the Cong Lab at Stanford and the Wang Lab at Princeton.[1]
Version 0.2.0-beta.2 (September 10, 2026) revised the agent architecture around typed AgentRuntime and VoiceAgent components and retired the in-tree MCP servers and the public Pipecat pipeline APIs. It also added lab-instrument-monitoring and tea-making samples, moved text-to-speech from Piper to Pocket-TTS, and made breaking package, API and configuration changes, including renaming the media hub to DeviceIOHub.[6][9]
| Version | Date | Notes |
|---|---|---|
| [PID] Alpha 0.1 (private) | May 6, 2026 | Private pre-release tag[7] |
| v0.1.0-beta | June 16, 2026 | First public release: XR Media Hub, agent SDK, GPU-backed VLM, LLM, STT and TTS services, six MCP servers, two end-to-end samples, CloudXR integration layer[3] |
| v0.2.0-beta.2 | September 10, 2026 | Revised agent architecture, two new samples, Pocket-TTS, versioned documentation, breaking changes including the DeviceIOHub rename[6] |
Architecture
NVIDIA's announcement groups the platform into four capabilities: ingesting video, audio, depth, pose and sensor data from AR and XR devices; connecting agents to tools such as NVIDIA Metropolis, its video search and summarization (VSS) blueprint and NeMo Retriever for retrieval-augmented generation; supporting models including Nemotron reasoning models and Cosmos Reason; and agent orchestration through NeMo Agent Toolkit, with DGX Spark, DGX Station and RTX PRO systems running inference in the cloud, data center or at the edge.[1]
In the code, a client device joins a session and publishes audio, video and data to the DeviceIOHub (called the XR Media Hub in the first beta). The hub tags each stream with the participant's identity and fans the events out to agent workers. Video frames stay in shared memory, and a worker fetches pixels only when a task needs an image. Replies are routed back only to the participant who asked. LiveKit carries the client connection, but workers talk to the hub over msgpack and ZeroMQ and never address LiveKit directly. A typical voice agent chains speech recognition, a language or vision-language model and text-to-speech, and can call tools in between.[10][2] Because several clients can join one hub and several agents can watch the same streams, NVIDIA presents the design as usable for multi-user and multi-agent sessions.[2]
At launch, enterprise data and XR functions were exposed to agents through the Model Context Protocol (MCP). The first beta shipped MCP servers for visual question answering, video queries, scene rendering, OpenXR spatial information, vector math and transcripts.[2][3] The xr-render-demo sample launches the hub, a CloudXR runtime, model services and an agent that creates and moves objects in the user's space by voice; CloudXR then streams the rendered scene from GPU servers to the device. CloudXR needs a separate license and is not bundled with XR AI.[2][3]
Bundled AI services
The reference stack runs NVIDIA open models locally by default, most of them served through vLLM. Developers can instead point the language and vision models at hosted or self-hosted NVIDIA NIM endpoints or other OpenAI-compatible APIs; the v0.2 documentation notes that speech recognition and text-to-speech cannot use hosted NIM speech services and stay on local or self-hosted servers.[11][2]
| Role | v0.1.0-beta model | v0.2.0-beta.2 model |
|---|---|---|
| Vision-language model | Cosmos-Reason1-7B[12] | Cosmos3 Nano Reasoner[11] |
| Speech-to-text | parakeet-tdt-0.6b-v3[12] | parakeet-tdt-0.6b-v3[11] |
| Text-to-speech | magpie_tts_multilingual_357m; Piper voices[12] | magpie_tts_multilingual_357m; kyutai/pocket-tts[11] |
| Fast language model | Llama-3.1-Nemotron-Nano-8B-v1[12] | Llama-3.1-Nemotron-Nano-8B-v1[11] |
| Tool-calling language model | NVIDIA-Nemotron-3-Nano-30B-A3B[12] | NVIDIA-Nemotron-3-Nano-30B-A3B[11] |
| Multimodal (text and video) model | Nemotron-3-Nano-Omni-30B-A3B-Reasoning[12] | Nemotron-3-Nano-Omni-30B-A3B-Reasoning[11] |
| Embeddings | None listed[12] | llama-nemotron-embed-1b-v2[11] |
The XR render demo uses a small model for quick acknowledgments and status updates while a larger model handles reasoning and tool calls in the background.[2]
Requirements
The server side runs on Linux, with Ubuntu 22.04 or 24.04 recommended, Python 3.11 or 3.12, NVIDIA driver 580 or later, Docker 24 or later and the NVIDIA Container Toolkit. WSL2 on Windows is not officially supported. The bundled GPU profiles target two 48 GB NVIDIA Ada GPUs, one RTX PRO 6000 Blackwell workstation GPU, or a DGX Spark; the full local model stack needs about 55 GB of GPU-visible memory. Developers who use cloud-hosted models still need an NVIDIA GPU with NVENC and NVDEC on the hub machine.[13]
Supported clients
The repository includes sample clients for each platform. They share a common StreamKit design in which one backend layer is the only code that imports a LiveKit SDK.[14]
| Client | Transport | Build |
|---|---|---|
| Web | LiveKit from CDN | None[14] |
| Web-XR | Local LiveKit and CloudXR bundles | Build script[14] |
| Android | LiveKit Android | Android Studio or Gradle[14] |
| iOS and visionOS | LiveKit Swift and CloudXRKit | Xcode[14] |
| Native C++ | LiveKit C++ | CMake[14] |
The first beta listed Android XR passthrough camera and immersive session APIs as not yet wired up.[3] NVIDIA's announcement names smart glasses from Meta, Rokid and Viture as compatible with the LabOS system built on XR AI.[1] In August 2026 the XR site The Ghost Howls reported running the XR AI client in a browser on the RayNeo X3 Pro glasses.[15]
Uses
NVIDIA's launch post listed these early users and demonstrations.[1]
| Organization | Field | Described use |
|---|---|---|
| Siemens | Manufacturing | Research into helping factory engineers wearing glasses find maintenance information, troubleshoot programmable logic controller issues, verify work and record shop-floor activity, using DGX Spark[1] |
| Rana (AutoBio) | Life sciences | LabOS, hands-free guidance for stem cell therapy and gene-editing experiments at Stanford and Princeton labs[1] |
| VITURE | Wearables | Integrated XR AI into a wearable interface for hands-free workplace guidance[1] |
| Surreality Lab, University of Pittsburgh Medical Center | Healthcare | Context-aware assistance for surgical teams, running on DGX Station[1] |
| Innoactive | Automotive design | Capturing information from design reviews, showrooms and digital twins, powered by DGX Spark[1] |
| Atlantic Studios | Immersive media | Voice-guided exploration of a scan of the Titanic wreck[1] |
VITURE introduced its Helix AI safety glasses at AWE 2026 as "the first AI safety eyewear platform built on NVIDIA's XR AI solution." In the arrangement VITURE describes, NVIDIA provides the AI infrastructure and VITURE supplies the hardware and edge software. VITURE said Helix would start shipping in Q1 2027 at prices from US$600.[4]
Reception
DEVELOP3D (July 2026) and Auganix (August 2026) reported on the public beta, summarizing NVIDIA's announcement and partner examples.[16][17] In a hands-on note, The Ghost Howls author Skarredghost described a backend that runs language models locally on a workstation so that company data does not leave it. He had trouble running it under WSL on Windows because it is made for Linux, but once running it "worked really well" and answered questions about what he was looking at through the RayNeo X3 Pro "out of the box."[15]
See also
References
- ↑ 1.00 1.01 1.02 1.03 1.04 1.05 1.06 1.07 1.08 1.09 1.10 1.11 1.12 David Chu (2026-06-16). "Hands Free, AIs Forward: NVIDIA XR AI Brings Agents to AR Glasses". NVIDIA Blog. NVIDIA. https://blogs.nvidia.com/blog/nvidia-xr-ai/. Retrieved 2026-09-29.
- ↑ 2.0 2.1 2.2 2.3 2.4 2.5 2.6 2.7 2.8 Greg Barbone (2026-06-16). "Building AI Agents for AR Glasses and XR Devices with NVIDIA XR AI". NVIDIA Technical Blog. NVIDIA. https://developer.nvidia.com/blog/building-ai-agents-for-ar-glasses-and-xr-devices-with-nvidia-xr-ai/. Retrieved 2026-09-29.
- ↑ 3.0 3.1 3.2 3.3 3.4 3.5 "XR AI v0.1.0 Beta Release". GitHub. NVIDIA. 2026-06-16. https://github.com/NVIDIA/xr-ai/releases/tag/v0.1.0-beta. Retrieved 2026-09-29.
- ↑ 4.0 4.1 4.2 4.3 "VITURE Unveils Helix, the First AI Safety Glasses Built on NVIDIA's XR AI Solution, at AWE 2026". PR Newswire. VITURE. 2026-06-16. https://www.prnewswire.com/news-releases/viture-unveils-helix-the-first-ai-safety-glasses-built-on-nvidias-xr-ai-solution-at-awe-2026-302802005.html. Retrieved 2026-09-29.
- ↑ 5.0 5.1 5.2 "NVIDIA/xr-ai: XR AI". GitHub. NVIDIA. https://github.com/NVIDIA/xr-ai. Retrieved 2026-09-29.
- ↑ 6.0 6.1 6.2 "XR AI v0.2.0 Beta". GitHub. NVIDIA. 2026-09-10. https://github.com/NVIDIA/xr-ai/releases/tag/v0.2.0-beta.2. Retrieved 2026-09-29.
- ↑ 7.0 7.1 "Releases - NVIDIA/xr-ai". GitHub. NVIDIA. https://github.com/NVIDIA/xr-ai/releases. Retrieved 2026-09-29.
- ↑ Le Cong, David Smerkous, Xiaotong Wang, Di Yin, Zaixi Zhang, et al., Mengdi Wang (2025-10-16). "LabOS: The AI-XR Co-Scientist That Sees and Works With Humans". arXiv preprint 2510.14861. https://arxiv.org/abs/2510.14861. Retrieved 2026-09-29.
- ↑ "Release migration". XR AI documentation. NVIDIA. https://nvidia.github.io/xr-ai/v0.2.0-beta.2/reference/migrations.html. Retrieved 2026-09-29.
- ↑ "Architecture". XR AI documentation. NVIDIA. https://nvidia.github.io/xr-ai/v0.2.0-beta.2/overview/architecture.html. Retrieved 2026-09-29.
- ↑ 11.0 11.1 11.2 11.3 11.4 11.5 11.6 11.7 "AI inference servers". XR AI documentation. NVIDIA. https://nvidia.github.io/xr-ai/v0.2.0-beta.2/components/ai-services.html. Retrieved 2026-09-29.
- ↑ 12.0 12.1 12.2 12.3 12.4 12.5 12.6 "AI inference servers". XR AI documentation (v0.1.0-beta). NVIDIA. https://nvidia.github.io/xr-ai/v0.1.0-beta/components/ai-services.html. Retrieved 2026-09-29.
- ↑ "Requirements". XR AI documentation. NVIDIA. https://nvidia.github.io/xr-ai/v0.2.0-beta.2/getting_started/requirements.html. Retrieved 2026-09-29.
- ↑ 14.0 14.1 14.2 14.3 14.4 14.5 "Connecting clients". XR AI documentation. NVIDIA. https://nvidia.github.io/xr-ai/v0.2.0-beta.2/getting_started/clients.html. Retrieved 2026-09-29.
- ↑ 15.0 15.1 Skarredghost (2026-08-10). "The XR Week Peek (2026.10.08): New rumors on Steam Frame, new VR mods, and more!". The Ghost Howls. https://skarredghost.com/2026/08/10/steam-frame-vr-mods/. Retrieved 2026-09-29.
- ↑ "NVIDIA XR AI developer library launched". DEVELOP3D. 2026-07-10. https://develop3d.com/ai/nvidia-xr-ai-developer-library-launched/. Retrieved 2026-09-29.
- ↑ Sam Sprigg (2026-08-04). "NVIDIA XR AI Enters Public Beta for AR Glasses and XR Devices". Auganix. https://www.auganix.org/xr-news-nvidia-xr-ai-public-beta/. Retrieved 2026-09-29.