Toggle menu
Toggle preferences menu
Toggle personal menu
Not logged in
Your IP address will be publicly visible if you make any edits.

Scene understanding is the ability of an augmented reality (AR), mixed reality (MR) or virtual reality (VR) device to build a structured, semantic model of the real environment around the user. Where spatial mapping produces raw geometry, usually an unlabeled triangle mesh, scene understanding identifies what that geometry is: floors, walls, ceilings, doors, windows, tables, seats and other objects, represented as labeled planes, bounding volumes and simplified meshes that applications can query. Microsoft, which uses the name for a HoloLens 2 runtime and SDK, describes it as "a structured, high-level environment representation designed to make developing for environmentally aware applications intuitive".[1]

In computer vision research the term covers recognizing a scene as a whole together with the objects in it.[2] In XR products it is exposed through platform APIs: Microsoft's Scene Understanding SDK for HoloLens 2, the scene model and Scene API on Meta Quest headsets, scene reconstruction with classification and RoomPlan in Apple's ARKit, the Scene Semantics API in Google's ARCore, the scene understanding permissions and extensions of Android XR, and, since 2025, cross-vendor OpenXR extensions. Applications use the model to place virtual content on real surfaces, hide it behind real objects, run physics against the room and plan paths for virtual characters.[1][3]

Reviewed 29 September 2026. Platform API details (Microsoft, Meta, Apple, Google, Khronos OpenXR registry), press dates and quotes, and the authors, venues and DOIs of the cited papers against the cited sources. About review dates.

Definition

Microsoft's documentation separates scene understanding from spatial mapping by what each trades away. Spatial mapping is "highly accurate but less structured"; scene understanding combines that data with "new AI driven runtimes" to produce representations similar to those developers know from game engines such as Unity or from ARKit and ARCore. According to Microsoft, the core difference "is a tradeoff of maximal accuracy and latency to structure and simplicity": an application that needs the lowest latency should read the spatial mapping mesh directly, while higher-level processing can use the scene model, which also includes a snapshot of the spatial mapping mesh.[1]

Meta describes its scene model as "a comprehensive, current representation of the physical world that can be indexed and queried", and compares it to a scene graph for the physical world. Its basic element is the scene anchor, which carries geometric components, such as planes or 3D bounding boxes, and semantic labels. Meta's example is a living room organized around anchors for the floor, the ceiling, walls, a table and a couch. Scene anchors are distinct from the spatial anchors that applications place themselves.[3]

A scene model usually combines several kinds of data:

Element What it represents Examples
Labeled planes Flat, bounded surfaces with a semantic class Microsoft SceneQuads; ARKit plane anchors classified as floor, wall or table; Android XR planes labeled WALL, FLOOR, CEILING or TABLE[4][5]
Bounding volumes 3D boxes around furniture and other objects Meta labels such as COUCH, TABLE, BED, STORAGE and SCREEN[6]
Classified meshes A triangle mesh whose faces or vertices carry a class ARKit mesh classification; Android XR scene meshing labels[7][8]
Watertight or parametric room models Simplified room structure, with unscanned areas inferred Microsoft watertight meshes; Apple RoomPlan parametric output; Meta room layout[1][9][10]
Per-pixel semantics A class label for every camera pixel ARCore Scene Semantics API for outdoor scenes[11]

History

Computer vision research

In 2009 Li-Jia Li, Richard Socher and Li Fei-Fei presented a model at CVPR, under the title "Towards total scene understanding", that classifies the overall scene in an image, recognizes and segments each object and annotates the image with a list of tags. The authors wrote that, to their knowledge, it was the first model to perform all three tasks in one framework.[2]

With affordable depth cameras, the problem moved into 3D. SLAM++, presented at CVPR 2013 by Renato Salas-Moreno, Richard Newcombe, Hauke Strasdat, Paul Kelly and Andrew Davison, performed simultaneous localization and mapping at the level of objects rather than points or surfaces: as a handheld depth camera moved through a cluttered room, real-time 3D object recognition fed an explicit graph of objects. The paper included a "context-aware augmented reality" demonstration in which virtual characters navigated the mapped room and found places to sit.[12]

SemanticFusion (McCormac, Handa, Davison and Leutenegger, ICRA 2017) combined convolutional neural networks with the ElasticFusion dense SLAM system to label a 3D map from RGB-D video at about 25 Hz. It found that fusing predictions from multiple viewpoints improved labeling even when measured in 2D, compared with single-frame predictions.[13] Training such systems required labeled 3D data. ScanNet (Dai et al., CVPR 2017) supplied 2.5 million RGB-D views of 1,513 indoor scenes with camera poses, surface reconstructions and instance-level semantic annotations; the scans were captured by novices using an app on an iPad fitted with a depth camera, and the labels were crowdsourced from about 500 workers.[14] In 2019 Armeni and colleagues proposed the 3D scene graph, a single structure that holds objects, rooms and cameras and the relationships between them, built semi-automatically for entire buildings.[15]

Platform timeline

Date Platform Scene understanding feature
iOS 12 Apple ARKit Plane anchors gain a classification (wall, floor, ceiling, table, seat, window, door)[16]
24 March 2020 Apple ARKit 3.5 Scene Geometry on the LiDAR-equipped iPad Pro: a 3D mesh classified into floors, walls, ceilings, windows, doors and seats[17][18]
iOS 16 Apple RoomPlan Guided room capture producing a parametric 3D model with walls, windows, openings, doors and furniture[10]
May 2023 Google ARCore Scene Semantics API labels each pixel of outdoor camera images[19]
September 2023 Meta Quest 3 Meta announces Space Setup with automated room layout detection[20]
v64 (April 2024) Meta Quest 3 Space Setup automatically recognizes and labels furniture, doors and windows[21]
June 2025 OpenXR Khronos releases the cross-vendor spatial entities extensions, including plane tracking[22]
visionOS 26 Apple visionOS Plane and mesh classifications renamed and unified as SurfaceClassification[23]
v83 Meta Horizon OS High-Fidelity Room API with multiple floor levels, columns and slanted ceilings; Meta later archived its documentation and says the feature is no longer supported[9]

Platform implementations

Microsoft HoloLens

Microsoft's Scene Understanding runtime is supported on HoloLens 2 and not on the first HoloLens or on Windows Mixed Reality immersive headsets.[1] An application asks a SceneObserver to compute a scene around the headset within a chosen radius; the runtime does the work in a separate process, the Mixed Reality driver, and returns the result to the application. Microsoft notes that the conversion is expensive, taking seconds for a medium space of about 10 by 10 m and minutes for a large space of about 50 by 50 m, so scenes are computed only when an application requests them.[4]

A computed scene is made of SceneObjects, each with a kind: Wall, Floor, Ceiling, Platform (large horizontal surfaces such as tables and countertops), Background, World (the unlabeled spatial mapping mesh) and Unknown. Objects reference SceneQuads, bounded 2D surfaces with helper functions for placing holograms, and SceneMeshes, including watertight meshes that infer the planar room structure without clutter.[4][1] With inference enabled, quads can extend placement areas into parts of a surface the device never scanned.[1] The same functions were exposed to OpenXR through Microsoft's XR_MSFT_scene_understanding vendor extension, which the OpenXR registry now lists as obsoleted by the cross-vendor XR_EXT_spatial_entity extension.[24] Microsoft ended HoloLens 2 production in October 2024 and told UploadVR that the device would receive critical security and regression updates until 31 December 2027.[25]

Meta Quest

On Meta's headsets the scene model is produced by Space Setup, a system flow run from the headset settings, and read by apps through the Scene API, the lower-level OVRAnchor API or the Mixed Reality Utility Kit, which Meta calls the preferred way of working with scene data.[3] Meta's design guidelines for MR are published under the heading "Scene understanding".[26] Meta's supported labels cover room structure (FLOOR, CEILING, WALL_FACE, INVISIBLE_WALL_FACE, DOOR_FRAME, WINDOW_FRAME, WALL_ART), room contents (COUCH, TABLE, BED, LAMP, PLANT, SCREEN, STORAGE), a GLOBAL_MESH captured during Space Setup and a general OTHER volume.[6] The OpenXR XR_FB_scene extension records how this vocabulary changed: from extension version 3 the runtime reports TABLE instead of DESK, and from version 4 it can report INVISIBLE_WALL_FACE for a wall used to separate spaces in an open-plan home where no real wall exists.[24]

For Meta Quest 3, Meta announced Space Setup in September 2023 as a feature that "provides automated room layout detection".[20] UploadVR reported that until v64 the headset could infer only walls, floor and ceiling from its room mesh; v64 added automatic bounding boxes for doors, windows, beds, tables, sofas, storage and screens, although in UploadVR's testing the storage category often produced inaccurate bounds.[21] Later changes extended the model across a home: according to UploadVR, Quest 3 can store up to 15 scene meshes, and SDK v65 added an API that lets apps read scene meshes from multiple rooms.[27] Meta's High-Fidelity Room API, which required Quest 3 or newer with software and SDK v83, described rooms with multiple floor levels, columns and slanted ceilings, where room data had previously been limited to a single floor, ceiling and set of walls; Space Setup on Quest 3 and Meta Quest 3S captures a scene mesh and estimates the room layout from it. Meta's documentation page for the API, last updated in July 2026, is marked as archived and states that the High-Fidelity Scene feature "is no longer supported", pointing developers to the Mixed Reality Utility Kit for current scene capabilities.[9]

Quest scans are captured at a point in time. In January 2025 UploadVR reported that changes such as moved furniture do not appear until the user scans again, and that the changelog for Meta XR Core SDK v72 said Meta was "on a path to remove the user's capability to edit the space settings in 2025".[28]

Apple

In ARKit on iOS and iPadOS, plane anchors can carry a classification (ceiling, door, floor, seat, table, wall or window), a feature introduced in iOS 12.[16] ARKit 3.5, released on 24 March 2020 for the iPad Pro with a LiDAR Scanner, added what Apple called "Scene Geometry for enhanced scene understanding and object occlusion".[17] VentureBeat described it as a 3D map of a space "differentiating between floors, walls, ceilings, windows, doors, and seats".[18] When scene reconstruction is enabled, ARKit provides a polygonal mesh that estimates the shape of the environment; if plane detection is also enabled it smooths the mesh where it detects a plane, and with people occlusion it removes parts of the mesh that overlap people in the camera feed.[29] Mesh faces can be classified as ceiling, door, floor, seat, table, wall or window.[7]

RoomPlan, introduced with iOS 16, guides a person through scanning a room and uses the camera feed, LiDAR readings and trained machine-learning models to identify walls, windows, openings and doors, together with furniture and appliances such as a fireplace, bed or refrigerator. It outputs the room as parametric data and in Universal Scene Description (USD) formats.[10]

On Apple Vision Pro, visionOS 1.0 provided plane anchors with the same seven classes and a scene reconstruction provider whose mesh faces can be classified into a wider set that adds bed, cabinet, home appliance, plant, stairs and TV.[30][31] visionOS 2.0 added a room tracking provider that reports the room the wearer is currently in, and visionOS 26 renamed both classification types to a single SurfaceClassification.[32][23]

Google ARCore and Android XR

Google announced the ARCore Scene Semantics API on 10 May 2023, alongside the Streetscape Geometry and Geospatial Depth features. It uses AI to give a class label to every pixel in an outdoor scene; at launch it offered twelve classes including sky, building, tree, road, sidewalk, vehicle, person and water.[19] The ARCore reference lists twelve label values: sky, building, tree, road, sidewalk, terrain, structure, water, vehicle, object, person and unlabeled.[33] The API returns a confidence value for each pixel label, is designed for outdoor scenes in the device's default portrait orientation, and shares its list of supported devices with the ARCore Depth API.[11]

On Android XR headsets such as the Samsung Galaxy XR, ARCore for Jetpack XR reports planes with semantic labels (WALL, FLOOR, CEILING and TABLE).[5] The OpenXR extension XR_ANDROID_scene_meshing supplies meshes of physical objects split into submeshes, with per-vertex labels for other, floor, ceiling, wall and table.[8] Android XR groups these features under two runtime permissions named after scene understanding: SCENE_UNDERSTANDING_COARSE covers light estimation, projecting passthrough onto mesh surfaces, raycasts against trackables in the environment, plane tracking, object tracking and persistent anchors, while SCENE_UNDERSTANDING_FINE covers the depth texture; the scene meshing extension requires the fine permission. Google classes these as dangerous permissions that apps must declare in the manifest and request at runtime.[34][8]

OpenXR

Until 2025, scene data in OpenXR came from vendor extensions such as XR_MSFT_scene_understanding and Meta's XR_FB_scene and XR_FB_scene_capture.[24] On 10 June 2025 the Khronos Group released a set of cross-vendor spatial entities extensions: a base spatial entity extension plus extensions for plane tracking, marker tracking, spatial anchors and two for persistence.[22] UploadVR reported that Meta, Google, Pico, Varjo, Unity, Godot and Collabora pledged support.[35] The plane tracking extension defines semantic labels for floor, wall, ceiling and table, plus an uncategorized value.[24] Khronos listed "the generation and processing of mesh-based models of the user's environment" as possible future work.[22]

Platform Interface Semantic classes (as documented)
Microsoft HoloLens 2 Scene Understanding SDK Wall, Floor, Ceiling, Platform, Background, World, Unknown[4]
Meta Quest Scene API / MRUK Floor, ceiling, wall face, invisible wall face, door frame, window frame, wall art, couch, table, bed, lamp, plant, screen, storage, global mesh, other[6]
Apple iOS and iPadOS ARKit mesh classification Ceiling, door, floor, seat, table, wall, window, none[7]
Apple visionOS ARKit SurfaceClassification Bed, cabinet, ceiling, door, floor, home appliance, plant, seat, stairs, table, TV, wall, window, none[23]
Google ARCore (outdoor) Scene Semantics API Sky, building, tree, road, sidewalk, terrain, structure, water, vehicle, object, person, unlabeled[33]
Android XR ARCore for Jetpack XR planes; XR_ANDROID_scene_meshing Wall, floor, ceiling, table (mesh also has other)[5][8]
OpenXR (cross-vendor) XR_EXT_spatial_plane_tracking Floor, wall, ceiling, table, uncategorized[24]

Applications in VR and AR

Microsoft groups the uses of scene data into placement, occlusion, physics, navigation and visualization. Quads let an app find a spot on a wall or table where a hologram of a given size fits; watertight meshes make sure physics ray casts always hit a surface; and floor meshes separated by semantic class simplify building navigation meshes for virtual characters, although apps still have to project furniture onto the floor to keep paths clear of clutter.[1] Meta's guidelines recommend grounding virtual objects on physical surfaces with correct alignment and shadows, keeping visualization of the scanned surfaces to a minimum, and using semantic labels to choose which virtual objects appear on which surface types.[26] UploadVR gave examples of what furniture labels allow on Quest 3: placing a tabletop game board on the largest table in the room, replacing windows with portals, or showing the user's TV inside a fully virtual game so they do not hit it.[21] Apple lists estimating the size of areas of a room, previewing catalog furniture and bringing a scanned room into a 3D game as uses of RoomPlan output.[10]

Research systems use scene semantics to make virtual characters behave plausibly in real rooms. Lang, Liang and Yu (IEEE VR 2019) reconstructed a room with the RGB-D cameras of a mixed reality headset such as HoloLens, detected relevant objects with the Mask R-CNN detector, and optimized the position of a virtual agent with a cost function based on visibility and spatial relations; users rated its placements higher than those of alternative approaches.[36] Li, Li, Huang and Yu (ACM Transactions on Graphics, 2022) used indoor scene semantics to populate a room with virtual characters and items that act out a story, adapting their behavior to the player's actions.[37]

Privacy

A scene model describes the layout and contents of a user's home, so platforms gate it behind permissions. Meta requires apps to request a runtime spatial data permission before reading the scene model,[3] Microsoft's SDK requires an access request before a scene can be computed,[4] and Android XR treats both scene understanding permissions as dangerous.[34] The Android XR scene meshing specification states that "scene meshing data is sensitive personal information and is closely linked to personal privacy and integrity", and recommends that apps storing or transferring it ask the user for active and specific acceptance.[8]

Research

Datasets captured with consumer AR hardware have become a basis for scene understanding research. Apple's ARKitScenes (NeurIPS 2021) was described by its authors as the first RGB-D dataset captured with a widely available depth sensor, the LiDAR Scanner in Apple devices, and as "the largest indoor scene understanding data released"; it pairs mobile captures with laser-scanned depth and manually annotated 3D bounding boxes for furniture.[38]

Meta Reality Labs Research's SceneScript (ECCV 2024) represents a room as a sequence of structured language commands generated by an autoregressive encoder-decoder model, rather than as a mesh or point cloud. It was trained on Aria Synthetic Environments, a dataset of 100,000 synthetic indoor scenes rendered with the sensor characteristics of Project Aria glasses, and reported state-of-the-art results in architectural layout estimation.[39] Meta's announcement contrasted the approach with current MR headsets such as Quest 3, which build their room representation from raw camera or 3D sensor data, and said the output describes walls, ceilings and doors in a form that takes only a few bytes.[40] Model weights were made available to academic researchers in September 2024.[41]

See also

References

  1. ↑ 1.0 1.1 1.2 1.3 1.4 1.5 1.6 1.7 "Scene understanding". Microsoft Learn. Microsoft. https://learn.microsoft.com/en-us/windows/mixed-reality/design/scene-understanding. Retrieved 2026-09-29.
  2. ↑ 2.0 2.1 Li-Jia Li, Richard Socher, Li Fei-Fei (2009). "Towards total scene understanding: Classification, annotation and segmentation in an automatic framework". 2009 IEEE Conference on Computer Vision and Pattern Recognition, pp. 2036-2043. https://doi.org/10.1109/CVPR.2009.5206718. Retrieved 2026-09-29.
  3. ↑ 3.0 3.1 3.2 3.3 "Scene Overview". Meta Horizon OS Developers. Meta Platforms. https://developers.meta.com/horizon/documentation/unity/unity-scene-overview/. Retrieved 2026-09-29.
  4. ↑ 4.0 4.1 4.2 4.3 4.4 "Scene understanding SDK". Microsoft Learn. Microsoft. https://learn.microsoft.com/en-us/windows/mixed-reality/develop/unity/scene-understanding-sdk. Retrieved 2026-09-29.
  5. ↑ 5.0 5.1 5.2 "Detect planes using ARCore for Jetpack XR". Android Developers. Google. https://developer.android.com/develop/xr/jetpack-xr-sdk/arcore/planes. Retrieved 2026-09-29.
  6. ↑ 6.0 6.1 6.2 "Semantic Classification for Scene". Meta Horizon OS Developers. Meta Platforms. https://developers.meta.com/horizon/documentation/unreal/unreal-scene-supported-semantic-labels/. Retrieved 2026-09-29.
  7. ↑ 7.0 7.1 7.2 "ARMeshClassification". Apple Developer Documentation. Apple. https://developer.apple.com/documentation/arkit/armeshclassification. Retrieved 2026-09-29.
  8. ↑ 8.0 8.1 8.2 8.3 8.4 "XR_ANDROID_scene_meshing OpenXR extension". Android Developers. Google. https://developer.android.com/develop/xr/openxr/extensions/XR_ANDROID_scene_meshing. Retrieved 2026-09-29.
  9. ↑ 9.0 9.1 9.2 "High-Fidelity Room". Meta Horizon OS Developers. Meta Platforms. https://developers.meta.com/horizon/documentation/unity/unity-scene-roommesh/. Retrieved 2026-09-29.
  10. ↑ 10.0 10.1 10.2 10.3 "RoomPlan". Apple Developer Documentation. Apple. https://developer.apple.com/documentation/roomplan. Retrieved 2026-09-29.
  11. ↑ 11.0 11.1 "Understand the user's environment with the Scene Semantics API". ARCore - Google for Developers. Google. https://developers.google.com/ar/develop/scene-semantics. Retrieved 2026-09-29.
  12. ↑ Renato F. Salas-Moreno, Richard A. Newcombe, Hauke Strasdat, Paul H. J. Kelly, Andrew J. Davison (2013). "SLAM++: Simultaneous Localisation and Mapping at the Level of Objects". 2013 IEEE Conference on Computer Vision and Pattern Recognition, pp. 1352-1359. https://doi.org/10.1109/CVPR.2013.178. Retrieved 2026-09-29.
  13. ↑ John McCormac, Ankur Handa, Andrew Davison, Stefan Leutenegger (2017). "SemanticFusion: Dense 3D semantic mapping with convolutional neural networks". 2017 IEEE International Conference on Robotics and Automation (ICRA), pp. 4628-4635. https://doi.org/10.1109/ICRA.2017.7989538. Retrieved 2026-09-29.
  14. ↑ Angela Dai, Angel X. Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, Matthias Niessner (2017). "ScanNet: Richly-Annotated 3D Reconstructions of Indoor Scenes". 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 2432-2443. https://doi.org/10.1109/CVPR.2017.261. Retrieved 2026-09-29.
  15. ↑ Iro Armeni, Zhi-Yang He, JunYoung Gwak, Amir R. Zamir, Martin Fischer, Jitendra Malik, Silvio Savarese (2019). "3D Scene Graph: A Structure for Unified Semantics, 3D Space, and Camera". 2019 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 5663-5672. https://doi.org/10.1109/ICCV.2019.00576. Retrieved 2026-09-29.
  16. ↑ 16.0 16.1 "ARPlaneAnchor.Classification". Apple Developer Documentation. Apple. https://developer.apple.com/documentation/arkit/arplaneanchor/classification-swift.enum. Retrieved 2026-09-29.
  17. ↑ 17.0 17.1 "ARKit 3.5 Now Available". Apple Developer News. Apple. 2020-03-24. https://developer.apple.com/news/?id=03242020a. Retrieved 2026-09-29.
  18. ↑ 18.0 18.1 Jeremy Horwitz (2020-03-24). "Apple releases ARKit 3.5, adding Scene Geometry API and lidar support". VentureBeat. https://venturebeat.com/mobile/apple-releases-arkit-3-5-adding-scene-geometry-api-and-lidar-support. Retrieved 2026-09-29.
  19. ↑ 19.0 19.1 Eric Lai (2023-05-10). "Build transformative augmented reality experiences with new ARCore and geospatial features". Google Developers Blog. Google. https://developers.googleblog.com/2023/05/build-transformative-augmented-reality-experiences-with-new-arcore-and-geospatial-features.html. Retrieved 2026-09-29.
  20. ↑ 20.0 20.1 "Building for Mixed Reality on Meta Quest 3". Meta Horizon OS Developers Blog. Meta Platforms. 2023-09-28. https://developers.meta.com/horizon/blog/building-mixed-reality-MR-meta-quest-3-connect-developers-presence-platform/. Retrieved 2026-09-29.
  21. ↑ 21.0 21.1 21.2 David Heaney (2024-04-16). "Quest 3's Latest Update Brought Two Undocumented Features". UploadVR. https://www.uploadvr.com/quest-v64-undocumented-features-furniture-recognition-multimodal/. Retrieved 2026-09-29.
  22. ↑ 22.0 22.1 22.2 "OpenXR Spatial Entities Extensions Released for Developer Feedback". Khronos Group Blog. Khronos Group. 2025-06-10. https://www.khronos.org/blog/openxr-spatial-entities-extensions-released-for-developer-feedback. Retrieved 2026-09-29.
  23. ↑ 23.0 23.1 23.2 "SurfaceClassification". Apple Developer Documentation. Apple. https://developer.apple.com/documentation/arkit/surfaceclassification. Retrieved 2026-09-29.
  24. ↑ 24.0 24.1 24.2 24.3 24.4 "The OpenXR Specification (1.1)". Khronos OpenXR Registry. Khronos Group. https://registry.khronos.org/OpenXR/specs/1.1/html/xrspec.html. Retrieved 2026-09-29.
  25. ↑ David Heaney (2024-10-01). "Microsoft Is Discontinuing HoloLens 2 As Production Ends". UploadVR. https://www.uploadvr.com/microsoft-discontinuing-hololens-2/. Retrieved 2026-09-29.
  26. ↑ 26.0 26.1 "Scene understanding". Meta Horizon OS Developers. Meta Platforms. https://developers.meta.com/horizon/design/mr-design-scene/. Retrieved 2026-09-29.
  27. ↑ David Heaney (2024-07-17). "Quest 3 Mixed Reality Apps Can Now Span Your Entire Home". UploadVR. https://www.uploadvr.com/quest-multi-room-mixed-reality-support/. Retrieved 2026-09-29.
  28. ↑ David Heaney (2025-01-28). "Meta Plans To Make Quest Scene Mesh Scans Update Automatically". UploadVR. https://www.uploadvr.com/meta-plans-quest-scene-mesh-scans-update-automatically/. Retrieved 2026-09-29.
  29. ↑ "sceneReconstruction". Apple Developer Documentation. Apple. https://developer.apple.com/documentation/arkit/arworldtrackingconfiguration/scenereconstruction. Retrieved 2026-09-29.
  30. ↑ "PlaneAnchor.Classification". Apple Developer Documentation. Apple. https://developer.apple.com/documentation/arkit/planeanchor/classification-swift.enum. Retrieved 2026-09-29.
  31. ↑ "MeshAnchor.MeshClassification". Apple Developer Documentation. Apple. https://developer.apple.com/documentation/arkit/meshanchor/meshclassification. Retrieved 2026-09-29.
  32. ↑ "RoomTrackingProvider". Apple Developer Documentation. Apple. https://developer.apple.com/documentation/arkit/roomtrackingprovider. Retrieved 2026-09-29.
  33. ↑ 33.0 33.1 "SemanticLabel". ARCore Java API Reference. Google. https://developers.google.com/ar/reference/java/com/google/ar/core/SemanticLabel. Retrieved 2026-09-29.
  34. ↑ 34.0 34.1 "Understand permissions for XR". Android Developers. Google. https://developer.android.com/develop/xr/permissions. Retrieved 2026-09-29.
  35. ↑ David Heaney (2025-06-17). "OpenXR Spatial Entities Extensions Standardize Surfaces, Markers, Anchors and Persistence". UploadVR. https://www.uploadvr.com/openxr-spatial-entities-extensions/. Retrieved 2026-09-29.
  36. ↑ Yining Lang, Wei Liang, Lap-Fai Yu (2019). "Virtual Agent Positioning Driven by Scene Semantics in Mixed Reality". 2019 IEEE Conference on Virtual Reality and 3D User Interfaces (VR), pp. 767-775. https://doi.org/10.1109/VR.2019.8798018. Retrieved 2026-09-29.
  37. ↑ Changyang Li, Wanwan Li, Haikun Huang, Lap-Fai Yu (2022). "Interactive augmented reality storytelling guided by scene semantics". ACM Transactions on Graphics, vol. 41, no. 4. https://doi.org/10.1145/3528223.3530061. Retrieved 2026-09-29.
  38. ↑ Gilad Baruch, Zhuoyuan Chen, Afshin Dehghan, et al. (2021-11). "ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data". Apple Machine Learning Research (NeurIPS 2021). Apple. https://machinelearning.apple.com/research/arkitscenes. Retrieved 2026-09-29.
  39. ↑ Armen Avetisyan, Christopher Xie, Henry Howard-Jenkins, et al. (2024). "SceneScript: Reconstructing Scenes with an Autoregressive Structured Language Model". Computer Vision - ECCV 2024, Lecture Notes in Computer Science, pp. 247-263. https://doi.org/10.1007/978-3-031-73030-6_14. Retrieved 2026-09-29.
  40. ↑ "Introducing SceneScript, a novel approach for 3D scene reconstruction". Meta AI Blog. Meta Platforms. 2024-03-20. https://ai.meta.com/blog/scenescript-3d-scene-reconstruction-reality-labs-research/. Retrieved 2026-09-29.
  41. ↑ "SceneScript". Project Aria. Meta Platforms. https://www.projectaria.com/scenescript/. Retrieved 2026-09-29.