Depth map
More actions
A depth map is an image in which each pixel stores the distance from a camera or viewpoint to the surface seen through that pixel, instead of a color or brightness value. Microsoft's documentation for the Azure Kinect depth camera defines it as "a set of Z-coordinate values for every pixel of the image, measured in units of millimeters".[1] Platform APIs also call it a depth image: Google's ARCore returns a "depth image" for each camera frame, and Apple's ARKit exposes a depthMap buffer.[2][3]
Depth maps come from two broad sources. Rendering engines produce one for every frame as the depth buffer (z-buffer) used to resolve which surfaces are visible, and cameras and algorithms estimate one for the real world through stereo matching, active depth sensors, depth from motion, or neural networks that predict depth from a single photograph. In augmented reality (AR) and mixed reality, a depth map of the real scene lets virtual objects be hidden behind real ones (occlusion) and interact physically with real surfaces.[4] In virtual reality (VR), compositors use the application's depth buffer to reproject frames for head movement, and the passthrough system Passthrough+ on the Oculus Quest used depth estimated from the headset cameras to warp the camera images to the viewpoint of the user's eyes.[5][6]
The hardware that captures real-world depth (stereo cameras, structured light, time of flight and LiDAR) is described in Depth sensing.
Definition and representation
What each pixel stores
A depth map holds one value per pixel, so it records only the first surface seen along each line of sight. The 1998 SIGGRAPH paper on layered depth images by Jonathan Shade, Steven Gortler, Li-wei He and Richard Szeliski describes this ordinary form as a "2D array of depth pixels" and extends it to a layered depth image, which stores several depth pixels along one line of sight so that surfaces hidden in the original view can be rendered when the viewpoint moves.[7]
On the major AR platforms the stored value is the distance along the camera's viewing axis, not the length of the ray to the surface. ARCore's documentation states that the value at a pixel equals the length of the line from the camera origin to the scene point "projected onto the principal axis", in other words the point's z-coordinate relative to the camera, and that it is "not the length of the ray" itself.[2] Apple describes ARKit's depth data as the distance to regions of the real world "from the plane of the camera", with depth-map values given in meters.[3] The Azure Kinect likewise reports Z-coordinate values.[1]
Units, encodings and confidence
Depth maps are stored in different numeric formats depending on the source:
| Source | Encoding | Notes |
|---|---|---|
| ARCore Depth API | 16-bit unsigned integers in millimeters[2] | Full Depth API gives a smoothed value for every pixel; the Raw Depth API adds a confidence image and leaves some pixels without an estimate[8] |
| ARKit scene depth | Depth in meters; Apple's documentation gives 32-bit floating point (kCVPixelFormatType_DepthFloat32) as an example format[9] |
Paired with a confidence map whose values are low, medium or high[10] |
| Azure Kinect | Z-coordinate per pixel in millimeters[1] | Invalid pixels are given a depth value of 0[1] |
| WebXR Depth Sensing Module | 16-bit unsigned integers ("luminance-alpha" or "unsigned-short") or "float32"; raw values are multiplied by rawValueToMeters to give meters[11] |
Applications may request "raw" or "smooth" depth[11] |
Meta XR_META_environment_depth (OpenXR) |
Follows OpenGL projection-matrix conventions; near and far plane values are supplied to convert pixels to metric distance[12] | Each depth image carries two views, each with its own field of view and pose[12] |
Google XR_ANDROID_depth_texture (Android XR) |
Depth and confidence images at 80x80, 160x160 or 320x320 pixels[13] | Exposes raw and smooth depth[13] |
Measured depth can be noisy or missing, and several platforms attach a per-pixel confidence value. Apple states that ARKit is less confident about the LiDAR Scanner's measurements on surfaces that are highly reflective or absorb much of the light, and its confidence map records that confidence for every depth pixel.[14] In ARCore's Raw Depth API, pixels with no depth estimate have a confidence of zero, and confidence tends to be higher in textured areas.[8] The Azure Kinect invalidates pixels for reasons including a saturated or weak infrared signal and multi-path interference.[1] Platforms also offer temporally smoothed variants: ARKit's smoothedSceneDepth smooths the depth data over time to lessen its frame-to-frame change.[15]
Disparity maps
Stereo algorithms often output a disparity map rather than a depth map. Disparity is the horizontal shift of a scene point between two rectified images, and in their 2002 survey of stereo algorithms Daniel Scharstein and Richard Szeliski note that in computer vision "disparity is often treated as synonymous with inverse depth".[16] The Middlebury stereo benchmark gives the conversion for its 2014 datasets as Z = baseline x f / (d + doffs), with depth Z and the camera baseline in millimeters, focal length f and disparity d in pixels, and doffs the horizontal offset between the two principal points.[17]
Depth buffers in rendering
The depth map produced by a graphics pipeline is generally not linear in distance. An NVIDIA developer article explains that GPU depth buffers "don't typically store a linear representation" of distance; the stored value is a linear remapping of 1/z, where z is the distance along the view axis; 1/z fits perspective projection and is linear in screen space, so it can be interpolated across a triangle during rasterization.[18] The same article describes reversed-Z, which maps the near plane to 1 and the far plane to 0 so that floating-point precision is spread more evenly with distance, and traces the technique to at least a SIGGRAPH 1999 paper by Eugene Lapidous and Guofang Jiao.[18] Oculus engineers make the same point for VR runtimes: raw depth-buffer values "do not directly represent the distance of a given pixel", so an application that hands its depth buffer to the compositor must also describe the projection used to create it.[5]
How depth maps are produced
Rendering: the z-buffer
In real-time 3D graphics, a depth map is a by-product of drawing each frame. Edwin Catmull's 1974 PhD thesis at the University of Utah introduced z-buffering, which the ACM's 2019 Turing Award announcement describes as managing "image depth coordinates" in computer graphics; the ACM notes that Wolfgang Strasser described the technique at the same time.[19] Lance Williams summarized the method in 1978: the z-buffer "resolves the visible surfaces in a scene by storing depth (Z) values at each point in the picture", and each new object's Z values are compared with the stored ones to decide visibility.[20]
Williams's paper also introduced the use of a second depth map rendered from a light's point of view: a view of the scene is first rendered from the light source storing only Z values, then each point in the observer's view is transformed into the light's view and tested against that depth map, and points not visible to the light are shaded as being in shadow. He reported a cost of roughly twice that of rendering the scene without shadows and named quantization and aliasing as the main drawbacks.[20]
Stereo matching
Passive stereo recovers depth by matching pixels between two images taken a known distance apart, and the output is a dense disparity map with one estimate per pixel. Scharstein and Szeliski's taxonomy divides most dense stereo algorithms into four steps: matching cost computation, cost aggregation, disparity computation or optimization, and disparity refinement.[16] Their evaluation compares algorithms against data sets with ground-truth disparity maps, and the Middlebury stereo site continues to publish such data sets.[16][17]
Stereo matching on a VR headset has tight power limits. In the Passthrough+ system for the original Oculus Quest, Facebook researchers re-purposed the hardware video encoder of the Qualcomm Snapdragon 835 to compute motion vectors between the left and right camera images, converted those correspondences into depth values, then densified and temporally filtered the result. They reported temporally stable passthrough rendering at 72 Hz using 200 mW.[6]
Active depth sensors
Active sensors emit their own light and produce depth maps directly. The first Microsoft Kinect used a structured-light approach, while the later Kinect One was based on the time-of-flight principle; Hamed Sarbolandi, Damien Lefloch and Andreas Kolb compared the range data of the two devices in a 2015 study.[21] The Azure Kinect's depth camera uses amplitude-modulated continuous-wave time of flight: it illuminates the scene with modulated near-infrared light, measures the travel time indirectly and processes the measurements into a depth map, alongside a "clean IR" image whose pixel values are proportional to the returned light.[1]
Sensor output is often fused with color images. Apple engineers explained at WWDC 2020 that ARKit combines the RGB image from the wide-angle camera with LiDAR depth readings using machine learning to create a dense depth map, delivered 60 times per second with each AR frame at a lower resolution than the camera image but with the same aspect ratio.[22]
Depth from motion
A single moving camera can also produce a depth map by comparing frames taken from different positions. Julien Valentin and colleagues at Google described a depth-from-motion pipeline in ACM Transactions on Graphics in 2018 that uses only a phone's existing monocular color camera and computes low-latency dense depth maps on a single CPU core of a range of medium- to high-end phones.[23] ARCore's Depth API uses depth-from-motion algorithms: it takes images from different angles as the user moves the phone, selectively uses machine learning so depth can be computed even with little motion, and merges data from a hardware sensor such as a time-of-flight camera when the device has one.[4]
Monocular depth estimation
Neural networks can predict a depth map from a single image. David Eigen, Christian Puhrsch and Rob Fergus presented a deep-network approach at NIPS 2014, using one network stack for a coarse global prediction and a second to refine it locally. They noted that the task is inherently ambiguous because of the unknown overall scale, and trained with a scale-invariant error.[24]
Later models were trained on larger and more varied data and evaluated on data sets not seen during training:
| Model | Authors and venue | Notes |
|---|---|---|
| MiDaS | René Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, Vladlen Koltun; IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022[25] | Training objective invariant to changes in depth range and scale; the original paper combined five data sets, including 3D films[25]. Later MiDaS models, up to the 3.1 release, were trained on up to 12 data sets and compute a relative (not metric) depth map[26] |
| DPT | René Ranftl, Alexey Bochkovskiy, Vladlen Koltun; 2021[27] | Vision-transformer backbone; the authors report up to 28% relative improvement in monocular depth estimation over a fully convolutional network[27] |
| Depth Anything | Lihe Yang et al.; CVPR 2024[28] | Scales training with about 62 million automatically annotated unlabeled images[28] |
| Depth Anything V2 | Lihe Yang et al.; NeurIPS 2024[29] | Trained from 595K synthetic labeled images and more than 62M real unlabeled images; models range from 25M to 1.3B parameters[29][30] |
| Depth Pro | Aleksei Bochkovskii et al.; ICLR 2025[31] | Predicts metric depth with absolute scale without camera intrinsics; the authors report a 2.25-megapixel depth map in 0.3 seconds on a standard GPU; code and weights were released in Apple's GitHub repository[31] |
Relative models such as MiDaS predict the ordering and shape of depth but not its absolute scale; the MiDaS project points users who need metric depth to ZoeDepth, which adds a metric depth module to MiDaS.[26]
History
| Year | Development |
|---|---|
| 1974 | Edwin Catmull's University of Utah PhD thesis introduces z-buffering; Wolfgang Strasser describes it at the same time[19] |
| 1978 | Lance Williams uses a depth map rendered from a light source to cast shadows[20] |
| 1998 | Shade, Gortler, He and Szeliski introduce layered depth images for image-based rendering[7] |
| 2002 | Scharstein and Szeliski publish their taxonomy and evaluation of dense two-frame stereo algorithms[16] |
| 2004 | Christoph Fehn describes depth-image-based rendering for 3D television within the European ATTEST project[32] |
| 2011 | KinectFusion fuses Kinect depth data into a single surface model in real time[33] |
| 2014 | Eigen, Puhrsch and Fergus predict depth maps from single images with deep networks[24] |
| 2017 | Oculus Dash lets PC VR apps submit depth buffers to the runtime through the ovrLayerEyeFovDepth layer[5]
|
| 2019 (April) | Oculus releases ASW 2.0 with depth-based Positional Timewarp for Rift[34] |
| 2019 (December) | Google previews the ARCore Depth API[35] |
| 2020 (June) | ARCore 1.18 makes the Depth API generally available; Apple announces ARKit 4 with a LiDAR-based Depth API[36][37] |
| 2023 (October) | Meta introduces the Depth API for Meta Quest 3[38] |
| 2026 (August) | The W3C publishes a new Working Draft of the WebXR Depth Sensing Module[11] |
Applications in VR and AR
Occlusion and interaction in AR
When Google announced the ARCore Depth API in December 2019, it described occlusion as "the ability for digital objects to accurately appear in front of or behind real world objects" and began rolling it out in Scene Viewer, the tool behind AR in Google Search, to more than 200 million ARCore-enabled Android devices; Houzz added it to its "View in My Room 3D" feature.[35] When the API became generally available in ARCore 1.18 in June 2020, Snap Inc. used it in Snapchat Lenses and offered a Depth API template to Lens creators.[36] ARCore's documentation lists occlusion, scene effects such as virtual snow or fog, distance measurement and depth-of-field effects, collision and physics interactions, and hit tests that work on non-planar and low-texture surfaces as uses, and says depth is most accurate between about 0.5 and 5 meters while textureless surfaces such as white walls give imprecise results.[4]
Ruofei Du and colleagues built a set of ARCore Depth API interactions called DepthLab, presented at UIST 2020. The open-source library groups depth use into localized depth, surface depth and dense depth, and implements geometry-aware rendering (occlusion and shadows), surface interaction (physics-based collisions and avatar path planning) and visual effects such as relighting and 3D-anchored focus and aperture effects.[39]
Apple's ARKit 4, announced with iOS 14 on 22 June 2020, added a Depth API that uses the LiDAR Scanner on iPad Pro to provide per-pixel depth; MacRumors reported that, combined with 3D mesh data, it makes virtual object occlusion more realistic.[37] At WWDC 2020 Apple showed depth pixels being unprojected into a 3D point cloud and accumulated across frames, with low-confidence pixels filtered out.[22]
On headsets, Meta's Depth API for Quest provides "a real-time depth map that represents the physical environment's depth as it's seen from the user's point-of-view", which supports dynamic occlusion behind fast-moving objects.[38] Meta's documentation contrasts it with the Scene Model, which supports room-scale occlusion but cannot handle objects moving within the user's view, such as hands, limbs, other people and pets.[12] The API requires a Meta Quest 3 or Meta Quest 3S, needs passthrough to be running, and can remove the user's hands from the depth map and fill in estimated background depth so that apps can use tracked hand models for hand occlusion instead.[40][12] Meta warns that real objects closer than 0.2 meters to the headset produce unreliable depth estimates.[12] Google's XR_ANDROID_depth_texture extension for Android XR similarly lets applications request depth maps of the environment around the headset for occlusion, hit tests and other tasks that need scene geometry.[13]
Reprojection and compositing in VR
Some VR runtimes use the application's rendered depth buffer to correct a frame for head motion that happened after it was rendered. Oculus introduced the ovrLayerEyeFovDepth layer type with Oculus Dash in 2017, to help with depth composition; the depth buffer produced an X-ray effect where app content intersected Dash content. The same data later drove Positional Timewarp, which corrects for head translation and parallax based on distance; if an app submits no depth, the Oculus runtime falls back to the earlier ASW 1.0 method.[5] Road to VR reported that ASW 2.0 combined Asynchronous Spacewarp with Positional Timewarp on Rift and quoted Oculus as saying that depth-based reprojection "works more accurately for lower frame rates than ASW 1.0 extrapolation".[34]
OpenXR standardizes the submission of depth. The XR_KHR_composition_layer_depth extension lets an application attach depth images to its projection layers, and "the XR runtime may use this information to perform more accurate reprojections taking depth into account"; the application supplies the window-space depth range and the near and far plane distances in meters.[41] Microsoft's OpenXR guidance for HoloLens 2 tells developers to always submit the depth buffer this way because hardware depth reprojection improves hologram stability, and recommends a narrow depth range and reversed-Z.[42] Vendor extensions also use depth for composition: XR_FB_composition_layer_depth_test enables depth-tested layer composition in the compositor, and the Varjo extension XR_VARJO_environment_depth_estimation lets the runtime's compositor estimate the depth of the real environment so that layers submitted with depth information are properly occluded.[41]
Passthrough
A headset's cameras are not located at the user's eyes, and the Passthrough+ system on the Oculus Quest used depth to compensate: it computed a coarse 3D scene proxy from the headset's stereo camera pair, densified and filtered the depth, and warped the camera images to the viewpoint of the user's eyes at display rate, textured stereoscopically. Oculus shipped it as a feature that let users see their surroundings while marking out a play area.[6] On Quest 3 and Quest 3S, Meta's Depth API documentation states that passthrough is required for receiving environment depth.[40]
3D reconstruction
Successive depth maps can be fused into a 3D model of a room. KinectFusion, presented at ISMAR 2011 by Richard Newcombe and colleagues, fused all the depth data streamed from a Kinect into a single global implicit surface model in real time, and tracked the sensor by aligning each live depth frame to that model with a coarse-to-fine iterative closest point algorithm.[33] ARKit's scene geometry mesh is built in a similar way: Apple says depth data from multiple frames is aggregated and processed to construct a 3D mesh.[22] See Point cloud and Scene understanding for these downstream representations.
View synthesis, 3D photos and video
A color image with a depth map can be re-rendered from nearby viewpoints. Christoph Fehn's 2004 SPIE paper, part of the European ATTEST ("Advanced Three-Dimensional Television System Technologies") project, proposed broadcasting monoscopic color video with associated per-pixel depth so that one or more "virtual" views of the scene could be synthesized in real time at the receiver, such as a 3D-TV set-top box, by depth-image-based rendering (DIBR).[32]
Johannes Kopf and colleagues' "One Shot 3D Photography" (ACM Transactions on Graphics, 2020) applies the idea to phone photos. It estimates depth from a single image with a monocular depth network optimized for mobile devices, lifts the result to a layered depth image, synthesizes color and geometry in regions revealed by parallax with an inpainting network, and converts it to a mesh. The resulting 3D photos can be shown with parallax on phones and computers, and on VR devices where viewing also includes stereo.[43] For 360-degree video, Jingwei Huang, Zhili Chen, Duygu Ceylan and Hailin Jin presented a method at IEEE VR 2017 that adapts structure-from-motion and dense reconstruction to monoscopic 360-degree video and warps the frames to synthesize a view for each eye as the headset rotates and translates, giving playback with six degrees of freedom at more than 120 frames per second.[44]
Privacy
A depth map of a user's surroundings reveals the geometry of their home or workplace, and platform specifications treat it as sensitive data. The W3C WebXR Depth Sensing Module notes that a depth buffer with high enough resolution and precision could let websites learn more than users are comfortable with, says user agents should seek user consent before enabling depth sensing, and suggests limiting the information exposed.[11] Google's Android XR extension exposes a downsampled depth texture to mitigate personally identifiable information concerns and requires the android.permission.SCENE_UNDERSTANDING_FINE permission, which Android treats as a dangerous permission requested at runtime.[13] On Meta Quest, apps must be granted the spatial data permission (com.oculus.permission.USE_SCENE) before calling the depth functions.[12]
See also
References
- ↑ 1.0 1.1 1.2 1.3 1.4 1.5 "Azure Kinect DK depth camera". Microsoft Learn. Microsoft. https://learn.microsoft.com/en-us/previous-versions/azure/kinect-dk/depth-camera. Retrieved 2026-10-06.
- ↑ 2.0 2.1 2.2 "Use Depth in your Android app". ARCore, Google for Developers. Google. https://developers.google.com/ar/develop/java/depth/developer-guide. Retrieved 2026-10-06.
- ↑ 3.0 3.1 "ARDepthData". Apple Developer Documentation. Apple. https://developer.apple.com/documentation/arkit/ardepthdata. Retrieved 2026-10-06.
- ↑ 4.0 4.1 4.2 "Depth adds realism". ARCore, Google for Developers. Google. https://developers.google.com/ar/develop/depth. Retrieved 2026-10-06.
- ↑ 5.0 5.1 5.2 5.3 Volga Aksoy, Dean Beeler (2019-08-09). "Developer Guide to ASW 2.0". Meta Horizon OS Developers Blog. Meta. https://developers.meta.com/horizon/blog/developer-guide-to-asw-20/. Retrieved 2026-10-06.
- ↑ 6.0 6.1 6.2 Gaurav Chaurasia, Arthur Nieuwoudt, Alexandru-Eugen Ichim, Richard Szeliski, Alexander Sorkine-Hornung (2020-05). "Passthrough+: Real-time Stereoscopic View Synthesis for Mobile Mixed Reality". Proceedings of the ACM on Computer Graphics and Interactive Techniques, vol. 3, no. 1, article 7. doi:10.1145/3384540. https://doi.org/10.1145/3384540. Retrieved 2026-10-06.
- ↑ 7.0 7.1 Jonathan Shade, Steven Gortler, Li-wei He, Richard Szeliski (1998). "Layered Depth Images". Proceedings of SIGGRAPH 98, pp. 231-242. ACM. doi:10.1145/280814.280882. https://doi.org/10.1145/280814.280882. Retrieved 2026-10-06.
- ↑ 8.0 8.1 "Use Raw Depth in your Android app". ARCore, Google for Developers. Google. https://developers.google.com/ar/develop/java/depth/raw-depth. Retrieved 2026-10-06.
- ↑ "depthMap". Apple Developer Documentation. Apple. https://developer.apple.com/documentation/arkit/ardepthdata/depthmap. Retrieved 2026-10-06.
- ↑ "ARConfidenceLevel". Apple Developer Documentation. Apple. https://developer.apple.com/documentation/arkit/arconfidencelevel. Retrieved 2026-10-06.
- ↑ 11.0 11.1 11.2 11.3 Alex Cooper (editor) (2026-08-25). "WebXR Depth Sensing Module (W3C Working Draft, 25 August 2026)". World Wide Web Consortium. W3C. https://www.w3.org/TR/webxr-depth-sensing-1/. Retrieved 2026-10-06.
- ↑ 12.0 12.1 12.2 12.3 12.4 12.5 "OpenXR Depth API Overview". Meta Horizon OS Developers. Meta. https://developers.meta.com/horizon/documentation/native/android/mobile-depth/. Retrieved 2026-10-06.
- ↑ 13.0 13.1 13.2 13.3 "XR_ANDROID_depth_texture". Android XR for OpenXR, Android Developers. Google. https://developer.android.com/develop/xr/openxr/extensions/XR_ANDROID_depth_texture. Retrieved 2026-10-06.
- ↑ "confidenceMap". Apple Developer Documentation. Apple. https://developer.apple.com/documentation/arkit/ardepthdata/confidencemap. Retrieved 2026-10-06.
- ↑ "smoothedSceneDepth". Apple Developer Documentation. Apple. https://developer.apple.com/documentation/arkit/arframe/smoothedscenedepth. Retrieved 2026-10-06.
- ↑ 16.0 16.1 16.2 16.3 Daniel Scharstein, Richard Szeliski (2002-04). "A Taxonomy and Evaluation of Dense Two-Frame Stereo Correspondence Algorithms". International Journal of Computer Vision, vol. 47, no. 1-3, pp. 7-42. doi:10.1023/A:1014573219977. https://vision.middlebury.edu/stereo/taxonomy-IJCV.pdf. Retrieved 2026-10-06.
- ↑ 17.0 17.1 Daniel Scharstein. "2014 Stereo Datasets". Middlebury Stereo Vision. Middlebury College. https://vision.middlebury.edu/stereo/data/scenes2014/. Retrieved 2026-10-06.
- ↑ 18.0 18.1 "Depth Precision Visualized". NVIDIA Developer. NVIDIA. https://developer.nvidia.com/content/depth-precision-visualized. Retrieved 2026-10-06.
- ↑ 19.0 19.1 "Pioneers of Modern Computer Graphics Recognized with ACM A.M. Turing Award". Association for Computing Machinery. ACM. https://web.archive.org/web/20250123200155/https://awards.acm.org/about/2019-turing. Retrieved 2026-10-06.
- ↑ 20.0 20.1 20.2 Lance Williams (1978-08). "Casting Curved Shadows on Curved Surfaces". Proceedings of SIGGRAPH 78, pp. 270-274. ACM. doi:10.1145/800248.807402. https://cseweb.ucsd.edu/~ravir/274/15/papers/p270-williams.pdf. Retrieved 2026-10-06.
- ↑ Hamed Sarbolandi, Damien Lefloch, Andreas Kolb (2015-10). "Kinect Range Sensing: Structured-Light versus Time-of-Flight Kinect". Computer Vision and Image Understanding, vol. 139, pp. 1-20. doi:10.1016/j.cviu.2015.05.006. https://arxiv.org/abs/1505.05459. Retrieved 2026-10-06.
- ↑ 22.0 22.1 22.2 "Explore ARKit 4". WWDC20, Apple Developer. Apple. 2020-06. https://developer.apple.com/videos/play/wwdc2020/10611/. Retrieved 2026-10-06.
- ↑ Julien Valentin, Adarsh Kowdle, Jonathan T. Barron, Neal Wadhwa et al. (2018). "Depth from motion for smartphone AR". ACM Transactions on Graphics, vol. 37, no. 6. Google Research. doi:10.1145/3272127.3275041. https://research.google/pubs/depth-from-motion-for-smartphone-ar/. Retrieved 2026-10-06.
- ↑ 24.0 24.1 David Eigen, Christian Puhrsch, Rob Fergus (2014). "Depth Map Prediction from a Single Image using a Multi-Scale Deep Network". Advances in Neural Information Processing Systems 27 (NIPS 2014). https://proceedings.neurips.cc/paper_files/paper/2014/hash/91c56ce4a249fae5419b90cba831e303-Abstract.html. Retrieved 2026-10-06.
- ↑ 25.0 25.1 René Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, Vladlen Koltun (2022). "Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-shot Cross-dataset Transfer". IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 3, pp. 1623-1637. doi:10.1109/TPAMI.2020.3019967. https://arxiv.org/abs/1907.01341. Retrieved 2026-10-06.
- ↑ 26.0 26.1 "MiDaS: robust monocular depth estimation". GitHub. Intel ISL. https://github.com/isl-org/MiDaS. Retrieved 2026-10-06.
- ↑ 27.0 27.1 René Ranftl, Alexey Bochkovskiy, Vladlen Koltun (2021-03-24). "Vision Transformers for Dense Prediction". arXiv. https://arxiv.org/abs/2103.13413. Retrieved 2026-10-06.
- ↑ 28.0 28.1 Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, Hengshuang Zhao (2024-01-19). "Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data". arXiv (CVPR 2024). https://arxiv.org/abs/2401.10891. Retrieved 2026-10-06.
- ↑ 29.0 29.1 Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xiaogang Xu, Jiashi Feng, Hengshuang Zhao. "Depth Anything V2". Depth Anything V2 project page. https://depth-anything-v2.github.io/. Retrieved 2026-10-06.
- ↑ Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xiaogang Xu, Jiashi Feng, Hengshuang Zhao (2024-06-13). "Depth Anything V2". arXiv. https://arxiv.org/abs/2406.09414. Retrieved 2026-10-06.
- ↑ 31.0 31.1 Aleksei Bochkovskii, Amaël Delaunoy, Hugo Germain, Marcel Santos, Yichao Zhou, Stephan R. Richter, Vladlen Koltun (2024-10-02). "Depth Pro: Sharp Monocular Metric Depth in Less Than a Second". arXiv (ICLR 2025). https://arxiv.org/abs/2410.02073. Retrieved 2026-10-06.
- ↑ 32.0 32.1 Christoph Fehn (2004-05-21). "Depth-image-based rendering (DIBR), compression, and transmission for a new approach on 3D-TV". Proceedings of SPIE, vol. 5291, Stereoscopic Displays and Virtual Reality Systems XI, pp. 93-104. SPIE. doi:10.1117/12.524762. https://web.archive.org/web/20250703053734/https://www.spiedigitallibrary.org/conference-proceedings-of-spie/5291/1/Depth-image-based-rendering-DIBR-compression-and-transmission-for-a/10.1117/12.524762.short. Retrieved 2026-10-06.
- ↑ 33.0 33.1 Richard A. Newcombe, Shahram Izadi, Otmar Hilliges, David Molyneaux, David Kim, Andrew J. Davison, Pushmeet Kohli, Jamie Shotton, Steve Hodges, Andrew Fitzgibbon (2011-10). "KinectFusion: real-time dense surface mapping and tracking". 10th IEEE International Symposium on Mixed and Augmented Reality (ISMAR 2011), pp. 127-136. IEEE. doi:10.1109/ISMAR.2011.6092378. https://research.lancaster-university.uk/en/publications/kinectfusion-real-time-dense-surface-mapping-and-tracking/. Retrieved 2026-10-06.
- ↑ 34.0 34.1 Ben Lang (2019-04-04). "Oculus Launches ASW 2.0 with Positional Timewarp to Reduce Latency, Improve Performance". Road to VR. https://www.roadtovr.com/oculus-launches-asw-2-0-asynchronous-spacewarp/. Retrieved 2026-10-06.
- ↑ 35.0 35.1 Shahram Izadi (2019-12-09). "Blending Realities with the ARCore Depth API". Google Developers Blog. Google. https://developers.googleblog.com/2019/12/blending-realities-with-arcore-depth-api.html. Retrieved 2026-10-06.
- ↑ 36.0 36.1 Rajat Paharia (2020-06-25). "A new wave of AR Realism with the ARCore Depth API". Google Developers Blog. Google. https://developers.googleblog.com/2020/06/a-new-wave-of-ar-realism-with-arcore-depth-api.html. Retrieved 2026-10-06.
- ↑ 37.0 37.1 Hartley Charlton (2020-06-22). "Apple Announces ARKit 4 with Location Anchors, Depth API, and Improved Face Tracking". MacRumors. https://www.macrumors.com/2020/06/22/apple-announces-arkit-4/. Retrieved 2026-10-06.
- ↑ 38.0 38.1 "Build Believable Mixed Reality Experiences with Mesh API and Depth API". Meta Horizon OS Developers Blog. Meta. 2023-10-12. https://developers.meta.com/horizon/blog/mesh-depth-api-meta-quest-3-developers-mixed-reality/. Retrieved 2026-10-06.
- ↑ Ruofei Du, Eric Turner, Maksym Dzitsiuk, Luca Prasso, Ivo Duarte, Jason Dourgarian, Joao Afonso, Jose Pascoal, Josh Gladstone, Nuno Cruces, Shahram Izadi, Adarsh Kowdle, Konstantine Tsotsos, David Kim (2020-10). "DepthLab: Real-Time 3D Interaction With Depth Maps for Mobile Augmented Reality". Proceedings of the 33rd Annual ACM Symposium on User Interface Software and Technology (UIST 2020), pp. 829-843. doi:10.1145/3379337.3415881. https://augmentedperception.github.io/depthlab/. Retrieved 2026-10-06.
- ↑ 40.0 40.1 "Depth API Overview". Meta Horizon OS Developers. Meta. 2026-04-29. https://developers.meta.com/horizon/documentation/unity/unity-depthapi-overview/. Retrieved 2026-10-06.
- ↑ 41.0 41.1 "The OpenXR 1.1.63 Specification (with all registered extensions)". Khronos Registry. Khronos Group. https://registry.khronos.org/OpenXR/specs/1.1/html/xrspec.html. Retrieved 2026-10-06.
- ↑ "OpenXR best practices". Microsoft Learn. Microsoft. https://learn.microsoft.com/en-us/windows/mixed-reality/develop/native/openxr-best-practices. Retrieved 2026-10-06.
- ↑ Johannes Kopf, Kevin Matzen, Suhib Alsisan, Ocean Quigley, Francis Ge, Yangming Chong, Josh Patterson, Jan-Michael Frahm, Shu Wu, Matthew Yu, Peizhao Zhang, Zijian He, Peter Vajda, Ayush Saraf, Michael Cohen (2020). "One Shot 3D Photography". ACM Transactions on Graphics, vol. 39, no. 4. doi:10.1145/3386569.3392420. https://arxiv.org/abs/2008.12298. Retrieved 2026-10-06.
- ↑ Jingwei Huang, Zhili Chen, Duygu Ceylan, Hailin Jin (2017). "6-DOF VR videos with a single 360-camera". 2017 IEEE Virtual Reality (VR), pp. 37-44. IEEE. doi:10.1109/VR.2017.7892229. https://doi.org/10.1109/VR.2017.7892229. Retrieved 2026-10-06.