Depth buffer
More actions
A depth buffer, also called a z-buffer, is a block of memory in a 3D graphics system that holds a depth value for every pixel (or sample) of the image being rendered. When a new fragment of a polygon is drawn, its depth is compared with the value already stored at that position, and the fragment is kept only if it passes the comparison, which by default means that it is closer to the viewer. This method, known as z-buffering or depth buffering, works out which surfaces are visible without sorting the scene beforehand. Microsoft's Direct3D documentation describes the depth buffer as the surface that stores the information telling Direct3D "how deep each visible pixel is in the scene".[1][2]
The technique is credited to two 1974 doctoral theses written independently: Edwin Catmull's at the University of Utah and Wolfgang Straßer's at TU Berlin. Microsoft's Direct3D documentation notes that nearly all accelerators support z-buffering.[3][4][1] In virtual reality (VR) and augmented reality (AR), the depth buffer has a second job after the frame is drawn. Headset runtimes can read the application's depth buffer to reproject a finished frame to a newer head position, correcting for head translation as well as rotation, to compose several layers in the right depth order, and to synthesize extra frames when an application runs below the display rate. OpenXR and other XR interfaces define ways for applications to hand their depth buffers to the compositor for these purposes.[5][6]
This article covers the rendering buffer and its uses. Depth images of the real world captured by sensors or estimated by cameras are covered in Depth map.
How it works
The depth test
At the start of a frame, every entry in the depth buffer is cleared to the largest possible depth value, and the color buffer is cleared to the background. As each polygon is rasterized, the depth of the polygon at each pixel it covers is compared with the stored value. In Microsoft's description of the default behavior, if the polygon's depth is smaller, it is written into the depth buffer and the polygon's color is written to the same pixel of the render target; if it is larger, the stored values are left unchanged and the next polygon is tested.[1] Lance Williams summarized the method in 1978: the z-buffer "resolves the visible surfaces in a scene by storing depth (Z) values at each point in the picture", and as objects are rendered their Z values are compared with the stored values to determine visibility.[2]
Modern graphics interfaces make each part of this test configurable. In Vulkan, the depth test compares the depth value already in the depth/stencil attachment at each sample's framebuffer coordinates with the depth of the incoming sample; if no depth attachment is bound, the test is skipped. An application chooses whether depth testing is enabled, whether passing fragments also write their depth, and which comparison operator is used.[7] Direct3D 9 likewise exposes the comparison through the D3DRS_ZFUNC render state, and Microsoft notes that on some hardware changing the comparison function may disable hierarchical z testing.[1]
Because the test is done one pixel at a time, objects can be drawn in any order. Williams wrote that objects "do not have to be sorted beforehand, so indefinitely complex scenes can be handled", and that the z-buffer was the only visible surface algorithm whose cost grew only linearly with the average depth complexity of the scene, that is, with the total screen area of every surface rendered, whether visible in the final image or not. He added that the algorithm is "extremely general and quite simple to implement but requires substantial memory".[2]
What is stored
The value in a depth buffer is usually not the distance to the surface. Microsoft's Direct3D 9 documentation explains that a depth buffer can store either a point's z coordinate or its homogeneous w coordinate from its position in projection space; a buffer that uses z values is called a z-buffer, and one that uses w values a w-buffer. Microsoft adds that nearly all accelerators support z-buffering, which makes it the most common type, while w-buffering is less widely supported in hardware.[1] An NVIDIA developer article explains that the standard depth value is a linear remapping of 1/z, the reciprocal of view-space depth. This mapping fits perspective projection and is linear in screen space, so it can be interpolated easily across a triangle; the article notes that this property also helps hierarchical z-buffers, early z-culling and depth compression.[8]
Since the stored values are not distances, any later stage that needs real distances must know how they were produced. Oculus engineers made this point for VR compositors: the raw values "do not directly represent the distance of a given pixel", so an application that submits its depth buffer must also pass the projection parameters used to create it.[6]
Precision
The 1/z mapping spends most of the available values close to the near clipping plane. Microsoft's Direct3D 9 documentation gives two examples: with a far plane to near plane distance ratio of 100, 90 percent of the depth buffer range is spent on the first 10 percent of the scene's depth range, and at a ratio of 1,000, 98 percent of the range is spent on the first 2 percent. The documentation notes that outdoor entertainment and simulation scenes often need ratios between 1,000 and 10,000, and that the uneven distribution can cause hidden surface artifacts on distant objects, especially with 16-bit depth buffers.[1] The NVIDIA article adds that moving the near plane closer makes the distribution more lopsided, while pushing the far plane to infinity has comparatively little effect.[8]
The common remedy is reversed-Z, which maps the near plane to 1 and the far plane to 0 and stores the result in a floating-point buffer. The quasi-logarithmic spacing of floating-point numbers then offsets the 1/z curve, giving precision near the camera similar to an integer buffer and much better precision farther away.[8] NVIDIA traces the idea at least to a paper by Eugene Lapidous and Guofang Jiao of Trident Microsystems at the 1999 SIGGRAPH/Eurographics Workshop on Graphics Hardware, while noting that it has probably been reinvented several times.[8] Lapidous and Jiao called their design a "complementary floating-point Z buffer", a combination of a reversed-direction Z buffer and a floating-point storage format whose non-linearities compensate for each other. They reported that it made depth errors much less dependent on eye-space distance and that it gave the best precision of the buffer types they compared at 32 bits per pixel.[9]
Surfaces that lie in the same plane, such as a shadow drawn on a wall, can produce visible depth artifacts because both have nearly the same depth. Direct3D addresses this with a depth bias added to a polygon's z values, which Microsoft describes as commonly used so that shadows display properly. The same setting reduces "shadow acne" in shadow maps, where a surface shadows itself because of small differences between depth computed in a shader and depth stored in the shadow buffer.[10]
Formats
Graphics interfaces offer depth buffers in several sizes, often packed together with an 8-bit stencil buffer. The Vulkan specification defines the following depth formats among others, and requires depth/stencil attachment support for at least one of X8_D24_UNORM_PACK32 and D32_SFLOAT and for at least one of D24_UNORM_S8_UINT and D32_SFLOAT_S8_UINT.[11]
| Vulkan format | Depth bits | Depth storage | Stencil bits |
|---|---|---|---|
| VK_FORMAT_D16_UNORM | 16 | Unsigned normalized integer | None |
| VK_FORMAT_X8_D24_UNORM_PACK32 | 24 | Unsigned normalized integer (8 unused bits) | None |
| VK_FORMAT_D32_SFLOAT | 32 | Signed floating point | None |
| VK_FORMAT_D16_UNORM_S8_UINT | 16 | Unsigned normalized integer | 8 |
| VK_FORMAT_D24_UNORM_S8_UINT | 24 | Unsigned normalized integer | 8 |
| VK_FORMAT_D32_SFLOAT_S8_UINT | 32 | Signed floating point | 8 |
The choice is a trade-off between precision and cost. For HoloLens 2, Microsoft recommends the 16-bit DXGI_FORMAT_D16_UNORM format for better frame rate and performance, while noting that 16-bit buffers have less depth resolution than 24-bit buffers and that on PC VR headsets 24-bit depth buffers have less of a performance impact.[12] For depth submitted to the Oculus PC runtime, Oculus recommended a 32-bit floating-point depth format without multisampling.[6]
History
In a survey published in ACM Computing Surveys in March 1974, Ivan Sutherland, Bob Sproull and Robert Schumacker compared ten hidden-surface algorithms and included a "brute-force" image-space method as a point of comparison. They wrote that if a memory large enough "to store a color and a depth at each point in the output picture raster" were available, a renderer could compare the depth already recorded at each point of a polygon with the depth of a new polygon and replace the stored data when the new polygon was closer. They noted that the cost of this approach depended only on what they called the depth number, and not otherwise on the complexity of the environment, and that it carried "some hazard with edge effects".[13] Ivan Sutherland was one of Catmull's advisors at Utah.[3]
Edwin Catmull's 1974 PhD thesis at the University of Utah, on displaying curved surface patches, is credited by the ACM with giving rise to z-buffering, which the ACM describes as managing "image depth coordinates", and to texture mapping; the ACM notes that Wolfgang Straßer also described z-buffering at the time. Catmull shared the 2019 ACM A.M. Turing Award with Pat Hanrahan.[3] According to the Eurographics Association, Straßer's 1974 dissertation at TU Berlin, "Schnelle Kurven- und Flächendarstellung auf graphischen Sichtgeräten" (fast generation of curves and surfaces on graphics displays), contained a description of the z-buffer algorithm written independently of Catmull's work.[4] Williams, writing in 1978, called the z-buffer algorithm "first published by Catmull" and said it was the first method that made shaded computer pictures of bicubic surface patches possible.[2]
Williams's own 1978 paper used depth buffers for shadows. A view of the scene is first rendered from the light source, keeping only depth; each point seen by the observer is then transformed into the light's view and tested against those stored depths, and points the light cannot see are drawn in shadow.[2]
In 1993 Ned Greene, Michael Kass and Gavin Miller of Apple Computer presented hierarchical z-buffer visibility at SIGGRAPH. They pointed out that a traditional z-buffer makes no use of object-space or temporal coherence and has to render "every polygon of every object in every drawer of every desk in a building even if the whole building cannot be seen". Their method combined an object-space octree with an image-space "Z pyramid" so that hidden geometry could be rejected quickly, and they reported speedups of orders of magnitude in some scenes with high depth complexity.[14] The broader family of such techniques is covered in occlusion culling.
Depth buffers also became an input for image-based rendering. In his 1999 dissertation at the University of North Carolina at Chapel Hill, William R. Mark described a system that rendered every Nth frame conventionally and generated the frames in between by warping rendered images to a new viewpoint and view direction using McMillan's 3D image warp, which unlike a plain perspective warp can correct for viewpoint changes "even for objects at different depths". Mark also noted that the technique could compensate for rendering and network latency.[15]
Applications in VR and AR
Depth-aware reprojection
A VR headset shows each frame some time after the application rendered it, and the user's head keeps moving in between. Orientation-only Timewarp corrects the image for head rotation only and treats the frame's content as if it were projected infinitely far away, so head translation, and the parallax it causes, is left uncorrected.[6] With a depth buffer, the compositor knows how far away each pixel is and can shift near and far content by different amounts, which is how Positional Timewarp corrects for both rotation and translation.[6] Oculus describes Positional Timewarp as reducing the latency of both rotation and translation of the headset, which makes depth-aware reprojection one of the ways runtimes reduce perceived motion-to-photon latency.[6]
Oculus introduced the ovrLayerEyeFovDepth layer type for its PC runtime in 2017, with Oculus Dash. Its first use was depth composition: Dash used the application's depth buffer to draw an X-ray effect where the app's content passed through Dash content. Oculus later used the same submitted depth for Positional Timewarp (PTW), which corrects for head translation as well as rotation, and for ASW 2.0, which combines PTW with Asynchronous Spacewarp. In ASW 2.0, PTW handles parallax caused by head movement and ASW's motion vectors handle animated content and the errors PTW leaves behind, such as disocclusion at object edges. To submit depth, an application uses ovrLayerEyeFovDepth together with an ovrTimewarpProjectionDesc describing its projection and an ovrViewScaleDesc giving the size of its world units in meters. The depth and color buffers must have the same resolution, color formats such as OVR_FORMAT_R32_FLOAT are not accepted as depth, and if an app submits no depth the runtime falls back to the ASW 1.0 method for that app. The Oculus engine integrations team used the layer type in Unreal Engine 4 and Unity.[6]
The OpenXR standard defines the same capability in the ratified XR_KHR_composition_layer_depth extension (registered extension number 11), whose contributors came from Oculus, Microsoft, Arm and Qualcomm. It lets an application attach a depth image to each view of a projection layer, and the specification states that "the XR runtime may use this information to perform more accurate reprojections taking depth into account". For each view the application supplies the depth swapchain image, the window-space depths that correspond to the near and far planes, and the distances to those planes in meters; giving a near distance larger than the far distance signals a reversed depth mapping.[5] Microsoft tells HoloLens 2 developers to always submit depth through this extension, because enabling hardware depth reprojection on HoloLens 2 improves hologram stability. It also recommends a narrow depth range (its sample uses 0.1 to 20 meters) and reversed-Z for more uniform depth resolution.[12]
Valve's OpenVR interface accepts depth as well. Its submit flags include Submit_TextureWithDepth, which passes a VRTextureWithDepth_t whose VRTextureDepthInfo_t member holds a depth texture handle, the projection matrix and the depth range.[16] On the web, the W3C WebXR Layers draft gives projection layers an ignoreDepthValues attribute: when it is false, the content of the depth buffer attachment will be used by the XR compositor and is expected to represent the scene rendered into the layer.[17]
On Apple Vision Pro, apps that draw fully immersive content with Metal receive a color texture and a depth texture for each drawable from Apple's Compositor Services framework. Apple's sample code stores the depth attachment at the end of the render pass and builds its projection matrix with reverse-Z, using near and far distances supplied by the drawable.[18]
Frame synthesis
Depth is also one of the inputs for generating whole frames. Meta's Application SpaceWarp (AppSW) lets a Quest app render at half rate, for example 36 frames per second instead of 72, while the runtime synthesizes the missing frames. Meta states that with AppSW it is the app's responsibility to generate motion vector and depth buffer data. Meta's Unity SDKs and Unreal integration render motion vectors in a dedicated low-resolution pass (368x400 by default on Meta Quest 2, against a default eye buffer of 1440x1584), collect the depth buffer of that pass at the same resolution, and pass it to the compositor, which also uses it for Positional Timewarp.[19] In OpenXR, Meta's XR_FB_space_warp extension submits this data with a dedicated depth image, usually of lower resolution, that is independent of XR_KHR_composition_layer_depth. That extension is now marked as deprecated by the multi-vendor XR_EXT_frame_synthesis, whose contributors include Meta, Varjo, Samsung Electronics, Microsoft, Collabora and ByteDance, and which uses application-generated motion vector and depth images to extrapolate and reproject new frames.[5] The WebXR Layers draft likewise sets ignoreDepthValues to false for projection layers in sessions created with the "space-warp" feature.[17]
Layer composition and occlusion
Depth lets a compositor combine separate layers correctly. With Meta's XR_FB_composition_layer_depth_test extension, the compositor maintains a depth buffer of its own, cleared to the infinitely far distance at the start of composition; each incoming layer's depths are transformed into the compositor's depth space and compared with the stored values, and failing fragments are discarded. Projection layers supply their depth through XR_KHR_composition_layer_depth, while the runtime computes depth directly for geometric layers such as quads.[5]
In mixed reality, virtual objects also need to be hidden behind real ones. Varjo's XR_VARJO_environment_depth_estimation extension lets the compositor estimate the depth of the real environment, so that layers submitted with depth information are properly occluded in the final image.[5] On phones, Google's ARCore documentation shows apps comparing the depth of a virtual object with the Depth API's depth image in the fragment shader; a comment in its sample shader describes letting the asset fade into the background "instead of a hard Z-buffer test".[20] How such real-world depth is measured or estimated is covered in Depth map.
Stereo rendering
A headset needs one image per eye, and each view has its own depth. OpenXR attaches a depth structure to every view of a projection layer.[5] For HoloLens 2, Microsoft recommends one color swapchain with two array slices for the left and right eyes and one swapchain for depth, rendered with instanced draw calls, and warns that alternatives such as double-wide rendering or separate swapchains per eye lead to runtime post-processing with a significant performance penalty.[12] See Stereoscopic rendering for the general techniques.
Limitations
A single depth value per pixel describes only the nearest surface. Greene, Kass and Miller noted in 1993 that z-buffers "do not handle partially transparent surfaces well".[14] The same problem affects depth-based frame synthesis in VR: Meta's AppSW documentation says that transparent objects are difficult because the algorithm supports only one direction of movement per pixel, and it usually does not recommend rendering transparent objects into the motion vector pass.[19] Reprojection from one depth buffer also cannot show surfaces that were hidden in the original frame; in Oculus's ASW 2.0 example, remaining corrections at the edges of geometry were attributed mostly to disocclusion and view-dependent shading.[6] Precision limits remain as well, which is why XR platform guidance recommends reversed-Z and a sensible near and far range.[12][8]
See also
References
- ↑ 1.0 1.1 1.2 1.3 1.4 1.5 "Depth Buffers (Direct3D 9)". Microsoft Learn. Microsoft. 2018-05-31. https://learn.microsoft.com/en-us/windows/win32/direct3d9/depth-buffers. Retrieved 2026-10-11.
- ↑ 2.0 2.1 2.2 2.3 2.4 Lance Williams (1978-08). "Casting Curved Shadows on Curved Surfaces". Proceedings of SIGGRAPH 78, pp. 270-274. ACM. doi:10.1145/800248.807402. https://cseweb.ucsd.edu/~ravir/274/15/papers/p270-williams.pdf. Retrieved 2026-10-11.
- ↑ 3.0 3.1 3.2 "Pioneers of Modern Computer Graphics Recognized with ACM A.M. Turing Award". Association for Computing Machinery. ACM. https://web.archive.org/web/20250123200155/https://awards.acm.org/about/2019-turing. Retrieved 2026-10-11.
- ↑ 4.0 4.1 "In Memoriam Wolfgang Straßer". Eurographics Association. Eurographics. https://www.eg.org/wp/obituaries/in-memoriam-wolfgang-straser/. Retrieved 2026-10-11.
- ↑ 5.0 5.1 5.2 5.3 5.4 5.5 "The OpenXR 1.1.63 Specification (with all registered extensions)". Khronos Registry. Khronos Group. https://registry.khronos.org/OpenXR/specs/1.1/html/xrspec.html. Retrieved 2026-10-11.
- ↑ 6.0 6.1 6.2 6.3 6.4 6.5 6.6 6.7 Volga Aksoy, Dean Beeler (2019-08-09). "Developer Guide to ASW 2.0". Meta Horizon OS Developers Blog. Meta. https://developers.meta.com/horizon/blog/developer-guide-to-asw-20/. Retrieved 2026-10-11.
- ↑ "Fragment Operations". Vulkan Documentation. Khronos Group. https://docs.vulkan.org/spec/latest/chapters/fragops.html. Retrieved 2026-10-11.
- ↑ 8.0 8.1 8.2 8.3 8.4 "Depth Precision Visualized". NVIDIA Developer. NVIDIA. https://developer.nvidia.com/content/depth-precision-visualized. Retrieved 2026-10-11.
- ↑ Eugene Lapidous, Guofang Jiao (1999-07). "Optimal depth buffer for low-cost graphics hardware". Proceedings of the ACM SIGGRAPH/Eurographics Workshop on Graphics Hardware (HWWS '99), pp. 67-73. ACM. doi:10.1145/311534.311579. https://doi.org/10.1145/311534.311579. Retrieved 2026-10-11.
- ↑ "Depth Bias". Microsoft Learn. Microsoft. 2018-05-31. https://learn.microsoft.com/en-us/windows/win32/direct3d11/d3d10-graphics-programming-guide-output-merger-stage-depth-bias. Retrieved 2026-10-11.
- ↑ "Formats". Vulkan Documentation. Khronos Group. https://docs.vulkan.org/spec/latest/chapters/formats.html. Retrieved 2026-10-11.
- ↑ 12.0 12.1 12.2 12.3 "OpenXR best practices". Microsoft Learn. Microsoft. https://learn.microsoft.com/en-us/windows/mixed-reality/develop/native/openxr-best-practices. Retrieved 2026-10-11.
- ↑ Ivan E. Sutherland, Robert F. Sproull, Robert A. Schumacker (1974-03). "A Characterization of Ten Hidden-Surface Algorithms". ACM Computing Surveys, vol. 6, no. 1, pp. 1-55. ACM. doi:10.1145/356625.356626. https://doi.org/10.1145/356625.356626. Retrieved 2026-10-11.
- ↑ 14.0 14.1 Ned Greene, Michael Kass, Gavin Miller (1993-09). "Hierarchical Z-buffer visibility". Proceedings of SIGGRAPH 93, pp. 231-238. ACM. doi:10.1145/166117.166147. https://doi.org/10.1145/166117.166147. Retrieved 2026-10-11.
- ↑ William R. Mark (1999-04-21). "Post-Rendering 3D Image Warping: Visibility, Reconstruction, and Performance for Depth-Image Warping". PhD dissertation, Technical Report TR99-022. University of North Carolina at Chapel Hill, Department of Computer Science. https://www.cs.unc.edu/techreports/99-022.pdf. Retrieved 2026-10-11.
- ↑ "openvr.h". ValveSoftware/openvr on GitHub. Valve. https://github.com/ValveSoftware/openvr/blob/master/headers/openvr.h. Retrieved 2026-10-11.
- ↑ 17.0 17.1 "WebXR Layers API Level 1". W3C Working Draft. World Wide Web Consortium. 2026-08-11. https://www.w3.org/TR/webxrlayers-1/. Retrieved 2026-10-11.
- ↑ "Drawing fully immersive content using Metal". Apple Developer Documentation. Apple. https://developer.apple.com/documentation/compositorservices/drawing-fully-immersive-content-using-metal. Retrieved 2026-10-11.
- ↑ 19.0 19.1 "Application SpaceWarp". Meta Horizon OS Developers. Meta. 2024-08-16. https://developers.meta.com/horizon/documentation/native/android/os-app-spacewarp. Retrieved 2026-10-11.
- ↑ "Use Depth in your Android app". ARCore, Google for Developers. Google. https://developers.google.com/ar/develop/java/depth/developer-guide. Retrieved 2026-10-11.