Draw call
More actions
A draw call is a command that an application running on the CPU issues through a graphics API, such as OpenGL, Direct3D, Metal or Vulkan, to make the graphics processing unit (GPU) draw a set of geometry. Each call submits a group of primitives, usually triangles, that share the same render state: the same shaders, textures, buffers and other settings. In OpenGL the basic draw calls are functions such as glDrawArrays; in Vulkan, "drawing commands (commands with Draw in the name) provoke work in a graphics pipeline".[1][2] The work submitted by one draw call is often called a "batch".[3]
Much of the overhead of a draw call falls on the CPU, which has to set up render state and pass the command through the graphics driver. When a scene is split into many small draw calls, the CPU can become the bottleneck while the GPU waits for work.[4][3] This matters more in virtual reality (VR) than in flat-screen games. A headset needs two views of the scene at a high, fixed frame rate, and a naive stereo renderer issues every draw call twice. Mobile and standalone headsets such as the Meta Quest 3 also run on battery-powered processors with limited heat dissipation.[5] Draw call counts are a standard performance budget in VR development guidelines, and techniques such as batching, instancing and single-pass stereo rendering are used to reduce them.[6][7]
How it works
Unity's documentation divides a draw call into two steps. First the CPU updates the render state: it uses the graphics API to prepare and send everything the GPU needs to draw the mesh, such as shader code, textures and buffers. Unity describes this as the most CPU-intensive step. Second, the CPU submits the draw call itself, telling the GPU what to draw. By default each mesh in a Unity scene needs its own draw call.[4] In NVIDIA's 2003 definition, every Direct3D DrawIndexedPrimitive() call was a batch: it submitted some number of triangles, the same render state applied to all of them, and the state-setting calls made before the draw counted as part of the batch. Changing state meant at least two batches.[3]
A draw call can draw more than one copy of its geometry. Vulkan's vkCmdDraw takes a vertex count and an instance count, and Vulkan also has indirect drawing commands that read their parameters from buffer memory instead of from the CPU-side call.[1] The OpenGL extension ARB_draw_instanced, approved by the OpenGL Architecture Review Board in 2008, introduced instanced draw calls that are "conceptually equivalent to a series of draw calls"; a vertex shader reads a built-in instance ID and can use it to fetch per-instance data such as a transform.[8] The later ARB_multi_draw_indirect extension, approved in 2012, lets an application assemble "large batches of drawing commands" in a GPU buffer and dispatch them with one function call.[9]
CPU-bound and GPU-bound rendering
Johansson, writing in the Journal of Computer Graphics Techniques and following Hillaire (2012), summarizes the distinction this way: if an application is CPU-bound, its performance depends mostly on the number of draw calls and related state changes per frame; if it is GPU-bound, it depends on the number of triangles drawn and the complexity of the shader stages.[2] A draw call count therefore says little about GPU load on its own. Meta's WebXR guidance gives an example: submitting 1,000 individual triangles as separate draw calls would likely keep an app below 72 frames per second, even though the Quest GPU can render several orders of magnitude more triangles.[10]
History
Batch counts in the early 2000s
Matthias Wloka of NVIDIA measured the problem in a talk at the Game Developers Conference in 2003. His test application drew about 100,000 triangles per frame in batches of different sizes, changing state between draws. Below an average of about 130 triangles per batch the application was "100% CPU limited", and throughput in batches per second depended on CPU speed rather than on the GPU. A profile of a 1 GHz Pentium 3 drawing two triangles per batch showed 78 percent of the time in the driver and 14 percent in the Direct3D runtime.[3]
From games he examined, Wloka derived a rule of thumb of about 25,000 batches per second using 100 percent of a 1 GHz CPU. His budget formula multiplied that figure by the CPU clock in GHz and by the share of CPU time available for draw and state calls, divided by the target frame rate: a 2 GHz CPU spending 20 percent of its time on submission at 30 frames per second gets about 333 batches per frame. He argued that the work grows linearly with the number of batches, that driver optimizations can reduce the constant but not the order, and that because GPUs were getting faster more quickly than CPUs, developers should put more triangles into each batch instead of making more batches. He suggested packing several textures into one surface and pre-transforming static geometry so that more objects could share a batch.[3]
In the 2005 book GPU Gems 2, Francesco Carucci of Lionhead Studios cited Wloka's finding that a 1 GHz CPU could submit only around 10,000 to 40,000 batches per second in Direct3D, and estimated that a newer CPU could manage around 30,000 to 120,000, or about 1,000 to 4,000 batches per frame at 30 frames per second. As an example, he noted that a forest drawn with one tree per batch could not contain more than 4,000 trees, with no CPU time left for the rest of the game. His chapter described geometry instancing as a way to render many copies of the same geometry in fewer calls.[11]
Low-overhead graphics APIs
Lower per-draw-call overhead was a headline goal of several graphics APIs introduced in the 2010s. In September 2013 AMD announced Mantle, and AMD's Raja Koduri said it would allow nine times as many draw calls per second as rival APIs by reducing CPU overhead.[12] In a March 2014 blog post introducing Direct3D 12, Microsoft said it would deliver "significantly reduced draw call overhead, and many more draw calls per frame".[13] Microsoft's documentation lists "vastly reduced CPU overhead" as its first benefit over Direct3D 11 and says the freed CPU time can be used to increase the number of draw calls.[14] At GDC 2014, engineers from AMD, Intel and NVIDIA presented "Approaching Zero Driver Overhead in OpenGL", a session on OpenGL techniques that the presenters said could reduce driver overhead "by up to 10x or more" across vendors.[15] In June 2014 Apple introduced Metal for iOS 8 and claimed a "10 times improvement in draw call speed".[16]
Vulkan, from the Khronos Group, followed the same approach. Google's Android team wrote in April 2016 that Vulkan's lower CPU overhead let some synthetic benchmarks reach as much as 10 times the draw call throughput on a single core compared with OpenGL ES, and that applications submit draw calls into command buffers that can be reused.[17] Benchmarks measure this directly: the 3DMark API Overhead test for Android raises the number of draw calls step by step and reports how many draw calls per second each API achieves before the frame rate falls below 30 frames per second.[18]
Peer-reviewed comparisons have found the gains depend on workload. In a 2019 study on a graphics rendering server, Lujan and colleagues reported that OpenGL's performance was unpredictable, that Vulkan saved a significant amount of energy at the same performance, and that OpenGL could not keep up with Vulkan when an extremely high frame rate was required.[19] A 2021 follow-up on an ARM processor with big and LITTLE cores, motivated by the battery drain of mobile games and VR, found that Vulkan could save up to 24 percent of energy on heavy workloads by running in parallel on the LITTLE cores and could render at a much higher frame rate once OpenGL ES reached its limit, but that the gains were negligible for light workloads.[20]
Reducing draw calls
Engines and graphics programmers use several methods to cut the number of draw calls or the cost of each one.
| Method | How it reduces draw call cost |
|---|---|
| Static batching | Combines static meshes that use the same material into shared vertex and index buffers in world space so they can be drawn with fewer render state updates. Unity's implementation allows up to 64,000 vertices per buffer.[21] |
| GPU instancing | Draws many copies of the same mesh and material in one draw call, with per-instance properties such as color or scale. Unity notes that the benefit depends on the platform and is larger on mobile than on desktop.[22][8] |
| Fewer render state changes | Unity's SRP Batcher reduces render state updates with techniques such as large GPU buffers.[4] Meta recommends sharing and instancing meshes, using global texture arrays and limiting an app to a small set of unified shaders.[6] |
| Texture packing | Packing several textures into one surface lets objects with different textures share a batch, at the cost of tool support and possibly wasted texture space.[3] |
| Culling and level of detail | Frustum and occlusion culling remove objects that cannot be seen, so they are never submitted; Meta's PC guidelines list level of detail, culling and batching as ways to stay within its draw call limit.[2][23] |
| Indirect and multi-draw commands | Draw parameters live in GPU buffers, and one call can dispatch many draws.[9][1] |
| Low-overhead APIs | Direct3D 12, Metal and Vulkan reduce the CPU and driver cost per draw call; Direct3D 12 is designed to make full use of multithreading, and Vulkan command buffers can be reused.[14][16][17] |
Reducing draw calls is not always worth its own cost. Johansson found that for one building model, once occlusion culling and stereo instancing were in use, the high level of occlusion in its exterior views meant that average frame times were actually lower with geometry instancing and batching of walls turned off.[2] Unity's engineers, comparing stereo rendering modes in 2017, found only a small CPU difference between single-pass and single-pass instanced rendering: instancing removes draw calls, but they wrote that this cost is quite low compared with processing the scene graph, and that issuing draw calls can be quite fast on the dispatching thread because most modern drivers are multithreaded.[24]
Draw calls in VR and AR
Stereo rendering
A VR frame has two views, one per eye. Meta's PC performance guidelines note that rendering both views typically means every draw call is issued twice, every mesh drawn twice and every texture bound twice.[23] Johansson made the same point: a two-pass stereo renderer doubles the number of occlusion tests, rasterized triangles and issued draw calls.[2] The Khronos OVR_multiview specification, whose contact is Cass Everitt of Oculus and whose contributors include John Carmack, states that sequential stereo rendering "typically incurs double the application and driver overhead, despite the fact that the command streams and render states are almost identical".[7]
Several methods remove the duplicate draw calls. In his GDC 2015 talk "Advanced VR Rendering", Valve's Alex Vlachos rated running the CPU code twice as "BAD", resubmitting command buffers as "GOOD" (Valve's solution at the time) and using instancing to double the geometry as "BETTER", with "half the API calls".[25] With stereo instancing, an application replaces each draw call with an instanced one, for example glDrawArraysInstanced in place of glDrawArrays, using an instance count of two, and the vertex shader uses the instance ID to project each vertex for the left or the right eye.[2] Unity's engineers wrote that this approach "literally halve[s] the number of draw calls" on the API side.[24] In multiview rendering, "draw calls are instanced into each corresponding element of the texture array", and the vertex shader reads a view ID to compute per-view values such as vertex position.[7] Vulkan's VK_KHR_multiview has the same goal, describing multiview as "a rendering technique originally designed for VR"; its functionality became part of core Vulkan 1.1.[26] Meta's RenderDoc guide for Unreal Engine developers says instanced stereo or multiview halves the number of draw calls in the base pass, with each draw call dispatching two instances per mesh.[27]
On mobile hardware the savings were measured early. Presenting at SIGGRAPH 2016, Everitt said that once devices with multiview support became available, real and synthetic apps showed CPU time reductions of 33 to 49 percent, and power numbers dropped by 27 to 33 percent; he added that driver overhead and validation with multiview cost about the same as monoscopic rendering.[5]
Other effects multiply draw calls further. In a 2019 Meta developer blog post, Trevor Dasch advised against a depth pre-pass on mobile VR because it doubles the draw calls, and noted that rendering a mirror or portal the naive way can triple them; he wrote that "draw calls are quite heavy on the cpu".[28]
Draw call budgets
Meta's performance guidelines for PC VR apps recommend limiting each frame to a maximum of 500 to 1,000 draw calls and 1 to 2 million triangles or vertices.[23] For its standalone headsets Meta publishes example ranges that depend on how much other work the app does on the CPU. Meta says the number of draw calls an app can afford depends on pipeline state changes, the capacity of the main and render threads and the graphics API, and that CPU work such as animation, skinning and networking can delay the render thread. It describes "busy" apps as those with heavy simulation such as many networked players or NPCs, "medium" as fitting most apps, and "light" as apps with minimal state changes such as puzzle games.[6]
| Headset | Busy simulation | Medium simulation | Light simulation |
|---|---|---|---|
| Oculus Quest (Quest 1) | 50-150 | 150-250 | 200-400 |
| Meta Quest 2, Meta Quest Pro | 80-200 | 200-300 | 400-600 |
| Meta Quest 3, Meta Quest 3S | 200-300 | 400-600 | 700-1000 |
Meta calls these figures examples and says actual results vary.[6] In January 2020 UploadVR reported that Unity had added Vulkan support for the Oculus Quest, noting that Vulkan's lower-level access to the hardware means less driver overhead for draw calls, so more draw calls can be used each frame.[29] Browser-based WebXR apps face the same limit. Meta's WebXR guide lists too many draw calls as a common cause of CPU-bound apps and recommends batching objects that share a material, merging small meshes, and using multiview or instanced mesh rendering, while still culling instanced meshes.[10]
Case studies and research
In a 2016 Intel case study, application engineer Finn Wong profiled Pangu, a PC VR game from Tencent built on Direct3D 11. Before optimization the game ran at an average of 36.4 frames per second on an Oculus Rift DK2 with 4,437 draw calls per frame and GPU load of 49.64 percent, because the GPU sat idle while the CPU prepared draw calls, culling and other work. Wong noted that Direct3D 11 rendering is single-threaded and has relatively high draw call overhead compared with Direct3D 12. After optimizations that included level of detail, Instanced Stereo Rendering, removal of dynamic shadows and an earlier start for the render thread, the game ran at 71.4 frames per second on an HTC Vive with 845 draw calls per frame, and the CPU-bound period in each frame fell from 7.37 ms to 2.62 ms.[30]
Johansson's 2016 study used building information models (BIM) of a dormitory, a hotel and an office, targeting the 90 Hz minimum frame rate (11.1 ms per frame) set for the consumer Oculus Rift and HTC Vive. With only view frustum culling, the models were mainly CPU-bound because of the large number of draw calls relative to the number of triangles. Stereo instancing cut average frame times on exterior camera paths by 27 to 52 percent compared with two-pass rendering on a GeForce GTX 980M. The exterior of the hotel model still had too many visible objects, and it reached 90 Hz only after the number of draw calls was cut further by combining occlusion culling and stereo instancing with batching of wall geometry and geometry instancing.[2]
Draw call counts also appear as a standard metric in applied VR research. In a 2025 study in the Bulletin of Electrical Engineering and Informatics, Miranto and colleagues measured frame rate, triangle count and draw calls in a Unity climate-education simulation running on a Meta Quest 3, treating triangle count as an indicator of GPU load and draw calls as an indicator of CPU load. They used static batching for terrain and infrastructure and GPU instancing for repeated trees; their baseline scene had 55,370 triangles and 65 draw calls. Trees with transparent, alpha-cutout foliage reduced batching efficiency and cost more GPU time than high-triangle opaque character models.[31]
See also
References
- ↑ 1.0 1.1 1.2 "Drawing Commands (Vulkan Specification)". Vulkan Documentation. Khronos Group. https://docs.vulkan.org/spec/latest/chapters/drawing.html. Retrieved 2026-10-11.
- ↑ 2.0 2.1 2.2 2.3 2.4 2.5 2.6 Mikael Johansson (2016). "Efficient Stereoscopic Rendering of Building Information Models (BIM)". Journal of Computer Graphics Techniques, vol. 5, no. 3. Chalmers University of Technology. https://jcgt.org/published/0005/03/01/. Retrieved 2026-10-11.
- ↑ 3.0 3.1 3.2 3.3 3.4 3.5 Matthias Wloka (2003). "Batch, Batch, Batch: What Does It Really Mean?". Game Developers Conference 2003. NVIDIA. https://web.archive.org/web/2010/http://developer.nvidia.com/docs/IO/8230/BatchBatchBatch.pdf. Retrieved 2026-10-11.
- ↑ 4.0 4.1 4.2 "Introduction to optimizing draw calls". Unity Manual (Unity 6.6). Unity Technologies. https://docs.unity3d.com/Manual/optimizing-draw-calls.html. Retrieved 2026-10-11.
- ↑ 5.0 5.1 Cass Everitt (2016). "Multiview Rendering (presentation slides with notes)". Moving Mobile Graphics course, SIGGRAPH 2016. Arm Community. https://community.arm.com/cfs-file/__key/communityserver-blogs-components-weblogfiles/00-00-00-20-66/5_2D00_mmg_2D00_siggraph2016_2D00_multiview_2D00_cass.pdf. Retrieved 2026-10-11.
- ↑ 6.0 6.1 6.2 6.3 6.4 "Testing and performance analysis". Meta Horizon OS Developers. Meta. https://developers.meta.com/horizon/documentation/unity/unity-perf/. Retrieved 2026-10-11.
- ↑ 7.0 7.1 7.2 "OVR_multiview (OpenGL and OpenGL ES extension specification, revision 6)". Khronos OpenGL Registry. Khronos Group. 2018-10-19. https://registry.khronos.org/OpenGL/extensions/OVR/OVR_multiview.txt. Retrieved 2026-10-11.
- ↑ 8.0 8.1 "ARB_draw_instanced (OpenGL extension specification)". Khronos OpenGL Registry. Khronos Group. 2011-04-08. https://registry.khronos.org/OpenGL/extensions/ARB/ARB_draw_instanced.txt. Retrieved 2026-10-11.
- ↑ 9.0 9.1 "ARB_multi_draw_indirect (OpenGL extension specification)". Khronos OpenGL Registry. Khronos Group. 2014-06-09. https://registry.khronos.org/OpenGL/extensions/ARB/ARB_multi_draw_indirect.txt. Retrieved 2026-10-11.
- ↑ 10.0 10.1 "WebXR performance optimization workflow". Meta Horizon OS Developers. Meta. https://developers.meta.com/horizon/documentation/web/webxr-perf-workflow/. Retrieved 2026-10-11.
- ↑ Francesco Carucci (2005). "Chapter 3. Inside Geometry Instancing". GPU Gems 2, ed. Matt Pharr. NVIDIA / Addison-Wesley. https://developer.nvidia.com/gpugems/gpugems2/part-i-geometric-complexity/chapter-3-inside-geometry-instancing. Retrieved 2026-10-11.
- ↑ Edward Chester (2013-09-25). "AMD unveils Mantle, a new high-speed gaming API". bit-tech. https://web.archive.org/web/20181109221545/https://www.bit-tech.net/news/tech/software/amd-unveils-mantle-a-new-high-speed-gaming/1/. Retrieved 2026-10-11.
- ↑ D3D Team (2014-03-20). "DirectX 12". DirectX Developer Blog. Microsoft. https://devblogs.microsoft.com/directx/directx-12/. Retrieved 2026-10-11.
- ↑ 14.0 14.1 "What is Direct3D 12". Microsoft Learn. Microsoft. https://learn.microsoft.com/en-us/windows/win32/direct3d12/what-is-directx-12-. Retrieved 2026-10-11.
- ↑ Graham Sellers, Tim Foley, Cass Everitt, John McDonald. "Approaching Zero Driver Overhead in OpenGL (Presented by NVIDIA)". GDC Vault (GDC 2014). https://gdcvault.com/play/1020791/Approaching-Zero-Driver-Overhead-in. Retrieved 2026-10-11.
- ↑ 16.0 16.1 "Apple Releases iOS 8 SDK With Over 4,000 New APIs". Apple Newsroom. Apple. 2014-06-02. https://www.apple.com/newsroom/2014/06/02Apple-Releases-iOS-8-SDK-With-Over-4-000-New-APIs/. Retrieved 2026-10-11.
- ↑ 17.0 17.1 Shannon Woods (2016-04-13). "Optimize, Develop, and Debug with Vulkan Developer Tools". Android Developers Blog. Google. https://android-developers.googleblog.com/2016/04/optimize-develop-and-debug-with-vulkan.html. Retrieved 2026-10-11.
- ↑ "New Vulkan API performance test for Android devices". UL Benchmarks. UL. 2017-08-17. https://benchmarks.ul.com/news/new-vulkan-api-performance-test-for-android-devices. Retrieved 2026-10-11.
- ↑ Michael Lujan, Michael Baum, Dayuan Chen, Ziliang Zong (2019-02). "Evaluating the Performance and Energy Efficiency of OpenGL and Vulkan on a Graphics Rendering Server". 2019 International Conference on Computing, Networking and Communications (ICNC). IEEE. doi:10.1109/ICCNC.2019.8685588. https://doi.org/10.1109/ICCNC.2019.8685588. Retrieved 2026-10-11.
- ↑ Michael Lujan, Michael McCrary, Blake W. Ford, Ziliang Zong (2021-10). "Vulkan vs OpenGL ES: Performance and Energy Efficiency Comparison on the big.LITTLE Architecture". 2021 IEEE International Conference on Networking, Architecture and Storage (NAS). IEEE. doi:10.1109/NAS51552.2021.9605447. https://doi.org/10.1109/NAS51552.2021.9605447. Retrieved 2026-10-11.
- ↑ "Introduction to static batching". Unity Manual (Unity 6.6). Unity Technologies. https://docs.unity3d.com/Manual/DrawCallBatching.html. Retrieved 2026-10-11.
- ↑ "Introduction to GPU instancing". Unity Manual (Unity 6.6). Unity Technologies. https://docs.unity3d.com/Manual/GPUInstancing.html. Retrieved 2026-10-11.
- ↑ 23.0 23.1 23.2 "Guidelines for VR Performance Optimization". Meta Horizon OS Developers. Meta. https://developers.meta.com/horizon/documentation/native/pc/dg-performance-guidelines/. Retrieved 2026-10-11.
- ↑ 24.0 24.1 Rob Srinivasiah (2017-11-21). "How to maximize AR and VR performance with advanced stereo rendering". Unity Blog. Unity Technologies. https://web.archive.org/web/20231201103727/https://blog.unity.com/technology/how-to-maximize-ar-and-vr-performance-with-advanced-stereo-rendering. Retrieved 2026-10-11.
- ↑ Alex Vlachos (2015). "Advanced VR Rendering". Game Developers Conference 2015. Valve. https://media.steampowered.com/apps/valve/2015/Alex_Vlachos_Advanced_VR_Rendering_GDC2015.pdf. Retrieved 2026-10-11.
- ↑ "VK_KHR_multiview". Vulkan Documentation. Khronos Group. 2016-10-28. https://docs.vulkan.org/refpages/latest/refpages/source/VK_KHR_multiview.html. Retrieved 2026-10-11.
- ↑ "Using RenderDoc Meta Fork to Optimize Your App - Part 1". Meta Horizon OS Developers. Meta. https://developers.meta.com/horizon/documentation/unreal/po-renderdoc-optimizations-1/. Retrieved 2026-10-11.
- ↑ Trevor Dasch (2019-11-19). "PC Rendering Techniques to Avoid when Developing for Mobile VR". Meta Horizon OS Developers Blog. Meta. https://developers.meta.com/horizon/blog/pc-rendering-techniques-to-avoid-when-developing-for-mobile-vr/. Retrieved 2026-10-11.
- ↑ David Heaney (2020-01-28). "Unity Now Supports Vulkan On Oculus Quest". UploadVR. https://www.uploadvr.com/unity-oculus-quest-vulkan/. Retrieved 2026-10-11.
- ↑ Finn Wong (2016). "Performance Analysis and Optimization for PC-Based VR Applications: From the CPU's Perspective". Intel Software. Intel Corporation. https://www.intel.cn/content/dam/develop/external/us/en/documents/performanceanalysisoptimizationforpcbasedvr-applicationscpuperspective-699994.pdf. Retrieved 2026-10-11.
- ↑ Cahya Miranto, Ardiman Firmanda, Hestiasari Rante, Sritrusta Sukaridhoto, Muhammad Agus Zainuddin, Haolia Rahman (2025-10). "Performance analysis of 3D assets in virtual reality simulations for climate change: a case study in sustainable energy systems". Bulletin of Electrical Engineering and Informatics, vol. 14, no. 5, pp. 3659-3670. doi:10.11591/eei.v14i5.9532. https://doi.org/10.11591/eei.v14i5.9532. Retrieved 2026-10-11.