Jump to content

Neural rendering

From VR & AR Wiki

Neural rendering is a family of rendering and image synthesis methods that combine machine learning, in particular deep neural networks, with physical knowledge from computer graphics. The 2020 Eurographics state-of-the-art report that first surveyed the field described it as "a new and rapidly emerging field that combines generative machine learning techniques with physical knowledge from computer graphics, e.g., by the integration of differentiable rendering into network training."[1] Typical tasks include novel view synthesis (generating views of a captured scene from camera positions that were never photographed), relighting, facial and body reenactment, free-viewpoint video and photorealistic avatars.[1]

The field matters to virtual reality and augmented reality for several reasons. The 2020 report lists "the creation of photo-realistic avatars for virtual and augmented reality telepresence" among its main use cases,[1] and its 2022 successor names augmented and virtual reality among the fields "starkly impacted" by neural volumetric scene representations because of the photorealism they add to rendered environments.[2] Meta's research division wrote in 2020 that "Neural rendering has great potential for AR/VR."[3] Neural and learned techniques now appear in headset avatar systems, in apps that turn scans of real rooms into places a user can walk around in VR, and in AI upscalers used by PC VR games.

Reviewed 27 September 2026. Checked every citation, quotation, paper author list and venue, product detail and date against the cited surveys, papers, company pages and UploadVR, CNET and NVIDIA articles. About review dates.

Definition

The 2020 report by Ayush Tewari, Ohad Fried, Justus Thies and 16 co-authors noted that neural rendering "has not yet a clear definition in the literature" and proposed one: "Deep image or video generation approaches that enable explicit or implicit control of scene properties such as illumination, camera parameters, pose, geometry, appearance, and semantic structure."[1] The authors contrasted two traditions. Classical graphics starts from physics, modelling geometry, surface properties and cameras, and its output quality depends on how physically correct those models are. Machine learning starts from statistics, learning from real-world examples, and its quality depends on the network design and the training data. Explicit reconstruction of scene properties is "hard and error prone", and image-based rendering, which blends captured photographs using heuristics, shows seams or ghosting in complex scenes; neural rendering tries to address both reconstruction and rendering by training deep networks to map captured images to new images.[1]

The 2022 follow-up report, "Advances in Neural Rendering", observed that the term is "often applied to what are two distinct concepts":[2]

Paradigm What the network learns Example
2D neural rendering (also called neural refinement, neural re-rendering or deferred neural rendering) The network takes a 2D input, such as a semantic label map or an image produced by a classical renderer from proxy geometry, and is trained to produce the output image. The network learns to render.[2] Deferred Neural Rendering (2019)[4]
3D neural rendering The network is trained to represent the shape or appearance of one particular scene in 3D; that representation is then drawn by a conventional, analytically defined graphics "engine" such as volume rendering. The network does not learn how to render.[2] Neural Radiance Fields (NeRF, 2020)[5]

The 2022 report focuses on the second paradigm, where learned 3D scene representations (which the authors say are "often now referred to as neural scene representations") are combined with classical rendering principles. Because the scene is represented in 3D, such methods are "3D-consistent by design", which is what makes novel view synthesis of a captured scene possible.[2] The report describes NeRF's physically motivated structure, density and radiance stored in 3D space and rendered by ray casting and volume integration, as the reason for its better generalization to new views, and credits NeRF's quality and simplicity with an "explosion" of follow-up work.[2]

History

Neural rendering grew out of two lines of work: deep generative models such as generative adversarial networks, which by 2019 could synthesize high-resolution portraits that are often indistinguishable from real faces, and earlier image-based rendering techniques for combining captured photographs into new views.[1] According to the 2020 report, one of the first publications to use the term "neural rendering" was the Generative Query Network (GQN).[1]

Year Work Authors and venue Contribution
2018 Generative Query Network S. M. Ali Eslami, Danilo Jimenez Rezende and colleagues at DeepMind, published in Science (vol. 360, no. 6394)[6] A representation network turns a few images of a scene into a compact description, and a generation network "predicts ('imagines') the scene from a previously unobserved viewpoint", learned with "no human labelling of the contents of scenes".[6]
2018 Deep Appearance Models Stephen Lombardi, Jason Saragih, Tomas Simon, Yaser Sheikh; ACM Transactions on Graphics vol. 37, no. 4 (SIGGRAPH 2018)[7] A variational autoencoder models facial geometry and view-specific texture; the authors describe it as "naturally suited to real-time interactive settings such as Virtual Reality (VR)".[7]
2018 DeepFocus Lei Xiao, Anton Kaplanyan, Alexander Fix, Matt Chapman, Douglas Lanman (Facebook Reality Labs); SIGGRAPH Asia 2018[8] A learned rendering system that synthesizes realistic defocus blur for varifocal headsets.[8]
2019 Deferred Neural Rendering Justus Thies, Michael Zollhöfer, Matthias Nießner; SIGGRAPH 2019[4] Introduced "neural textures", learned feature maps trained during scene capture and interpreted by a deferred neural rendering pipeline, allowing photorealistic output from imperfect 3D geometry.[4]
2020 State of the Art on Neural Rendering Tewari, Fried, Thies et al.; Eurographics 2020[1] First survey of the field, with the working definition quoted above.[1]
2020 NeRF Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, Ren Ng; ECCV 2020 (oral)[5] A fully connected network maps a 5D coordinate (3D position plus 2D viewing direction) to volume density and view-dependent radiance; images are formed with standard, differentiable volume rendering, so only posed photographs are needed for training.[5]
2020 Neural supersampling for real-time rendering Xiao, Nouri, Chapman, Fix, Lanman, Kaplanyan (Facebook Reality Labs); SIGGRAPH 2020[9] Learned 16x supersampling of rendered game content in real time.[3]
2022 Instant Neural Graphics Primitives Thomas Müller, Alex Evans, Christoph Schied, Alexander Keller (NVIDIA); ACM Transactions on Graphics vol. 41, no. 4 (SIGGRAPH 2022)[10] A small network combined with multiresolution hash tables of learnable feature vectors; the authors report training "in a matter of seconds" and rendering in tens of milliseconds at 1920 x 1080.[10]
2022 Advances in Neural Rendering Tewari, Thies, Mildenhall, Srinivasan et al.; Eurographics 2022[2] Second survey, centred on neural scene representations such as NeRF.[2]
2023 3D Gaussian Splatting Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George Drettakis; ACM Transactions on Graphics vol. 42, no. 4[11] Represents a radiance field with explicit 3D Gaussians and a fast splatting rasterizer, giving real-time novel view synthesis at 1080p without neural network inference at render time.[11]
2023 VR-NeRF Linning Xu, Vasu Agrawal, William Laney and colleagues at Meta and the Chinese University of Hong Kong; SIGGRAPH Asia 2023[12] Capture, reconstruction and real-time VR rendering of walkable spaces with neural radiance fields.[12]

Relationship to Gaussian splatting

3D Gaussian splatting is closely tied to neural rendering but is not a neural network method in the strict sense. Kerbl and colleagues framed their 2023 paper as a response to radiance field methods in which "achieving high visual quality still requires neural networks that are costly to train and render". Their method starts from the sparse points produced during camera calibration, represents the scene with 3D Gaussians that keep the useful properties of continuous volumetric radiance fields while "avoiding unnecessary computation in empty space", optimizes the Gaussians' anisotropic covariance, and renders them with a visibility-aware splatting algorithm.[11][13] A survey by Guikun Chen and Wenguan Wang, accepted by ACM Computing Surveys, describes the technique as a radiance field method that, "Unlike mainstream implicit neural models", uses "millions of learnable 3D Gaussians for an explicit scene representation."[14]

The two approaches share the same problem setting (novel view synthesis from photographs) and the same idea of optimizing a scene representation through differentiable rendering,[11][14] and press coverage of consumer products usually describes Gaussian splat capture as a machine learning technique. UploadVR, for example, wrote that Varjo's Teleport service "uses Gaussian splatting, leveraging advances in machine learning to 'train' a high quality output based on a sequence of image views of the scene."[15] Apple has said that its Gaussian splat based avatars also depend on a set of neural networks (see below).[16]

Applications in VR and AR

Avatars and telepresence

Photorealistic avatars for telepresence are one of the VR and AR uses that the 2020 state-of-the-art report covers in detail. The 2020 state-of-the-art report describes how Deep Appearance Models produce "high-quality, high-resolution viewpoint-dependent renderings of the face" that run at 90 Hz in virtual reality. The same system animates the face from cameras inside a VR headset: two look at the eyes and one, at the bottom, looks at the mouth. Conditioning the model on the camera viewpoint lets the decoder correct geometric tracking errors. A follow-up by Wei et al. used a "training" headset with six additional cameras and unsupervised image-to-image translation to close the gap between rendered avatars and infrared headset images, matching lip shapes and complex expressions more precisely.[1]

The Deep Appearance Models authors, including Yaser Sheikh, worked at Facebook Reality Labs,[7] whose Pittsburgh lab develops Meta's Codec Avatars; in 2019 Meta identified Sheikh, the lab's director of research, as the leader of that project. Meta describes Codec Avatars as "Highly realistic digital representations that accurately reflect how a person looks and moves in real life, and in real-time", built with computer vision and machine learning. An encoder uses cameras and microphones on the headset to build "a unique code, a numeric representation of the state of a person's body and environment", and a decoder turns that code back into audio and images.[17]

Apple's Personas on Apple Vision Pro are a shipping example. Apple said in June 2025 that the new Personas in visionOS 26 take "advantage of industry-leading volumetric rendering and machine learning technology".[18] In an October 2025 interview with CNET, Jeff Norris, senior director of Apple's Vision Products Group, explained that Personas use Gaussian splatting to create the 3D facial scans, and said: "There's machine learning involved, but not many people really realize that it's a concert of networks that come together." He added that Apple had counted "over a dozen" networks and reduced the number for the new version, and that the company also applies Gaussian splatting to the 3D conversion of photos called Spatial Scenes.[16]

Captured real-world spaces

Neural and radiance field reconstruction lets a real room be captured with a camera and revisited in a headset with correct parallax. The VR-NeRF system from Meta researchers (SIGGRAPH Asia 2023) is an end-to-end pipeline for "the high-fidelity capture, model reconstruction, and real-time rendering of walkable spaces in virtual reality using neural radiance fields". It uses a custom multi-camera rig called the Eyeful Tower to take dense, high dynamic range images, and a multi-GPU renderer that draws the neural radiance field at full VR resolution, two 2K by 2K views at 36 Hz on the authors' demo machine.[12]

Consumer capture tools that followed mostly use Gaussian splatting:

Product Company Capture device Viewing Notes
Teleport Varjo iPhone, 5 to 10 minutes of walking around the space iPhone, web browsers and, through Varjo's Windows software, any OpenXR PC VR headset, including Meta Quest headsets over Quest Link, Steam Link or Virtual Desktop Launched November 2024; 30 minutes to 24 hours of cloud processing; US$30 per month for up to 15 scans.[15]
Horizon Hyperscape Capture (Beta) Meta The headset itself: Meta Quest 3 or Meta Quest 3S Cloud streamed from Meta's servers using technology internally codenamed Avalanche Began rolling out in the US in September 2025; a scan is ready 1 to 8 hours after upload, depending on the size of the space.[19]

In its hands-on, UploadVR noted the distortions on fine details such as small text that are typical of Gaussian splats.[19] At WWDC 2026, Apple added Gaussian splat rendering to RealityKit: its "Explore advances in RealityKit" session introduces GaussianSplatResource and GaussianSplatComponent, which take app-supplied buffers of splat position, scale, rotation, opacity and spherical harmonics, and demonstrates them on Apple Vision Pro.[20]

Real-time rendering, upscaling and display effects

A second group of applications uses neural networks inside the real-time rendering loop of a headset or GPU, rather than to reconstruct a scene. Facebook Reality Labs presented DeepFocus in December 2018 as "a new AI-powered rendering system that works with Half Dome", its prototype varifocal headset with eye tracking, to produce natural defocus blur outside the point of focus. The team wrote that "All previous methods of rendering highly realistic blur were based on hand-crafted mathematical models, with corner cases and limitations that lead to low-quality results and artifacts." The demonstration ran on a four-GPU machine, trained on 196,000 images, and the code and training data were open-sourced; the stated goal was real-time blur at VR resolutions on a single GPU.[8]

Its 2020 neural supersampling work targeted the cost of rendering at high resolution and refresh rate. The team wrote that "Real-time rendering in virtual reality presents a unique set of challenges", and described the method as "the first learned supersampling method that achieves significant 16x supersampling of rendered content with high spatial and temporal fidelity". The network takes color, depth and dense motion vectors from the current and several previous frames; for a 3840 x 2160 display it starts from a 960 x 540 image rendered by the game engine.[3]

On PC VR, NVIDIA's DLSS upscaling, which runs on the Tensor Cores of GeForce RTX GPUs, reached its first VR titles in May 2021: No Man's Sky, Into the Radius and Wrench. NVIDIA claimed that DLSS "doubles VR performance at the Ultra graphics preset" in No Man's Sky, maintaining 90 FPS on an Oculus Quest 2 with a GeForce RTX 3080, and accelerates Wrench by up to 80 percent.[21] NVIDIA also applies the neural rendering name to its newer GPU features. Announcing the GeForce RTX 50 series on 6 January 2025, it said that "By integrating neural networks into the rendering process, we can take dramatic leaps forward in performance, image quality, and interactivity", claimed that DLSS 4 can "multiply frame rates by up to 8X over traditional brute-force rendering", and said that RTX Neural Shaders can "compress textures by up to 7X".[22] Its RTX Kit lets developers "train and deploy AI directly within shaders", and Microsoft's DirectX 12 Agility SDK preview added Cooperative Vectors, which give shaders direct access to the Tensor Cores.[23]

Research challenges

The 2020 report called true generalization to unseen settings "one of the grand challenges for the future". Learned reenactment methods can fail for poses they were not trained on, and capturing training data for every scenario, or for every possible user of an avatar system, is not realistic. It also suggested that immersive VR and AR experiences may need neural rendering to produce multimodal output, such as spatial audio, tactile and haptic signals, rather than images alone.[1] The 2022 report listed further open problems for neural scene representations: large-scale scenes that are only partly visible in each frame, generalization to deforming scenes, training time, and fewer input views. It noted that telepresence and augmented reality "would highly benefit" from methods that render novel views of interacting, talking people together with matching sound.[2]

Both reports discuss social implications. The 2020 report warns that the technology can lower the barrier to manipulating images and video, for example through talking-head synthesis, and reviews methods for detecting synthetic imagery.[1] The 2022 report adds that training neural volumetric representations on GPU clusters uses a sizable amount of energy, and that high GPU demand may keep some research groups from contributing on equal terms.[2]

See also

References

  1. ↑ 1.00 1.01 1.02 1.03 1.04 1.05 1.06 1.07 1.08 1.09 1.10 1.11 Ayush Tewari, Ohad Fried, Justus Thies, Vincent Sitzmann, Stephen Lombardi, Kalyan Sunkavalli, Ricardo Martin-Brualla, Tomas Simon, Jason Saragih, Matthias Nießner, Rohit Pandey, Sean Fanello, Gordon Wetzstein, Jun-Yan Zhu, Christian Theobalt, Maneesh Agrawala, Eli Shechtman, Dan B Goldman, Michael Zollhöfer (2020). "State of the Art on Neural Rendering". Computer Graphics Forum, vol. 39, no. 2 (Eurographics 2020 state-of-the-art report), pp. 701-727. doi:10.1111/cgf.14022. https://arxiv.org/abs/2004.03805. Retrieved 2026-09-27.
  2. ↑ 2.00 2.01 2.02 2.03 2.04 2.05 2.06 2.07 2.08 2.09 Ayush Tewari, Justus Thies, Ben Mildenhall, Pratul Srinivasan, Edgar Tretschk, Yifan Wang, Christoph Lassner, Vincent Sitzmann, Ricardo Martin-Brualla, Stephen Lombardi, Tomas Simon, Christian Theobalt, Matthias Nießner, Jonathan T. Barron, Gordon Wetzstein, Michael Zollhöfer, Vladislav Golyanik (2022). "Advances in Neural Rendering". Computer Graphics Forum, vol. 41, no. 2 (Eurographics 2022 state-of-the-art report), pp. 703-735. doi:10.1111/cgf.14507. https://arxiv.org/abs/2111.05849. Retrieved 2026-09-27.
  3. ↑ 3.0 3.1 3.2 Lei Xiao, Salah Nouri, Matt Chapman, Alexander Fix, Douglas Lanman, Anton Kaplanyan (2020-07-01). "Introducing neural supersampling for real-time rendering". Meta Research. Facebook Reality Labs. https://research.facebook.com/blog/2020/07/introducing-neural-supersampling-for-real-time-rendering/. Retrieved 2026-09-27.
  4. ↑ 4.0 4.1 4.2 Justus Thies, Michael Zollhöfer, Matthias Nießner (2019-04-28). "Deferred Neural Rendering: Image Synthesis using Neural Textures". ACM Transactions on Graphics (SIGGRAPH 2019). https://arxiv.org/abs/1904.12356. Retrieved 2026-09-27.
  5. ↑ 5.0 5.1 5.2 Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, Ren Ng (2020-03-19). "NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis". European Conference on Computer Vision (ECCV 2020). https://arxiv.org/abs/2003.08934. Retrieved 2026-09-27.
  6. ↑ 6.0 6.1 Ali Eslami, Danilo Jimenez Rezende (2018-06-14). "Neural scene representation and rendering". Google DeepMind. https://deepmind.google/discover/blog/neural-scene-representation-and-rendering/. Retrieved 2026-09-27.
  7. ↑ 7.0 7.1 7.2 Stephen Lombardi, Jason Saragih, Tomas Simon, Yaser Sheikh (2018). "Deep Appearance Models for Face Rendering". ACM Transactions on Graphics, vol. 37, no. 4, article 68 (SIGGRAPH 2018). https://arxiv.org/abs/1808.00362. Retrieved 2026-09-27.
  8. ↑ 8.0 8.1 8.2 "Introducing DeepFocus: The AI Rendering System Powering Half Dome". Meta Quest Blog. Meta. 2018-12-19. https://www.meta.com/blog/introducing-deepfocus-the-ai-rendering-system-powering-half-dome/. Retrieved 2026-09-27.
  9. ↑ "Neural Supersampling for Real-time Rendering by Xiao, Nouri, Chapman, Fix, Lanman, et al.". ACM SIGGRAPH History Archives. ACM SIGGRAPH. https://history.siggraph.org/learning/neural-supersampling-for-real-time-rendering-by-xiao-nouri-chapman-fix-lanman-et-al/. Retrieved 2026-09-27.
  10. ↑ 10.0 10.1 Thomas Müller, Alex Evans, Christoph Schied, Alexander Keller (2022-07). "Instant Neural Graphics Primitives with a Multiresolution Hash Encoding". ACM Transactions on Graphics, vol. 41, no. 4, article 102 (SIGGRAPH 2022). https://arxiv.org/abs/2201.05989. Retrieved 2026-09-27.
  11. ↑ 11.0 11.1 11.2 11.3 Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George Drettakis (2023-07). "3D Gaussian Splatting for Real-Time Radiance Field Rendering". ACM Transactions on Graphics, vol. 42, no. 4 (SIGGRAPH 2023). https://arxiv.org/abs/2308.04079. Retrieved 2026-09-27.
  12. ↑ 12.0 12.1 12.2 Linning Xu, Vasu Agrawal, William Laney, Tony Garcia, Aayush Bansal, Changil Kim, Samuel Rota Bulò, Lorenzo Porzi, Peter Kontschieder, Aljaž Božič, Dahua Lin, Michael Zollhöfer, Christian Richardt (2023). "VR-NeRF: High-Fidelity Virtualized Walkable Spaces". SIGGRAPH Asia 2023 project page. https://vr-nerf.github.io/. Retrieved 2026-09-27.
  13. ↑ Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George Drettakis. "3D Gaussian Splatting for Real-Time Radiance Field Rendering". Inria GraphDeco project page. https://repo-sam.inria.fr/fungraph/3d-gaussian-splatting/. Retrieved 2026-09-27.
  14. ↑ 14.0 14.1 Guikun Chen, Wenguan Wang (2026-04-09). "A Survey on 3D Gaussian Splatting". ACM Computing Surveys (accepted; arXiv version 9). https://arxiv.org/abs/2401.03890. Retrieved 2026-09-27.
  15. ↑ 15.0 15.1 David Heaney (2024-11-25). "Varjo Teleport Lets You Easily Capture Scenes With An iPhone To View In PC VR". UploadVR. https://www.uploadvr.com/varjo-teleport-launch/. Retrieved 2026-09-27.
  16. ↑ 16.0 16.1 Scott Stein (2025-10-28). "Apple Vision Pro's Persona Avatars Are Evolving Fast". CNET. https://www.cnet.com/tech/computing/apple-talks-to-me-about-vision-pro-personas-where-is-our-virtual-presence-headed/. Retrieved 2026-09-27.
  17. ↑ "Facebook is building the future of connection with lifelike avatars". Tech at Meta. Meta. 2019-03-12. https://tech.facebook.com/ar-vr/2019/03/codec-avatars-facebook-reality-labs/. Retrieved 2026-09-27.
  18. ↑ "visionOS 26 introduces powerful new spatial experiences for Apple Vision Pro". Apple Newsroom. Apple. 2025-06-09. https://www.apple.com/newsroom/2025/06/visionos-26-introduces-powerful-new-spatial-experiences-for-apple-vision-pro/. Retrieved 2026-09-27.
  19. ↑ 19.0 19.1 David Heaney (2025-09-17). "Hands-On: Meta Horizon Hyperscape Captures Photorealistic VR Scenes On Quest 3". UploadVR. https://www.uploadvr.com/meta-horizon-hyperscape-photorealistic-scene-capture-quest-3/. Retrieved 2026-09-27.
  20. ↑ "Explore advances in RealityKit - WWDC26". Apple Developer. Apple. 2026. https://developer.apple.com/videos/play/wwdc2026/279/. Retrieved 2026-09-27.
  21. ↑ "NVIDIA DLSS: No Man's Sky And 8 Other Games, Including The First VR Titles, Add Performance-Accelerating Tech This Month". NVIDIA GeForce News. NVIDIA. 2021-05-18. https://www.nvidia.com/en-gb/geforce/news/may-2021-rtx-dlss-game-update. Retrieved 2026-09-27.
  22. ↑ "New GeForce RTX 50 Series Graphics Cards and Laptops Powered By NVIDIA Blackwell Bring Game-Changing AI and Neural Rendering Capabilities To Gamers and Creators". NVIDIA GeForce News. NVIDIA. 2025-01-06. https://www.nvidia.com/en-us/geforce/news/rtx-50-series-graphics-cards-gpu-laptop-announcements/. Retrieved 2026-09-27.
  23. ↑ Ike Nnoli (2025-08-18). "Announcing the Latest NVIDIA Gaming AI and Neural Rendering Technologies". NVIDIA Technical Blog. NVIDIA. https://developer.nvidia.com/blog/announcing-the-latest-nvidia-gaming-ai-and-neural-rendering-technologies/. Retrieved 2026-09-27.