Voxel
More actions
A voxel is the basic unit of a three-dimensional digital representation of an image or object; the American Heritage Dictionary gives the word's origin as a blend of "volumetric" and "pixel".[1] Where a pixel is one sample on a 2D image grid, a voxel is one sample on a regular 3D grid. The value it holds may be a color, density, heat or pressure, or even a vector such as velocity, and the array that stores these samples is also called a volume buffer, cubic frame buffer or 3D raster.[2]
Voxel data comes from scanning, simulation or modeling. CT and MRI scanners produce stacks of slices that are reconstructed into voxel volumes, and depth cameras can be fused into voxel grids that describe the shape of a room.[2][3] In virtual and augmented reality, voxels appear in several places: medical volumes viewed in a head-mounted display, dense room reconstruction for spatial mapping, the meshing settings of AR software development kits, voxel-based sculpting and game worlds, and some recent scene representations for novel-view synthesis.
Definition
In Kaufman's formulation, a volume dataset is a set of samples, each giving the value of some measured property at a location (x, y, z). When the samples lie on a regular grid, the position of each sample is implied by its place in a 3D array rather than stored with it. Spacing can be the same along all three axes (isotropic) or use three different constants (anisotropic).[2] The samples define values only at grid points; a value elsewhere comes from interpolation. With the simplest, nearest-neighbor interpolation, each sample is surrounded by a region of constant value, and that region "is known as a voxel", a rectangular cuboid with six faces, twelve edges and eight corners. Trilinear interpolation instead lets the value vary linearly between neighboring samples.[2]
In 3D discrete topology a voxel is treated as the unit cube centered on an integer grid point. Two voxels that share a face are 6-adjacent, those that share a face or an edge are 18-adjacent, and those that share a face, edge or vertex are 26-adjacent; these adjacency rules decide, for example, whether a voxelized surface lets a discrete ray slip through it. A binary voxel stores only 1 (opaque object) or 0 (transparent background). Multivalued voxels hold a value between 0 and 1 for partial coverage, density or opacity, which allows 3D antialiasing and higher-quality rendering.[2]
Voxels compared with polygons
Kaufman compares the move from surface (polygon) graphics to volume graphics with the earlier move from vector to raster graphics. In volume graphics every object is converted, or voxelized, into one uniform primitive, the voxel. Rendering then depends mainly on the fixed resolution of the volume buffer rather than on the number or complexity of objects in the scene, and the representation supports Boolean and block operations, interior structure, and amorphous phenomena.[2]
The same notes list the weaknesses. Memory is the first: at a "medium resolution" of 512 x 512 x 512 voxels and two bytes per voxel, one volume buffer occupies 256 MB. Because the scene is stored in discrete form, low-resolution volumes show aliasing, especially when the viewer zooms in, and rotations by angles other than 90 degrees degrade the data. A voxelized object also keeps no record of its original geometric definition, so surface normals have to be estimated from neighboring voxels.[2] Sparse and hierarchical data structures, described below, were developed largely to reduce the memory cost.
History
According to Kaufman, the volumetric approach first took hold in applications whose data is already volumetric, such as 3D medical imaging and scientific visualization.[2]
| Year | Work | Contribution |
|---|---|---|
| 1982 | Donald Meagher, "Geometric Modeling Using Octree Encoding" (Computer Graphics and Image Processing) | Represented arbitrary 3D objects to any chosen resolution in a hierarchical 8-ary tree, the octree, with memory on the order of the object's surface area[4] |
| 1987 | William E. Lorensen and Harvey E. Cline, marching cubes (SIGGRAPH 1987) | Built triangle models of constant-density surfaces from CT, MR and SPECT data, using a case table that defines triangle topology[5] |
| 1988 | Marc Levoy, "Display of Surfaces from Volume Data" (IEEE Computer Graphics and Applications) | Direct volume rendering: shading and opacity computed at every voxel and composited along viewing rays, with no geometric primitives fitted to the data[6] |
| 1993 | Arie Kaufman, D. Cohen and R. Yagel, "Volume Graphics" (IEEE Computer, vol. 26, no. 7, July 1993) | The paper Kaufman's later survey cites for volume graphics: the synthesis, modeling, manipulation and rendering of objects stored in a volume buffer of voxels[2] |
| 1996 | Brian Curless and Marc Levoy (SIGGRAPH 1996) | Merged range scans into a voxel grid holding a cumulative weighted signed distance function; integrated as many as 70 range images into models of up to 2.6 million triangles[7] |
| 2009 | Cyril Crassin, Fabrice Neyret, Sylvain Lefebvre and Elmar Eisemann, GigaVoxels (I3D 2009) | Ray-guided streaming of large volumes into GPU memory, reaching interactive to real-time rates for several billion voxels[8] |
| 2010 | Samuli Laine and Tero Karras, "Efficient Sparse Voxel Octrees" (I3D 2010) | A compact sparse voxel octree with an efficient GPU ray-cast algorithm, aimed at using voxels as a generic geometry representation[9] |
| 2011 | Richard Newcombe, Shahram Izadi and colleagues, KinectFusion (ISMAR 2011) | Real-time fusion of a moving depth camera's data into a voxel grid of truncated signed distances on the GPU[3] |
| 2012 | OpenVDB | Released as open source by DreamWorks Animation on 3 August 2012; a hierarchical data structure for sparse volumetric data on 3D grids[10][11] |
| 2013 | Matthias Nießner, Michael Zollhöfer, Shahram Izadi and Marc Stamminger, voxel hashing (ACM Transactions on Graphics) | Stored surface data densely only where depth measurements exist, using a spatial hash instead of a regular or hierarchical grid[12] |
| 2022 | Sara Fridovich-Keil, Alex Yu and colleagues, Plenoxels (CVPR 2022) | Radiance fields stored in a sparse voxel grid with density and spherical harmonic coefficients, optimized without a neural network[13] |
Rendering and data structures
Voxel data is shown in one of two ways. Surface extraction converts the grid into triangles. Marching cubes creates triangle models of constant-density surfaces: it processes the data in scan-line order, computes triangle vertices by linear interpolation and shades the result with the normalized gradient of the original data.[5] Direct volume rendering skips the triangles. In Levoy's 1988 method, shading is computed at every voxel with the local gradient serving as the surface normal, a classification step assigns each voxel a partial opacity, and colors and opacities are composited from back to front along viewing rays. Levoy demonstrated it on molecular graphics and medical imaging.[6]
Dense grids grow with the cube of their resolution, so most real-time systems store voxels sparsely. In Meagher's octree, a node whose value completely describes its region is a leaf; only a node that remains ambiguous is divided into eight child octants, so storage concentrates near the object's surface.[4] GigaVoxels adapts its data representation to the current view and occlusion, and uses information extracted during ray casting to decide which data to produce and stream into GPU memory.[8] Laine and Karras at NVIDIA reported that ray casting their sparse voxel octree was competitive with triangle ray tracing while allowing much greater geometric detail; their format adds contour information to each voxel, a compressed normal format and a post-process filter.[9] Crassin, Neyret, Sainz, Green and Eisemann used a voxel octree that is built and updated on the fly from an ordinary triangle mesh to approximate indirect lighting by cone tracing, handling two light bounces at 25 to 70 frames per second.[14] In film production, OpenVDB, originally created at DreamWorks Animation and now maintained by the Academy Software Foundation, is an Academy Award-winning C++ library for storing and manipulating sparse volumes; its tree structure was described by Ken Museth in ACM Transactions on Graphics in 2013.[10][11]
Stereoscopic display doubles the rendering load, because a separate image is needed for each eye. A Stony Brook group (Ming Wan, Nan Zhang, Arie Kaufman and Huamin Qu) presented a stereoscopic renderer for voxel-based terrain at IEEE Virtual Reality 2000 for their virtual flythrough system. It ray casts the left-eye image and builds most of the right-eye image by reprojecting pixels from the left one, ray casting only the remaining pixels.[15]
Applications in VR and AR
Scene reconstruction and spatial mapping
Voxel grids are a common way to turn the depth data of a moving headset or phone into a model of the real environment. KinectFusion, presented at the IEEE International Symposium on Mixed and Augmented Reality in 2011, fused the depth stream of a Microsoft Kinect into a single global surface model in real time on commodity graphics hardware. The model is a discretized truncated signed distance function held in GPU memory, and each voxel stores a truncated signed distance value and a weight. The authors reported that their GPU could update more than 65 gigavoxels per second, about 2 ms for a full update of a 512 x 512 x 512 voxel volume, and evaluated resolutions from 64 x 64 x 64 to 512 x 512 x 512 voxels inside a 3 m3 volume. They noted that the system worked well for rooms of up to 7 m3, and that the dense volume would need too much memory for a whole building.[3] A companion paper at UIST 2011 used the same reconstruction for geometry-aware augmented reality, physics-based interaction and multi-touch input on any reconstructed surface.[16] The signed-distance idea goes back to Curless and Levoy's 1996 volumetric method,[7] and voxel hashing later replaced the regular grid with a spatial hash, storing surface data densely only where measurements are observed and streaming it in and out of the hash table as the sensor moves.[12] In robotics, OctoMap applies an octree to probabilistic 3D occupancy mapping, modeling occupied, free and unknown space.[17]
For two early games on the first Microsoft HoloLens, Fragments and Young Conker, Microsoft's development partner Asobo Studio created a spatial understanding technology that identifies surfaces such as floors, walls, tables and places where a character could sit. According to Microsoft's developer case study, the resulting module stores the scanned playspace as a grid of 8 cm voxel cubes aligned with the room's main axes, generates a mesh about once per second by extracting the isosurface from the voxel volume, and answers ray casts against this voxel representation, with each voxel holding surface elements (surfels) that carry topology labels. Microsoft and Asobo released the code as open source in the Mixed Reality Toolkit. Microsoft HoloLens 2 provides a separate Scene Understanding runtime.[18]
Phone AR toolkits expose the voxel size as a developer setting. In Niantic Spatial's NSDK 3.17.0, the meshing extension's VoxelSize property sets the size, in meters, of the voxel elements in the scene representation; larger values reduce memory use but also reduce the precision of the surface. The mesh block size used to build colliders is rounded to a multiple of the voxel size, and a separate option fuses only depth keyframes into the mesh.[19]
Medical and scientific visualization
CT and MRI data are voxel volumes, and VR viewers can render them directly without first converting them to meshes. David Shattuck's multiuser VR environment for neuroimaging, published in 2018, was developed for the HTC Vive with the OpenVR SDK; it stores volumetric data on the GPU as a 3D texture and draws it with slice-based volume rendering, intersecting the volume with a series of user-facing planes rendered with translucency.[20] A 2024 study in the International Journal of Computer Assisted Radiology and Surgery used Specto VR version 4.0, which renders imaging data in real time with ray casting at 90 frames per second per eye, on HTC Vive and HP Mixed Reality headsets. The abdominal CT scans had in-plane voxel dimensions of about 0.8 mm and a mean slice thickness of about 2.4 mm. When 9 surgeons and 3 radiologists judged the resectability of pancreatic cancer, the median number of correct assessments out of six CT scans was 5.5 in VR and 3 with standard 2D PACS viewing.[21]
Sculpting and content creation
Oculus Medium, Oculus's VR sculpting application, stores the sculpture as a signed distance field rather than as a triangle mesh. In a 2017 developer post, David Farrell explained that Medium keeps the distance values in blocks of 8 x 8 x 8 voxels inside a narrow band two voxels wide on each side of the surface, stores those blocks sparsely to save memory, and renders triangle meshes generated with Transvoxel, a technique similar to marching cubes that produces level-of-detail meshes which stitch together without seams.[22]
Games
Some games that have been playable in VR build their worlds from voxels. At GDC 2017, Innes McKendrick of Hello Games described the world generation pipeline of No Man's Sky, which runs from voxel-based generation through polygonization and texturing to population and simulation.[23] The game became playable in VR with the Beyond update on 14 August 2019, on Steam headsets including the HTC Vive, Valve Index and Oculus Rift, and on PlayStation VR.[24] Minecraft worlds are built from blocks; Mojang ended VR and mixed reality headset support in Minecraft: Bedrock Edition after March 2025, although the community-made ViveCraft mod still adds VR support to the Java edition.[25] GANcraft (ICCV 2021) converts a semantically labeled Minecraft-style block world into photorealistic imagery with a voxel-bounded neural radiance field, assigning a learnable feature vector to every corner of the blocks and using trilinear interpolation inside each voxel.[26]
Neural scene representations
Voxel grids also appear in methods that reconstruct photographed scenes for free-viewpoint viewing. Plenoxels replaces the neural network of a neural radiance field with a sparse voxel grid that stores a density and spherical harmonic coefficients at each voxel, optimized through differentiable volume rendering. The authors report a typical optimization time of 11 minutes on a single GPU, which they describe as a speedup of two orders of magnitude over NeRF, at comparable image quality.[13]
Volumetric displays
In display engineering the word refers to a physical point of light. A 2023 review in Current Optics and Photonics describes volumetric 3D displays as generating voxels so that viewers can see a three-dimensional virtual object from various angles, and notes that such displays avoid the vergence-accommodation conflict that affects other types of 3D display.[27]
See also
References
- ↑ "voxel". The American Heritage Dictionary of the English Language. HarperCollins. https://ahdictionary.com/word/search.html?q=voxel. Retrieved 2026-10-06.
- ↑ 2.0 2.1 2.2 2.3 2.4 2.5 2.6 2.7 2.8 Arie E. Kaufman. "Introduction to Volume Graphics". Center for Visual Computing, State University of New York at Stony Brook (course reading hosted by Duke University). https://courses.cs.duke.edu/spring03/cps296.8/papers/KaufmanVolumeGraphics.pdf. Retrieved 2026-10-06.
- ↑ 3.0 3.1 3.2 Richard A. Newcombe, Shahram Izadi, Otmar Hilliges, David Molyneaux, David Kim, Andrew J. Davison, Pushmeet Kohli, Jamie Shotton, Steve Hodges, Andrew Fitzgibbon (2011-10). "KinectFusion: Real-Time Dense Surface Mapping and Tracking". 10th IEEE International Symposium on Mixed and Augmented Reality (ISMAR 2011), pp. 127-136. Microsoft Research. doi:10.1109/ISMAR.2011.6092378. https://www.microsoft.com/en-us/research/wp-content/uploads/2016/11/ismar_2011.pdf. Retrieved 2026-10-06.
- ↑ 4.0 4.1 Donald Meagher (1982). "Geometric Modeling Using Octree Encoding". Computer Graphics and Image Processing, vol. 19, no. 2, pp. 129-147. doi:10.1016/0146-664X(82)90104-6. http://fab.cba.mit.edu/classes/S62.12/docs/Meagher_octree.pdf. Retrieved 2026-10-06.
- ↑ 5.0 5.1 William E. Lorensen, Harvey E. Cline (1987). "Marching cubes: A high resolution 3D surface construction algorithm". SIGGRAPH 1987 (ACM SIGGRAPH History Archives). doi:10.1145/37401.37422. https://history.siggraph.org/learning/marching-cubes-a-high-resolution-3d-surface-construction-algorithm-by-lorensen-and-cline/. Retrieved 2026-10-06.
- ↑ 6.0 6.1 Marc Levoy (1988-05). "Display of Surfaces from Volume Data". IEEE Computer Graphics and Applications, vol. 8, no. 3, pp. 29-37. https://graphics.stanford.edu/papers/volume-cga88/. Retrieved 2026-10-06.
- ↑ 7.0 7.1 Brian Curless, Marc Levoy (1996). "A volumetric method for building complex models from range images". SIGGRAPH 1996 (ACM SIGGRAPH History Archives). doi:10.1145/237170.237269. https://history.siggraph.org/learning/a-volumetric-method-for-building-complex-models-from-range-images-by-curless-and-levoy/. Retrieved 2026-10-06.
- ↑ 8.0 8.1 Cyril Crassin, Fabrice Neyret, Sylvain Lefebvre, Elmar Eisemann (2009-02). "GigaVoxels: Ray-Guided Streaming for Efficient and Detailed Voxel Rendering". ACM SIGGRAPH Symposium on Interactive 3D Graphics and Games (I3D). Inria. https://artis.inrialpes.fr/Publications/2009/CNLE09/. Retrieved 2026-10-06.
- ↑ 9.0 9.1 Samuli Laine, Tero Karras (2010-02). "Efficient Sparse Voxel Octrees". ACM SIGGRAPH Symposium on Interactive 3D Graphics and Games (I3D). NVIDIA Research. doi:10.1145/1730804.1730814. https://research.nvidia.com/publication/2010-02_efficient-sparse-voxel-octrees. Retrieved 2026-10-06.
- ↑ 10.0 10.1 "OpenVDB Documentation". OpenVDB. Academy Software Foundation. https://www.openvdb.org/documentation/. Retrieved 2026-10-06.
- ↑ 11.0 11.1 "OpenVDB". OpenVDB. Academy Software Foundation. https://www.openvdb.org/. Retrieved 2026-10-06.
- ↑ 12.0 12.1 Matthias Nießner, Michael Zollhöfer, Shahram Izadi, Marc Stamminger (2013). "Real-time 3D Reconstruction at Scale using Voxel Hashing". ACM Transactions on Graphics, vol. 32, no. 6. doi:10.1145/2508363.2508374. https://niessnerlab.org/projects/niessner2013hashing.html. Retrieved 2026-10-06.
- ↑ 13.0 13.1 Sara Fridovich-Keil, Alex Yu, Matthew Tancik, Qinhong Chen, Benjamin Recht, Angjoo Kanazawa (2022). "Plenoxels: Radiance Fields without Neural Networks". IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2022). https://alexyu.net/plenoxels/. Retrieved 2026-10-06.
- ↑ Cyril Crassin, Fabrice Neyret, Miguel Sainz, Simon Green, Elmar Eisemann (2011-09). "Interactive Indirect Illumination Using Voxel Cone Tracing". Computer Graphics Forum, vol. 30, no. 7 (Proceedings of Pacific Graphics 2011). Inria. https://artis.inrialpes.fr/Publications/2011/CNSGE11b/. Retrieved 2026-10-06.
- ↑ Ming Wan, Nan Zhang, Arie Kaufman, Huamin Qu (2000-03). "Interactive Stereoscopic Rendering of Voxel-based Terrain". Proceedings IEEE Virtual Reality 2000, pp. 197-206. doi:10.1109/VR.2000.840499. https://cse.hkust.edu.hk/~huamin/vr_00.pdf. Retrieved 2026-10-06.
- ↑ Shahram Izadi, David Kim, Otmar Hilliges, David Molyneaux, Richard Newcombe, Pushmeet Kohli, Jamie Shotton, Steve Hodges, Dustin Freeman, Andrew Davison, Andrew Fitzgibbon (2011-10). "KinectFusion: Real-time 3D Reconstruction and Interaction Using a Moving Depth Camera". Proceedings of the 24th Annual ACM Symposium on User Interface Software and Technology (UIST 2011). Microsoft Research. https://www.microsoft.com/en-us/research/publication/kinectfusion-real-time-3d-reconstruction-and-interaction-using-a-moving-depth-camera/. Retrieved 2026-10-06.
- ↑ Armin Hornung, Kai M. Wurm, Maren Bennewitz, Cyrill Stachniss, Wolfram Burgard (2013). "OctoMap: An Efficient Probabilistic 3D Mapping Framework Based on Octrees". Autonomous Robots. doi:10.1007/s10514-012-9321-0. https://octomap.github.io/. Retrieved 2026-10-06.
- ↑ Jeff Evertt (2018-03-21). "Case study - Expanding the spatial mapping capabilities of HoloLens". Microsoft Learn. Microsoft. https://learn.microsoft.com/en-us/windows/mixed-reality/out-of-scope/case-study-expanding-the-spatial-mapping-capabilities-of-hololens. Retrieved 2026-10-06.
- ↑ "class LightshipMeshingExtension". Niantic Spatial SDK 3.17.0 documentation. Niantic Spatial. https://www.nianticspatial.com/docs/nsdk/3.17.0/apiref/Niantic/Lightship/AR/Meshing/LightshipMeshingExtension/. Retrieved 2026-10-06.
- ↑ David W. Shattuck (2018). "Multiuser virtual reality environment for visualising neuroimaging data". Healthcare Technology Letters, vol. 5, no. 5, pp. 183-188. doi:10.1049/htl.2018.5077. https://pmc.ncbi.nlm.nih.gov/articles/PMC6222246/. Retrieved 2026-10-06.
- ↑ J. M. Kunz, P. Maloca, A. Allemann, et al. (2024). "Assessment of resectability of pancreatic cancer using novel immersive high-performance virtual reality rendering of abdominal computed tomography and magnetic resonance imaging". International Journal of Computer Assisted Radiology and Surgery, vol. 19, no. 9. doi:10.1007/s11548-023-03048-0. https://pmc.ncbi.nlm.nih.gov/articles/PMC11365822/. Retrieved 2026-10-06.
- ↑ David Farrell (2017-10-03). "Medium Under the Hood: Part 2 - Move Tool Implementation". Meta for Developers blog. Meta. https://developers.meta.com/vr/blog/medium-under-the-hood-part-2-move-tool-implementation/. Retrieved 2026-10-06.
- ↑ Innes McKendrick (2017). "Continuous World Generation in 'No Man's Sky'". GDC Vault. Game Developers Conference. https://www.gdcvault.com/play/1024265/Continuous_World_Generation_in__No_Man_s_Sky_. Retrieved 2026-10-06.
- ↑ Gabriel Moss (2019-08-19). "'No Man's Sky' VR Review - A Wonderful, Deeply Flawed Space Odyssey". Road to VR. https://www.roadtovr.com/no-mans-sky-vr-review-wonderful-deeply-flawed-space-odyssey/. Retrieved 2026-10-06.
- ↑ Laurent Giret (2025-05-07). "Minecraft Bedrock Edition Drops Virtual and Mixed Reality Support". Thurrott.com. https://www.thurrott.com/games/320608/minecraft-bedrock-edition-drops-virtual-and-mixed-reality-support. Retrieved 2026-10-06.
- ↑ Zekun Hao, Arun Mallya, Serge Belongie, Ming-Yu Liu (2021). "GANcraft: Unsupervised 3D Neural Rendering of Minecraft Worlds". IEEE/CVF International Conference on Computer Vision (ICCV 2021). NVIDIA Research. https://nvlabs.github.io/GANcraft/. Retrieved 2026-10-06.
- ↑ Joonku Hahn, Woonchan Moon, Hosung Jeon, Minwoo Jung, Seongju Lee, Gunhee Lee, Muhan Choi (2023). "Volumetric 3D Display: Features and Classification". Current Optics and Photonics, vol. 7, no. 6, pp. 597-607. doi:10.3807/COPP.2023.7.6.597. https://koreascience.or.kr/article/JAKO202336161705871.page. Retrieved 2026-10-06.