Equirectangular projection
More actions
The equirectangular projection is a map projection in which meridians and parallels are drawn as equally spaced straight lines crossing at right angles, so that longitude and latitude map linearly to the horizontal and vertical axes of a rectangle.[1] Cartographers also call it the equidistant cylindrical or rectangular projection, and the form whose standard parallel is the equator is known as the plate carrée or simple cylindrical projection.[1][2] It is one of the oldest projections; Claudius Ptolemy credited its invention to Marinus of Tyre around AD 100, although it probably originated with Eratosthenes.[1]
In virtual reality the projection is the usual way to store a full sphere of imagery in a flat frame. A 360-degree photo or video in equirectangular form covers 360 degrees of longitude across its width and 180 degrees of latitude down its height, so frames are typically twice as wide as they are tall.[3][4] Coors, Condurache and Geiger describe it as "the most popular representation of 360° images".[5] Its sampling is uneven: the regions near the poles receive far more pixels than their area on the sphere warrants, while the equator, where the important content of a spherical video usually lies, receives relatively few.[6] That trade-off has driven alternative layouts such as the cubemap and research on coding and analysing equirectangular content.
Geometry
Map projection
In Snyder's description for the U.S. Geological Survey, the equidistant cylindrical projection is neither equal-area nor conformal, its meridians and parallels are equidistant straight lines intersecting at right angles, and its poles are shown as lines rather than points.[1] For a sphere of radius R, central meridian λ0 and standard parallel φ1, the forward formulas are:[1]
- x = R (λ - λ0) cos φ1
- y = R φ
The scale along every meridian is true (h = 1), while the scale along a parallel is k = cos φ1 / cos φ. Distances are therefore correct along any meridian and along the standard parallels, and east-west stretching grows with distance from them.[1][2] If the equator is the standard parallel, meridians are spaced at the same distance as parallels and the graticule appears square; this is the plate carrée.[1] For the plate carrée, k = 1 / cos φ, so a parallel at latitude 60 degrees is drawn at twice its true length, and each pole, a single point on the globe, becomes a line as long as the equator.[1][2]
The PROJ cartographic library calls it "the simplest of all projections" and notes that, although its distortions limit its use in navigation and cadastral mapping, the plate carrée "has become a standard for global raster datasets" because of the simple relationship between a pixel's position and its geographic location.[7]
Image coordinates
Applied to a viewing sphere instead of the Earth, the projection maps viewing directions to pixels. Apple's documentation for its Apple Projected Media Profile states that "the pixel coordinates of an enclosing sphere are expressed as angles of latitude and longitude and projected equally into the rows and columns of a rectangular video frame", with the horizontal axis covering longitude from -180 to +180 degrees and the vertical axis covering latitude from -90 to +90 degrees.[3] Regensky, Herglotz and Kaup give the equivalent pixel mapping: the azimuthal angle (0 to 2π) is scaled to the frame width U and the polar angle (0 to π) to the frame height V, and typically U = 2V because the azimuth spans twice the angular range of the polar angle.[4]
Google's Spherical Video V2 metadata specification fixes the orientation convention for video: the default pose has the forward vector at the centre of the frame, the up vector at the top and the right vector toward the right of the frame.[8] Because longitude wraps around, the left and right edges of an equirectangular frame are neighbours on the sphere, while the whole top edge and the whole bottom edge each represent a single point on the sphere, one of the poles.[9][1]
History
Cartography
Snyder writes that the projection "originated probably with Eratosthenes (275?-195? B.C.)". Ptolemy credited Marinus of Tyre with the invention about AD 100 and criticised his choice as "a manner of representing the distances which gives the worst results of all". On Marinus's world map only the parallel of Rhodes (latitude 36 degrees north) was true to scale, so the meridians were spaced at about four-fifths of the spacing of the parallels.[1] Ptolemy approved the projection for maps of smaller areas, with meridians spaced to give correct scale along the central parallel; all the Greek manuscript maps of the Geographia, dating from the 13th century, use that modification. The projection was used for some maps until the 18th century; Snyder notes that it is now used mainly for a few maps on which distortion matters less than the ease of displaying special information. Nordenskiöld called it the projection of Marinus in 1889, and other names include La Carte Parallelogrammatique, Die Rechteckige Plattkarte and "Equirectangular".[1] Some sources state the attribution more simply: Esri's documentation says it "was invented by Marinus of Tyre about A.D. 100".[2]
Computer graphics
Early environment mapping in computer graphics also indexed images by angle. In their 1976 paper Texture and Reflection in Computer Generated Images, James Blinn and Martin Newell of the University of Utah modelled the environment as "a two-dimensional intensity map indexed by the polar coordinate angles of the ray reflected", showing such maps "with azimuthal angle plotted as abscissa and polar angle plotted as ordinate"; when a reflection direction was computed, it was converted to polar coordinates and the reflected intensity read from the map.[10] Ned Greene's 1986 paper Environment Mapping and Other Applications of World Projections later used cubic environment maps.[11] Current engines offer both approaches; Unity's Panoramic skybox shader, for example, has a "Latitude Longitude Layout" mapping, which uses a cylindrical wrapping method to map a texture to the skybox, alongside a six-sided layout, with 360-degree and 180-degree image types.[12]
360-degree media
Early open metadata for 360-degree video assumed the equirectangular layout. Version 1 of Google's open Spherical Video metadata scheme for MP4 and Matroska/WebM files required the ProjectionType field to be equirectangular; the scheme is now superseded by Spherical Video V2.[13] Facebook's engineers wrote in October 2015 that they converted uploaded equirectangular 360 videos to a cube map layout, and in January 2016 that a pyramid-shaped encoding cut file size by 80 percent against the equirectangular original.[14][15] In March 2017 Google described the equi-angular cubemap, developed by YouTube and the Daydream team, as an alternative that spreads pixels more evenly than the equirectangular projection.[6]
Use in VR and 360-degree media
File formats and metadata
Several formats signal equirectangular content explicitly so that players know how to display it:
| Format or standard | How equirectangular content is signalled |
|---|---|
| Spherical Video V1 (Google) | XML metadata in which ProjectionType must be equirectangular; StereoMode may be mono, left-right or top-bottom[13]
|
| Spherical Video V2 (Google) | An equi box inside the projection box, with four projection_bounds fields giving the proportion cropped from each edge (all zero for an uncropped frame); the same scheme defines cbmp (cubemap) and mshp (mesh) projections[8]
|
| Photo Sphere XMP metadata (Google) | GPano:ProjectionType, for which equirectangular is the only value supported by Google products; FullPanoWidthPixels and FullPanoHeightPixels record the size of the full panorama when an image is cropped[16]
|
| Apple Projected Media Profile (visionOS 26) | Projection kind "equirectangular" for 360-degree video and "half-equirectangular" for 180-degree video, in which the horizontal axis covers -90 to +90 degrees[3] |
| MPEG OMAF (ISO/IEC 23090-2) | The HEVC-based viewport-independent video profiles use equirectangular projection; the HEVC-based and AVC-based viewport-dependent profiles allow equirectangular or cubemap projection[17] |
| 3GPP TS 26.118 VR video operation points | The Basic H.264/AVC operation point allows only equirectangular projection with full 360-degree coverage; the Main and Main 8K H.265/HEVC points allow equirectangular with any coverage, and the Flexible H.265/HEVC point allows equirectangular or cubemap[17] |
Capture devices and services also produce or require the format. Ricoh's THETA API documentation states that stitched still images are saved in equirectangular format and unstitched ones in dual-fisheye format.[18] Ricoh's API lists still-image sizes with the same 2:1 shape, for example 6720 x 3360 pixels for the THETA Z1 and 11008 x 5504 for the THETA X.[19] YouTube's help page for 360-degree live streams states that "YouTube only supports equirectangular projection for 360 videos at this time".[20] Apple's documentation says the projection is widely supported by editing applications such as Final Cut Pro, which reads and writes the Apple Projected Media Profile for 360-degree formats.[3]
Uploading in equirectangular form does not mean a service stores or streams it that way. In a January 2020 analysis, Paul Bourke noted that 360 video uploaded to YouTube is generally equirectangular but that YouTube "does not retain this format but instead remaps the footage" into a cube map layout.[21] The VR180 format also departs from it: Google's VR180 specification requires a mesh projection, which lets cameras keep raw fisheye pixels and avoid reprojection.[22]
Playback and rendering
To display an equirectangular image in a head-mounted display, the image is mapped onto the inside of a sphere.[23] XR runtimes provide this as a compositor feature. The OpenXR standard from the Khronos Group has two ratified extensions, XR_KHR_composition_layer_equirect and XR_KHR_composition_layer_equirect2, whose layer structures contain "the information needed to render an equirectangular image onto a sphere"; the application supplies the sphere's pose and radius (zero or infinity meaning an infinite sphere) and either a texture-coordinate scale and bias or the horizontal and vertical angles that the image covers.[24][25] The W3C WebXR Layers API draft defines an equivalent XREquirectLayer for browsers, in which the XR compositor "MUST map an equirectangular coded data onto the inside of a sphere".[23]
Engines can also produce equirectangular output. Unity's RenderTexture.ConvertToEquirect converts a cubemap render texture to "equirectangular format (both stereoscopic or monoscopic equirect)"; in the stereoscopic version "the left eye will occupy the top half and the right eye will occupy the bottom".[26] Stereoscopic 360-degree content can be stored in this projection by packing the two eye views into one frame, top-bottom or left-right; Spherical Video V1 states that cropping, initial view and projection properties are shared by the two eyes, with each eye's region treated as a separate frame.[13]
Sampling and distortion
Because every row of an equirectangular frame has the same number of pixels, while the circumference of a parallel shrinks toward the poles, the projection oversamples high latitudes. Google's Chip Brown summarised the problem in 2017: "the poles get a lot of pixels, and the equator gets relatively few", although spherical videos usually have their important content around the equator, the viewer's horizon, and the projection "has high distortion, which makes existing video compression technology work harder".[6] Facebook's engineers likewise described the layout as containing "redundant information" at the poles, comparing it to the way Antarctica is stretched into a line on a world map.[14][15] The oversampling is greatest where viewers look least: studies of head and eye movement summarised by Xu, Li, Zhang and Le Callet found that subjects view regions near the equator more frequently, a pattern called equator bias.[9]
The distortion also affects image analysis. Coors and colleagues note that the equirectangular representation "suffers from heavy distortions in the polar regions which implies that an object will appear differently depending on its latitudinal position", and Regensky and colleagues show that blocks of constant size in an equirectangular frame become increasingly distorted on the sphere with distance from the equator.[5][4]
Other projections trade the equirectangular format's simplicity for more even sampling:
| Layout | Reported comparison with equirectangular |
|---|---|
| Cube map | Facebook reported that its cube map conversion used 25 percent fewer pixels per frame than the equirectangular input,[14] and Xu et al. note that a cubemap has less geometric distortion[9] |
| Equi-angular cubemap | Google reported pixel density "significantly more uniform" than either equirectangular or a standard cubemap[6] |
| Pyramid (Facebook, 2016) | Viewport-dependent layout that Facebook said reduced file size by 80 percent against the equirectangular original[15] |
| Mesh projection | Generic mapping defined in Spherical Video V2 and required by VR180 for fisheye frames[8][22] |
Research
Quality measurement
Ordinary PSNR treats every pixel equally, so on an equirectangular frame it over-weights errors near the poles.[9] Several metrics correct for this. Yu, Lakshman and Girod's spherical PSNR (S-PSNR), presented at the 2015 IEEE International Symposium on Mixed and Augmented Reality, computes PSNR over a set of points distributed uniformly on the sphere; its official implementation uses 655,362 points. Zakharchenko and colleagues instead compute PSNR after reprojecting to the Craster parabolic projection (CPP-PSNR), which has uniform sampling density.[9] Sun, Lu and Yu's weighted-to-spherically-uniform PSNR (WS-PSNR), published in IEEE Signal Processing Letters in 2017, keeps the computation on the projection plane but multiplies each pixel's error by a weight that compensates for the projection's stretching, so that pixels covering equal areas of the sphere have equal influence; for equirectangular frames the weight is the cosine of the pixel row's latitude.[9] The same cosine weight map has been applied to SSIM in the W-SSIM and WS-SSIM variants.[9]
Video coding
Coding research has targeted both problems of the projection: oversampled poles and the artificial left and right borders. To reduce wasted bits, Budagavi and colleagues applied Gaussian smoothing to the top and bottom regions of equirectangular video, Youvalari and colleagues split frames into strips down-sampled by latitude, and Tang and colleagues added a latitude-dependent multiplier to the quantization parameter to give higher quality near the equator.[9] For the borders, Li and colleagues proposed padding the left edge with pixels from near the right edge and vice versa, reflecting the cyclic nature of the format.[9] According to Regensky, Herglotz and Kaup, a geometrically correct wrap-around padding for equirectangular video proposed by He and colleagues was integrated into the H.266/VVC standard because of its low computational complexity.[4] Their own motion-plane-adaptive inter prediction, which performs motion compensation on planes in 3D space instead of on the projected image, reported average Bjøntegaard Delta rate savings of 1.72 percent (PSNR) and 1.56 percent (WS-PSNR) over the VVC VTM-14.2 reference software.[4]
Computer vision
Neural networks trained on perspective images do not transfer directly to equirectangular input. Su and Grauman proposed learning a spherical convolutional network that "translates a planar CNN to process 360° imagery directly in its equirectangular projection", reproducing the outputs of flat filters while saving orders of magnitude in computation compared with reprojecting to tangent planes.[27] SphereNet, by Coors, Condurache and Geiger at ECCV 2018, adapts the sampling locations of convolutional filters to undo the distortion and wraps the filters around the sphere, so that existing perspective models can be transferred to omnidirectional images.[5] In saliency prediction for 360-degree images, some methods run 2D models on rotated equirectangular versions of an image to reduce border artifacts, while others work on extracted viewports or cubemap faces to avoid the projection's geometric distortion.[9]
See also
References
- ↑ 1.00 1.01 1.02 1.03 1.04 1.05 1.06 1.07 1.08 1.09 1.10 John P. Snyder (1987). "Map Projections: A Working Manual (Professional Paper 1395), chapter 12, Equidistant Cylindrical projection, pp. 90-91". U.S. Geological Survey. doi:10.3133/pp1395. https://pubs.usgs.gov/pp/1395/report.pdf. Retrieved 2026-10-06.
- ↑ 2.0 2.1 2.2 2.3 "Equidistant cylindrical". ArcMap documentation. Esri. https://desktop.arcgis.com/en/arcmap/latest/map/projections/equidistant-cylindrical.htm. Retrieved 2026-10-06.
- ↑ 3.0 3.1 3.2 3.3 "Learn about the Apple Projected Media Profile (WWDC25 session 297)". Apple Developer. Apple. 2025-06. https://developer.apple.com/videos/play/wwdc2025/297/. Retrieved 2026-10-06.
- ↑ 4.0 4.1 4.2 4.3 4.4 Andy Regensky, Christian Herglotz, André Kaup (2022). "Motion-Plane-Adaptive Inter Prediction in 360-Degree Video Coding". arXiv preprint 2202.03323 (Friedrich-Alexander University Erlangen-Nürnberg). https://arxiv.org/abs/2202.03323. Retrieved 2026-10-06.
- ↑ 5.0 5.1 5.2 Benjamin Coors, Alexandru Paul Condurache, Andreas Geiger (2018). "SphereNet: Learning Spherical Representations for Detection and Classification in Omnidirectional Images". Computer Vision - ECCV 2018, Lecture Notes in Computer Science, pp. 525-541. doi:10.1007/978-3-030-01240-3_32. https://www.cvlibs.net/publications/Coors2018ECCV.pdf. Retrieved 2026-10-06.
- ↑ 6.0 6.1 6.2 6.3 Chip Brown (2017-03-14). "Bringing pixels front and center in VR video". The Keyword. Google. https://blog.google/products/google-ar-vr/bringing-pixels-front-and-center-vr-video/. Retrieved 2026-10-06.
- ↑ "Equidistant Cylindrical (Plate Carrée)". PROJ documentation. PROJ contributors. https://proj.org/en/stable/operations/projections/eqc.html. Retrieved 2026-10-06.
- ↑ 8.0 8.1 8.2 "Spherical Video V2 RFC". google/spatial-media on GitHub. Google. https://github.com/google/spatial-media/blob/master/docs/spherical-video-v2-rfc.md. Retrieved 2026-10-06.
- ↑ 9.00 9.01 9.02 9.03 9.04 9.05 9.06 9.07 9.08 9.09 Mai Xu, Chen Li, Shanyi Zhang, Patrick Le Callet (2020). "State-of-the-art in 360° Video/Image Processing: Perception, Assessment and Compression". IEEE Journal of Selected Topics in Signal Processing, vol. 14, no. 1, pp. 5-26. doi:10.1109/JSTSP.2020.2966864. https://arxiv.org/abs/1905.00161. Retrieved 2026-10-06.
- ↑ James F. Blinn, Martin E. Newell (1976-10). "Texture and Reflection in Computer Generated Images". Communications of the ACM, vol. 19, no. 10, pp. 542-547. doi:10.1145/360349.360353. https://doi.org/10.1145/360349.360353. Retrieved 2026-10-06.
- ↑ Paul Debevec. "Reflection Mapping History". pauldebevec.com. https://www.pauldebevec.com/ReflectionMapping/. Retrieved 2026-10-06.
- ↑ "Panoramic Skybox Shader". Unity Manual. Unity Technologies. https://docs.unity3d.com/Manual/shader-skybox-panoramic.html. Retrieved 2026-10-06.
- ↑ 13.0 13.1 13.2 "Spherical Video RFC". google/spatial-media on GitHub. Google. https://github.com/google/spatial-media/blob/master/docs/spherical-video-rfc.md. Retrieved 2026-10-06.
- ↑ 14.0 14.1 14.2 David Pio, Evgeny Kuzyakov (2015-10-15). "Under the hood: Building 360 video". Engineering at Meta. Meta. https://engineering.fb.com/2015/10/15/video-engineering/under-the-hood-building-360-video/. Retrieved 2026-10-06.
- ↑ 15.0 15.1 15.2 Evgeny Kuzyakov, David Pio (2016-01-21). "Next-generation video encoding techniques for 360 video and VR". Engineering at Meta. Meta. https://engineering.fb.com/2016/01/21/virtual-reality/next-generation-video-encoding-techniques-for-360-video-and-vr/. Retrieved 2026-10-06.
- ↑ "Photo Sphere XMP Metadata". Google for Developers. Google. https://developers.google.com/streetview/spherical-metadata. Retrieved 2026-10-06.
- ↑ 17.0 17.1 Sachin Deshpande, Miska M. Hannuksela (2021). "Omnidirectional MediA Format (OMAF): Toolbox for Virtual Reality Services". 2021 IEEE Conference on Standards for Communications and Networking (CSCN), pp. 20-25. doi:10.1109/CSCN53733.2021.9686150. https://arxiv.org/abs/2203.01183. Retrieved 2026-10-06.
- ↑ "0xD834 Image Stitching". RICOH THETA API documentation. Ricoh. https://docs-theta-api.ricoh360.com/usb-api/property/image_stitching.html. Retrieved 2026-10-06.
- ↑ "fileFormat (THETA Web API v2.1 options)". ricohapi/theta-api-specs on GitHub. Ricoh. https://github.com/ricohapi/theta-api-specs/blob/main/theta-web-api-v2.1/options/file_format.md. Retrieved 2026-10-06.
- ↑ "Encoder settings for 360-degree live streams". YouTube Help. Google. https://support.google.com/youtube/answer/6396222?hl=en. Retrieved 2026-10-06.
- ↑ Paul Bourke (2020-01). "YouTube 360 video format". paulbourke.net. https://paulbourke.net/panorama/youtubeformat/. Retrieved 2026-10-06.
- ↑ 22.0 22.1 "VR180 Video Format". google/spatial-media on GitHub. Google. https://github.com/google/spatial-media/blob/master/docs/vr180.md. Retrieved 2026-10-06.
- ↑ 23.0 23.1 "WebXR Layers API Level 1 (W3C Working Draft)". World Wide Web Consortium. W3C. https://www.w3.org/TR/webxrlayers-1/. Retrieved 2026-10-06.
- ↑ "XrCompositionLayerEquirectKHR(3)". OpenXR Registry. Khronos Group. https://registry.khronos.org/OpenXR/specs/1.1/man/html/XrCompositionLayerEquirectKHR.html. Retrieved 2026-10-06.
- ↑ "XrCompositionLayerEquirect2KHR(3)". OpenXR Registry. Khronos Group. https://registry.khronos.org/OpenXR/specs/1.1/man/html/XrCompositionLayerEquirect2KHR.html. Retrieved 2026-10-06.
- ↑ "RenderTexture.ConvertToEquirect". Unity Scripting API. Unity Technologies. https://docs.unity3d.com/ScriptReference/RenderTexture.ConvertToEquirect.html. Retrieved 2026-10-06.
- ↑ Yu-Chuan Su, Kristen Grauman (2017). "Learning Spherical Convolution for Fast Features from 360° Imagery". Advances in Neural Information Processing Systems 30 (NIPS 2017). https://arxiv.org/abs/1708.00919. Retrieved 2026-10-06.