Ambisonics
More actions
Ambisonics is a full-sphere surround sound technique that records, stores and reproduces a sound field as a set of directional components based on spherical harmonics, rather than as a set of signals for particular loudspeakers. In a 2025 developer session, Apple describes it as "a technique for recording, mixing, and playing back full-sphere Spatial Audio" whose recordings "are not tied to a specific speaker layout as the sound field is encoded mathematically using a set of spherical harmonic basis functions."[1] The technique was developed in the United Kingdom in the 1970s by Michael Gerzon, Peter Fellgett and Peter Craven, with support from the National Research Development Corporation (NRDC).[2][3]
An ambisonic signal can be decoded to almost any loudspeaker arrangement or, by way of head-related impulse responses, to headphones, and the whole sound scene can be rotated with a simple matrix operation. The rotation property is what head-mounted displays use: the scene is counter-rotated as the listener turns, so sounds stay fixed in the virtual world.[3] Zotter and Frank, in their 2019 open-access textbook, write that first-order Ambisonics "is nowadays strongly revived by internet technology supported by Google/YouTube, Facebook 360°, 360° audio and video recording and rendering, as well as VR in games."[3] It is the spatial audio format of YouTube's 360-degree and VR videos, and it is supported by the Meta XR Audio SDK, Resonance Audio, Steam Audio, Unity and Apple's spatial audio codec.[4][5][6]
This article covers the ambisonic format and its use in VR and AR. Spatial sound in headsets more generally, including object-based spatializers, room acoustics and output hardware, is covered in 3D audio and VR audio; the filters used for headphone rendering are covered in Head-related transfer function.
How it works
Spherical harmonic components and order
Ambisonics describes the sound arriving at a single point as a weighted sum of spherical harmonic patterns. Each order n adds 2n + 1 patterns, so a signal set truncated at order N has (N + 1)² channels.[3] Apple's documentation summarizes the channel counts: "1st-order is 4 components, or channels, and correspond to 1 omnidirectional channel and 3 channels representing front-back, left-right, and up-down directionally oriented audio," while "2nd-order ambisonics uses 9 components, while 3rd order ambisonics use 16."[1]
| Order | Channels | Common name |
|---|---|---|
| 0 | 1 | Omnidirectional (W) only[3] |
| 1 | 4 | First-order Ambisonics (FOA)[1] |
| 2 | 9 | Higher-order Ambisonics (HOA)[1] |
| 3 | 16 | Higher-order Ambisonics (HOA)[1] |
| 4 | 25 | Higher-order Ambisonics (HOA)[3] |
| 5 | 36 | Higher-order Ambisonics (HOA)[3] |
Higher orders give a sharper directional image. According to Zotter and Frank, the perception of spatial depth improves strongly when the order is raised from 1 to 3 for a listener at the center of a loudspeaker array; away from that sweet spot, orders above 3, such as 5, improve it further, and the area of plausible playback grows with order. At fifth order nearly the whole area inside the horizontal loudspeakers of the IEM CUBE, a 12 by 10 m concert space at their institute in Graz, became a valid listening area.[3] A perceptual study by Bertet, Daniel, Parizet and Warusfel (2013) that compared microphones of order one to four found that localization improved with higher-order ambisonic microphones and that accuracy depended on both the order and the direction of the source.[7]
B-format and A-format
The first-order signal set is traditionally called B-format. It consists of a signal W with an omnidirectional pickup pattern and three signals X, Y and Z with figure-of-eight patterns aligned to the front-back, left-right and up-down axes.[3]
Practical first-order microphones usually place four cardioid or subcardioid capsules as close together as possible in a tetrahedral arrangement, an approach described in Craven and Gerzon's patent.[8][3][9] Their raw outputs (A-format) are converted to B-format by a matrix that sums all four capsules for W and takes front-minus-back, left-minus-right and up-minus-down differences for X, Y and Z.[3][9] Because the capsules are not truly coincident, high frequencies need correction; Zotter and Frank give a high-shelf boost of roughly 3 dB above the frequency at which capsule spacing exceeds half a wavelength, for example 5 kHz for a 3.4 cm spacing.[3] Higher orders require larger spherical arrays. Boaz Rafaely's 2005 paper in IEEE Transactions on Speech and Audio Processing set out a spherical-harmonics framework for designing such arrays, including the errors caused by a finite number of microphones, spatial aliasing and positioning inaccuracies.[10]
Channel ordering and normalization
Ambisonic data long lacked a single exchange convention: the 2011 ambiX paper notes that, although the technique dated from the 1970s, there were disagreements about how to normalize, store and exchange it.[11] In that proposal, Nachbar, Zotter, Deleflie and Sontacchi called the B-format ".amb" file, with Furse-Malham (FuMa) weighting, "the only de-facto existing Ambisonic format", but noted that it allowed at most third order (16 channels) and inherited a 4 GB file-size limit from the WAV container.[11] Their proposal, ambiX ("Ambisonics exchangeable"), orders channels by Ambisonic Channel Number (ACN = n² + n + m) and uses SN3D normalization. The authors chose SN3D because "the peak amplitude of single point sources will never exceed the level of the 0th order signal", which avoids clipping in the higher-order channels.[11] The paper proposed Apple's Core Audio Format as the container and an "extended" variant that carries an adaptor matrix for mixed-order or reduced channel sets.[11]
The VR platforms and tools described below use the AmbiX convention. In first order it puts the channels in the order W, Y, Z, X, which is the layout YouTube requires.[4] Google's Resonance Audio, the Meta XR Audio SDK, Unity and Apple's spatial audio capture API all use ACN ordering with SN3D normalization.[12][13][6][14]
Decoding to loudspeakers and headphones
A decoder converts the ambisonic channels into loudspeaker feeds. Early decoder design was addressed by Gerzon, David Malham and Jérôme Daniel; Zotter and Frank describe All-Round Ambisonic Decoding (AllRAD), developed around 2010, as "the most robust and flexible higher-order decoding method known today."[3]
For headphones, the classic approach renders a small set of virtual loudspeakers and convolves each feed with the head-related impulse responses (HRIRs) for its direction.[3] At low orders this causes problems. Zaunschirm, Schörkhuber and Höldrich (2018) wrote that order reduction of the HRIRs leads to "limited externalization, localization accuracy, and altered timbre."[15] Ben-Hur and colleagues (2017) traced part of the timbral loss to a rapid high-frequency roll-off in order-truncated binaural signals and proposed an equalization filter for it.[16] The time-alignment (TAC) and magnitude least squares (MagLS) decoders of Zaunschirm, Schörkhuber and Höldrich remove HRTF delays or optimize HRTF phases above about 3 kHz in favor of accurate spectral magnitude; in a listening test by Schörkhuber and colleagues, most listeners could not hear the removal of the HRTF linear-phase trend above 3 kHz.[3][15] Meta states that its SDK renders ambisonics with a spherical harmonic binaural renderer that gives "better frequency response, better externalization, less spatial smearing, and also uses less compute resources" than virtual-speaker methods.[5]
Rotation and head tracking
Rotating a first-order scene only requires mixing the figure-of-eight channels with a rotation matrix; W is unchanged. Zotter and Frank note that this "is important for head-tracked headphone playback to render the VR/360° audio scene static around the listener," and that rotation updates can run at high control rates while the HRIR convolution filters stay constant.[3] Unity's manual likewise says the sound field can be rotated "based on the listener's orientation (such as the user's head rotation in VR)."[6] In the Meta XR Audio SDK, rotating either the headset or the audio source changes the ambisonic orientation.[5]
History
Ambisonics builds on coincident stereo microphone techniques that go back to Alan Blumlein's patents of the 1930s. After D. H. Cooper and T. Shiga described surround panning as a directional Fourier series in 1972, the concept and technology of Ambisonics were developed by Fellgett, Gerzon and Craven, who also worked on a matching recording method.[3]
In "Periphony: With-Height Sound Reproduction" (Journal of the Audio Engineering Society, February 1973), Gerzon, then at the Mathematical Institute of the University of Oxford, described how periphony, "sound reproduction in both vertical and horizontal directions around a listener", could be recorded with practical two-, four- and nine-channel systems, giving matrix parameters and microphone techniques for 19 systems.[17] Peter Fellgett's article "Ambisonic reproduction of directionality in surround-sound systems" appeared in Nature in December 1974.[18] According to the University of Oxford's Into The Soundfield project, the ambisonic system was supported by the NRDC and involved Fellgett, then a professor at the University of Reading; once Gerzon joined the project, funds were made available to patent and develop a new microphone, and the Soundfield microphone made its first appearance in a February 1975 recording of the Schola Cantorum of Oxford.[2] Craven and Gerzon's tetrahedral microphone patent, US 4,042,779, claimed priority from July 1974, was filed in July 1975 and granted on 16 August 1977; it was assigned to the National Research Development Corporation (NRDC).[8]
The microphone became the Soundfield, which Gerzon and Craven developed with Ken Farrar at the British manufacturer Calrec; Calrec brought it to market in 1978. According to Sound On Sound, Calrec's were for many years the only tetrahedral ambisonic microphones on sale. Calrec stopped making microphones in 1993, and in 2016 the Soundfield business was acquired by the Australian maker Røde, whose NT-SF1 was the first jointly branded model.[9]
In 1995 David Malham and Anthony Myatt described Ambisonics in the Computer Music Journal as "a powerful technique for sound spatialization" and listed virtual reality, multimedia computing, film, video and computer games among new applications for spatial sound; their article included Csound implementations of ambisonic encoding and decoding.[19] Around the early 2000s, Jeffrey Bamford, Malham, Poletti, Jean-Marc Jot and Jérôme Daniel led the development of higher-order panning and decoding.[3] MPEG-H 3D Audio, the ISO/MPEG coding standard for immersive audio described by Herre and colleagues in 2015, carries higher-order ambisonic input alongside channel-based and object-based audio and can render it to loudspeakers or binaurally.[20] Zotter and Frank write that ITU, MPEG-H and ETSI standards "firmly fix it into the production and media broadcasting world."[3]
| Year | Event |
|---|---|
| 1973 | Gerzon publishes "Periphony: With-Height Sound Reproduction" in the Journal of the Audio Engineering Society[17] |
| 1974 | Fellgett's Nature article on ambisonic reproduction; priority date of the Craven and Gerzon microphone patent[18][8] |
| 1977 | US patent 4,042,779 granted, assigned to the NRDC[8] |
| 1978 | Calrec commercializes the Soundfield microphone[9] |
| 2011 | AmbiX (ACN/SN3D) proposed at the Ambisonics Symposium in Lexington, Kentucky[11] |
| 2016 | YouTube launches spatial audio for on-demand videos (18 April); Facebook acquires Two Big Ears, whose Spatial Workstation becomes a free download (May)[21][22] |
| 2017 | Google releases Resonance Audio, built on higher-order Ambisonics (6 November)[23] |
| 2018 | Resonance Audio open-sourced under the Apache 2.0 license (14 March)[12] |
| 2025 | Apple recommends APAC, with first- to third-order ambisonic encoding, for Apple Projected Media Profile video, new in visionOS 26[1] |
Applications in VR and AR
360-degree and VR video
YouTube launched spatial audio for on-demand videos on 18 April 2016, in the same announcement as 360-degree live streaming; it said it was working with VideoStitch and Two Big Ears on compatible software.[21] TechCrunch reported that at launch the feature was limited to Android smartphones used with headphones and was not available for live streams.[24] YouTube's current specification for 360-degree and VR uploads accepts first-order Ambisonics in AmbiX format (ACN ordering, SN3D normalization) as a 4-channel W, Y, Z, X track at 48 kHz, or a 6-channel track that adds head-locked stereo, which "doesn't change when a viewer moves their head." YouTube tells creators to run the latest version of its spatial-media metadata tool on the video before uploading.[4]
Facebook acquired the British VR audio company Two Big Ears in May 2016, after which Two Big Ears released its Spatial Workstation authoring tools as a free download.[22] In a February 2017 engineering post, Facebook's Hans Fugal and Varun Nair explained that "a first-order sound field results in four channels of audio data, while a third-order sound field results in 16 channels," and that the Spatial Workstation instead output eight channels of what Facebook called "hybrid higher-order ambisonics," plus two optional channels of head-locked audio for narration and music.[25]
Apple described ambisonics support in its 2025 developer sessions. The Apple Projected Media Profile, new in visionOS 26, signals 180-degree, 360-degree and wide-FOV video, and Apple recommends the Apple Positional Audio Codec (APAC) for its spatial audio, "including ambisonics." The built-in APAC encoder on iOS, macOS and visionOS supports first-, second- and third-order ambisonics, and APAC decodes on every Apple platform except watchOS.[1] In a 2025 developer session on audio recording, Apple explained that its spatial audio capture API, introduced in iOS 18, transforms the signals of a microphone array, such as an iPhone's, into first-order Ambisonics stored as "4 channels of the ambisonic layout HOA - ACN - SN3D"; apps can write these recordings with AVCaptureMovieFileOutput or assemble them with AVAssetWriter, and they support playback features such as head tracking on AirPods.[14] Apple's stereoscopic Apple Immersive Video format is a separate topic.
Game engines and audio SDKs
Unity's manual describes ambisonics as "an audio skybox for distant ambient sounds" that is "particularly useful for 360-degree videos and applications."[6]
| Engine or SDK | Developer | Ambisonic support |
|---|---|---|
| Meta XR Audio SDK | Meta | 4-channel AmbiX (first-order) clips, decoded by a spherical harmonic binaural renderer; the FMOD plugin supports first order only[5][13] |
| Resonance Audio | Higher-order Ambisonics used internally to spatialize "hundreds of simultaneous 3D sound sources"; includes a reference implementation of YouTube's ambisonic decoder (AmbiX ACN/SN3D) and an ambisonic recording tool for Unity[23][12] | |
| Steam Audio | Valve | Renders physics-based audio into first- or higher-order Ambisonics and spatializes ambisonic audio using HRTFs[26] |
| Unity | Unity Technologies | Up to third order (4, 9 or 16 channels); preferred format B-format WAV with ACN ordering and SN3D normalization; no built-in decoder, so a third-party decoder (often from a VR hardware SDK) or custom plugin is required[6] |
Resonance Audio is based on technology from Google's earlier VR Audio SDK; Google had added spatial audio to the Cardboard SDK in January 2016.[23] When Google open-sourced it in March 2018, it said developers could "easily render Ambisonic content in their VR media and other applications."[12]
Recording hardware
In explaining the revival of first-order Ambisonics for 360-degree media and VR, Zotter and Frank cite its compact microphone arrays that capture the whole sound scene in four channels, naming the Zoom H3-VR, the Oktava A-format microphone, the Røde NT-SF1 and the Sennheiser AMBEO VR Mic; they name the Eigenmike and Zylia as higher-order main microphone arrays.[3] The Zoom H3-VR combines a four-capsule ambisonic microphone with a recorder and decoder; Zoom describes it as a solution "for capturing and processing spatial audio for VR, AR and mixed-reality content" and it records A-format or B-format (FuMa or AmbiX) files, as well as binaural stereo.[27]
Research
Zaunschirm, Schörkhuber and Höldrich describe binaural rendering of Ambisonic signals as "of great interest in the fields of virtual reality, immersive media, and virtual acoustics."[15] They proposed a renderer that time-aligns HRIRs by frequency and constrains the diffuse-field behavior, and reported significant improvements over existing methods in listening tests.[15] Ben-Hur and colleagues found that their equalization filter restored timbre in order-truncated binaural signals while preserving, and in some respects improving, spatial properties.[16]
Reverberation is computationally expensive to render with high fidelity, and Ambisonics-based methods can render a reverberant field more efficiently by limiting its spatial resolution. Engel, Henry, Amengual Garí, Robinson and Picinali, a group from Imperial College London and Facebook Reality Labs, tested a "hybrid Ambisonics" method that renders the direct sound with a dense HRIR set and the reverberation in Ambisonics. Perceived quality of their auralizations of two measured rooms stopped improving beyond third order, a lower threshold than earlier studies had found when the direct sound was not processed separately.[28]
Streaming is another research topic. Narbutt and colleagues, including Google researchers, measured the effect of Opus 1.2 compression on first- and third-order ambisonic audio rendered to headphones and proposed AMBIQUAL, an objective metric for listening quality and localization accuracy, noting that services such as YouTube transcode content into several bit rates.[29] Hold, McCormack, Politis and Pulkki (2024) proposed a parametric codec based on higher-order Directional Audio Coding that downmixes higher-order input into fewer signals and synthesizes the components that cannot be fully recovered from transmitted scene parameters; in their listening test with fifth-order signals it improved on the underlying perceptual coder, especially at low to medium-high bit rates.[30]
For audiovisual perception research with head-mounted displays, Robotham and colleagues published a database of twelve real-world scenes recorded as 7680 by 3840 360-degree video at 60 frames per second with fourth-order Ambisonics audio, presented at QoMEX 2022.[31]
See also
References
- ↑ 1.0 1.1 1.2 1.3 1.4 1.5 1.6 "Learn about the Apple Projected Media Profile (WWDC25 session 297)". Apple Developer. Apple. 2025. https://developer.apple.com/videos/play/wwdc2025/297/. Retrieved 2026-10-06.
- ↑ 2.0 2.1 "Ambisonics". Into The Soundfield. University of Oxford, Faculty of Music. https://intothesoundfield.music.ox.ac.uk/ambisonics. Retrieved 2026-10-06.
- ↑ 3.00 3.01 3.02 3.03 3.04 3.05 3.06 3.07 3.08 3.09 3.10 3.11 3.12 3.13 3.14 3.15 3.16 3.17 3.18 3.19 Franz Zotter, Matthias Frank (2019). "Ambisonics: A Practical 3D Audio Theory for Recording, Studio Production, Sound Reinforcement, and Virtual Reality". Springer Topics in Signal Processing, vol. 19. Springer. https://doi.org/10.1007/978-3-030-17207-7. Retrieved 2026-10-06.
- ↑ 4.0 4.1 4.2 "Use spatial audio in 360-degree and VR videos". YouTube Help. Google. https://support.google.com/youtube/answer/6395969?hl=en. Retrieved 2026-10-06.
- ↑ 5.0 5.1 5.2 5.3 "Play Ambisonic Audio in Unity". Meta Horizon Developers. Meta. https://developers.meta.com/horizon/documentation/unity/meta-xr-audio-sdk-unity-ambisonic/. Retrieved 2026-10-06.
- ↑ 6.0 6.1 6.2 6.3 6.4 "Introduction to ambisonic audio". Unity Manual (Unity 6.6). Unity Technologies. https://docs.unity3d.com/Manual/AmbisonicAudio.html. Retrieved 2026-10-06.
- ↑ Stéphanie Bertet, Jérôme Daniel, Étienne Parizet, Olivier Warusfel (2013). "Investigation on Localisation Accuracy for First and Higher Order Ambisonics Reproduced Sound Sources". Acta Acustica united with Acustica, vol. 99, no. 4. pp. 642-657. https://doi.org/10.3813/AAA.918643. Retrieved 2026-10-06.
- ↑ 8.0 8.1 8.2 8.3 Peter Graham Craven, Michael Anthony Gerzon (1977-08-16). "US4042779A - Coincident microphone simulation covering three dimensional space and yielding various directional outputs". Google Patents. https://patents.google.com/patent/US4042779A/en. Retrieved 2026-10-06.
- ↑ 9.0 9.1 9.2 9.3 Sam Inglis (2018-12). "Rode NT-SF1 Ambisonic Microphone". Sound On Sound. https://www.soundonsound.com/reviews/rode-nt-sf1. Retrieved 2026-10-06.
- ↑ Boaz Rafaely (2005). "Analysis and design of spherical microphone arrays". IEEE Transactions on Speech and Audio Processing, vol. 13, no. 1. pp. 135-143. https://doi.org/10.1109/TSA.2004.839244. Retrieved 2026-10-06.
- ↑ 11.0 11.1 11.2 11.3 11.4 Christian Nachbar, Franz Zotter, Etienne Deleflie, Alois Sontacchi (2011-06). "AMBIX - A Suggested Ambisonics Format". Proceedings of the Ambisonics Symposium 2011, Lexington, KY. Institute of Electronic Music and Acoustics, Graz. https://ambisonics.iem.at/proceedings-of-the-ambisonics-symposium-2011/ambix-a-suggested-ambisonics-format. Retrieved 2026-10-06.
- ↑ 12.0 12.1 12.2 12.3 Eric Mauskopf (2018-03-14). "Open sourcing Resonance Audio". Google Open Source Blog. Google. https://opensource.googleblog.com/2018/03/resonance-audio-open-source.html. Retrieved 2026-10-06.
- ↑ 13.0 13.1 "Meta XR Audio Plugin for FMOD - Ambisonic". Meta Horizon Developers. Meta. https://developers.meta.com/horizon/documentation/unreal/meta-xr-audio-sdk-fmod-ambisonic/. Retrieved 2026-10-06.
- ↑ 14.0 14.1 "Enhance your app's audio recording capabilities (WWDC25 session 251)". Apple Developer. Apple. 2025. https://developer.apple.com/videos/play/wwdc2025/251/. Retrieved 2026-10-06.
- ↑ 15.0 15.1 15.2 15.3 Markus Zaunschirm, Christian Schörkhuber, Robert Höldrich (2018). "Binaural rendering of Ambisonic signals by head-related impulse response time alignment and a diffuseness constraint". The Journal of the Acoustical Society of America, vol. 143, no. 6. pp. 3616-3627. https://doi.org/10.1121/1.5040489. Retrieved 2026-10-06.
- ↑ 16.0 16.1 Zamir Ben-Hur, Fabian Brinkmann, Jonathan Sheaffer, Stefan Weinzierl, Boaz Rafaely (2017). "Spectral equalization in binaural signals represented by order-truncated spherical harmonics". The Journal of the Acoustical Society of America, vol. 141, no. 6. pp. 4087-4096. https://doi.org/10.1121/1.4983652. Retrieved 2026-10-06.
- ↑ 17.0 17.1 Michael A. Gerzon (1973-02). "Periphony: With-Height Sound Reproduction". Journal of the Audio Engineering Society, vol. 21, no. 1. pp. 2-10. https://www.aes.org/e-lib/browse.cfm?elib=2012. Retrieved 2026-10-06.
- ↑ 18.0 18.1 P. B. Fellgett (1974-12). "Ambisonic reproduction of directionality in surround-sound systems". Nature, vol. 252, no. 5484. pp. 534-538. https://doi.org/10.1038/252534b0. Retrieved 2026-10-06.
- ↑ David G. Malham, Anthony Myatt (1995). "3-D Sound Spatialization using Ambisonic Techniques". Computer Music Journal, vol. 19, no. 4. pp. 58-70. https://doi.org/10.2307/3680991. Retrieved 2026-10-06.
- ↑ Jürgen Herre, Johannes Hilpert, Achim Kuntz, Jan Plogsties (2015). "MPEG-H 3D Audio - The New Standard for Coding of Immersive Spatial Audio". IEEE Journal of Selected Topics in Signal Processing, vol. 9, no. 5. pp. 770-779. https://doi.org/10.1109/JSTSP.2015.2411578. Retrieved 2026-10-06.
- ↑ 21.0 21.1 Neal Mohan (2016-04-18). "One step closer to reality: introducing 360-degree live streaming and spatial audio on YouTube". YouTube Official Blog. YouTube. https://blog.youtube/news-and-events/one-step-closer-to-reality-introducing. Retrieved 2026-10-06.
- ↑ 22.0 22.1 Paul James (2016-05-24). "VR Audio Specialists 'Two Big Ears' Acquired by Facebook". Road to VR. https://www.roadtovr.com/vr-audio-specialists-two-big-ears-acquired-by-facebook/. Retrieved 2026-10-06.
- ↑ 23.0 23.1 23.2 "Google Releases 'Resonance Audio', a New Multi-Platform Spatial Audio SDK". Road to VR. 2017-11-06. https://www.roadtovr.com/resonance-audio-googles-new-multi-platform-spatial-audio-sdk/. Retrieved 2026-10-06.
- ↑ Sarah Perez (2016-04-18). "YouTube rolls out support for 360-degree live streams and spatial audio". TechCrunch. https://techcrunch.com/2016/04/18/youtube-rolls-out-support-for-360-degree-live-streams-and-spatial-audio/. Retrieved 2026-10-06.
- ↑ Hans Fugal, Varun Nair (2017-02-22). "Spatial audio - bringing realistic sound to 360 video". Engineering at Meta. Meta. https://engineering.fb.com/2017/02/22/virtual-reality/spatial-audio-bringing-realistic-sound-to-360-video/. Retrieved 2026-10-06.
- ↑ "Steam Audio". Steam Audio. Valve. https://valvesoftware.github.io/steam-audio/. Retrieved 2026-10-06.
- ↑ "H3-VR 360° Audio Recorder". Zoom Corporation. https://zoomcorp.com/en/us/handheld-recorders/handheld-recorders/h3-vr-360-audio-recorder/. Retrieved 2026-10-06.
- ↑ Isaac Engel, Craig Henry, Sebastià V. Amengual Garí, Philip W. Robinson, Lorenzo Picinali (2021). "Perceptual implications of different Ambisonics-based methods for binaural reverberation". The Journal of the Acoustical Society of America, vol. 149, no. 2. pp. 895-910. https://doi.org/10.1121/10.0003437. Retrieved 2026-10-06.
- ↑ Miroslaw Narbutt, Jan Skoglund, Andrew Allen, Michael Chinen, Dan Barry, Andrew Hines (2020). "AMBIQUAL: Towards a Quality Metric for Headphone Rendered Compressed Ambisonic Spatial Audio". Applied Sciences, vol. 10, no. 9. pp. 3188. https://doi.org/10.3390/app10093188. Retrieved 2026-10-06.
- ↑ Christoph Hold, Leo McCormack, Archontis Politis, Ville Pulkki (2024). "Perceptually-Motivated Spatial Audio Codec for Higher-Order Ambisonics Compression". 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 1121-1125. https://doi.org/10.1109/ICASSP48485.2024.10447577. Retrieved 2026-10-06.
- ↑ Thomas Robotham, Ashutosh Singla, Olli S. Rummukainen, Alexander Raake, Emanuël A. P. Habets (2022-12-27). "Audiovisual Database with 360 Video and Higher-Order Ambisonics Audio for Perception, Cognition, Behavior, and QoE Evaluation Research". arXiv (2022 14th International Conference on Quality of Multimedia Experience, QoMEX). doi:10.1109/QoMEX55416.2022.9900893. https://arxiv.org/abs/2212.13442. Retrieved 2026-10-06.