Toggle menu
Toggle preferences menu
Toggle personal menu
Not logged in
Your IP address will be publicly visible if you make any edits.

Microphone is a transducer that converts sound into an electrical signal.[1] In a head-mounted display, microphones provide audio for voice input, dictation and communication with other people.[2] A microphone array combines signals from multiple microphones to use spatial information when capturing a desired sound and suppressing interference.[3]

Operating principles

Dynamic and condenser microphones use a diaphragm that responds to sound pressure. In a moving-coil dynamic microphone, the diaphragm moves a coil within a magnetic field, generating an electrical signal. In a condenser microphone, a diaphragm and backplate form a capacitor; diaphragm movement changes their spacing and capacitance. Condenser microphones require electrical power for their associated circuitry.[1]

Capacitive microelectromechanical systems (MEMS) microphones use a moving silicon membrane and fixed backplate. An integrated circuit converts changing capacitance into an electrical output. STMicroelectronics' tutorial describes analog and digital pulse-density modulation (PDM) outputs.[4]

Transducer construction, output format and directionality describe different properties. An omnidirectional element can be part of an array whose combined output is directional: the array obtains that response by processing its separate microphone signals.[3] A microphone's polar pattern describes its sensitivity to sounds arriving from different directions, while its frequency response describes sensitivity across frequencies.[1]

Microphone specifications

Sensitivity measures the electrical output produced by a specified acoustic input. A common reference is a 1 kHz tone at 94 dB sound pressure level (SPL), equivalent to 1 pascal. Analog sensitivity is expressed using voltage units, such as millivolts per pascal or dBV, while digital sensitivity is expressed relative to digital full scale, in dBFS. Sensitivity alone does not establish microphone quality; noise, clipping and distortion also affect suitability for an application.[5]

Property Meaning
Frequency response Variation in sensitivity with frequency. It describes which parts of the captured sound are emphasized or attenuated.[1]
Signal-to-noise ratio (SNR) Ratio of reference-sound output to residual microphone noise; ST uses a 1 Pa, 1 kHz reference.[4]
Equivalent input noise The microphone's output noise expressed as an equivalent acoustic input level in dB SPL.[4]
Acoustic overload point (AOP) Upper acoustic level at a specified distortion threshold; ST's tutorial allows 10% distortion at this point.[4]
Dynamic range Range between the microphone's residual noise and its upper usable signal level.[4]
Sensitivity and headroom Higher sensitivity can reduce the required preamplifier gain, but it can also leave less headroom before clipping. The appropriate balance depends on the sound level at the microphone.[5]

These transducer specifications are distinct from an application's audio-stream format. The W3C's Media Capture and Streams specification defines sample rate as samples per second, sample size as the number of bits in a linear sample, and channel count as independent sound channels in the delivered data. It defines capture latency as the interval from the start of processing to availability of data at the next processing step; the actual latency can vary around the configured target.[6]

Headset placement and acoustics

Compact headsets can place microphones farther from the mouth than a close-talking boom microphone. Tashev, Seltzer and Acero's 2005 study treated the resulting loss of speech level relative to environmental noise as a reason to combine a three-microphone array with noise suppression.[7]

Its design accounted for mouth directivity and diffraction around the head. These affect speech reaching each sensor, so a distant-source model for unobstructed space did not adequately describe the study's headset configuration.[7]

Head rotation changes the array's orientation relative to external talkers and noise. Hafezi and colleagues describe how rapidly changing noise fields can undermine an adaptive beamformer's assumption of stable noise statistics over a short estimation interval.[8]

A microphone port can also be obstructed by hair, skin or clothing. Middelberg and colleagues studied how this changes the acoustic transfer functions of the wearer's voice and background noise, causing filters designed for unobstructed microphones to lose effectiveness. Their evaluation used recorded speech and noise with simulated occlusion, rather than a field trial of a finished headset.[9]

Microphone arrays and beamforming

Sound reaches spatially separated microphones at different times. Beamforming combines their signals with selected delays and weights to favor a desired source or direction. In delay-and-sum beamforming, the channels are aligned for a chosen arrival direction and added together. Its directional response has a main lobe and side lobes, so sounds outside the intended direction can still pass through.[10] For a receiving microphone array, "beam" describes its directional sensitivity; no physical beam of sound is emitted by the microphones.[3]

Adaptive methods adjust their weights using estimates of the captured signals. A minimum variance distortionless response (MVDR) beamformer minimizes estimated interference while constraining its response to the desired source.[10] An earlier foundation is Frost's 1972 linearly constrained adaptive array algorithm, which minimizes output noise power while preserving a specified response in the desired direction. Such constraints aim to prevent the adaptation from suppressing the wanted signal along with the noise.[11]

Array geometry affects the available spatial information. If microphones are too close together, the acoustic differences between their signals can be small relative to self-noise and mismatch. Excessive spacing can introduce spatial aliasing, making different arrival directions difficult to distinguish at affected frequencies. Matching microphone sensitivity and frequency response is particularly important when forming directional responses from a compact array.[3]

Steering errors and incorrect noise estimates can distort or attenuate desired speech. Hafezi and colleagues combine predefined noise models and subsequent denoising to address changing conditions around a moving head.[8]

Processing for speech and communication

Acoustic echo cancellation addresses sound played by the device that returns to the microphone. In hands-free communication research, an echo canceller can operate together with a beamformer and a spectral postfilter, with the latter stages suppressing residual echo and background noise.[12]

Processing operation Role
Beamforming Uses differences among microphone channels to produce a response favoring a desired sound source.[10]
Acoustic echo cancellation Attempts to remove device playback from microphone capture. A model-based canceller can remove an early echo component before further echo and noise suppression.[12]
Noise suppression Reduces unwanted components in the captured signal. The 2005 headset study distinguished stationary noise from spatially varying interference and processed them separately.[7]

The W3C specification exposes echo cancellation, automatic gain control and noise suppression as audio-capture properties. Support and configurability depend on the source's reported capabilities; an application cannot assume that every feature can be enabled or disabled. The specification also recognizes cases in which processing should be disabled to preserve unaltered audio.[6]

Beamforming produces audio for subsequent speech recognition. Kumatani and colleagues evaluated recognition through word errors after array processing, measuring its effect on the recognizer separately from the beamformer's acoustic response.[10]

Uses in VR and AR

Commands and dictation

On HoloLens, commands can be combined with head or eye gaze to identify an object. Microsoft's "see it, say it" approach uses visible button labels as voice commands; the keyboard's microphone control enables dictation. Its guidance recommends distinguishable commands, recognition feedback and actions that can easily be undone if nearby speech accidentally triggers a command. Accents, noise and ambiguity can affect recognition, and continuous fine control can be difficult to express verbally.[2]

Conversation capture

Head-worn microphone research includes both enhancing the wearer's own voice and capturing another person's speech for the wearer to hear. EasyCom was released to study the latter task: conversational focus in noisy environments. It includes more than five hours of conversations, with microphone-array audio, egocentric video, participant pose and speech annotations.[13] Its released data includes six-channel head-mounted recordings and separate close-microphone recordings of participants, supporting the comparison of enhanced audio with reference speech.[14]

Documented headset implementations

Headset Officially documented microphone feature
Apple Vision Pro (M5) Six microphones with directional beamforming. Apple's specifications also list voice as an input method.[15]
Microsoft HoloLens 2 A microphone array specified as five channels.[16]
PlayStation VR2 Built-in microphone. Its function button can be configured to mute or unmute the microphone.[17]

The number of physical microphones is distinct from the channel count of a processed application stream. Microsoft's HoloLens voice-input documentation specifies different formats for different capture purposes:[2][6]

HoloLens audio category Documented stream Intended use
Communications 16 kHz, 24-bit mono Calls and recording the wearer's voice.[2]
Speech 16 kHz, 24-bit mono Speech-recognition engines.[2]
Other 48 kHz, 24-bit stereo Ambient sound recording.[2]

Selected head-worn array research

Study Setup and contribution Limits of the evidence
Tashev, Seltzer and Acero (2005) Three-microphone headset with fixed beamforming and spatial and stationary-noise suppression; the design accounted for the mouth and head's acoustic effects.[7] Reported office, cafe and car-noise tests apply to the study's headset and processing configuration.[7]
Donley and colleagues, EasyCom (2021) Conversations recorded around a table with restaurant-like background noise, four microphones on the glasses and two at the ears, alongside video and pose data.[13] A dataset and baseline for algorithm evaluation. The paper explicitly states that the illustrated microphone positions do not depict a future product.[13]
Hafezi and colleagues (2023) Hybrid MVDR selects among predefined noise-field models, followed by two-channel principal-components denoising. Evaluation used EasyCom recordings and reported improvements over a superdirective baseline in noise suppression and estimated intelligibility.[8] The evaluation used supplied target-direction information. Its results concern the tested recordings, metrics and baseline.[8]
Middelberg and colleagues (2025) Compared adaptive, precomputed switching and combined switching-adaptive approaches to an occluded microphone in a five-channel glasses array.[9] Occlusion was simulated using measured transfer functions. Evaluation assumed knowledge of the occlusion state and examined deliberately introduced voice-activity-detection errors.[9]

Permissions, indicators and mute controls

Microphone access and communication controls operate at several levels. Apple Vision Pro allows users to change each requesting app's microphone access in Privacy & Security settings. An orange status dot indicates microphone use without the Persona Virtual Camera. A green dot appears for view sharing, screen recording or Persona Virtual Camera use, including when that camera and microphone are used together.[18]

In VRChat, the safety system's Voice setting controls whether another user's voice is heard, while its Audio setting controls sounds from that user's avatar.[19] VRChat's microphone troubleshooting guidance separately checks operating-system input, the selected recording device, microphone level and the user's mute state. These settings can prevent voice transmission even when a working microphone is present.[20]

Reviewed 12 October 2026. Independent verification of all 20 cited sources: transducer principles and units, official headset and audio-stream documentation, permissions and application controls, and the cited research claims against original papers and their stated evaluation limits. About review dates.

References

  1. ↑ 1.0 1.1 1.2 1.3 "Microphone Techniques for Recording". Shure. https://www.shure.com/en-US/docs/education/Microphone-Techniques-for-Recording. Retrieved 2026-10-12.
  2. ↑ 2.0 2.1 2.2 2.3 2.4 2.5 "Voice input". Microsoft Learn. 2026-01-06. https://learn.microsoft.com/en-us/windows/mixed-reality/design/voice-input. Retrieved 2026-10-12.
  3. ↑ 3.0 3.1 3.2 3.3 Ken Waurin, Dietmar Ruwisch and Yu Du. "How A2B Technology and Digital Microphones Enable Superior Performance in Emerging Automotive Applications". Analog Dialogue. https://www.analog.com/en/resources/analog-dialogue/articles/how-a2b-technology-and-digital-microphones-enable-superior-performance-in-automotive-applications.html. Retrieved 2026-10-12.
  4. ↑ 4.0 4.1 4.2 4.3 4.4 "Tutorial for MEMS microphones". STMicroelectronics, application note AN4426. https://www.st.com/resource/en/application_note/an4426-tutorial-for-mems-microphones-stmicroelectronics.pdf. Retrieved 2026-10-12.
  5. ↑ 5.0 5.1 Jerad Lewis. "Understanding Microphone Sensitivity". Analog Dialogue. https://www.analog.com/en/resources/analog-dialogue/articles/understanding-microphone-sensitivity.html. Retrieved 2026-10-12.
  6. ↑ 6.0 6.1 6.2 "Media Capture and Streams". World Wide Web Consortium. https://www.w3.org/TR/mediacapture-streams/. Retrieved 2026-10-12.
  7. ↑ 7.0 7.1 7.2 7.3 7.4 Ivan Tashev, Michael L. Seltzer and Alex Acero (2005-09). "Microphone Array for Headset with Spatial Noise Suppressor". International Workshop on Acoustic Echo and Noise Control (IWAENC). https://www.microsoft.com/en-us/research/wp-content/uploads/2016/02/Tashev_MAforHeadset_IWAENC_05.pdf. Retrieved 2026-10-12.
  8. ↑ 8.0 8.1 8.2 8.3 Sina Hafezi, Alastair H. Moore, Pierre Guiraud, Patrick A. Naylor, Jacob Donley, Vladimir Tourbabin and Thomas Lunner (2023-03-15). "Subspace Hybrid Beamforming for Head-worn Microphone Arrays". arXiv, author manuscript accepted for ICASSP 2023. https://arxiv.org/abs/2303.08967. Retrieved 2026-10-12.
  9. ↑ 9.0 9.1 9.2 Wiebke Middelberg, Jung-Suk Lee, Saeed Bagheri Sereshki, Ali Aroudi, Vladimir Tourbabin and Daniel D. E. Wong (2025-07-12). "Microphone Occlusion Mitigation for Own-Voice Enhancement in Head-Worn Microphone Arrays Using Switching-Adaptive Beamforming". arXiv, author manuscript accepted for WASPAA 2025. https://arxiv.org/abs/2507.09350. Retrieved 2026-10-12.
  10. ↑ 10.0 10.1 10.2 10.3 Kenichi Kumatani, Takayuki Arakawa, Kazumasa Yamamoto, John McDonough, Bhiksha Raj, Rita Singh and Ivan Tashev (2012-12). "Microphone Array Processing for Distant Speech Recognition: Towards Real-World Deployment". APSIPA Annual Summit and Conference. https://www.microsoft.com/en-us/research/wp-content/uploads/2016/02/Microphone20Array20Processing20for20Distant20Speech20Recognition20-20Towards20Real-World20Deployment.pdf. Retrieved 2026-10-12.
  11. ↑ Otis Lamont Frost III (1972-08). "An Algorithm for Linearly Constrained Adaptive Array Processing". Proceedings of the IEEE, volume 60, number 8, pages 926-935. doi:10.1109/PROC.1972.8817. https://calebrascon.info/AR/Topic7/addresources/LCMV.pdf. Retrieved 2026-10-12.
  12. ↑ 12.0 12.1 Thomas Haubner and Walter Kellermann (2022-08-10). "Deep Learning-Based Joint Control of Acoustic Echo Cancellation, Beamforming and Postfiltering". arXiv, author manuscript accepted for EUSIPCO 2022. https://arxiv.org/abs/2203.01793. Retrieved 2026-10-12.
  13. ↑ 13.0 13.1 13.2 Jacob Donley, Vladimir Tourbabin, Jung-Suk Lee, Mark Broyles, Hao Jiang, Jie Shen, Maja Pantic, Vamsi Krishna Ithapu and Ravish Mehra (2021-10-18). "EasyCom: An Augmented Reality Dataset to Support Algorithms for Easy Communication in Noisy Environments". arXiv. https://arxiv.org/abs/2107.04174. Retrieved 2026-10-12.
  14. ↑ "EasyComDataset". Facebook Research. https://github.com/facebookresearch/EasyComDataset. Retrieved 2026-10-12.
  15. ↑ "Apple Vision Pro - Technical Specifications". Apple. https://www.apple.com/apple-vision-pro/specs/. Retrieved 2026-10-12.
  16. ↑ "About HoloLens 2". Microsoft Learn. https://learn.microsoft.com/en-us/hololens/hololens2-hardware. Retrieved 2026-10-12.
  17. ↑ Sid Shuman (2023-02-06). "PlayStation VR2: The ultimate FAQ". PlayStation Blog. https://blog.playstation.com/2023/02/06/playstation-vr2-the-ultimate-faq/. Retrieved 2026-10-12.
  18. ↑ "Control access to hardware features, your surroundings, and more on Apple Vision Pro". Apple Support. https://support.apple.com/guide/apple-vision-pro/control-access-to-hardware-features-tane3312ebe/visionos. Retrieved 2026-10-12.
  19. ↑ "VRChat Safety and Trust System". VRChat Documentation. https://docs.vrchat.com/docs/vrchat-safety-and-trust-system. Retrieved 2026-10-12.
  20. ↑ Spark (2025-10-30). "My microphone isn't working". VRChat Support. https://help.vrchat.com/hc/en-us/articles/1500002248141-My-microphone-isn-t-working. Retrieved 2026-10-12.