Toggle menu
Toggle preferences menu
Toggle personal menu
Not logged in
Your IP address will be publicly visible if you make any edits.

Facial motion capture (also called facial capture or facial performance capture) is the recording of a person's facial movements and expressions so that they can be reproduced on a computer-generated face. Systems either track markers or makeup applied to the skin, or analyze unmarked video and depth images of the face, and the result is converted into animation data for a 3D character.[1][2] It is used to animate characters in films and video games, to build digital humans, and, in real time, to drive avatars in video calls, live streaming and social VR.[3][4] When the face and body are recorded together, the combined process is usually called performance capture, a branch of motion capture.[5][6]

Lance Williams described an early system of this kind in a 1990 SIGGRAPH paper, which he framed as an "electronic mask" for bringing actors' performances into digital animation.[1] Film studios adopted marker-based, image-based and head-mounted camera methods during the 2000s, and depth sensors and phone cameras brought markerless capture to consumer devices in the 2010s.[5][7] For virtual reality, the main difficulty is that a head-mounted display hides much of the wearer's face from outside cameras, which led researchers to build capture sensors into the headset itself.[3] Real-time expression sensing inside consumer headsets is covered in more detail at Face tracking.

Reviewed 6 October 2026. Checked paper metadata (Crossref, OpenAlex, DataCite) and the cited Weta, Polar Express, fxguide, AWN, Vicon, TechCrunch, CG Channel, EPFL, Eurographics, Apple, ARKit, MacRumors, MetaHuman Animator, Meta, FRL, Ekman and VTube Studio pages for the dates, numbers, quotes and attributions in this article. About review dates.

Definition and scope

In Williams' description, facial motion capture means "acquiring the expressions of real faces, and applying them to computer-generated faces."[1] A capture pipeline typically has three stages: recording the face with one or more cameras or other sensors; tracking or reconstructing the face's shape over time; and solving that motion onto the controls of a character's facial rig, a step often called retargeting when the character's face differs from the performer's.[4][8]

The term usually refers to capturing a performance as animation data for film, games or other media, whether offline or live. Face tracking (also called facial tracking) is the closely related task of estimating a user's expression in real time, for example from cameras inside a VR headset, so that an avatar can mirror it. The two use similar methods, and research papers on headset-based systems describe their goal as real-time facial motion capture of a person wearing a head-mounted display.[9]

How it works

Marker-based capture

Marker-based systems place small dots on the performer's face and track them with calibrated cameras, giving the 3D positions of a set of points on the skin.[10] In the 1998 SIGGRAPH paper "Making Faces", Brian Guenter and colleagues used "a large set of sampling points on the face" to track the face's 3D deformation while multiple high-resolution video cameras recorded color and shading, which they turned into a sequence of texture maps for a polygonal face model.[11] Markers give only sparse points, so finer detail such as wrinkles needs other data. Bickel and colleagues (2007) added two synchronized video cameras to "a traditional marker-based facial motion-capture system" to track expression wrinkles, then combined the marker data, the wrinkles and a high-resolution static scan into a multi-scale face model.[12]

Markerless and image-based capture

Markerless systems work from the natural features and texture of the face. For the visual effects of The Matrix Reloaded, whose research and development stage began in January 2000, George Borshukov's team rejected muscle deformers and blendshapes and built "Universal Capture", which combined optical flow and photogrammetry to make a 3D recording of an actor's performance that could be replayed from new angles and under new lighting.[13] Image Metrics' Faceware product line combined what the company called "marker-less video analysis technology" with an "artist-driven performance transfer toolset".[14]

Dense capture rigs surround the head with many cameras. The Mova Contour rig used on The Curious Case of Benjamin Button held 28 cameras around the actor and tracked a pattern of phosphorescent makeup in 3D. According to fxguide, Digital Domain used the rig to record Brad Pitt's facial shapes, while his performance in the role was filmed with four HD cameras on a soundstage and converted into animation curves with Image Metrics' image-analysis technology.[15] Depth Analysis' MotionScan system, used for the 2011 game L.A. Noire, recorded seated actors with 32 calibrated 2-megapixel HD cameras and no dots or markers on the face.[16]

Academic work in 2010 and 2011 showed that high-resolution capture is possible with passive cameras alone. Bradley, Heidrich, Popa and Sheffer used an array of video cameras with "no template facial geometry, no special makeup or markers, and no active lighting", reconstructing each frame with multi-view stereo and tracking skin texture detail between frames.[2] Beeler and colleagues (2011) proposed "anchor frames", frames whose expression resembles a chosen reference expression, which let a single mesh be propagated through a whole performance in parallel while limiting tracker drift.[17]

Head-mounted cameras

A head-mounted camera (HMC) rig is worn on the performer's head and holds one or more cameras in front of the face, so the face stays in view while the actor moves. Weta Digital introduced HMCs on Avatar (2009), "giving actors better freedom of movement on the stage" and letting body and face be captured together; on Alita: Battle Angel (2019) it used a stereo camera rig in the HMC for the first time.[5] Vicon's Cara, launched in July 2013, carried four synchronized high-speed cameras, and its CaraPost software produced 3D points from the tracked facial dots.[10] Faceware and Dynamixyz have sold head-mounted cameras together with their facial capture software, and Cubic Motion's Persona product bundled software with capture hardware.[8][18][19]

Depth cameras and phones

Consumer depth sensors made real-time markerless capture possible outside the studio. Weise, Bouaziz, Li and Pauly (2011) recorded users with "a non-intrusive, commercially available 3D sensor" and combined geometry and texture registration with pre-recorded animation priors, so that "any user" could control a digital avatar's expressions in real time without markers or special lighting.[7] Weise founded faceshift, a spin-off of the Computer Graphics and Geometry Laboratory at EPFL; in 2012 its software worked with Kinect-style or Asus Xtion depth cameras and needed about ten minutes of training on a new user's face.[20] Apple acquired faceshift in 2015; it confirmed the deal only with its standard statement that it buys smaller technology companies from time to time.[21] According to the Eurographics citation for his 2018 Young Researcher Award, co-founder Sofien Bouaziz later developed and productized at Apple the real-time face tracking algorithm behind the iPhone X's Animoji.[22]

The iPhone X, announced on 12 September 2017, has a TrueDepth camera made up of a dot projector, an infrared camera and a flood illuminator; Apple said it "captures and analyzes over 50 different facial muscle movements" to animate Animoji characters.[23] Apple's ARKit exposes the same kind of data to developers: a face-tracking configuration detects faces within 3 meters of the front camera and reports the face's position, orientation, topology and expressions, the last as coefficients for 52 named blend shape locations covering the eyes, brows, cheeks, nose, mouth, jaw and tongue.[24][25] Face tracking required a TrueDepth camera up to iOS 13; from iOS 14 it also runs on devices with the Apple Neural Engine.[24]

Monocular methods recover a 3D face model and its motion from a single ordinary RGB camera, a problem that a 2018 Eurographics state-of-the-art report describes as under-constrained and addressed with face priors and optimization-based reconstruction.[4] Learning-based methods such as DECA (2021) regress shape, expression, pose and expression-dependent wrinkle detail from a single image.[26]

Solving and retargeting

Captured motion is usually solved onto a blendshape rig. A 2014 Eurographics state-of-the-art report by J. P. Lewis and co-authors calls blendshapes, "a simple linear model of facial expression", the prevalent approach to realistic facial animation, and notes that the technique originated in industry before becoming an academic research subject.[27] The Facial Action Coding System (FACS) of Paul Ekman and Wallace Friesen, first published in 1978 and revised in 2002, is an anatomically based system that describes visible facial movement as combinations of action units.[28] In the 2009 Digital Emily project, a collaboration between Image Metrics and the USC Institute for Creative Technologies, an actress was scanned in 33 expressions to build a blendshape rig, which a semi-automatic video-based facial animation system then animated to match her filmed performance.[29]

History

Williams' 1990 paper built a high-resolution head model with photographic texture mapping and demonstrated animating it "by tracking and applying the expressions of a human performer".[1] Guenter and colleagues' "Making Faces" (1998) captured both 3D geometry and video texture of expressions and reconstructed them as photorealistic animation.[11]

In film, Sony Pictures Imageworks built a capture volume for The Polar Express (2004) that recorded facial and body motion together for up to four performers, with the aim of preserving each actor's performance in a single session; it let Tom Hanks play several characters.[6] On King Kong (2005), Weta Digital used facial motion capture to drive a facial performance for the first time, with Andy Serkis staying inside a one-meter cube so that he could be captured accurately; on Avatar (2009) it moved to head-mounted cameras.[5]

Commercial facial capture grew into a separate business. Image Metrics sold its Faceware product line to a new spin-off, Faceware Technologies, in January 2012, and game companies later bought two other specialists.[14] Epic Games acquired Cubic Motion, a UK company whose facial animation technology was used in God of War and Marvel's Spider-Man, in March 2020; Cubic Motion had launched Persona, a package of software and capture hardware, the previous year.[18] Take-Two Interactive acquired Dynamixyz, founded in 2010, in July 2021. Its markerless Performer software, used on Red Dead Redemption 2 and by Framestore for the Hulk in Avengers: Endgame, was withdrawn from sale and made exclusive to Take-Two's own studios.[8]

Milestones

Year Milestone
1978 Ekman and Friesen publish the Facial Action Coding System[28]
1990 Lance Williams presents "Performance-driven facial animation" at SIGGRAPH[1]
1998 Guenter et al. present "Making Faces" at SIGGRAPH[11]
2000 Research begins on Universal Capture for The Matrix Reloaded[13]
2004 The Polar Express records face and body together for up to four performers[6]
2005 Weta Digital first uses facial motion capture, on King Kong[5]
2009 Weta Digital introduces head-mounted cameras on Avatar[5]
2011 Real-time markerless capture with a consumer depth sensor (Weise et al.)[7]; L.A. Noire uses 32-camera MotionScan[16]
2012 Faceware Technologies spun off from Image Metrics[14]
2013 Vicon launches the four-camera Cara head rig[10]
2015 Facial capture built into a VR headset (Li et al.)[3]; Apple acquires faceshift[21]
2017 iPhone X introduces the TrueDepth camera[23]
2020 Epic Games acquires Cubic Motion[18]; Live Link Face app released[30]
2021 Take-Two Interactive acquires Dynamixyz[8]
2022 Meta Quest Pro announced with inward-facing sensors for facial expressions[31]
2023 Epic Games releases MetaHuman Animator[32]

Real-time engine tools

Game engine makers also supply their own capture tools. Epic Games released the iOS app Live Link Face on 9 July 2020; it uses ARKit and the TrueDepth camera to capture a performer's facial expressions and stream them to characters in Unreal Engine for real-time rendering. The app can track head and neck rotation as well as facial expressions, works either with a single actor at a desk or as part of a larger motion capture stage setup, and supports timecode so that it stays synchronized with other recording devices.[30] MetaHuman Animator, released on 15 June 2023, records an actor with an iPhone 12 or later or with a stereo head-mounted camera, then uses a "4D solver" that combines video and depth data with a MetaHuman model of the performer to produce facial animation for any MetaHuman character in minutes on local GPU hardware.[32] Epic had shown it at GDC in March 2023.[33] Dynamixyz's Performer could likewise stream to Unreal Engine or Unity as well as export to animation packages such as Maya and MotionBuilder.[8]

Applications in VR and AR

The headset occlusion problem

A headset blocks a large part of the wearer's face, which prevents effective capture with conventional external cameras. Li and colleagues stated in 2015 that there were then "no solutions for enabling direct face-to-face interaction" between VR users wearing head-mounted displays.[3] Later research systems, and later consumer headsets, placed cameras inside or below the headset and used the resulting expression data to animate the wearer's avatar.[34][35]

Consumer headsets

Meta described the Meta Quest Pro (2022) as its first headset with inward-facing sensors to capture natural facial expressions and eye tracking, so that an avatar raises an eyebrow or smiles when the wearer does.[31] Meta states that the images of the face captured by the headset never leave it and are deleted after processing.[36] On Apple Vision Pro, a FaceTime caller appears as a Persona, which Apple describes as a digital representation made with machine learning that "reflects face and hand movements in real time".[37] Devices, add-on trackers and software support for this kind of in-headset sensing are listed in the face tracking article.

Avatars, streaming and production

Phone and webcam capture is also used for streamed 2D characters. VTube Studio, an app for iPhone, iPad, Android, macOS and Windows, uses a smartphone or webcam to track the user's face and animate a Live2D model; its developer ranks tracking quality as iOS first, then webcam, then Android, and supports iPhones and iPads with Face ID as well as some A12-based devices without it.[38] On motion capture stages, face and body are often captured together, and timecode keeps the facial data synchronized with the other recording devices.[30][5]

Because face data is personal, Apple requires apps that use ARKit face tracking to include a privacy policy explaining how they use face tracking and face data.[24]

Research

Research on facial capture for VR has concentrated on headset-mounted sensing and on photorealistic avatars.

Year Paper Approach
2015 Li et al., "Facial performance sensing head-mounted display" (SIGGRAPH) Ultra-thin flexible strain sensors on the headset's foam liner measure the upper face; a head-mounted RGB-D camera tracks the mouth; one offline training session per person plus a short calibration before each use[3]
2016 Olszewski et al., "High-fidelity facial and speech animation for VR HMDs" (SIGGRAPH Asia) A convolutional neural network maps images from a mouth camera attached to the headset to avatar animation parameters; an internal infrared camera covers the eye region; no user-specific calibration[34]
2018 Thies et al., "FaceVR" (ACM TOG) Real-time facial motion capture of a headset wearer plus monocular eye tracking, used to re-render a prerecorded stereo video of the person for VR teleconferencing[9]
2018 Lombardi et al., "Deep appearance models for face rendering" (SIGGRAPH) A variational autoencoder learns face geometry and view-dependent texture from a multiview capture rig, aimed at real-time use in VR[39]
2019 Wei et al., "VR facial animation via multiview image translation" (SIGGRAPH) Separate training and tracking headset-mounted camera designs; self-supervised image translation links headset images to a photorealistic avatar for two-way VR conversation[35]
2022 Cao et al., "Authentic volumetric avatars from a phone scan" (SIGGRAPH) A universal prior trained on multiview capture of hundreds of people turns a short phone scan into a drivable 3D head avatar[40]

Yaser Sheikh, who was director of research at Facebook Reality Labs Pittsburgh when Facebook described the project in 2019, is a co-author of the 2018, 2019 and 2022 papers. That lab worked on Meta's Codec Avatars and built prototype head-mounted capture systems fitted with cameras, accelerometers, gyroscopes, magnetometers and microphones for them.[41] Wei and colleagues argued that stylized avatars had limited the adoption of social VR for professional and intimate conversations, and that animating avatars of each user's full likeness from consumer-friendly headset cameras would address this.[35]

See also

References

  1. ↑ 1.0 1.1 1.2 1.3 1.4 Lance Williams (1990). "Performance-driven facial animation". SIGGRAPH '90, Computer Graphics, vol. 24, no. 4, pp. 235-242. ACM. https://doi.org/10.1145/97879.97906. Retrieved 2026-10-06.
  2. ↑ 2.0 2.1 Derek Bradley, Wolfgang Heidrich, Tiberiu Popa, Alla Sheffer (2010). "High resolution passive facial performance capture". ACM Transactions on Graphics, vol. 29, no. 4 (SIGGRAPH 2010). https://doi.org/10.1145/1778765.1778778. Retrieved 2026-10-06.
  3. ↑ 3.0 3.1 3.2 3.3 3.4 Hao Li, Laura Trutoiu, Kyle Olszewski, Lingyu Wei, Tristan Trutna, Pei-Lun Hsieh, Aaron Nicholls, Chongyang Ma (2015). "Facial performance sensing head-mounted display". ACM Transactions on Graphics, vol. 34, no. 4 (SIGGRAPH 2015). https://doi.org/10.1145/2766939. Retrieved 2026-10-06.
  4. ↑ 4.0 4.1 4.2 Michael Zollhöfer, Justus Thies, Pablo Garrido, Derek Bradley, Thabo Beeler, Patrick Pérez, Marc Stamminger, Matthias Nießner, Christian Theobalt (2018). "State of the Art on Monocular 3D Face Reconstruction, Tracking, and Applications". Computer Graphics Forum, vol. 37, no. 2, pp. 523-550 (Eurographics 2018 STAR). https://doi.org/10.1111/cgf.13382. Retrieved 2026-10-06.
  5. ↑ 5.0 5.1 5.2 5.3 5.4 5.5 5.6 Ian Failes (2019-08-21). "A visual history of performance capture at Weta Digital". befores & afters. https://beforesandafters.com/2019/08/21/a-visual-history-of-performance-capture-at-weta-digital/. Retrieved 2026-10-06.
  6. ↑ 6.0 6.1 6.2 "The Polar Express by Zemeckis". ACM SIGGRAPH History Archive. ACM SIGGRAPH. https://history.siggraph.org/animation-video-pod/the-polar-express-by-zemeckis/. Retrieved 2026-10-06.
  7. ↑ 7.0 7.1 7.2 Thibaut Weise, Sofien Bouaziz, Hao Li, Mark Pauly (2011). "Realtime performance-based facial animation". ACM Transactions on Graphics, vol. 30, no. 4 (SIGGRAPH 2011). https://doi.org/10.1145/2010324.1964972. Retrieved 2026-10-06.
  8. ↑ 8.0 8.1 8.2 8.3 8.4 Jim Thacker (2021-07-05). "Take-Two Interactive Software acquires Dynamixyz". CG Channel. https://www.cgchannel.com/2021/07/take-two-interactive-acquires-dynamixyz/. Retrieved 2026-10-06.
  9. ↑ 9.0 9.1 Justus Thies, Michael Zollhöfer, Marc Stamminger, Christian Theobalt, Matthias Nießner (2018). "FaceVR". ACM Transactions on Graphics, vol. 37, no. 2. https://doi.org/10.1145/3182644. Retrieved 2026-10-06.
  10. ↑ 10.0 10.1 10.2 "OMG's Vicon launches Cara, the new face of film and games production". Oxford Metrics. 2013-07-24. https://oxfordmetrics.com/news/2013-07-24/omgs-vicon-launches-cara-the-new-face-of-film-and-games-production. Retrieved 2026-10-06.
  11. ↑ 11.0 11.1 11.2 Brian Guenter, Cindy Grimm, Daniel Wood, Henrique Malvar, Fredric Pighin (1998). "Making faces". SIGGRAPH '98, Proceedings of the 25th Annual Conference on Computer Graphics and Interactive Techniques, pp. 55-66. ACM. https://doi.org/10.1145/280814.280822. Retrieved 2026-10-06.
  12. ↑ Bernd Bickel, Mario Botsch, Roland Angst, Wojciech Matusik, Miguel Otaduy, Hanspeter Pfister, Markus Gross (2007). "Multi-scale capture of facial geometry and motion". ACM SIGGRAPH 2007 Papers, article 33. https://doi.org/10.1145/1275808.1276419. Retrieved 2026-10-06.
  13. ↑ 13.0 13.1 George Borshukov, Dan Piponi, Oystein Larsen, J. P. Lewis, Christina Tempelaar-Lietz (2005). "Universal capture: image-based facial animation for "The Matrix Reloaded"". ACM SIGGRAPH 2005 Courses. ACM. https://doi.org/10.1145/1198555.1198596. Retrieved 2026-10-06.
  14. ↑ 14.0 14.1 14.2 "Image Metrics Launches Spin-Off Company Faceware Technologies". Animation World Network. 2012-01-19. https://www.awn.com/news/image-metrics-launches-spin-company-faceware-technologies. Retrieved 2026-10-06.
  15. ↑ Mike Seymour (2009-01-01). "The Curious Case of Aging Visual Effects". fxguide. https://www.fxguide.com/fxfeatured/the_curious_case_of_aging_visual_effects/. Retrieved 2026-10-06.
  16. ↑ 16.0 16.1 Bill Desowitz (2011-06-17). "Raising the Bar with 'L.A. Noire'". Animation World Network. https://www.awn.com/vfxworld/raising-bar-la-noire. Retrieved 2026-10-06.
  17. ↑ Thabo Beeler, Fabian Hahn, Derek Bradley, Bernd Bickel, Paul Beardsley, Craig Gotsman, Robert W. Sumner, Markus Gross (2011). "High-quality passive facial performance capture using anchor frames". ACM Transactions on Graphics, vol. 30, no. 4 (SIGGRAPH 2011). https://doi.org/10.1145/2010324.1964970. Retrieved 2026-10-06.
  18. ↑ 18.0 18.1 18.2 Lucas Matney (2020-03-12). "Epic Games buys UK facial mapping startup Cubic Motion". TechCrunch. https://techcrunch.com/2020/03/12/epic-games-buys-uk-facial-mapping-startup-cubic-motion/. Retrieved 2026-10-06.
  19. ↑ "Faceware Technologies". Faceware Technologies. https://facewaretech.com/. Retrieved 2026-10-06.
  20. ↑ "Software Enables Avatar to Reproduce Our Emotions in Real Time". EPFL News. EPFL. 2012-11-19. https://actu.epfl.ch/news/software-enables-avatar-to-reproduce-our-emotions-/. Retrieved 2026-10-06.
  21. ↑ 21.0 21.1 Ingrid Lunden, Natasha Lomas (2015-11-24). "Apple Has Acquired Faceshift, Maker Of Motion Capture Tech Used In Star Wars". TechCrunch. https://techcrunch.com/2015/11/24/apple-faceshift/. Retrieved 2026-10-06.
  22. ↑ "Young Researcher Award 2018: Sofien Bouaziz". Eurographics Association. https://www.eg.org/wp/eurographics-awards-programme/the-young-researcher-award/young-researcher-award-2018-sofien-bouaziz/. Retrieved 2026-10-06.
  23. ↑ 23.0 23.1 "The future is here: iPhone X". Apple Newsroom. Apple. 2017-09-12. https://www.apple.com/newsroom/2017/09/the-future-is-here-iphone-x/. Retrieved 2026-10-06.
  24. ↑ 24.0 24.1 24.2 "ARFaceTrackingConfiguration". Apple Developer Documentation. Apple. https://developer.apple.com/documentation/arkit/arfacetrackingconfiguration. Retrieved 2026-10-06.
  25. ↑ "ARFaceAnchor.BlendShapeLocation". Apple Developer Documentation. Apple. https://developer.apple.com/documentation/arkit/arfaceanchor/blendshapelocation. Retrieved 2026-10-06.
  26. ↑ Yao Feng, Haiwen Feng, Michael J. Black, Timo Bolkart (2021). "Learning an animatable detailed 3D face model from in-the-wild images". ACM Transactions on Graphics, vol. 40, no. 4 (SIGGRAPH 2021). https://doi.org/10.1145/3450626.3459936. Retrieved 2026-10-06.
  27. ↑ J. P. Lewis, Ken Anjyo, Taehyun Rhee, Mengjie Zhang, Fred Pighin, Zhigang Deng (2014). "Practice and Theory of Blendshape Facial Models". Eurographics 2014, State of the Art Reports, pp. 199-218. Eurographics Association. https://doi.org/10.2312/egst.20141042. Retrieved 2026-10-06.
  28. ↑ 28.0 28.1 "Facial Action Coding System". Paul Ekman Group. https://www.paulekman.com/facial-action-coding-system/. Retrieved 2026-10-06.
  29. ↑ Oleg Alexander, Mike Rogers, William Lambeth, Matt Chiang, Paul Debevec (2009). "Creating a Photoreal Digital Actor: The Digital Emily Project". 2009 Conference for Visual Media Production (CVMP). IEEE. https://doi.org/10.1109/CVMP.2009.29. Retrieved 2026-10-06.
  30. ↑ 30.0 30.1 30.2 Eric Slivka (2020-07-09). "Epic's 'Live Link Face' Leverages ARKit and iPhone's TrueDepth Camera for Real-Time Facial Animation". MacRumors. https://www.macrumors.com/2020/07/09/epic-unreal-engine-live-link-face/. Retrieved 2026-10-06.
  31. ↑ 31.0 31.1 "Meta Connect 2022: Meta Quest Pro, More Social VR and a Look Into the Future". Meta Newsroom. Meta Platforms. 2022-10-11. https://about.fb.com/news/2022/10/meta-quest-pro-social-vr-connect-2022/. Retrieved 2026-10-06.
  32. ↑ 32.0 32.1 Jose Antunes (2023-06-15). "MetaHuman Animator is now available". ProVideo Coalition. https://www.provideocoalition.com/metahuman-animator-is-now-available/. Retrieved 2026-10-06.
  33. ↑ Amid Amidi (2023-03-24). "Epic Games Unveils Metahuman Animator, A Potential Game Changer For Automated Facial Animation And Lip Sync". Cartoon Brew. https://www.cartoonbrew.com/tools/epic-games-unveils-metahuman-animator-a-potential-gamechanger-for-automated-facial-animation-and-lip-sync-227250.html. Retrieved 2026-10-06.
  34. ↑ 34.0 34.1 Kyle Olszewski, Joseph J. Lim, Shunsuke Saito, Hao Li (2016). "High-fidelity facial and speech animation for VR HMDs". ACM Transactions on Graphics, vol. 35, no. 6 (SIGGRAPH Asia 2016). https://doi.org/10.1145/2980179.2980252. Retrieved 2026-10-06.
  35. ↑ 35.0 35.1 35.2 Shih-En Wei, Jason Saragih, Tomas Simon, Adam W. Harley, Stephen Lombardi, Michal Perďoch, Alexander Hypes, Dawei Wang, Hernán Badino, Yaser Sheikh (2019). "VR facial animation via multiview image translation". ACM Transactions on Graphics, vol. 38, no. 4 (SIGGRAPH 2019). https://doi.org/10.1145/3306346.3323030. Retrieved 2026-10-06.
  36. ↑ "Learn about Natural Facial Expressions on Meta Quest Pro". Meta Quest Help. Meta Platforms. https://www.meta.com/help/quest/402982851992067/. Retrieved 2026-10-06.
  37. ↑ "Introducing Apple Vision Pro: Apple's first spatial computer". Apple Newsroom. Apple. 2023-06-05. https://www.apple.com/newsroom/2023/06/introducing-apple-vision-pro/. Retrieved 2026-10-06.
  38. ↑ "Introduction & Requirements". VTube Studio Wiki. DenchiSoft. https://github.com/DenchiSoft/VTubeStudio/wiki/Introduction-&-Requirements. Retrieved 2026-10-06.
  39. ↑ Stephen Lombardi, Jason Saragih, Tomas Simon, Yaser Sheikh (2018). "Deep appearance models for face rendering". ACM Transactions on Graphics, vol. 37, no. 4 (SIGGRAPH 2018). https://doi.org/10.1145/3197517.3201401. Retrieved 2026-10-06.
  40. ↑ Chen Cao, Tomas Simon, Jin Kyu Kim, Gabe Schwartz, Michael Zollhöfer, Shunsuke Saito, Stephen Lombardi, Shih-En Wei, Danielle Belko, Shoou-I Yu, Yaser Sheikh, Jason Saragih (2022). "Authentic volumetric avatars from a phone scan". ACM Transactions on Graphics, vol. 41, no. 4 (SIGGRAPH 2022). https://doi.org/10.1145/3528223.3530143. Retrieved 2026-10-06.
  41. ↑ "Facebook is building the future of connection with lifelike avatars". Tech at Meta. Meta Platforms. 2019-03-12. https://tech.facebook.com/ar-vr/2019/03/codec-avatars-facebook-reality-labs/. Retrieved 2026-10-06.