Toggle menu
Toggle preferences menu
Toggle personal menu
Not logged in
Your IP address will be publicly visible if you make any edits.
MediaPipe
Information
Type Framework
Industry Machine learning, computer vision
Developer Google
Written In C++
Operating System Android, iOS, Linux, macOS, Windows, web browsers
License Apache License 2.0
Supported Devices Smartphones, desktop computers, web browsers, edge and IoT devices
Release Date July 2019
Website https://ai.google.dev/edge/mediapipe


MediaPipe is an open-source framework from Google for building on-device machine learning pipelines that process live and streaming media such as camera video and audio. It is best known in VR and AR development for its ready-made perception models, such as a hand tracker that estimates 21 3D landmarks per hand from a single frame of an ordinary RGB camera, and similar models for the face and body.[1][2]

Google researchers introduced the framework in June 2019 in a paper presented at the Third Workshop on Computer Vision for AR/VR at CVPR 2019, and published the code on GitHub.[3][4] Several of its perception models were also presented at that workshop series, including the face mesh in 2019 and the hand tracker in 2020, whose paper describes the system as a hand tracking pipeline "for AR/VR applications".[5][6]

MediaPipe is now part of Google's AI Edge developer tools. Since 2023 Google has offered MediaPipe Solutions, a set of task libraries for Android, iOS, the web and Python, with tools for retraining and testing models.[7][8] The project reached version 1.0.0 in July 2026.[9]

Reviewed 6 October 2026. Checked GitHub repository and release dates, PyPI 1.0.1, the six arXiv papers, the Reimer et al. 2023 paper, Google Research and Developers blog posts, current MediaPipe documentation and the cited third-party repositories and Next Reality article. About review dates.

History

Origins and open-source release

The 2019 framework paper presents MediaPipe as a way for developers to combine existing and new perception components into prototypes, turn them into cross-platform applications, and measure performance and resource use on target devices.[4] Google's research page for the CVPR 2019 workshop paper lists earlier Google projects as example MediaPipe projects: real-time AR self-expression effects, Motion Photos on the Pixel 2, mobile real-time video segmentation, the instant motion tracking behind Motion Stills AR, and the Motion Stills apps.[10] The AR self-expression work, described by Google in March 2019, is the face geometry technology behind creator effects in YouTube Stories and the Augmented Faces API in ARCore.[11]

The GitHub repository was created on 13 June 2019 under the google organization and is now hosted at google-ai-edge/mediapipe.[10][1][12] Google demonstrated the framework at the Google booth of the CVPR 2019 expo in Long Beach on 18 and 20 June 2019.[10] The first tagged GitHub release, v0.5.0, followed on 10 July 2019.[13]

Perception solutions, 2019-2020

On 19 August 2019 Google Research published its on-device hand tracking pipeline and open-sourced it in MediaPipe, the same day as release v0.6.0.[2][13] The blog post said the approach could support sign language understanding and hand gesture control, and could "enable the overlay of digital content and information on top of the physical world in augmented reality".[2] Writing in Next Reality, Tommy Palladino noted that the approach could run on Android or iOS devices without dedicated motion or depth sensors, and suggested that it could let makers of lightweight smartglasses offer hand gesture input of the kind used on HoloLens 2 and Magic Leap One without the bulk and cost of those sensors.[14]

In January 2020 the MediaPipe team brought the framework to web browsers. The C++ code was compiled to WebAssembly with Emscripten, TensorFlow Lite inference was accelerated with the XNNPACK library, and most OpenGL-based calculators ran through WebGL. The first browser demos covered edge detection, face detection, hair segmentation and hand tracking.[15] Later in 2020 Google released several more pipelines: MediaPipe Objectron for 3D object detection in March, MediaPipe Iris for iris tracking and camera-to-eye distance estimation in August, and MediaPipe Holistic, which combines face, hand and pose tracking, in December.[16][17][18] In October 2020 Google stated that the background blur and replacement features of Google Meet in the browser were developed with MediaPipe.[19]

MediaPipe Solutions and Tasks, 2022-2023

Google launched the first five MediaPipe Tasks in December 2022: gesture recognition, hand landmarker, image classification, object detection and text classification. At Google I/O in May 2023 it introduced MediaPipe Solutions, made up of MediaPipe Studio, MediaPipe Tasks and MediaPipe Model Maker, along with nine further tasks including the Face Landmarker.[7]

Support for the older "MediaPipe Legacy Solutions" ended on 1 March 2023. Face Detection, Face Mesh, Iris, Hands, Pose, Holistic, Selfie segmentation, Hair segmentation and Object detection were upgraded to new Solutions. Support ended without a replacement for Box tracking, Instant motion tracking, Objectron, KNIFT, AutoFlip, MediaSequence and YouTube 8M; their code and prebuilt binaries remain available on an as-is basis.[8] On 3 April 2023 the primary documentation moved from the GitHub repository to Google's developer site.[1]

Google AI Edge and version 1.0

When Google renamed TensorFlow Lite to LiteRT in September 2024, it described LiteRT as part of the Google AI Edge suite and encouraged developers to use MediaPipe Tasks for future development instead of the TensorFlow Lite Task libraries.[20] MediaPipe also gained an LLM Inference API for running large language models on device; Google's documentation now lists that API as maintenance-only and recommends LiteRT-LM instead.[21]

Version 1.0.0 was published on GitHub on 28 July 2026. Its release notes list changes such as the removal of macOS prebuilts, new text tasks for summarization and proofreading, and the removal of the LLM inference engine from the iOS and Android release.[9] As of October 2026 the newest Python package on PyPI is 1.0.1, uploaded on 14 August 2026, with wheels for macOS on Apple silicon, Linux (x86-64 and AArch64) and Windows (x64 and Arm64).[22]

Architecture

MediaPipe Framework

The low-level MediaPipe Framework runs processing inside a graph that defines how data flows between nodes. Nodes, which the documentation also calls calculators, consume and produce packets. A packet holds a numeric timestamp and a shared pointer to an immutable payload of any C++ type. A connection between two nodes that carries a sequence of packets with increasing timestamps is a stream; a side packet carries a single value that stays constant, such as a setting. By default, a node receives all inputs that share a timestamp together, in timestamp order, which keeps streams such as video frames and model outputs synchronized.[23] According to the 2019 paper, an application built with MediaPipe can be configured to manage CPU and GPU resources for low latency and to synchronize time-series data such as audio and video frames.[3]

MediaPipe Solutions

MediaPipe Solutions sits on top of the framework and has four parts:[8]

Component Purpose
MediaPipe Tasks Cross-platform APIs and libraries for deploying solutions
MediaPipe Models Pre-trained, ready-to-run models for each solution
MediaPipe Model Maker Customizing a solution's model with the developer's own data
MediaPipe Studio Visualizing, evaluating and benchmarking solutions in a web browser

MediaPipe Tasks has APIs for Android (Java), Python, Web/JavaScript and iOS. Google's privacy notice for Tasks states that input data such as images, video and text is processed on the device and is not sent to Google servers; the APIs do send performance and usage metrics.[24] The solutions list in the October 2026 documentation covers audio classification, face detection, face landmark detection, gesture recognition, hand landmark detection, holistic landmark detection, image classification, image embedding, image segmentation, interactive segmentation, language detection, LLM inference, object detection, pose landmark detection, text classification, text embedding, text proofreading and text summarization.[8]

Body and object tracking solutions

The vision tasks most used in XR work estimate the position of the user's hands, face and body from a single camera image.

Solution Output Notes
Hand Landmarker 21 hand landmarks per hand in image and world coordinates, plus handedness A palm detection model finds hands; a landmark model then locates the 21 hand-knuckle coordinates. In video and live-stream modes the palm detector re-runs only when tracking is lost.[25]
Gesture Recognizer Hand landmarks plus gesture categories Built-in gestures are Closed_Fist, Open_Palm, Pointing_Up, Thumb_Down, Thumb_Up, Victory and ILoveYou (plus None); custom gestures can be trained with Model Maker.[26]
Face Landmarker 478 3D face landmarks, 52 blendshape scores and facial transformation matrices Uses the BlazeFace short-range face detector, a face mesh model and a blendshape model. Google lists facial filters and effects and virtual avatars among its uses.[27]
Pose Landmarker 33 body landmarks in image and 3D world coordinates, optional segmentation mask Lite, Full and Heavy model bundles; Google describes the model as a BlazePose variant that uses GHUM, a 3D human shape modeling pipeline, to estimate full 3D body pose.[28]
Holistic Landmarker 553 landmarks: 33 pose, 478 face (including 10 iris landmarks) and 21 per hand Combines the pose, face and hand models; face blendshapes and pose segmentation masks are optional outputs.[29]
Objectron (legacy) Oriented 3D bounding boxes of everyday objects First released for shoes and chairs; ran at 26 FPS on an Adreno 650 mobile GPU; support ended in March 2023.[16][8]

The legacy Face Mesh produced 468 landmarks, the figure given in the 2019 face geometry paper and in the 2020 Holistic announcement.[6][18] According to Google, the hand landmark model was trained on about 30,000 manually annotated real-world images plus rendered synthetic hand models. Its palm detector reached an average precision of 95.7 percent.[2][25]

Research

Google researchers published several of MediaPipe's models as short papers at the CVPR Workshop on Computer Vision for Augmented and Virtual Reality:

Paper Authors Venue Key claim
MediaPipe: A Framework for Building Perception Pipelines Lugaresi, Tang, Nash, McClanahan, Uboweja, Hays, Zhang, Chang, Yong, Lee, Chang, Hua, Georg, Grundmann Third Workshop on Computer Vision for AR/VR, CVPR 2019 Framework for building and benchmarking cross-platform perception pipelines[3]
BlazeFace: Sub-millisecond Neural Face Detection on Mobile GPUs Bazarevsky, Kartynnik, Vakunov, Raveendran, Grundmann CVPR Workshop on Computer Vision for AR/VR, Long Beach, 2019 Face detector running at 200-1000+ FPS on flagship devices, intended as the first stage of AR face pipelines[30]
Real-time Facial Surface Geometry from Monocular Video on Mobile GPUs Kartynnik, Ablavatski, Grishchenko, Grundmann CVPR Workshop on Computer Vision for AR/VR, Long Beach, 2019 468-vertex face mesh for AR effects, 100-1000+ FPS on mobile GPUs depending on device and model variant[6]
MediaPipe Hands: On-device Real-time Hand Tracking Zhang, Bazarevsky, Vakunov, Tkachenka, Sung, Chang, Grundmann CVPR Workshop on Computer Vision for AR/VR, Seattle, 2020 Palm detector plus hand landmark model predicting a hand skeleton from a single RGB camera in real time on mobile GPUs[5]
BlazePose: On-device Real-time Body Pose tracking Bazarevsky, Grishchenko, Raveendran, Zhu, Zhang, Grundmann CVPR Workshop on Computer Vision for AR/VR, Seattle, 2020 33 body keypoints for one person at over 30 FPS on a Pixel 2[31]
Attention Mesh: High-fidelity Face Mesh Prediction in Real-time Grishchenko, Ablavatski, Kartynnik, Raveendran, Grundmann CVPR Workshop on Computer Vision for AR/VR, Seattle, 2020 Face mesh with attention to eye and lip regions, over 50 FPS on a Pixel 2, for AR makeup, eye tracking and AR puppeteering[32]

Outside Google, researchers have used MediaPipe as a camera-based alternative to headset hand tracking. In a 2023 study in Frontiers in Virtual Reality, Dennis Reimer, Iana Podkosova, Daniel Scherzer and Hannes Kaufmann built an RGB camera hand tracking method on MediaPipe and compared it with the hand tracking of the Oculus Quest and the Leap Motion controller. Their method had median errors below 1.75 cm at distances under 75 cm, slightly less accurate than the two commercial systems. At 2.75 m its median error was 4.7 cm, a distance at which the Quest and Leap Motion either lost tracking or became very inaccurate. The median distance at which MediaPipe acquired tracking was 324.94 cm, compared with 40.33 cm for the Oculus Quest and 46.05 cm for Leap Motion.[33]

Use in VR and AR

The documented XR uses of MediaPipe's models involve camera-based hand tracking, face tracking, body tracking and avatar animation on phones, PCs and in web browsers, using ordinary cameras rather than headset sensors.

Face effects and AR filters

The face mesh technology behind YouTube Stories creator effects and ARCore's Augmented Faces API appears on Google's list of MediaPipe example projects.[11][10] The current Face Landmarker outputs facial transformation matrices so that effects can be rendered on a detected face.[27]

MediaPipe Iris estimated the metric distance from the camera to the user from the size of the iris, based on the observation that the horizontal iris diameter is roughly constant at 11.7 mm, plus or minus 0.5 mm, across a wide population. Google listed AR effects such as virtual avatars and the virtual try-on of glasses and hats among its uses.[17]

3D object detection from AR data

To train Objectron, Google annotated video captured during mobile AR sessions on ARCore and ARKit devices, using the recorded camera poses, sparse 3D point clouds and detected planes. It also generated synthetic training data by placing virtual objects into scenes that had AR session data, which improved accuracy by about 10 percent.[16]

Avatars

In a July 2023 Google Developers Blog guest post, the XR development team at KDDI and Alpha-U described using the Face Landmarker's 52 blendshapes, extracted in real time with MediaPipe's Python package, to drive the face of the virtual human "Metako" in Unreal Engine. The scene was streamed to phones and browsers through Google Cloud's Immersive Stream for XR.[34]

Smartphone AR and game engines

Portal-ble, a library from Brown University's HCI group for manipulating virtual objects in smartphone AR with bare hands, relies on MediaPipe for hand tracking in its Android release, which runs as a Unity application.[35] The third-party MediaPipe Unity Plugin by the developer homuler ports MediaPipe's C++ API to C#, so that MediaPipe solutions and custom calculators can run in Unity on Android, iOS, Linux, macOS and Windows.[36]

The web version runs in the browser through WebAssembly and WebGL, so the tracking models can run inside a web page.[15][24]

See also

References

  1. ↑ 1.0 1.1 1.2 "google-ai-edge/mediapipe: Cross-platform, customizable ML solutions for live and streaming media". GitHub. Google. https://github.com/google-ai-edge/mediapipe. Retrieved 2026-10-06.
  2. ↑ 2.0 2.1 2.2 2.3 Valentin Bazarevsky, Fan Zhang (2019-08-19). "On-Device, Real-Time Hand Tracking with MediaPipe". Google Research Blog. https://research.google/blog/on-device-real-time-hand-tracking-with-mediapipe/. Retrieved 2026-10-06.
  3. ↑ 3.0 3.1 3.2 Camillo Lugaresi, Jiuqiang Tang, Hadon Nash, Chris McClanahan, Esha Uboweja, Michael Hays, Fan Zhang, Chuo-Ling Chang, Ming Yong, Juhyun Lee, Wan-Teh Chang, Wei Hua, Manfred Georg, Matthias Grundmann (2019). "MediaPipe: A Framework for Perceiving and Processing Reality". Third Workshop on Computer Vision for AR/VR at CVPR 2019. Google Research. https://research.google/pubs/mediapipe-a-framework-for-perceiving-and-processing-reality/. Retrieved 2026-10-06.
  4. ↑ 4.0 4.1 Camillo Lugaresi, Jiuqiang Tang, Hadon Nash, Chris McClanahan, Esha Uboweja, Michael Hays, Fan Zhang, Chuo-Ling Chang, Ming Guang Yong, Juhyun Lee, Wan-Teh Chang, Wei Hua, Manfred Georg, Matthias Grundmann (2019-06-14). "MediaPipe: A Framework for Building Perception Pipelines". arXiv. arXiv:1906.08172. https://arxiv.org/abs/1906.08172. Retrieved 2026-10-06.
  5. ↑ 5.0 5.1 Fan Zhang, Valentin Bazarevsky, Andrey Vakunov, Andrei Tkachenka, George Sung, Chuo-Ling Chang, Matthias Grundmann (2020-06-18). "MediaPipe Hands: On-device Real-time Hand Tracking". CVPR Workshop on Computer Vision for Augmented and Virtual Reality, Seattle, 2020 (arXiv). arXiv:2006.10214. https://arxiv.org/abs/2006.10214. Retrieved 2026-10-06.
  6. ↑ 6.0 6.1 6.2 Yury Kartynnik, Artsiom Ablavatski, Ivan Grishchenko, Matthias Grundmann (2019-07-15). "Real-time Facial Surface Geometry from Monocular Video on Mobile GPUs". CVPR Workshop on Computer Vision for Augmented and Virtual Reality, Long Beach, 2019 (arXiv). arXiv:1907.06724. https://arxiv.org/abs/1907.06724. Retrieved 2026-10-06.
  7. ↑ 7.0 7.1 Paul Ruiz, Kris Tonthat (2023-05-11). "Introducing MediaPipe Solutions for On-Device Machine Learning". Google Developers Blog. https://developers.googleblog.com/introducing-mediapipe-solutions-for-on-device-machine-learning/. Retrieved 2026-10-06.
  8. ↑ 8.0 8.1 8.2 8.3 8.4 "MediaPipe Solutions guide". Google AI Edge. Google for Developers. 2026-09-28. https://developers.google.com/edge/mediapipe/solutions/guide. Retrieved 2026-10-06.
  9. ↑ 9.0 9.1 "MediaPipe v1.0.0". GitHub. Google. 2026-07-28. https://github.com/google-ai-edge/mediapipe/releases/tag/v1.0.0. Retrieved 2026-10-06.
  10. ↑ 10.0 10.1 10.2 10.3 "MediaPipe - Google Research Perception - CV4AR/VR". Google Research Perception. Google. 2019. https://sites.google.com/view/perception-cv4arvr/mediapipe. Retrieved 2026-10-06.
  11. ↑ 11.0 11.1 Artsiom Ablavatski, Ivan Grishchenko (2019-03-08). "Real-Time AR Self-Expression with Machine Learning". Google AI Blog. https://research.google/blog/real-time-ar-self-expression-with-machine-learning/. Retrieved 2026-10-06.
  12. ↑ "google-ai-edge/mediapipe repository metadata". GitHub REST API. GitHub. https://api.github.com/repos/google-ai-edge/mediapipe. Retrieved 2026-10-06.
  13. ↑ 13.0 13.1 "Releases - google-ai-edge/mediapipe". GitHub. Google. https://github.com/google-ai-edge/mediapipe/releases. Retrieved 2026-10-06.
  14. ↑ Tommy Palladino (2019-08-21). "Google's AI Solution for Hand & Finger Tracking Could Be Huge for Smartglasses". Next Reality. https://next.reality.news/news/googles-ai-solution-for-hand-finger-tracking-could-be-huge-for-smartglasses-0203914/. Retrieved 2026-10-06.
  15. ↑ 15.0 15.1 Michael Hays, Tyler Mullen (2020-01-28). "MediaPipe on the Web". Google Developers Blog. https://developers.googleblog.com/en/mediapipe-on-the-web/. Retrieved 2026-10-06.
  16. ↑ 16.0 16.1 16.2 Adel Ahmadyan, Tingbo Hou (2020-03-11). "Real-Time 3D Object Detection on Mobile Devices with MediaPipe". Google Research Blog. https://research.google/blog/real-time-3d-object-detection-on-mobile-devices-with-mediapipe/. Retrieved 2026-10-06.
  17. ↑ 17.0 17.1 Andrey Vakunov, Dmitry Lagun (2020-08-06). "MediaPipe Iris: Real-time Iris Tracking and Depth Estimation". Google Research Blog. https://research.google/blog/mediapipe-iris-real-time-iris-tracking-depth-estimation/. Retrieved 2026-10-06.
  18. ↑ 18.0 18.1 Ivan Grishchenko, Valentin Bazarevsky (2020-12-10). "MediaPipe Holistic - Simultaneous Face, Hand and Pose Prediction, on Device". Google Research Blog. https://research.google/blog/mediapipe-holistic-simultaneous-face-hand-and-pose-prediction-on-device/. Retrieved 2026-10-06.
  19. ↑ Tingbo Hou, Tyler Mullen (2020-10-30). "Background Features in Google Meet, Powered by Web ML". Google Research Blog. https://research.google/blog/background-features-in-google-meet-powered-by-web-ml/. Retrieved 2026-10-06.
  20. ↑ Google AI Edge team (2024-09-04). "TensorFlow Lite is now LiteRT". Google Developers Blog. https://developers.googleblog.com/tensorflow-lite-is-now-litert/. Retrieved 2026-10-06.
  21. ↑ "LLM Inference guide". Google AI Edge. Google for Developers. https://developers.google.com/edge/mediapipe/solutions/genai/llm_inference. Retrieved 2026-10-06.
  22. ↑ "mediapipe". PyPI. Python Software Foundation. 2026-08-14. https://pypi.org/project/mediapipe/. Retrieved 2026-10-06.
  23. ↑ "Framework concepts". Google AI Edge. Google for Developers. https://developers.google.com/edge/mediapipe/framework/framework_concepts/overview. Retrieved 2026-10-06.
  24. ↑ 24.0 24.1 "MediaPipe Tasks". Google AI Edge. Google for Developers. https://developers.google.com/edge/mediapipe/solutions/tasks. Retrieved 2026-10-06.
  25. ↑ 25.0 25.1 "Hand landmarks detection guide". Google AI Edge. Google for Developers. https://developers.google.com/edge/mediapipe/solutions/vision/hand_landmarker. Retrieved 2026-10-06.
  26. ↑ "Gesture recognition task guide". Google AI Edge. Google for Developers. https://developers.google.com/edge/mediapipe/solutions/vision/gesture_recognizer. Retrieved 2026-10-06.
  27. ↑ 27.0 27.1 "Face landmark detection guide". Google AI Edge. Google for Developers. https://developers.google.com/edge/mediapipe/solutions/vision/face_landmarker. Retrieved 2026-10-06.
  28. ↑ "Pose landmark detection guide". Google AI Edge. Google for Developers. https://developers.google.com/edge/mediapipe/solutions/vision/pose_landmarker. Retrieved 2026-10-06.
  29. ↑ "Holistic landmarks detection task guide". Google AI Edge. Google for Developers. https://developers.google.com/edge/mediapipe/solutions/vision/holistic_landmarker. Retrieved 2026-10-06.
  30. ↑ Valentin Bazarevsky, Yury Kartynnik, Andrey Vakunov, Karthik Raveendran, Matthias Grundmann (2019-07-11). "BlazeFace: Sub-millisecond Neural Face Detection on Mobile GPUs". CVPR Workshop on Computer Vision for Augmented and Virtual Reality, Long Beach, 2019 (arXiv). arXiv:1907.05047. https://arxiv.org/abs/1907.05047. Retrieved 2026-10-06.
  31. ↑ Valentin Bazarevsky, Ivan Grishchenko, Karthik Raveendran, Tyler Zhu, Fan Zhang, Matthias Grundmann (2020-06-17). "BlazePose: On-device Real-time Body Pose tracking". CVPR Workshop on Computer Vision for Augmented and Virtual Reality, Seattle, 2020 (arXiv). arXiv:2006.10204. https://arxiv.org/abs/2006.10204. Retrieved 2026-10-06.
  32. ↑ Ivan Grishchenko, Artsiom Ablavatski, Yury Kartynnik, Karthik Raveendran, Matthias Grundmann (2020-06-19). "Attention Mesh: High-fidelity Face Mesh Prediction in Real-time". CVPR Workshop on Computer Vision for Augmented and Virtual Reality, Seattle, 2020 (arXiv). arXiv:2006.10962. https://arxiv.org/abs/2006.10962. Retrieved 2026-10-06.
  33. ↑ Dennis Reimer, Iana Podkosova, Daniel Scherzer, Hannes Kaufmann (2023-07-19). "Evaluation and improvement of HMD-based and RGB-based hand tracking solutions in VR". Frontiers in Virtual Reality, vol. 4, art. 1169313. https://doi.org/10.3389/frvir.2023.1169313. Retrieved 2026-10-06.
  34. ↑ XR Development team at KDDI and Alpha-U (2023-07-10). "MediaPipe: Enhancing Virtual Humans to be more realistic". Google Developers Blog. https://developers.googleblog.com/mediapipe-enhancing-virtual-humans-to-be-more-realistic/. Retrieved 2026-10-06.
  35. ↑ "Portal-ble". GitHub. Brown University HCI. https://github.com/brownhci/Portalble. Retrieved 2026-10-06.
  36. ↑ homuler. "MediaPipeUnityPlugin". GitHub. https://github.com/homuler/MediaPipeUnityPlugin. Retrieved 2026-10-06.