Position and orientation
More actions
Position and orientation (P&O, also known as PnO) is a quantity that holds position and orientation.
It is used in PnO tracking. It is used in magnetic tracking systems, such as systems from Polhemus. The Polhemus VIPER manual, for example, describes a system in which sensors detect the emitted magnetic field and the electronics unit "calculates Position and Orientation (P&O) of each of the Sensors".[1]
PnO is substantially the same concept as "pose". It is the rotation and translation of an object relative to a root transform. In terms of a transform matrix with SRT order, it is just the rotation (R) and translation (T).
Together the two parts make up the six degrees of freedom of a rigid body: three coordinates for position and three angles for orientation.[2][3] In VR and AR, trackers estimate this quantity for the user's head so that a head-mounted display can render the correct view, and report the same data for hand-held controllers and other tracked objects; simpler 3DoF devices supply only the orientation part.[3][4] Ivan Sutherland described the job of the head position sensor in his 1968 head-mounted display as measuring and reporting "the position and orientation of the user's head".[5] Modern XR interfaces such as OpenXR and WebXR define a pose type made of a position vector and a unit quaternion.[6][4]
Definition
In his textbook Virtual Reality, Steven M. LaValle writes that "for convenience, we will refer to the position and orientation of a body as its pose", and describes tracking of all six degrees of freedom of a moving rigid body, with head tracking as the most important case.[7] Polhemus describes the same six values for its electromagnetic trackers: the object's position within X, Y and Z coordinates of a space, and its orientation, given as pitch, roll and yaw.[2]
The WebXR specification uses the P&O split to classify hardware. A 3DoF device "is one that can only track rotational movement", which the specification says is common in devices that rely only on accelerometer and gyroscope readings; such devices do not respond to translational movement, although they may estimate it with a model of the neck or arms. A 6DoF device "is one that can track both rotation and translation, enabling precise 1:1 tracking in space".[4] The first kind of device has rotational tracking only; the second adds positional tracking.
Applying a pose to an object means rotating it and then translating it. LaValle shows that a rotation matrix R and a translation (xt, yt, zt) can be combined in one 4 by 4 homogeneous transformation matrix, which he calls a rigid body transform "meaning that it does not distort objects"; the matrix represents rotation followed by translation, "not the other way around", because the two operations do not commute.[8] OpenXR states the same rule for its pose structure: "the rotation described by orientation is always applied before the translation described by position".[6] The WebXR specification says that when interpreting a rigid transform "the orientation is always applied prior to the position".[4]
Representations
Position is almost always stored as a three-component vector. Orientation has several common representations, each with trade-offs:
| Representation | Description | Notes |
|---|---|---|
| Yaw, pitch and roll (Euler angles) | Three successive rotations about coordinate axes | LaValle calls this "one of the simplest ways to parameterize 3D rotations", but warns that the order matters and that the angles have kinematic singularities: at a pitch of 90 degrees the viewpoint can "spin uncontrollably", the effect known as gimbal lock in mechanical gimbals.[8] In the 1979 Polhemus paper, sensor orientation is azimuth, then elevation, then roll, and "the order of the rotations cannot be interchanged".[9] |
| Rotation matrix | A 3 by 3 matrix | LaValle notes that the inverse of a rotation matrix is its transpose, which makes rigid transforms simple to invert.[8] |
| Unit quaternion | A four-component vector of unit length that encodes an axis and an angle | Introduced by William Rowan Hamilton in 1843. LaValle describes quaternions as the representation that handles "all problems with 3D rotations" except the unavoidable fact that q and -q encode the same rotation.[8] The Oculus Rift tracking team chose them because they allow "singularity-free manipulation of rotations with few parameters".[10] |
| Homogeneous matrix | Rotation and translation together in one 4 by 4 (or 3 by 4) matrix | Sutherland's 1968 display used a 4 by 4 matrix of 18-bit fixed-point numbers so that "the single four-by-four matrix multiplication provides for both translation and rotation".[5] |
Software interfaces pick different combinations:
| Interface | Pose type | Contents |
|---|---|---|
| OpenXR | XrPosef |
An XrQuaternionf orientation and an XrVector3f position in meters. A runtime must return XR_ERROR_POSE_INVALID if the orientation's norm deviates from unit length by more than 1%.[6]
|
| WebXR | XRRigidTransform, inside XRPose |
A position point in meters and a unit quaternion orientation; XRPose adds optional linear and angular velocity and an emulatedPosition flag.[4][11]
|
| OpenVR (SteamVR) | TrackedDevicePose_t |
A 3 by 4 matrix (HmdMatrix34_t mDeviceToAbsoluteTracking), velocity in m/s, angular velocity, a tracking result and a validity flag.[12]
|
| Unity | Pose |
The engine describes it as a "Representation of a Position, and a Rotation in 3D Space", with a position component and a rotation component.[13] |
| Polhemus VIPER | P&O data frame | Position in inches, feet, centimeters or meters; orientation as Euler angle degrees, Euler angle radians or a quaternion. The default is inches and Euler degrees.[1] |
Reference frames
A position and orientation only has meaning relative to a reference frame. In the Polhemus VIPER, P&O is calculated at the electromagnetic center of each sensor relative to a Cartesian origin that by default is the electromagnetic center of a VIPER source, although the user can define another reference.[1] Raab and colleagues likewise specified the sensor position "relative to the source coordinate frame", and noted that the result could be converted mathematically to any other frame.[9]
XR runtimes expose several frames. OpenXR's VIEW reference space follows the viewer's view origin; LOCAL is a world-locked origin "gravity-aligned to exclude pitch and roll"; LOCAL_FLOOR is the same with its origin at an estimate of floor level; and STAGE is a rectangular walkable area with its origin on the floor at the center. Each uses +Y up, and VIEW, LOCAL and LOCAL_FLOOR use +X to the right and -Z forward.[14] WebXR uses the same axis convention ("+X is considered 'Right', +Y is considered 'Up', and -Z is considered 'Forward'") for the viewer, local, local-floor, bounded-floor and unbounded reference spaces.[4] OpenVR offers a seated universe, relative to the seated zero pose, and a standing universe, relative to the chaperone soft bounds.[15]
Runtimes also report how much of a pose is real. OpenXR's xrLocateSpace returns the location of a space in a base space at a given time, and separate flags state whether the orientation and position are valid and whether each is actively tracked. When tracking is lost, the runtime should keep supplying inferred or last-known values, for example from a neck model or dead reckoning, with the position still marked valid but no longer marked tracked.[16][17] In WebXR, emulatedPosition is false for "an actively tracked 6DoF pose based on sensor readings" and true when the position includes a computed offset such as one from a neck or arm model.[4]
History
Sutherland's head-mounted display
Sutherland's 1968 system measured "only the position and orientation of the optical system fastened to the user's head", since the perspective picture only had to change when the head moved, not the eyes. A computer used the measured head position to compute the elements of a rotation and translation matrix for each viewing position.[5] The initial equipment allowed a working volume of head motion about six feet in diameter and three feet high, in which the user could turn completely around and tilt the head up or down about forty degrees. Sutherland's target resolution was 1/100 of an inch and "one part in 10,000 of rotation".[5]
Two sensors were built. The mechanical one was an arm hanging from the ceiling with two universal joints and a sliding center section "to provide the six motions required to measure both translation and rotation", each joint read by a digital shaft encoder; Sutherland called it "rather heavy and uncomfortable to use" but built it "to have a sure method of measuring head position". The second was an ultrasonic sensor.[5]
Magnetic P&O tracking
Jack Kuipers patented a magnetic method for tracking an object and determining its orientation in 1975 (US patent 3,868,565, filed in 1973).[18] In 1979, Frederick Raab, Ernest Blood, Terry Steiner and Herbert Jones of Polhemus Navigation Sciences published "Magnetic Position and Orientation Tracking System", a detailed analysis of the concept "invented by J. Kuipers".[9] Their abstract states that "three-axis generation and sensing of quasi-static magnetic-dipole fields provide information sufficient to determine both the position and orientation of the sensor relative to the source". The source was driven through three sequential excitation states, and the three resulting output vectors of a three-axis sensor carried enough information to determine all six degrees of freedom; the system tracked them by computing small changes and adding them to the previous estimate.[9]
The paper's worked example was a helmet-mounted sight for fighter pilots: the sensor on the helmet, the source above and behind the head, typical separations of a few centimeters to one meter, carrier frequencies of 7 to 14 kHz, and updates "typically at 30- to 120-Hz rates". The authors gave a rule of thumb for metal in the environment: an object at least twice as far from the source as the sensor produces a scattered field of 1 percent or less of the desired field.[9]
Welch and Foxlin's 2002 survey describes magnetic tracking as a popular method for interactive graphics for many years. The user-worn part can be small, the fields pass through the human body so no line of sight is needed, and one source can excite several sensors. The drawbacks are field distortion from ferromagnetic and conductive objects, which is why some projection displays were built from wood or plastic, and an inverse cubic falloff of field strength that makes positional jitter grow "as the fourth power of the separation distance".[3] Current Polhemus hardware follows the same principle. The company states that its trackers deliver "true six degrees of freedom" because position and orientation are both measured natively.[2] The VIPER manual (2020) specifies update rates of up to 960 Hz for the VIPER 8 and VIPER 16 (240 Hz for the VIPER 4), with 1 millisecond latency at 960 Hz, and a static accuracy of 0.015 in. RMS for position and 0.10 degrees RMS for orientation in a magnetically clean environment.[1]
Consumer headsets
The 2014 paper on head tracking for the Oculus Rift Development Kit dealt mainly with orientation. Its authors described "efficiently maintaining human head orientation using low-cost MEMS sensors", with a gyroscope, accelerometer and magnetometer on one board, sensor observations reported at 1000 Hz, and the orientation kept as a quaternion. Gravity was used to correct tilt drift and the magnetic field to correct yaw drift.[10] For position, the Oculus SDK moved the rotation center to the base of the neck so that the eyes shifted as the head turned. The authors tested double integration of the accelerometer and reported that after one minute its position error exceeded 500 m, compared with 1.1 m for a kinematically constrained method. They wrote that this kinematic method was only effective over a few seconds and was insufficient as a standalone technique, and that their most recent prototypes combined their methods with position data from a camera observing infrared LEDs on the headset.[10]
In 2019, Oculus Insight shipped in the Oculus Quest and Oculus Rift S. Meta's engineers described it as "the first time that fully untethered six-degree-of-freedom (6DoF) headset and controller tracking has shipped in a consumer AR/VR device". It computes "an accurate and real-time position for the headset and controllers every millisecond" from IMU data, headset camera images and infrared LEDs in the controllers, using visual-inertial SLAM for the headset and constellation tracking for the controllers.[19]
Measuring position and orientation
Welch and Foxlin describe an imaginary "tracker-on-a-chip" that would track "all six degrees of freedom (position and orientation)" with resolution better than 1 mm in position and 0.1 degree in orientation, at 1,000 Hz with latency under 1 ms. They wrote that every tracker of the time fell short on at least seven of these ten characteristics, and that motion trackers most often derive pose estimates from mechanical, inertial, acoustic, magnetic, optical and radio frequency sensors.[3]
Orientation is the easier half with inertial sensors. A strapdown inertial system gets orientation by integrating three gyroscopes, and gets position by rotating the accelerometer readings into navigation coordinates, removing gravity and integrating twice.[3] The second integration makes position drift much faster: Welch and Foxlin calculate that an accelerometer bias of 1 milli-g would put the position estimate 4.5 meters off after 30 seconds, and that an orientation error of 1 milliradian from the gyroscopes produces exactly that bias in the gravity compensation.[3] LaValle makes the same point: a calibration error leads to linearly growing drift for a gyroscope but quadratically growing drift after the double integration needed for position, which "becomes unbearable in practice after a fraction of a second". An IMU also cannot tell constant velocity from standing still. LaValle identifies visibility, locating features along lines of sight from a known place, as "the most powerful paradigm for 6-DOF tracking", while IMUs stay in these systems because of their high sampling rates and good handling of rotation.[7] Oculus Insight's visual-inertial approach, described above, is one example of sensor fusion of this kind: it combines inertial measurement unit data with camera images.[19]
Latency and prediction
A measured pose is already old when it reaches the display. Ronald Azuma's 1997 survey of augmented reality defines the end-to-end system delay as the time between the moment "the tracking system measures the position and orientation of the viewpoint" and the moment the matching images appear in the displays. Delays of 100 ms were then typical; with a head rotating at 50 degrees per second, that lag gives an angular error of 5 degrees, or almost 60 mm at a 68 cm arm length. Azuma called system delay "the largest single source of registration error in existing AR systems".[20]
Systems therefore predict the pose for the time the image will be seen (see Motion-to-photon latency and Predictive tracking). The Oculus Rift team compared no prediction, constant angular velocity and constant angular acceleration over a 20 ms interval; average errors were 1.46302, 0.19395 and 0.07596 degrees respectively, and with a 40 ms interval the third method averaged 0.17 degrees.[10] OpenVR's GetDeviceToAbsoluteTrackingPose takes a number of seconds to predict ahead, and Valve's documentation advises applications to pass the time until photons leave the display.[15] OpenXR specifies that when a space is located at a future time, the runtime should base the result on its most up-to-date prediction of how the world will be at that time.[16] Oculus Insight uses "an extrapolation function with dynamic damping" to predict where the user's head and hands will move in the milliseconds ahead.[19]
Applications in VR and AR
In virtual reality, the tracked P&O of the head sets the virtual camera. Welch and Foxlin list "view control" among the purposes of motion tracking: "position and orientation control of a virtual camera for rendering computer graphics in a head-mounted display".[3] Tracking position as well as orientation strengthens the depth cue of parallax as the user moves the head from side to side, and lets the user approach an object and look at it from any viewpoint.[7] The same poses place hand-held controllers and tracked objects in the scene.
Augmented reality is stricter. Azuma writes that it "demands much more accurate registration than Virtual Environments", because errors that show up as visual-kinesthetic conflicts in VR become visual-visual conflicts in a see-through display, where "even tiny offsets" between real and virtual objects are easy to see. For comparison he notes that the full moon spans only about 0.5 degrees of arc, so the angular accuracy needed is "a small fraction of a degree".[20]
Outside consumer XR, the 1979 Polhemus paper describes magnetic P&O trackers in helmet-mounted sights, where the helmet's orientation gives the wearer's line of sight, so that a target sighted by one crew member can be cued to another, or a range-measuring radar can be slaved to that line of sight.[9] Polhemus states that because its sensors need no line of sight they can be embedded inside almost anything, which the company says makes them "the top choice for many training and simulation applications".[2]
See also
References
- ↑ 1.0 1.1 1.2 1.3 "VIPER User Manual for Models VIPER 4, VIPER 8, VIPER 16 (URM18PH392 Rev. A)". Polhemus. Alken, Inc. dba Polhemus. 2020-05. https://www.creact.co.jp/measure/mocap/viper/file/VIPER_User_Manual_URM18PH392-A.pdf. Retrieved 2026-09-27.
- ↑ 2.0 2.1 2.2 2.3 "Our Technology". Polhemus. https://polhemus.com/technology. Retrieved 2026-09-27.
- ↑ 3.0 3.1 3.2 3.3 3.4 3.5 3.6 Greg Welch, Eric Foxlin (2002). "Motion Tracking: No Silver Bullet, but a Respectable Arsenal". IEEE Computer Graphics and Applications, vol. 22, no. 6. pp. 24-38. doi:10.1109/MCG.2002.1046626. https://doi.org/10.1109/MCG.2002.1046626. Retrieved 2026-09-27.
- ↑ 4.0 4.1 4.2 4.3 4.4 4.5 4.6 "WebXR Device API (W3C Candidate Recommendation Draft)". World Wide Web Consortium. 2026-06-09. https://www.w3.org/TR/webxr/. Retrieved 2026-09-27.
- ↑ 5.0 5.1 5.2 5.3 5.4 Ivan E. Sutherland (1968). "A head-mounted three dimensional display". Proceedings of the December 9-11, 1968, Fall Joint Computer Conference, Part I (AFIPS '68). pp. 757-764. doi:10.1145/1476589.1476686. https://doi.org/10.1145/1476589.1476686. Retrieved 2026-09-27.
- ↑ 6.0 6.1 6.2 "XrPosef(3)". OpenXR 1.1 Reference Pages. The Khronos Group. https://registry.khronos.org/OpenXR/specs/1.1/man/html/XrPosef.html. Retrieved 2026-09-27.
- ↑ 7.0 7.1 7.2 Steven M. LaValle (2020). "Virtual Reality, Chapter 9: Tracking". Virtual Reality (online edition). University of Oulu. http://lavalle.pl/vr/vrch9.pdf. Retrieved 2026-09-27.
- ↑ 8.0 8.1 8.2 8.3 Steven M. LaValle (2020). "Virtual Reality, Chapter 3: The Geometry of Virtual Worlds". Virtual Reality (online edition). University of Oulu. http://lavalle.pl/vr/vrch3.pdf. Retrieved 2026-09-27.
- ↑ 9.0 9.1 9.2 9.3 9.4 9.5 Frederick H. Raab, Ernest B. Blood, Terry O. Steiner, Herbert R. Jones (1979-09). "Magnetic Position and Orientation Tracking System". IEEE Transactions on Aerospace and Electronic Systems, vol. AES-15, no. 5. pp. 709-718. doi:10.1109/TAES.1979.308860. https://doi.org/10.1109/TAES.1979.308860. Retrieved 2026-09-27.
- ↑ 10.0 10.1 10.2 10.3 Steven M. LaValle, Anna Yershova, Max Katsev, Michael Antonov (2014). "Head Tracking for the Oculus Rift". 2014 IEEE International Conference on Robotics and Automation (ICRA). pp. 187-194. https://msl.cs.illinois.edu/~lavalle/papers/LavYerKatAnt14.pdf. Retrieved 2026-09-27.
- ↑ "XRPose". MDN Web Docs. Mozilla. https://developer.mozilla.org/en-US/docs/Web/API/XRPose. Retrieved 2026-09-27.
- ↑ "openvr.h". ValveSoftware/openvr on GitHub. Valve Corporation. https://github.com/ValveSoftware/openvr/blob/master/headers/openvr.h. Retrieved 2026-09-27.
- ↑ "Pose". Unity Scripting API. Unity Technologies. https://docs.unity3d.com/ScriptReference/Pose.html. Retrieved 2026-09-27.
- ↑ "XrReferenceSpaceType(3)". OpenXR 1.1 Reference Pages. The Khronos Group. https://registry.khronos.org/OpenXR/specs/1.1/man/html/XrReferenceSpaceType.html. Retrieved 2026-09-27.
- ↑ 15.0 15.1 "IVRSystem::GetDeviceToAbsoluteTrackingPose". ValveSoftware/openvr wiki on GitHub. Valve Corporation. https://github.com/ValveSoftware/openvr/wiki/IVRSystem::GetDeviceToAbsoluteTrackingPose. Retrieved 2026-09-27.
- ↑ 16.0 16.1 "xrLocateSpace(3)". OpenXR 1.1 Reference Pages. The Khronos Group. https://registry.khronos.org/OpenXR/specs/1.1/man/html/xrLocateSpace.html. Retrieved 2026-09-27.
- ↑ "XrSpaceLocationFlagBits(3)". OpenXR 1.1 Reference Pages. The Khronos Group. https://registry.khronos.org/OpenXR/specs/1.1/man/html/XrSpaceLocationFlagBits.html. Retrieved 2026-09-27.
- ↑ Jack Kuipers (1975-02-25). "US3868565A - Object tracking and orientation determination means, system and process". Google Patents. https://patents.google.com/patent/US3868565A/en. Retrieved 2026-09-27.
- ↑ 19.0 19.1 19.2 Joel Hesch, Anna Kozminski, Oskar Linde (2019-08-22). "Powered by AI: Oculus Insight". Meta AI Blog. Meta Platforms. https://ai.meta.com/blog/powered-by-ai-oculus-insight/. Retrieved 2026-09-27.
- ↑ 20.0 20.1 Ronald T. Azuma (1997-08). "A Survey of Augmented Reality". Presence: Teleoperators and Virtual Environments, vol. 6, no. 4. pp. 355-385. doi:10.1162/pres.1997.6.4.355. https://doi.org/10.1162/pres.1997.6.4.355. Retrieved 2026-09-27.