Perspective-n-Point
More actions
Perspective-n-Point (PnP) is the problem of estimating the position and orientation (the pose) of a calibrated camera from a set of n points whose 3D coordinates are known and whose 2D projections in the camera image have been identified. Lepetit, Moreno-Noguer and Fua define the aim of the problem as determining "the position and orientation of a camera given its intrinsic parameters and a set of n correspondences between 3D points and their 2D projections", and list computer vision, robotics and augmented reality among its applications.[1] The name was introduced in 1981 by Martin Fischler and Robert Bolles of SRI International in the paper that also introduced the RANSAC (random sample consensus) algorithm.[2]
The output of a PnP solver is a rotation and a translation, a full six degrees of freedom (6DoF) pose. In VR and AR systems, PnP solvers recover the pose of a camera relative to a printed fiducial marker, compute the pose of a headset or controller from the image positions of its infrared LEDs, and relocalize a device against a stored map of 3D feature points.[3][4][5] The smallest useful case, with three points, is called the Perspective-three-point (P3P) problem. It can have up to four solutions, so practical systems use additional points to choose among them.[2][1]
Definition
PnP assumes that the camera's intrinsic parameters (focal length and principal point, and in software libraries usually the lens distortion coefficients as well) are already known from a prior calibration. Fischler and Bolles stated the problem in this form: the principal point and focal length are known, so the angle between the viewing rays to any pair of control points can be computed, and the task is to find the lengths of the rays ("legs") from the camera's center of perspective to each control point. Once the distances to three control points are known, the camera position and the orientation of its image plane follow directly.[2]
The OpenCV library describes the same problem as finding the rotation and translation "that minimizes the reprojection error from 3D-2D point correspondences". In its formulation, a world point is transformed into the camera frame by a 4x4 rigid transformation (a 3x3 rotation matrix and a translation vector), then projected onto the image with the perspective projection model and the 3x3 camera intrinsic matrix. The library's convention puts the camera X axis to the right, the Y axis downward and the Z axis forward.[6] The reprojection error is the distance in the image between where each known 3D point is observed and where the estimated pose predicts it should appear; OpenCV's iterative method adjusts the pose to make the sum of these squared distances as small as possible.[6]
PnP is related to, but distinct from, two neighboring problems. When the intrinsic parameters are unknown as well, they must be estimated together with the pose; Schönberger and Frahm note that this extended version is needed for Internet photographs without reliable calibration data.[7] In structure from motion and bundle adjustment the 3D points themselves are unknown and are estimated jointly with the camera poses, while in PnP the 3D points are given.[7]
Number of points and solutions
The number of correspondences determines whether the pose is defined at all and how many poses fit the data:
| Case | Result | Source |
|---|---|---|
| Two points (P2P) | Infinitely many solutions: the camera can lie anywhere on a circle rotated about the line joining the two points | Fischler and Bolles (1981)[2] |
| Three points (P3P) | At most four positive solutions; Fischler and Bolles gave an example in which all four occur | Fischler and Bolles (1981)[2] |
| Four coplanar points (no three collinear) | A unique solution | Fischler and Bolles (1981)[2] |
| Four or five points in general position | Multiple physically real solutions are possible; at least two for some non-planar four-point configurations | Fischler and Bolles (1981)[2] |
| Six points in general position (P6P) | A unique solution | Fischler and Bolles (1981)[2] |
In Fischler and Bolles' derivation, the three P3P equations follow from the law of cosines applied to the triangle formed by each pair of control points and the camera center. A system of three independent second-degree equations in three unknowns can have at most eight solutions, and because every term is constant or of second degree, each positive solution has a mirror-image negative one, which leaves at most four physically meaningful poses.[2] Most P3P solvers reduce the problem to the roots of a quartic polynomial, and a fourth correspondence is used to pick the correct root.[1][8] For that reason OpenCV's P3P methods take exactly four points and return a unique pose.[6]
Planar targets such as square markers have a further, practical ambiguity. Although four coplanar, non-collinear points determine the pose in theory, Gerald Schweighofer and Axel Pinz showed in 2006 that the error function used by real-time trackers can have two distinct local minima, even with wide-angle lenses and nearby targets. They derived an analytical way to locate the second minimum and built a pose estimator for planar targets around it.[9]
History
The three-point case was studied long before digital cameras. According to Persson and Nordberg, the first P3P solution was published in 1841 by J. A. Grunert, who showed that the problem has up to four feasible solutions.[8] In photogrammetry and cartography, recovering where a photograph was taken from known ground control points was known as determining the elements of exterior camera orientation, and Fischler and Bolles noted in 1981 that it was "routinely solved using a least-squares technique" with a human operator matching image points to control points by hand. They added that the photogrammetric literature of the time offered the iterative Church method for P3P "without any indication that more than one physically real solution is possible".[2]
Fischler and Bolles' 1981 paper in Communications of the ACM framed the task for automated systems as the "Location Determination Problem": given an image of landmarks with known locations, determine the point in space from which the image was taken. They named the reduced geometric problem the "perspective-n-point" problem, gave a closed-form P3P solution based on a quartic polynomial, and combined it with RANSAC so that wrong landmark matches from imperfect feature detectors would not corrupt the result.[2] Haralick, Lee, Ottenberg and Nölle published a review and analysis of the classical three-point solutions in the International Journal of Computer Vision in 1994.[10]
Later research concentrated on speed, numerical stability and the use of many points at once. Lu, Hager and Mjolsness (2000) minimized an error measured in object space rather than in the image, which gave an iterative algorithm that is globally convergent and produces orthogonal rotation matrices directly.[11] Gao, Hou, Tang and Cheng (2003) used Wu-Ritt zero decomposition to produce what they described as the first complete analytical solution to P3P, with explicit criteria for when the problem has one, two, three or four solutions.[12] EPnP, published by Lepetit, Moreno-Noguer and Fua at EPFL, is a non-iterative solution whose cost grows linearly with the number of points, O(n), whereas the state-of-the-art non-iterative methods the authors compared against were O(n5) or O(n8). The authors cited real-time feature-point camera tracking, which has to handle hundreds of noisy points, as the motivating use.[1] Later work produced faster minimal solvers, methods specialized for planar targets, and solvers that guarantee the global optimum.
Solution methods
PnP solvers fall into a few families. Minimal solvers handle exactly three points and are designed to be fast enough to run many times inside a RANSAC loop. Non-iterative solvers for n points use all correspondences at once. Iterative solvers refine an initial guess by minimizing reprojection error, and planar solvers exploit the geometry of a flat target.[1][6] The table lists notable published methods.
| Method | Authors | Venue and year | Points | Approach |
|---|---|---|---|---|
| First P3P solution | J. A. Grunert | 1841 | 3 | Showed that P3P has up to four feasible solutions[8] |
| RANSAC-based location determination | Martin A. Fischler, Robert C. Bolles | Communications of the ACM, 1981 | 3 and more | Closed-form P3P via a quartic polynomial inside a random-sampling loop that rejects gross errors[2] |
| Orthogonal iteration | C.-P. Lu, Gregory D. Hager, Eric Mjolsness | IEEE TPAMI, 2000 | n | Iterative minimization of object-space collinearity error; globally convergent[11] |
| Complete solution classification | Xiao-Shan Gao, Xiao-Rong Hou, Jianliang Tang, Hang-Fei Cheng | IEEE TPAMI, 2003 | 3 | Triangular decomposition of the P3P equations; criteria for one to four solutions; algorithm named CASSC[12] |
| Robust planar pose | Gerald Schweighofer, Axel Pinz | IEEE TPAMI, 2006 | 4 or more coplanar | Analytically locates the second local minimum of planar pose[9] |
| EPnP | Vincent Lepetit, Francesc Moreno-Noguer, Pascal Fua | International Journal of Computer Vision, 2008 (online) | 4 or more | Expresses all points as a weighted sum of four virtual control points; O(n); optional Gauss-Newton refinement[1] |
| Kneip P3P | Laurent Kneip, Davide Scaramuzza, Roland Siegwart | CVPR 2011 | 3 | Computes the camera pose directly in one stage through intermediate reference frames; reported as 15 times faster than a standard approach[13] |
| DLS | Joel A. Hesch, Stergios I. Roumeliotis | ICCV 2011 | 3 or more | Direct least squares; finds all minima from a system of three third-order polynomials, whose size does not depend on n[14] |
| IPPE | Toby Collins, Adrien Bartoli | International Journal of Computer Vision, 2014 | 4 or more coplanar | Analytic solution from the local transformation of a homography at one well-estimated point on the plane[15] |
| AP3P | Tong Ke, Stergios I. Roumeliotis | CVPR 2017 | 3 | Solves for camera attitude first from trigonometric equations, then position[16] |
| Lambda Twist | Mikael Persson, Klas Nordberg | ECCV 2018 | 3 | Diagonalization that needs one real root of a cubic instead of all roots of a quartic; never produces invalid or duplicate solutions[8] |
| SQPnP | George Terzakis, Manolis Lourakis | ECCV 2020 | 3 or more | Non-linear quadratic program solved by sequential quadratic programming; always finds the global minimum, including for coplanar points[17] |
| Revisiting P3P | Yaqing Ding, Jian Yang, Viktor Larsson, Carl Olsson, Kalle Åström | CVPR 2023 | 3 | Formulates P3P as the intersection of two conics, with a tailored strategy for each case to handle configurations where earlier solvers fail[18] |
Robust estimation
Real correspondences contain mismatches, either from a feature matcher pairing the wrong points or from an LED being identified incorrectly. A standard approach, going back to Fischler and Bolles, is to draw small random subsets of correspondences, solve each subset with a minimal or fast solver, count how many of the other correspondences agree with the resulting pose, and keep the hypothesis with the largest consensus.[2] Kneip, Scaramuzza and Siegwart describe the speed of their P3P solver as "particularly suitable for any RANSAC-outlier-rejection step, which is always recommended before applying PnP or non-linear optimization of the final solution".[13] After the inliers are found, the pose is usually refined by non-linear least squares; EPnP, for example, can pass its closed-form result to a Gauss-Newton step that improves accuracy at little extra cost.[1]
Software implementations
OpenCV provides PnP through the solvePnP family in its calib3d module. solveP3P takes exactly three correspondences and can return up to four solutions; solvePnPGeneric returns all solutions for the methods that can produce several; solvePnPRansac wraps the solvers in a RANSAC scheme; and solvePnPRefineLM (Levenberg-Marquardt) and solvePnPRefineVVS (Gauss-Newton with an exponential-map rotation update) refine an existing estimate.[6] The method is selected with a flag:
| OpenCV flag | Based on | Input requirement |
|---|---|---|
SOLVEPNP_ITERATIVE |
Levenberg-Marquardt minimization of reprojection error, initialized by DLT (non-planar) or homography decomposition (planar) | At least 6 points (non-planar) or 4 (planar) for the initial solution |
SOLVEPNP_P3P |
Ding et al. (2023) from OpenCV 4.13.0; Gao et al. (2003) in 4.12.0 and earlier | Exactly 4 points |
SOLVEPNP_AP3P |
Ke and Roumeliotis (2017) | Exactly 4 points |
SOLVEPNP_EPNP |
Lepetit, Moreno-Noguer and Fua | 4 or more points |
SOLVEPNP_IPPE |
Collins and Bartoli (2014) | 4 or more coplanar points |
SOLVEPNP_IPPE_SQUARE |
Collins and Bartoli (2014), for square markers | Exactly 4 coplanar points in a fixed corner order |
SOLVEPNP_SQPNP |
Terzakis and Lourakis (2020) | 3 or more points |
The table follows the OpenCV documentation source.[6] The change of P3P method appears between the 4.12.0 and 4.13.0 release tags of that file.[19][20] The documentation marks two older flags, SOLVEPNP_DLS and SOLVEPNP_UPNP, as broken implementations that fall back to EPnP.[6]
The AprilTag fiducial library takes a different route for its square tags. Its estimate_tag_pose_orthogonal_iteration function computes an initial pose from the tag's homography, refines it with the Lu, Hager and Mjolsness orthogonal iteration, then uses the Schweighofer and Pinz method to find a possible second local minimum, and returns one or two candidate poses with their object-space errors.[21]
Applications in VR and AR
Marker-based AR
A square fiducial marker gives a tracker four known corner points, which is enough for a pose. The OpenCV ArUco tutorial states that "a single marker provides enough correspondences (its four corners) to obtain the camera pose", that the camera's intrinsic matrix and distortion coefficients must be calibrated first, and that the camera pose relative to each marker is computed with solvePnP().[3] Collins and Bartoli list the use of planar markers to perform AR among the main applications of plane-based pose estimation, and cite mobile applications, where the run-time budget is critical, as the reason a faster solution was needed.[15] OpenCV's SOLVEPNP_IPPE_SQUARE flag is documented as "suitable for marker pose estimation".[6]
The square-marker approach predates these libraries. ARToolKit was developed in 1999 by Hirokazu Kato and Mark Billinghurst at the University of Washington's HIT Lab.[22] In their 1999 paper on marker tracking for an AR conferencing system, Kato and Billinghurst found the four vertices of a square marker of known size in the image, estimated the marker-to-camera transformation from the marker's edge lines and vertices, and then optimized it so that the projected marker corners matched the measured image coordinates as closely as possible.[23] Related marker systems include ARTag.
LED constellation tracking
Outside-in tracking systems that watch infrared LEDs on a headset or controller have to solve a PnP problem: the 3D positions of the LEDs on the device are known from its design, and the camera sees a set of bright blobs. Persson and Nordberg motivated Lambda Twist with a related scenario. They described vision-based AR/VR localization that places "a few markers/beacons on a target" observed by a high-speed camera, and wrote that because "both latency and localization errors independently not only break immersion, but also cause nausea, accurate solutions and minimal latency are crucial".[8]
The open-source Monado OpenXR runtime shows how this works in practice. Its Oculus Rift driver, which handles both the Oculus Rift DK2 and Oculus Rift CV1, reads the headset's LED positions from the device and passes them to a constellation tracker.[24] The tracker performs a brute-force search over groups of four LEDs and four observed blobs. The first three points are passed to a Lambda Twist P3P solver to produce a hypothesis pose; if the fourth point matches, the code checks how many other LEDs project onto observed blobs and how small their reprojection error is. The search walks nearest-neighbor lists, because nearby blobs are likely to belong to nearby LEDs, and, when available, it rejects poses that disagree with the gravity direction extracted from the device's IMU. A "strong" match is defined as seven or more LEDs matching blobs within 1.5 pixels of reprojection error.[4] The correspondence search and the Lambda Twist solver were imported into Monado from OpenHMD on 8 April 2026; the files carry copyright notices from 2020 and 2019 respectively.[25][26][4] Oculus's own system for these headsets is described in Constellation.
Relocalization and SLAM
Camera-based SLAM systems use PnP to recover after tracking is lost. ORB-SLAM, a monocular SLAM system published by Mur-Artal, Montiel and Tardós in 2015, relocalizes by matching features from the current frame to the map points of candidate keyframes. For each candidate it runs RANSAC iterations with EPnP to find a camera pose, then optimizes that pose and searches for further matches before tracking resumes.[5] In incremental structure from motion, as in the COLMAP pipeline described by Schönberger and Frahm, each new image is registered to the existing reconstruction by solving PnP from 2D-3D feature correspondences, usually with RANSAC and a minimal pose solver.[7]
Large-scale visual localization
Visual positioning systems apply the same idea at city scale. In a 2020 paper in The International Journal of Robotics Research, Simon Lynen of Google and his co-authors describe feature-based localization as matching 2D image features to 3D map points and then "applying a perspective-n-point-pose (PnP) solver, e.g., the 3-point-pose (P3P) solver for calibrated cameras", inside a RANSAC loop. Their system runs visual-inertial tracking on the device for real-time rendering and asynchronously fuses in global poses computed on a server against maps built from Street View imagery; the authors list virtual and augmented reality among the target applications and report query latencies "in the 200 ms range" in a proof-of-concept that localized 2.5 million images in four cities.[27]
Research
Recent work connects PnP with deep learning for 6DoF object pose estimation from a single RGB image. PVNet (Peng et al., CVPR 2019) follows a two-stage approach: a network predicts the image locations of object keypoints, using per-pixel vectors that vote for each keypoint so that occluded or truncated keypoints can still be located, and a PnP solver then computes the pose, using the predicted keypoint uncertainties.[28] EPro-PnP (Chen et al., CVPR 2022) treats PnP as a probabilistic layer that outputs a distribution of poses, so that the 2D-3D correspondences and their weights can be learned end to end. The authors report that this narrowed the gap between PnP-based methods and task-specific leaders on the LineMOD 6DoF pose and nuScenes 3D object detection benchmarks.[29]
Work on minimal solvers also continues. Ding, Yang, Larsson, Olsson and Åström noted in 2023 that although P3P solvers are "a critical component of many vision systems (e.g. in localization and Structure-from-Motion)" and state-of-the-art solvers are fast and stable, "there still exist configurations where they break down"; their conic-intersection solver is the basis of OpenCV's SOLVEPNP_P3P method from version 4.13.0.[18][20]
See also
References
- ↑ 1.0 1.1 1.2 1.3 1.4 1.5 1.6 Vincent Lepetit, Francesc Moreno-Noguer, Pascal Fua (2008-07-19). "EPnP: An Accurate O(n) Solution to the PnP Problem". International Journal of Computer Vision, vol. 81, no. 2, pp. 155-166. doi:10.1007/s11263-008-0152-6. https://doi.org/10.1007/s11263-008-0152-6. Retrieved 2026-10-06.
- ↑ 2.00 2.01 2.02 2.03 2.04 2.05 2.06 2.07 2.08 2.09 2.10 2.11 2.12 Martin A. Fischler, Robert C. Bolles (1981-06). "Random Sample Consensus: A Paradigm for Model Fitting with Applications to Image Analysis and Automated Cartography". Communications of the ACM, vol. 24, no. 6, pp. 381-395. Association for Computing Machinery. doi:10.1145/358669.358692. https://doi.org/10.1145/358669.358692. Retrieved 2026-10-06.
- ↑ 3.0 3.1 Sergio Garrido, Alexander Panov. "Detection of ArUco Markers (OpenCV tutorial source)". OpenCV (GitHub). https://github.com/opencv/opencv/blob/4.x/doc/tutorials/objdetect/aruco_detection/aruco_detection.markdown. Retrieved 2026-10-06.
- ↑ 4.0 4.1 4.2 Jan Schmidt, Beyley Cardellio. "correspondence_search.c: Ab-initio blob/LED correspondence search". Monado source repository (freedesktop.org GitLab). https://gitlab.freedesktop.org/monado/monado/-/blob/main/src/xrt/tracking/constellation/correspondence_search.c. Retrieved 2026-10-06.
- ↑ 5.0 5.1 Raúl Mur-Artal, J. M. M. Montiel, Juan D. Tardós (2015-10). "ORB-SLAM: A Versatile and Accurate Monocular SLAM System". IEEE Transactions on Robotics, vol. 31, no. 5, pp. 1147-1163. doi:10.1109/TRO.2015.2463671. https://doi.org/10.1109/TRO.2015.2463671. Retrieved 2026-10-06.
- ↑ 6.0 6.1 6.2 6.3 6.4 6.5 6.6 6.7 "Perspective-n-Point (PnP) pose computation". OpenCV documentation source (GitHub, 4.x branch). OpenCV. https://github.com/opencv/opencv/blob/4.x/modules/calib3d/doc/solvePnP.markdown. Retrieved 2026-10-06.
- ↑ 7.0 7.1 7.2 Johannes L. Schönberger, Jan-Michael Frahm (2016-06). "Structure-from-Motion Revisited". 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 4104-4113. doi:10.1109/CVPR.2016.445. https://doi.org/10.1109/CVPR.2016.445. Retrieved 2026-10-06.
- ↑ 8.0 8.1 8.2 8.3 8.4 Mikael Persson, Klas Nordberg (2018). "Lambda Twist: An Accurate Fast Robust Perspective Three Point (P3P) Solver". Computer Vision - ECCV 2018, Lecture Notes in Computer Science, pp. 334-349. Springer. doi:10.1007/978-3-030-01225-0_20. https://doi.org/10.1007/978-3-030-01225-0_20. Retrieved 2026-10-06.
- ↑ 9.0 9.1 Gerald Schweighofer, Axel Pinz (2006). "Robust Pose Estimation from a Planar Target". IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 28, no. 12, pp. 2024-2030. doi:10.1109/TPAMI.2006.252. https://doi.org/10.1109/TPAMI.2006.252. Retrieved 2026-10-06.
- ↑ Bert M. Haralick, Chung-Nan Lee, Karsten Ottenberg, Michael Nölle (1994-12). "Review and analysis of solutions of the three point perspective pose estimation problem". International Journal of Computer Vision, vol. 13, no. 3, pp. 331-356. doi:10.1007/BF02028352. https://doi.org/10.1007/BF02028352. Retrieved 2026-10-06.
- ↑ 11.0 11.1 C.-P. Lu, G. D. Hager, E. Mjolsness (2000-06). "Fast and globally convergent pose estimation from video images". IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 22, no. 6, pp. 610-622. doi:10.1109/34.862199. https://doi.org/10.1109/34.862199. Retrieved 2026-10-06.
- ↑ 12.0 12.1 Xiao-Shan Gao, Xiao-Rong Hou, Jianliang Tang, Hang-Fei Cheng (2003-08). "Complete solution classification for the perspective-three-point problem". IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 25, no. 8, pp. 930-943. doi:10.1109/TPAMI.2003.1217599. https://doi.org/10.1109/TPAMI.2003.1217599. Retrieved 2026-10-06.
- ↑ 13.0 13.1 Laurent Kneip, Davide Scaramuzza, Roland Siegwart (2011-06). "A novel parametrization of the perspective-three-point problem for a direct computation of absolute camera position and orientation". 2011 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 2969-2976. doi:10.1109/CVPR.2011.5995464. https://doi.org/10.1109/CVPR.2011.5995464. Retrieved 2026-10-06.
- ↑ Joel A. Hesch, Stergios I. Roumeliotis (2011-11). "A Direct Least-Squares (DLS) method for PnP". 2011 International Conference on Computer Vision (ICCV), pp. 383-390. doi:10.1109/ICCV.2011.6126266. https://doi.org/10.1109/ICCV.2011.6126266. Retrieved 2026-10-06.
- ↑ 15.0 15.1 Toby Collins, Adrien Bartoli (2014-07-24). "Infinitesimal Plane-Based Pose Estimation". International Journal of Computer Vision, vol. 109, no. 3, pp. 252-286. doi:10.1007/s11263-014-0725-5. https://doi.org/10.1007/s11263-014-0725-5. Retrieved 2026-10-06.
- ↑ Tong Ke, Stergios I. Roumeliotis (2017-07). "An Efficient Algebraic Solution to the Perspective-Three-Point Problem". 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 4618-4626. doi:10.1109/CVPR.2017.491. https://doi.org/10.1109/CVPR.2017.491. Retrieved 2026-10-06.
- ↑ George Terzakis, Manolis Lourakis (2020). "A Consistently Fast and Globally Optimal Solution to the Perspective-n-Point Problem". Computer Vision - ECCV 2020, Lecture Notes in Computer Science, pp. 478-494. Springer. doi:10.1007/978-3-030-58452-8_28. https://doi.org/10.1007/978-3-030-58452-8_28. Retrieved 2026-10-06.
- ↑ 18.0 18.1 Yaqing Ding, Jian Yang, Viktor Larsson, Carl Olsson, Kalle Åström (2023-06). "Revisiting the P3P Problem". 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 4872-4880. doi:10.1109/CVPR52729.2023.00472. https://doi.org/10.1109/CVPR52729.2023.00472. Retrieved 2026-10-06.
- ↑ "Perspective-n-Point (PnP) pose computation (OpenCV 4.12.0)". OpenCV (GitHub). https://github.com/opencv/opencv/blob/4.12.0/modules/calib3d/doc/solvePnP.markdown. Retrieved 2026-10-06.
- ↑ 20.0 20.1 "Perspective-n-Point (PnP) pose computation (OpenCV 4.13.0)". OpenCV (GitHub). https://github.com/opencv/opencv/blob/4.13.0/modules/calib3d/doc/solvePnP.markdown. Retrieved 2026-10-06.
- ↑ "apriltag_pose.h". AprilTag (AprilRobotics, GitHub). https://github.com/AprilRobotics/apriltag/blob/master/apriltag_pose.h. Retrieved 2026-10-06.
- ↑ "ARToolKit Documentation (History)". ARToolKit, Human Interface Technology Laboratory, University of Washington. http://www.hitl.washington.edu/artoolkit/documentation/history.htm. Retrieved 2026-10-06.
- ↑ Hirokazu Kato, Mark Billinghurst (1999). "Marker tracking and HMD calibration for a video-based augmented reality conferencing system". Proceedings of the 2nd IEEE and ACM International Workshop on Augmented Reality (IWAR '99), pp. 85-94. doi:10.1109/IWAR.1999.803809. https://doi.org/10.1109/IWAR.1999.803809. Retrieved 2026-10-06.
- ↑ "rift_driver.c". Monado source repository (freedesktop.org GitLab). https://gitlab.freedesktop.org/monado/monado/-/blob/main/src/xrt/drivers/rift/rift_driver.c. Retrieved 2026-10-06.
- ↑ Beyley Cardellio (2026-04-08). "t/constellation: Import correspondence search from OpenHMD". Monado source repository (freedesktop.org GitLab). https://gitlab.freedesktop.org/monado/monado/-/commit/7f47872f47d5269cd8aec96141905ce14d8376b2. Retrieved 2026-10-06.
- ↑ Beyley Cardellio (2026-04-08). "t/constellation: Import lambdatwist p3p solver from OpenHMD". Monado source repository (freedesktop.org GitLab). https://gitlab.freedesktop.org/monado/monado/-/commit/5297bb845cad00db19ef6361c8d84301463491f9. Retrieved 2026-10-06.
- ↑ Simon Lynen, Bernhard Zeisl, Dror Aiger, Michael Bosse, Joel Hesch, Marc Pollefeys, Roland Siegwart, Torsten Sattler (2020-07-07). "Large-scale, real-time visual-inertial localization revisited". The International Journal of Robotics Research, vol. 39, no. 9, pp. 1061-1084. doi:10.1177/0278364920931151. https://doi.org/10.1177/0278364920931151. Retrieved 2026-10-06.
- ↑ Sida Peng, Yuan Liu, Qixing Huang, Xiaowei Zhou, Hujun Bao (2019-06). "PVNet: Pixel-Wise Voting Network for 6DoF Pose Estimation". 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 4556-4565. doi:10.1109/CVPR.2019.00469. https://doi.org/10.1109/CVPR.2019.00469. Retrieved 2026-10-06.
- ↑ Hansheng Chen, Pichao Wang, Fan Wang, Wei Tian, Lu Xiong, Hao Li (2022-06). "EPro-PnP: Generalized End-to-End Probabilistic Perspective-n-Points for Monocular Object Pose Estimation". 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 2771-2780. doi:10.1109/CVPR52688.2022.00280. https://doi.org/10.1109/CVPR52688.2022.00280. Retrieved 2026-10-06.