Structure from motion
More actions
Structure from motion (SfM) is the computer vision problem of recovering the three-dimensional structure of a stationary scene, together with the positions and orientations of the cameras that photographed it, from a set of overlapping two-dimensional images taken from different viewpoints. The term also covers the family of algorithms that solve it. In their 2017 survey in Acta Numerica, Onur Özyeşil, Vladislav Voroninski, Ronen Basri and Amit Singer define SfM as "the problem of recovering the three-dimensional (3D) structure of a stationary scene from a set of projective measurements, represented as a collection of two-dimensional (2D) images, via estimation of motion of the cameras corresponding to these images."[1] A typical pipeline matches local image features across photos, estimates the camera poses, triangulates the matched features into a sparse point cloud, and refines the whole solution with bundle adjustment.[2]
In virtual reality and augmented reality, SfM is used for both content creation and device localization. It underlies SfM-based photogrammetry, it supplies the camera poses and starting points for neural radiance fields and Gaussian splatting, and SfM maps are a classic basis for visual localization of AR devices.[3][4][5] The survey above treats simultaneous localization and mapping (SLAM), the problem of building a map while tracking a moving camera within it, as "effectively a special case of the SfM problem", specific to robotics and AR/VR applications.[1]
How it works
The input is a collection of photographs or video frames of a static scene with enough overlap that the same surface points appear in several images. The output is an estimated pose (position and orientation) for each registered camera, often its intrinsic calibration, and a set of 3D scene points.[2][1] Schönberger and Frahm divide the widely used incremental pipeline into a correspondence-search stage and an incremental reconstruction stage.[2]
Correspondence search
- Feature extraction. Local features are detected in every image and described so that the same point can be recognized under changes of viewpoint and lighting. Schönberger and Frahm describe SIFT, its derivatives and learned features as "the gold standard in terms of robustness", with binary features as a faster but less robust alternative.[2] SIFT is described in David Lowe's 2004 paper in the International Journal of Computer Vision.[6]
- Matching. Images that see the same part of the scene are found by comparing feature descriptors. Testing every image pair is prohibitively expensive for large collections, so many systems use scalable matching strategies.[2]
- Geometric verification. Because appearance matches can be wrong, each candidate pair is checked by fitting a geometric model with a robust estimator such as RANSAC: a homography for a purely rotating camera or a planar scene, or epipolar geometry (the essential matrix for calibrated cameras, the fundamental matrix for uncalibrated ones). The verified pairs form a "scene graph" with images as nodes.[2]
Reconstruction
- Initialization. The model is seeded with a carefully chosen two-view reconstruction. Schönberger and Frahm note that "the reconstruction may never recover from a bad initialization".[2]
- Image registration. New images are added by solving the Perspective-n-Point problem from matches between their features and already triangulated points.[2]
- Triangulation. Points seen from at least two registered viewpoints are triangulated and added to the model, which in turn provides more 2D-3D matches for registering further images.[2]
- Bundle adjustment. Bundle adjustment is the joint non-linear refinement of all camera parameters and point positions that minimizes the reprojection error, the distance between where each 3D point projects into an image and where its feature was observed. Without it, Schönberger and Frahm write, "SfM usually drifts quickly to a non-recoverable state."[2] The cost function is not convex, so it needs good initial estimates of the cameras and structure to avoid converging to a poor local minimum.[1]
Incremental SfM repeats registration, triangulation, outlier filtering and bundle adjustment as images are added one at a time or in small groups.[2] Global SfM instead estimates the rotations and positions of all cameras at once before a final refinement. The survey notes that the incremental approach relies on greedy steps that may not reach an optimal solution and can accumulate errors on large, weakly connected image sets.[1] In 2024 Linfei Pan, Dániel Baráth, Marc Pollefeys and Johannes Schönberger wrote that the most popular systems had followed the incremental paradigm "due to its superior accuracy and robustness, while global approaches are drastically more scalable and efficient"; their GLOMAP system reported accuracy on par with or better than COLMAP while running "orders of magnitude faster".[7]
Image measurements alone do not fix the absolute frame of the result. In the Photo Tourism paper, Noah Snavely, Steven Seitz and Richard Szeliski note that the estimated camera locations are related to true locations by a similarity transform (a global translation, rotation and uniform scale); where absolute positions were needed they aligned models to overhead maps or images, to GPS-tagged control points, or to digital elevation maps.[8] The 3D structure SfM produces is sparse: one point per matched feature track. A dense surface is usually computed afterwards with multi-view stereo (MVS), and tools such as COLMAP combine both stages.[1][9]
History
Shimon Ullman examined "the interpretation of structure from motion" from a computational point of view, asking how the 3D structure and motion of objects can be inferred from the changing 2D projections of moving objects. His analysis rests on the "structure from motion theorem", which states that the structure of four non-coplanar points is recoverable from three orthographic projections. The work circulated as MIT Artificial Intelligence Laboratory memo AIM-476 in October 1976 and was published in the Proceedings of the Royal Society of London, Series B in January 1979; the paper also discusses the scheme's psychological relevance.[10][11]
The Özyeşil survey describes H. C. Longuet-Higgins's 1981 Nature paper "A computer algorithm for reconstructing a scene from two projections" as seminal work that introduced the first linear method for a pair of cameras based on point correspondences, later named the eight-point algorithm.[1][12] In 1992 Carlo Tomasi and Takeo Kanade introduced a factorization method that recovers shape and motion from image streams under an orthographic camera model, intended for objects that are distant relative to their size.[1][13]
Photo Tourism was presented at SIGGRAPH 2006 by Snavely and Seitz of the University of Washington and Szeliski of Microsoft Research, who called it "the first successful demonstration of SfM techniques being applied to the kinds of real-world image sets found on Google and Flickr". It detected SIFT keypoints, matched them between image pairs, estimated fundamental matrices with RANSAC and the eight-point algorithm, and then reconstructed cameras and points incrementally with sparse bundle adjustment, starting from a well-separated initial pair. Run times ranged from a few hours for a 120-photo set of the Great Wall of China (82 registered) to about two weeks for 2,635 Flickr photos of Notre Dame Cathedral in Paris (597 registered).[8] The survey credits the paper with demonstrating accurate reconstruction from "hundreds, or even thousands of independently captured photographs", which sparked wide interest in SfM for large, unordered image sets.[1] Its SfM software was released as the open-source package Bundler.[2][1]
At ICCV 2009 the Building Rome in a Day project (Sameer Agarwal, Snavely, Ian Simon, Seitz and Szeliski) scaled the approach to entire cities of Internet photos using parallel computation. Its Rome reconstruction processed 150,000 images on 496 compute cores, with 13 hours of matching and 8 hours of reconstruction; the project also reconstructed Venice (250,000 images) and Dubrovnik (57,845 images).[14] A 2011 Communications of the ACM version of the work added Yasutaka Furukawa and Brian Curless as authors.[15] Schönberger and Frahm's 2016 CVPR paper "Structure-from-Motion Revisited" addressed robustness, accuracy, completeness and scalability in incremental SfM and released its pipeline as the open-source program COLMAP.[2]
Milestones
| Year | Work | Contribution |
|---|---|---|
| 1979 | S. Ullman, "The interpretation of structure from motion" (memo 1976) | Computational analysis of structure from motion in vision; four non-coplanar points recoverable from three orthographic views[10] |
| 1981 | H. C. Longuet-Higgins, Nature | Linear two-view reconstruction from point correspondences, later called the eight-point algorithm[1] |
| 1992 | C. Tomasi and T. Kanade, IJCV | Factorization method for shape and motion under orthography[13] |
| 2004 | D. Nistér, IEEE TPAMI | Efficient solution to the five-point relative pose problem[16] |
| 2006 | Snavely, Seitz and Szeliski, Photo Tourism (SIGGRAPH) | Incremental SfM on unordered Internet photo collections[8] |
| 2009 | Agarwal et al., Building Rome in a Day (ICCV) | City-scale reconstruction from Internet photos on hundreds of cores[14] |
| 2016 | Schönberger and Frahm, COLMAP (CVPR) | Improved incremental SfM released as open source[2] |
| 2024 | Pan et al., GLOMAP (ECCV) | Global SfM matching incremental accuracy at much higher speed[7] |
| 2024 | Brachmann et al., ACE0 (ECCV) | SfM recast as repeated neural relocalization[17] |
| 2025 | Wang et al., VGGT (CVPR) | Feed-forward network predicting cameras, depth and point maps; CVPR 2025 Best Paper[18] |
| 2026 | COLMAP 4.0 | GLOMAP global pipeline integrated into COLMAP[19] |
Applications in VR and AR
Photogrammetry and 3D capture
In a 2012 paper in Geomorphology, M. J. Westoby and colleagues contrasted SfM with traditional softcopy photogrammetry, which requires the 3D location and pose of the cameras, or the 3D locations of ground control points, to be known: "the SfM method solves the camera pose and scene geometry simultaneously and automatically, using a highly redundant bundle adjustment based on matching features in multiple overlapping, offset images." Using a consumer-grade digital camera, they reported decimetre-scale vertical accuracy against a terrestrial laser scan, and described SfM as "a major advancement in the field of photogrammetry for geoscience applications".[20]
Radiance fields and Gaussian splatting
Novel-view synthesis methods that let a viewer move freely through a captured scene depend on SfM for their camera parameters. The original NeRF paper (Mildenhall et al., ECCV 2020) states that for real data it uses "the COLMAP structure-from-motion package" to estimate camera poses, intrinsics and scene bounds.[3] 3D Gaussian splatting (Kerbl et al., 2023) starts from the same input, cameras calibrated by SfM, and initializes its 3D Gaussians from "the sparse point cloud produced for free as part of the SfM process"; the method targets real-time rendering at 30 frames per second or more at 1080p.[4] The open-source Nerfstudio framework likewise uses COLMAP in its data-processing script to obtain the camera pose of every image, and warns that "COLMAP can be finicky", advising users to capture overlapping, non-blurry images. Its documentation lists alternatives that skip COLMAP, such as poses from the Polycam app or from an iPhone's LiDAR in Record3D.[21]
Localization and tracking for AR
SfM maps are a long-standing way to localize an AR device in a known place. LaMAR, a 2022 benchmark from ETH Zurich and the Microsoft Mixed Reality & AI Lab, describes image-based localization as "classically tackled" by estimating a camera pose from correspondences between local features and a 3D SfM map of the scene. The benchmark itself contains more than 100 hours of sensor recordings from head-mounted HoloLens 2 and hand-held iPhone and iPad devices, covering 45,000 square meters over one year, and its authors call localization and mapping "the foundational technology for augmented reality (AR) that enables sharing and persistence of digital content in the real world".[5] At ISMAR 2012, Jonathan Ventura and Tobias Höllerer of UC Santa Barbara built point-cloud models of wide outdoor areas offline from panoramas, using an incremental SfM pipeline, and then tracked a handheld device against the model in real time; with a reconstruction made from less than five minutes of video they reported translational error below 25 cm and rotational error of 0.5 degrees for over 80% of test images.[22] Google's visual positioning system (VPS), offered through the ARCore Geospatial API, also localizes devices against a point-cloud model built from images. Google describes it as combining features from tens of billions of Street View images "to compute a 3D point cloud of the global environment", then matching recognizable parts of the user's surroundings to that model and computing the device's position and orientation.[23]
Techniques from SfM also moved into real-time tracking. Georg Klein and David Murray's Parallel Tracking and Mapping (PTAM) system, presented at ISMAR 2007 for small AR workspaces, split camera tracking and map building into separate threads so that the mapping thread could run bundle adjustment, which they described as having "long been a proven method for offline Structure-from-Motion (SfM)".[24]
Image-based virtual tourism
Photo Tourism was built as an interactive 3D browser for photo collections of landmarks, with smooth image-based transitions between photos, and its introduction cites the hope that image-based rendering will one day allow "virtual tourism of the world's interesting and important sites". The authors described their system as a way to automatically create a version of the Aspen Movie Map from unorganized photos, and compared its transfer of annotations between photos to augmented reality work such as that of Steven Feiner and colleagues.[8] Microsoft Live Labs released the technology publicly as Photosynth on 20 August 2008, describing it as a way to turn ordinary digital photos into "a three-dimensional, 360-degree experience".[25] A 2014 Microsoft Research blog post quotes Eric Stollnitz tracing Photosynth to Photo Tourism, "a collaboration between Noah Snavely and Steve Seitz at the University of Washington and my colleague Richard Szeliski."[26]
Relationship to SLAM
SfM and SLAM estimate the same quantities, camera poses and scene structure, from images. The Özyeşil survey calls SLAM "effectively a special case of the SfM problem", specific to robotics and augmented and virtual reality, and notes that a SLAM pipeline without a prebuilt map begins like SfM by establishing feature correspondences and estimating camera orientations.[1] The practical difference is timing. SfM is typically run offline over a whole, often unordered, image collection, while SLAM and visual-inertial odometry must update the pose in real time as the device moves.[1][24] PTAM showed one way to bring the two together, by running batch bundle adjustment on selected keyframes in a background thread while tracking continued at frame rate.[24]
Software
| Software | Origin | Approach | Notes |
|---|---|---|---|
| Bundler | Noah Snavely, from Photo Tourism (2006) | Incremental | Open-source SfM from Photo Tourism, built on a modified sparse bundle adjustment (SBA) package[1][2] |
| VisualSfM | Changchang Wu (2011) | Incremental | Closed-source; multicore CPU and GPU bundle adjustment, GPU SIFT and a graphical interface[1][2] |
| COLMAP | Schönberger and Frahm (2016) | Incremental, hierarchical and global | Open source; "general-purpose Structure-from-Motion (SfM) and Multi-View Stereo (MVS) pipeline with a graphical and command-line interface"[9][19] |
| GLOMAP | Pan et al. (2024) | Global | Open source; the separate repository is marked deprecated after migration into COLMAP 4.0[27][19] |
Both the original NeRF paper and the Nerfstudio framework use COLMAP for camera poses.[3][21] COLMAP version 4.0.0, released on 15 March 2026, integrated GLOMAP as a "first-class alternative" to the incremental and hierarchical mappers, maintained in the COLMAP repository from then on.[19] Version 4.1.0 (26 June 2026) added Caspar, a GPU-accelerated bundle adjustment backend that the release notes describe as often one to two orders of magnitude faster than the Ceres CUDA backend for medium and large problems, and spherical (equirectangular) camera models for native reconstruction of 360-degree panoramic images.[28] As of 4 October 2026 the latest release is 4.2.1, published on 29 September 2026.[29]
Research
Several learning-based methods aim to replace parts of the pipeline or all of it. ACE0 (ECCV 2024), from researchers at Niantic and the University of Oxford, reinterprets incremental SfM as "an iterated application and refinement of a visual relocalizer" and builds an implicit neural scene representation from unposed images; the authors report camera poses that in many cases come close to feature-based SfM in accuracy.[17] DUSt3R, from Naver Labs Europe and Aalto University (CVPR 2024), regresses 3D "pointmaps" from image pairs without prior camera calibration or poses and recovers cameras from them.[30] VGGT, from the University of Oxford's Visual Geometry Group and Meta AI, is a feed-forward network that infers camera parameters, depth maps, point maps and 3D point tracks from one to hundreds of views, reconstructing in under one second according to its authors.[18] It received the Best Paper Award at CVPR 2025, chosen from more than 13,000 submissions; Oxford's announcement lists augmented reality among its potential uses.[31] Work on classical optimization has continued alongside these methods, for example GLOMAP and the GPU bundle adjustment backend added to COLMAP in 2026.[7][28]
See also
References
- ↑ 1.00 1.01 1.02 1.03 1.04 1.05 1.06 1.07 1.08 1.09 1.10 1.11 1.12 1.13 1.14 Onur Özyeşil, Vladislav Voroninski, Ronen Basri, Amit Singer (2017). "A survey of structure from motion". Acta Numerica, vol. 26, pp. 305-364. Cambridge University Press. doi:10.1017/S096249291700006X; preprint arXiv:1701.08493. https://doi.org/10.1017/S096249291700006X. Retrieved 2026-10-04.
- ↑ 2.00 2.01 2.02 2.03 2.04 2.05 2.06 2.07 2.08 2.09 2.10 2.11 2.12 2.13 2.14 2.15 Johannes L. Schönberger, Jan-Michael Frahm (2016). "Structure-from-Motion Revisited". 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 4104-4113. doi:10.1109/CVPR.2016.445. https://openaccess.thecvf.com/content_cvpr_2016/papers/Schonberger_Structure-From-Motion_Revisited_CVPR_2016_paper.pdf. Retrieved 2026-10-04.
- ↑ 3.0 3.1 3.2 Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, Ren Ng (2020). "NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis". European Conference on Computer Vision (ECCV) 2020. arXiv:2003.08934. https://arxiv.org/abs/2003.08934. Retrieved 2026-10-04.
- ↑ 4.0 4.1 Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George Drettakis (2023-07). "3D Gaussian Splatting for Real-Time Radiance Field Rendering". ACM Transactions on Graphics, vol. 42, no. 4. arXiv:2308.04079. https://arxiv.org/abs/2308.04079. Retrieved 2026-10-04.
- ↑ 5.0 5.1 Paul-Edouard Sarlin, Mihai Dusmanu, Johannes L. Schönberger, Pablo Speciale, Lukas Gruber, Viktor Larsson, Ondrej Miksik, Marc Pollefeys (2022). "LaMAR: Benchmarking Localization and Mapping for Augmented Reality". European Conference on Computer Vision (ECCV) 2022. arXiv:2210.10770. https://arxiv.org/abs/2210.10770. Retrieved 2026-10-04.
- ↑ David G. Lowe (2004). "Distinctive Image Features from Scale-Invariant Keypoints". International Journal of Computer Vision, vol. 60, no. 2, pp. 91-110. doi:10.1023/B:VISI.0000029664.99615.94. https://doi.org/10.1023/B:VISI.0000029664.99615.94. Retrieved 2026-10-04.
- ↑ 7.0 7.1 7.2 Linfei Pan, Dániel Baráth, Marc Pollefeys, Johannes L. Schönberger (2024). "Global Structure-from-Motion Revisited". European Conference on Computer Vision (ECCV) 2024. arXiv:2407.20219. https://arxiv.org/abs/2407.20219. Retrieved 2026-10-04.
- ↑ 8.0 8.1 8.2 8.3 Noah Snavely, Steven M. Seitz, Richard Szeliski (2006-07). "Photo Tourism: Exploring Photo Collections in 3D". ACM Transactions on Graphics, vol. 25, no. 3 (SIGGRAPH 2006), pp. 835-846. doi:10.1145/1141911.1141964. https://phototour.cs.washington.edu/Photo_Tourism.pdf. Retrieved 2026-10-04.
- ↑ 9.0 9.1 "COLMAP documentation". COLMAP. https://colmap.github.io/. Retrieved 2026-10-04.
- ↑ 10.0 10.1 S. Ullman (1979-01-15). "The interpretation of structure from motion". Proceedings of the Royal Society of London. Series B. Biological Sciences, vol. 203, no. 1153, pp. 405-426. doi:10.1098/rspb.1979.0006. https://doi.org/10.1098/rspb.1979.0006. Retrieved 2026-10-04.
- ↑ S. Ullman (1976-10-01). "The Interpretation of Structure From Motion (AIM-476)". MIT Artificial Intelligence Laboratory Memos, DSpace@MIT. Massachusetts Institute of Technology. https://dspace.mit.edu/handle/1721.1/6298. Retrieved 2026-10-04.
- ↑ H. C. Longuet-Higgins (1981-09). "A computer algorithm for reconstructing a scene from two projections". Nature, vol. 293, no. 5828, pp. 133-135. doi:10.1038/293133a0. https://doi.org/10.1038/293133a0. Retrieved 2026-10-04.
- ↑ 13.0 13.1 Carlo Tomasi, Takeo Kanade (1992-11). "Shape and motion from image streams under orthography: a factorization method". International Journal of Computer Vision, vol. 9, no. 2, pp. 137-154. doi:10.1007/BF00129684. https://doi.org/10.1007/BF00129684. Retrieved 2026-10-04.
- ↑ 14.0 14.1 "Building Rome in a Day". GRAIL, University of Washington. https://grail.cs.washington.edu/rome/. Retrieved 2026-10-04.
- ↑ Sameer Agarwal, Yasutaka Furukawa, Noah Snavely, Ian Simon, Brian Curless, Steven M. Seitz, Richard Szeliski (2011-10). "Building Rome in a day". Communications of the ACM, vol. 54, no. 10, pp. 105-112. doi:10.1145/2001269.2001293. https://doi.org/10.1145/2001269.2001293. Retrieved 2026-10-04.
- ↑ D. Nistér (2004-06). "An efficient solution to the five-point relative pose problem". IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 26, no. 6, pp. 756-770. doi:10.1109/TPAMI.2004.17. https://doi.org/10.1109/TPAMI.2004.17. Retrieved 2026-10-04.
- ↑ 17.0 17.1 Eric Brachmann, Jamie Wynn, Shuai Chen, Tommaso Cavallari, Áron Monszpart, Daniyar Turmukhambetov, Victor Adrian Prisacariu (2024). "Scene Coordinate Reconstruction: Posing of Image Collections via Incremental Learning of a Relocalizer". European Conference on Computer Vision (ECCV) 2024. arXiv:2404.14351. https://arxiv.org/abs/2404.14351. Retrieved 2026-10-04.
- ↑ 18.0 18.1 Jianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi, Christian Rupprecht, David Novotny (2025). "VGGT: Visual Geometry Grounded Transformer". IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2025. arXiv:2503.11651. https://arxiv.org/abs/2503.11651. Retrieved 2026-10-04.
- ↑ 19.0 19.1 19.2 19.3 "COLMAP 4.0.0 release notes". GitHub. COLMAP. 2026-03-15. https://github.com/colmap/colmap/releases/tag/4.0.0. Retrieved 2026-10-04.
- ↑ M. J. Westoby, J. Brasington, N. F. Glasser, M. J. Hambrey, J. M. Reynolds (2012-12). "'Structure-from-Motion' photogrammetry: A low-cost, effective tool for geoscience applications". Geomorphology, vol. 179, pp. 300-314. Aberystwyth University Research Portal. doi:10.1016/j.geomorph.2012.08.021. https://research.aber.ac.uk/en/publications/structure-from-motion-photogrammetry-a-low-cost-effective-tool-fo/. Retrieved 2026-10-04.
- ↑ 21.0 21.1 "Using custom data". Nerfstudio documentation. https://docs.nerf.studio/quickstart/custom_dataset.html. Retrieved 2026-10-04.
- ↑ Jonathan Ventura, Tobias Höllerer (2012-11). "Wide-Area Scene Mapping for Mobile Visual Tracking". 2012 IEEE International Symposium on Mixed and Augmented Reality (ISMAR). doi:10.1109/ISMAR.2012.6402531. https://sites.cs.ucsb.edu/~holl/pubs/Ventura-2012-ISMAR.pdf. Retrieved 2026-10-04.
- ↑ "Geospatial API overview". ARCore - Google for Developers. Google. https://developers.google.com/ar/develop/geospatial. Retrieved 2026-10-04.
- ↑ 24.0 24.1 24.2 Georg Klein, David Murray (2007-11). "Parallel Tracking and Mapping for Small AR Workspaces". 6th IEEE and ACM International Symposium on Mixed and Augmented Reality (ISMAR 2007). pp. 1-10. doi:10.1109/ISMAR.2007.4538852. https://www.robots.ox.ac.uk/~gk/publications/KleinMurray2007ISMAR.pdf. Retrieved 2026-10-04.
- ↑ "Microsoft Live Labs Introduces Photosynth, a Breakthrough Visual Medium". Microsoft Source. Microsoft. 2008-08-20. https://news.microsoft.com/source/2008/08/20/microsoft-live-labs-introduces-photosynth-a-breakthrough-visual-medium/. Retrieved 2026-10-04.
- ↑ "A New Spin for Photosynth". Microsoft Research Blog. Microsoft. 2014-01-07. https://www.microsoft.com/en-us/research/blog/new-spin-photosynth/. Retrieved 2026-10-04.
- ↑ "GLOMAP: Global Structure-from-Motion Revisited". GitHub. COLMAP. https://github.com/colmap/glomap. Retrieved 2026-10-04.
- ↑ 28.0 28.1 "COLMAP 4.1.0 release notes". GitHub. COLMAP. 2026-06-26. https://github.com/colmap/colmap/releases/tag/4.1.0. Retrieved 2026-10-04.
- ↑ "COLMAP 4.2.1 release notes". GitHub. COLMAP. 2026-09-29. https://github.com/colmap/colmap/releases/tag/4.2.1. Retrieved 2026-10-04.
- ↑ Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, Jerome Revaud (2024). "DUSt3R: Geometric 3D Vision Made Easy". IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2024. arXiv:2312.14132. https://arxiv.org/abs/2312.14132. Retrieved 2026-10-04.
- ↑ "Oxford researchers awarded Best Paper at CVPR 2025". Department of Engineering Science, University of Oxford. 2025-06-30. https://eng.ox.ac.uk/news/oxford-researchers-awarded-best-paper-at-cvpr-2025. Retrieved 2026-10-04.