Ray-casting selection
More actions
Ray-casting selection (also written raycasting selection) is a pointing technique for choosing objects in a three-dimensional virtual environment: a virtual ray is cast from a tracked hand, controller or other origin into the scene, and the object the ray intersects becomes the target, which the user then confirms with a button press, gesture or other trigger. It is used above all to select objects beyond arm's reach in virtual reality (VR) and augmented reality (AR) interfaces. In their 2013 survey of 3D selection techniques, Ferran Argelaguet and Carlos Andujar call raycasting selection "one of the most popular techniques for 3D object selection tasks", and a 2019 CHI paper by Baloup, Pietrzak and Casiez describes it as "the most common target pointing technique in virtual reality environments".[1][2]
The technique was described for immersive virtual environments in the mid-1990s, notably in Mark Mine's 1995 University of North Carolina technical report on virtual environment interaction techniques, which proposed "laser beams or spotlights which project out from the user's hand and intersect with the objects in the virtual world" for selecting distant objects.[3] Today it is built into the system software of consumer headsets and into the main cross-platform standards: the OpenXR specification defines an "aim" pose for pointing with a hand or controller, and the WebXR Device API gives every input source a "target ray".[4][5]
How it works
Selection is one of the basic tasks of a 3D user interface. Mine's 1995 report names movement, selection, manipulation and scaling as the fundamental forms of interaction in a virtual world,[3] and Doug Bowman argued that most virtual environment interactions fall into three categories: viewpoint motion control, selection and manipulation.[6] A selection technique has to let the user indicate an object, confirm the selection, and receive feedback while doing so.[1] In ray-casting, the indication step uses a ray (a half-line) as the selection tool. In the classic form, the ray starts at the position of the tracked hand or hand-held device and points along its orientation, much like a laser pointer; the system intersects the ray with the scene, either at the moment of confirmation or every frame if the indicated object is to be highlighted continuously, and an object hit by the ray becomes the candidate target. When the ray passes through several objects, the technique needs a rule or extra input to decide between them.[1] Mine's 1995 report states the general requirement for any selection method: some mechanism to identify the object, and "some signal or command to indicate the actual act of selection", typically "a button press, a gesture, or some kind of voice command".[3] In Bowman's description of the technique, "a light ray emanates from the user's virtual hand. To select an object, the user intersects the object with the light ray and performs a 'grab' action (usually by pressing a button)."[6]
A ray controlled by hand position and orientation has five degrees of freedom: three for its origin and two for its direction. Argelaguet and Andujar point out that, in the absence of a visibility mismatch, any object can be indicated by changing only the two orientation values (yaw and pitch) with the origin held fixed; in practice ray-casting is governed mainly by the yaw and pitch of the hand.[1] The same property defines its reach. With a virtual hand, only objects inside the user's physical working space can be touched; with a virtual pointer, the space the user can act on matches the visual space, so anything visible can in principle be selected.[1]
Feedback is part of the technique. Mine wrote that the user "must know when he has chosen an object for selection (perhaps by highlighting the object or its bounding box)" and when a selection has succeeded, through both visual and audible cues.[3] Current systems draw the ray as a visible line and attach a cursor where it meets a surface. Microsoft's design guidance for HoloLens 2, for example, places a donut-shaped cursor at the end of the hand ray; while pointing, the ray is drawn as a dashed line, and in the commit state it becomes a solid line and the cursor shrinks to a dot.[7]
History
Pointing at a distant display with a tracked hand was demonstrated more than a decade before Mine's report. Richard Bolt's "Put-That-There" system, presented at SIGGRAPH 1980, combined speech recognition with a Polhemus space-sensing cube worn on a watchband on the user's wrist; seated in front of the large screen of the Media Room at MIT's Architecture Machine Group, the user pointed at locations while a small white "x" cursor gave running visual feedback, and spoken words such as "there" fixed the moment the pointing position was read.[8]
In the JDCAD 3D modeling system, Jiandong Liang and Mark Green used a cone instead of a thin ray. Their "spotlight" has a cross-section proportional to the distance from the user, which reduces accuracy errors when selecting at a distance; because several objects can fall inside the cone, the object closest to the cone's axis is chosen.[3][1][9] The paper appeared in Computers & Graphics in July 1994.[10]
Mine's technical report, dated 5 May 1995, divided selection into local techniques (moving a hand-attached cursor into an object's bounding box) and at-a-distance techniques using laser beams or spotlights. It also described gaze-directed selection, which in the absence of reliable eye tracking "can be approximated using the current orientation of the user's head", and recommended hysteresis to counter noise in the tracking system.[3] In 1996 Andrew Forsberg, Kenneth Herndon and Robert Zeleznik presented aperture selection at UIST, a variant of the spotlight in which the user adjusts the apex angle of the selection cone.[1][11]
Doug Bowman and Larry Hodges compared ray-casting with arm-extension techniques such as Go-Go at the 1997 Symposium on Interactive 3D Graphics.[12] On the project page for that work, Bowman summarized that ray-casting techniques "made it easy to grab virtual objects, but manipulation was difficult", while arm-extension made manipulation natural but made it hard to get the hand to the object. Their answer was HOMER (Hand-Centered Object Manipulation Extending Ray-Casting): the user selects with a ray, the virtual hand immediately moves to the object, and the object is manipulated with ordinary hand motion.[13][6] The same year, Pierce and colleagues described image-plane techniques, in which the selection ray is cast from the user's viewpoint through the hand.[1][9]
In a 1998 study, Ivan Poupyrev and colleagues compared virtual hand and virtual pointer metaphors. As summarized by Steinicke, Ropinski and Hinrichs, Go-Go and simple ray-casting performed comparably for objects within reach regardless of object size, but as distance grew, and especially when higher accuracy was required, Go-Go had a significant advantage.[9] The classification by Poupyrev and Ichikawa, which divides egocentric selection techniques into virtual hand and virtual pointer metaphors, is reproduced in Argelaguet and Andujar's 2013 survey.[1]
Variants
Most variants change the shape of the selection tool, how it is controlled, or how the system chooses between several objects that the ray or cone touches.[1]
| Technique | Proposed by | Selection tool | Key idea |
|---|---|---|---|
| Spotlight (flashlight) | Liang and Green (1994) | Cone | The cone widens with distance; the object nearest the cone's axis is selected.[1][9] |
| Aperture selection | Forsberg, Herndon and Zeleznik (1996) | Adjustable cone | The selection tool is a cone whose apex angle the user adjusts manually.[1] |
| Image-plane or occlusion selection | Pierce et al. (1997) | Ray from the eye | The ray runs from the user's viewpoint through the hand.[1] |
| HOMER | Bowman and Hodges (1997) | Ray, then virtual hand | Ray-casting selects; the virtual hand moves to the object for manipulation.[13] |
| Flexible pointer | Olwal and Feiner (2003) | Curved ray | A two-handed Bezier curve that the user bends to reach partly occluded objects.[1][9] |
| Improved virtual pointer and sticky ray | Steinicke, Ropinski and Hinrichs (2004) | Bendable ray | A curved ray points to the selectable object closest to the pointing direction; in the sticky variant the last object hit stays active until the ray hits another.[9] |
| IntenSelect | de Haan, Koutek and Post (2005) | Ray with scoring | Objects are ranked continuously over time, which helps with moving targets.[1] |
| Depth ray, lock ray, flower ray, smart ray | Grossman and Balakrishnan (2006) | Ray with depth marker or menu | Different ways to pick one of several intersected targets: a movable depth marker, a two-step lock, a marking menu, or a history-weighted guess.[14] |
| Raycasting from the eye | Argelaguet, Andujar and Trueba (2008) | Ray from the eye | The origin is at the eye but wrist rotation steers the ray, so visible objects and selectable objects match.[1] |
| SQUAD | Kopper, Bacim and Bowman (2011) | Ray and sphere, then menu | Progressive refinement: the indicated objects are split across a four-part menu until one remains.[1] |
| RayCursor | Baloup, Pietrzak and Casiez (2019) | Filtered ray with cursor | Filtering the ray and adding a controllable cursor that selects the nearest target.[2] |
| Bubble Ray | Lu, Yu and Shi (2020) | Ray with bubble mechanism | Selects the target nearest the ray, so the ray does not have to pass exactly through it.[15] |
Disambiguation in dense scenes
A single thin ray can pass through several objects in a cluttered scene, while volumetric tools such as cones are "prone to indicate more than one object at once". Argelaguet and Andujar group the remedies into manual, heuristic and behavioral mechanisms. Manual methods let the user choose among the candidates (for example by cycling with a button, or with a second menu step); heuristic methods rank candidates by a rule such as distance from the cone's axis; behavioral methods, such as IntenSelect, accumulate evidence while the user is pointing.[1]
Tovi Grossman and Ravin Balakrishnan's 2006 UIST study was carried out for volumetric displays. In a first experiment they found a ray cursor superior to a 3D point cursor in a single-target environment. They then designed four ray techniques for dense environments: the depth ray, which adds a depth marker that the user moves along the ray by moving the hand forward and back, selecting the intersected target closest to the marker; the lock ray, which separates pointing and depth adjustment into two steps; the flower ray, which spreads intersected targets into a marking menu; and the predictive smart ray. The depth ray "performed particularly well, significantly reducing movement time, error rate, and input device footprint in comparison to the 3D point cursor".[14]
Curved and parabolic rays
Rays do not have to be straight. Besides the two-handed flexible pointer and the bendable ray described above, development frameworks offer curved rays as standard options. Unity's XR Interaction Toolkit can cast a straight line, a projectile curve that samples "the trajectory of a projectile", or a quadratic Bezier curve; its documentation also describes how to configure the ray interactor for teleportation.[16]
Performance and limitations
Pointing performance is commonly modeled with Fitts' law, which estimates the time needed to acquire a target from the amplitude of the movement and the size of the target.[1] For distal pointing, Regis Kopper, Doug Bowman, Mara Silva and Ryan McMahan found in a 2010 study that movement time "is best described as a function of the angular amplitude of movement and the angular size of the target"; contrary to Fitts' law, angular size had a much larger effect than angular amplitude, and task difficulty grew quadratically rather than linearly.[17] In a 2025 paper in IEEE Transactions on Visualization and Computer Graphics, Logan Lane, Feiyu Lu, Shakiba Davari, Robert J. Teather and Doug Bowman revisited the question with current VR technology and a new method for collecting distal pointing data; they reported that "the best model used a simple Fitts-Law-style index of difficulty with angular measures of amplitude and width".[18]
Because the ray is steered mostly by hand rotation, its precision is limited by the angular accuracy and steadiness of the hand, and the further away an object is, the more accuracy its selection requires. Selecting small or distant objects therefore remains difficult.[1][2] Other known problems include:
- Tracker noise and hand tremor. Noise from tracking devices and the lack of physical support for the hand both reduce accuracy on small targets.[1] Remedies include filtering the ray, which in the RayCursor studies reduced the error rate "in a drastic way",[2] and velocity-dependent control-display gain such as the PRISM technique and Adaptive Pointing, which scale down slow corrective movements.[1]
- The Heisenberg effect. Pressing a button to confirm a selection can itself change the device's orientation and cause a wrong selection. Triggering the selection on button release may be less sensitive to this effect, and gestures have been designed specifically to minimize hand movement during confirmation.[1]
- Eye-hand visibility mismatch. When the ray starts at the hand rather than the eye, an object can be visible but unreachable by the ray, or hidden from the eye but hit by it. Casting the ray from the eye and steering it with the wrist removes this conflict.[1]
- Occlusion and density. In cluttered scenes the ray may intersect many objects, which led to the disambiguation techniques described above.[1][14]
- Fatigue. Argelaguet and Andujar note that virtual pointing needs less arm effort than virtual hand techniques but more wrist effort, and that eye-aligned variants such as occlusion selection increase arm effort because the arm must be held up in line with the eye.[1]
Ray-casting is also used with bare hands. In a study published in IEEE Transactions on Visualization and Computer Graphics in 2023, Tiffany Luong, Yi Fei Cheng, Max Moebus, Andreas Fender and Christian Holz compared controllers and free-hand input for touch and raycast interaction. In the raycast setting, participants "reported less physical exertion, felt more in control, and were faster and more accurate when using VR controllers compared to free-hand interaction", and they preferred controllers for raycast but free hands for mid-air touch.[19]
Applications in VR and AR
Standards
Among the standard poses that the OpenXR 1.1 specification defines for tracked hands and motion controllers are a "grip" pose for rendering a virtual object held in the hand and an "aim" pose, "a pose that allows applications to point in the world using the input source, according to the platform's conventions". For controllers the aim ray "will often emerge from the frontmost tip of a motion controller"; for tracked hands it is runtime-dependent, "often a ray emerging from the hand at a target pointed by moving the forearm".[4]
The W3C WebXR Device API gives each input source a target ray space and a target ray mode. The "gaze" mode means the ray "will originate at the viewer and follow the direction it is facing", what is commonly called gaze input on head-mounted displays. The "tracked-pointer" mode is for handheld devices and hand tracking; when a platform has no ergonomics convention, the ray should point "in the same direction as the user's index finger if it was outstretched". The "screen" mode is for interaction with the canvas element of an inline session, such as a mouse click or touch event, and the "transient-pointer" mode covers inputs generated from an operating-system intent, such as gaze, that are too sensitive to expose directly.[5]
Headsets and toolkits
On Meta Quest headsets, users of hand tracking select system interface items by pointing and pinching: "When the cursor appears, point your hand at what you want to select. Then, pinch your thumb and finger together to select."[20] Meta's design guidelines recommend the standard pinch for hand-based selection and the trigger for controllers, a shoulder or hip point as a secondary origin "to stabilize the ray and minimize the impact of natural hand tremors", and techniques such as "filtering rotational noise, casting conical frustums, and snapping to interactables"; they advise against using a head ray as the primary pointer.[21] In Meta's Interaction SDK, a RayInteractor "defines the origin and direction of raycasts for a ray interaction, as well as a max distance for the interaction" and is paired with a selector, which can be "a button, a gesture, or a voice command".[22]
Microsoft calls its version "point and commit with hands". On HoloLens 2 the hand ray "shoots out from the center of the user's palm", and the user commits with an air tap of thumb and index finger. Rays switch off automatically when an object is within arm's length (roughly 50 cm), encouraging near interaction, and switch on for objects farther away. Microsoft states that point and commit for far interaction was created for the Mixed Reality Portal of Windows Mixed Reality, where users of immersive headsets point with rays from motion controllers, and that the same ray model was then attached to the hands so that users of both need only one interaction model; the Mixed Reality Toolkit supplies a hand ray prefab matching the system's.[7] Unity's XR Interaction Toolkit provides a comparable component, the XR Ray Interactor, "used for interacting with Interactables at a distance".[16]
Gaze plus pinch
Apple's visionOS for Apple Vision Pro targets its "indirect" gestures with the eyes, through eye tracking, rather than with a pointing hand: "a person can look at a button to focus it and select it by quickly tapping their finger and thumb together." Apple recommends indirect gestures for interface elements and common components such as buttons, and reserves direct gestures for objects that invite close-up interaction.[23] A similar combination had been studied in research: in 2017 Ken Pfeuffer, Benedikt Mayer, Diako Mardanbegi and Hans Gellersen described "Gaze + Pinch" for VR, a technique that "integrates eye gaze to select targets, and indirect freehand gestures to manipulate them".[24] When web content uses WebXR on Vision Pro, a pointing ray exists only while the user pinches: in WebKit's description, the transient-pointer's target ray "begins with its origin between the user's eyes and points to what the user was looking at the start of the gesture", and it then follows the movement of the hand rather than the eyes. The mode was introduced behind a flag in Safari 17.4 on visionOS 1.1.[25]
See also
References
- ↑ 1.00 1.01 1.02 1.03 1.04 1.05 1.06 1.07 1.08 1.09 1.10 1.11 1.12 1.13 1.14 1.15 1.16 1.17 1.18 1.19 1.20 1.21 1.22 1.23 1.24 1.25 Ferran Argelaguet, Carlos Andujar (2013-05). "A survey of 3D object selection techniques for virtual environments". Computers & Graphics, vol. 37, no. 3, pp. 121-136. Elsevier. https://doi.org/10.1016/j.cag.2012.12.003. Retrieved 2026-10-06.
- ↑ 2.0 2.1 2.2 2.3 Marc Baloup, Thomas Pietrzak, Géry Casiez (2019-05-02). "RayCursor: A 3D Pointing Facilitation Technique based on Raycasting". Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, pp. 1-12. Association for Computing Machinery. https://doi.org/10.1145/3290605.3300331. Retrieved 2026-10-06.
- ↑ 3.0 3.1 3.2 3.3 3.4 3.5 Mark R. Mine (1995-05-05). "Virtual Environment Interaction Techniques". Department of Computer Science, University of North Carolina at Chapel Hill, Technical Report TR95-018. https://www.cs.unc.edu/techreports/95-018.pdf. Retrieved 2026-10-06.
- ↑ 4.0 4.1 "The OpenXR Specification (version 1.1) - Standard pose identifiers". Khronos Group. The Khronos Group. https://registry.khronos.org/OpenXR/specs/1.1/html/xrspec.html. Retrieved 2026-10-06.
- ↑ 5.0 5.1 "WebXR Device API - W3C Candidate Recommendation Draft". World Wide Web Consortium (W3C). 2026-06-09. https://www.w3.org/TR/webxr/. Retrieved 2026-10-06.
- ↑ 6.0 6.1 6.2 Doug A. Bowman. "Interaction Techniques for Immersive Virtual Environments: Design, Evaluation, and Application". Graphics, Visualization, and Usability Center, Georgia Institute of Technology. https://people.cs.vt.edu/bowman/papers/hcic.pdf. Retrieved 2026-10-06.
- ↑ 7.0 7.1 "Point and commit with hands". Microsoft Learn (Mixed Reality documentation). Microsoft. https://learn.microsoft.com/en-us/windows/mixed-reality/design/point-and-commit. Retrieved 2026-10-06.
- ↑ Richard A. Bolt (1980). ""Put-That-There": Voice and Gesture at the Graphics Interface". SIGGRAPH '80: Proceedings of the 7th Annual Conference on Computer Graphics and Interactive Techniques, pp. 262-270. Association for Computing Machinery. https://www.media.mit.edu/speech/papers/1980/bolt_SIGGRAPH80_put-that-there.pdf. Retrieved 2026-10-06.
- ↑ 9.0 9.1 9.2 9.3 9.4 9.5 Frank Steinicke, Timo Ropinski, Klaus Hinrichs. "Object Selection in Virtual Environments with an Improved Virtual Pointer Metaphor". Computer Vision and Graphics (ICCVG 2004), Computational Imaging and Vision, pp. 320-326. Springer. https://www.uni-ulm.de/fileadmin/website_uni_ulm/iui.inst.100/institut/Papers/viscom/2006/steinicke2006object.pdf. Retrieved 2026-10-06.
- ↑ Jiandong Liang, Mark Green (1994-07). "JDCAD: A highly interactive 3D modeling system". Computers & Graphics, vol. 18, no. 4, pp. 499-506. Elsevier. https://doi.org/10.1016/0097-8493(94)90062-0. Retrieved 2026-10-06.
- ↑ Andrew Forsberg, Kenneth Herndon, Robert Zeleznik (1996). "Aperture based selection for immersive virtual environments". UIST '96: Proceedings of the 9th Annual ACM Symposium on User Interface Software and Technology, pp. 95-96. Association for Computing Machinery. https://doi.org/10.1145/237091.237105. Retrieved 2026-10-06.
- ↑ Doug A. Bowman, Larry F. Hodges (1997). "An evaluation of techniques for grabbing and manipulating remote objects in immersive virtual environments". Proceedings of the 1997 Symposium on Interactive 3D Graphics (I3D '97). Association for Computing Machinery. https://doi.org/10.1145/253284.253301. Retrieved 2026-10-06.
- ↑ 13.0 13.1 Doug A. Bowman. "Evaluation of Virtual Grabbing/Manipulation Techniques". Georgia Institute of Technology (project page, hosted at Virginia Tech). https://people.cs.vt.edu/bowman/grab.html. Retrieved 2026-10-06.
- ↑ 14.0 14.1 14.2 Tovi Grossman, Ravin Balakrishnan (2006-10-15). "The Design and Evaluation of Selection Techniques for 3D Volumetric Displays". UIST '06: Proceedings of the 19th Annual ACM Symposium on User Interface Software and Technology, pp. 3-12. Association for Computing Machinery. doi:10.1145/1166253.1166257. https://www.dgp.toronto.edu/~ravin/papers/uist2006_volumetricselection.pdf. Retrieved 2026-10-06.
- ↑ Yiqin Lu, Chun Yu, Yuanchun Shi (2020). "Investigating Bubble Mechanism for Ray-Casting to Improve 3D Target Acquisition in Virtual Reality". 2020 IEEE Conference on Virtual Reality and 3D User Interfaces (IEEE VR), pp. 35-43. IEEE. https://doi.org/10.1109/VR46266.2020.00021. Retrieved 2026-10-06.
- ↑ 16.0 16.1 "XR Ray Interactor". Unity Documentation (XR Interaction Toolkit 3.0). Unity Technologies. https://docs.unity3d.com/Packages/[email protected]/manual/xr-ray-interactor.html. Retrieved 2026-10-06.
- ↑ Regis Kopper, Doug A. Bowman, Mara G. Silva, Ryan P. McMahan (2010-10). "A human motor behavior model for distal pointing tasks". International Journal of Human-Computer Studies, vol. 68, no. 10, pp. 603-615. Elsevier. https://doi.org/10.1016/j.ijhcs.2010.05.001. Retrieved 2026-10-06.
- ↑ Logan Lane, Feiyu Lu, Shakiba Davari, Robert J. Teather, Doug A. Bowman (2025-10). "Revisiting Performance Models of Distal Pointing Tasks in Virtual Reality". IEEE Transactions on Visualization and Computer Graphics, vol. 31, no. 10, pp. 8283-8296. IEEE. arXiv:2505.03027. https://doi.org/10.1109/TVCG.2025.3567078. Retrieved 2026-10-06.
- ↑ Tiffany Luong, Yi Fei Cheng, Max Moebus, Andreas Fender, Christian Holz (2023). "Controllers or Bare Hands? A Controlled Evaluation of Input Techniques on Interaction Performance and Exertion in Virtual Reality". IEEE Transactions on Visualization and Computer Graphics (project page, Sensing, Perception and Interaction Lab). https://siplab.org/projects/Controllers_or_Bare_Hands. Retrieved 2026-10-06.
- ↑ "Learn about Hand and Body Tracking on Meta Quest". Meta Quest Help Center. Meta. https://www.meta.com/help/quest/290147772643252/. Retrieved 2026-10-06.
- ↑ "Ray casting best practices". Meta Horizon Developers (design guidelines). Meta. https://developers.meta.com/vr/design/raycasting_bp/. Retrieved 2026-10-06.
- ↑ "Ray Interactions". Meta Horizon Developers (Interaction SDK documentation). Meta. https://developers.meta.com/horizon/documentation/unity/unity-isdk-ray-interaction/. Retrieved 2026-10-06.
- ↑ "Gestures - Human Interface Guidelines". Apple Developer. Apple. https://developer.apple.com/design/human-interface-guidelines/gestures. Retrieved 2026-10-06.
- ↑ Ken Pfeuffer, Benedikt Mayer, Diako Mardanbegi, Hans Gellersen (2017-10-16). "Gaze + pinch interaction in virtual reality". SUI '17: Proceedings of the 5th Symposium on Spatial User Interaction, pp. 99-108. Association for Computing Machinery. https://doi.org/10.1145/3131277.3132180. Retrieved 2026-10-06.
- ↑ Ada Rose Cannon, Brandel Zachernuk (2024-03-19). "Introducing Natural Input for WebXR in Apple Vision Pro". WebKit Blog. Apple. https://webkit.org/blog/15162/introducing-natural-input-for-webxr-in-apple-vision-pro/. Retrieved 2026-10-06.