Understanding 3D Mid-Air Spatial Manipulation in Immersive Environments: A Taxonomy and Visual Representation

In this project, we propose a novel taxonomy for mid-air spatial manipulation techniques in 3D virtual environments (VEs). Mid-air spatial manipulation refers to the application of mathematical transformations that alter an object’s position or orientation in virtual space, performed through gestures or actions directly in 3D space and closely resembling real-world interactions with physical objects.

Over the years, a wide range of mid-air manipulation techniques have been proposed, and the field has shifted from one-size-fits-all solutions to a diverse set of strategies tailored to specific use cases. Designers now face a complex design space with numerous considerations, including input devices and mechanisms, interaction constraints, and strategies to map physical movement to virtual motion. Critically, these decisions must account for contextual factors such as task demands, user capabilities, system constraints, and environmental conditions.

To support reasoning within this design space, several taxonomies of 3D interaction have been proposed, but none provides a complete view of the spatial manipulation design space or its complexity. Our novel taxonomy characterizes 3D spatial manipulation techniques across five key orthogonal components, facilitating the description, comparison, and design of mid-air spatial manipulation techniques.

Fig. 1: Novel Taxonomy for Classifying 3D Mid-Air Spatial Manipulation Techniques: Five components uniquely describe each operating mode of a technique. The blue boxes indicate the options within a component or subcomponent, from which a single option should be chosen. The separation method subcomponent only applies when a level of separation exists (partial or total). * [T][R] in the Transfer Function Mappings component indicates that an operating mode can include multiple spatial transformations, each with one or more associated transfer functions. Thus, this component should be repeated for each.

We also propose a visual representation that enables techniques and their operating modes to be visualized in a compact, easy-to-read format, facilitating comparison of operating modes and techniques.

Finally, to demonstrate the expressiveness and practical utility of our framework, we present a curated representative set of existing manipulation techniques and discuss how they are classified and illustrated.

We aim for our taxonomy and visual representation to facilitate the entire research and design process of 3D manipulation techniques, starting from identifying limitations in the existing methods, pinpointing potential areas that could be enhanced to address a specific limitation, designing innovative approaches based on these identified shortcomings, and finally, comparing those different approaches to assess their effectiveness.

, ,