All episodes
Cover art for Self-Confrontation in VR: How Seeing Yourself Shapes Motor Skill Reflection

Self-Confrontation in VR: How Seeing Yourself Shapes Motor Skill Reflection

— VRST 2025 — episode 58

0:00 14:26

What happens when you watch yourself perform in VR—and then have to critique that performance? This study explored self-reflection in motor skill learning using a Karate training task. Participants were embodied as a "trainer" avatar and asked to give verbal feedback on a "trainee"—which was either their own 3D-scanned appearance or a stranger, performing either their own recorded movements or an expert's. The results revealed a psychological "role conflict": participants felt split between being the evaluator and being evaluated. Seeing their own appearance triggered deeper, more emotional reflection, while recognizing their own movements created bodily connection even in a stranger's avatar. The findings suggest VR embodiment isn't binary but multi-faceted, with implications for training and therapy.

Dennis Dietz, Samuel Benjamin Rogers, Julian Rasch, Sophia Sakel, Nadine Wagener, Andreas Martin Butz, and Matthias Hoppe. 2025. The 2×2 of Being Me and You: How the Combination of Self and Other Avatars and Movements Alters How We Reflect on Ourselves in VR. In Proceedings of the 31st ACM Symposium on Virtual Reality Software and Technology (VRST '25). Association for Computing Machinery, New York, NY, USA, 11 pages. https://doi.org/10.1145/3756884.3765986

Download the audio (28 MB)
Transcript 2,492 words

Welcome to the Deep Dive. Today we're immersing ourselves in the cognitive science of learning. We're focusing on how we master complex motor skills when, well, when the teacher, the critic, and the subject are all the same person, all inside virtual reality. It's a fascinating area. And effective self-reflection is just so critical for skill mastery, isn't it? It's the engine that drives that whole cycle of what we call self-regulated learning.

It is. And that engine often stalls with the traditional methods we use. Oh, how? Well, if you think about how we usually do self-observation, say, analyzing your performance on a standard 2D video, there are significant drawbacks. Okay. They struggle to diagnose the subtle three-dimensional posture errors you find in complex movements. And crucially, they enforce a kind of fundamental disembodiment. They separate you, the observer, from the performance you're supposed to be reviewing.

I see. So there's a disconnect. Let's unpack that challenge then. The mission for this deep dive is to examine a study that proposes a genuine paradigm shift. It does. Moving from that passive, flat self-observation to something active. Embodied self-evaluation, all within a virtual environment. And the core goal is to investigate precisely how that combination of visual appearance, seeing yourself and observed movement, seeing your own performance,

how that alters a user's sense itself, and of course, the quality of their reflection. This research touches on some profound psychological territory. It examines the emergence of a complex psychological duality, a role conflict, as the paper calls it. A role conflict. It's what surfaces when you are simultaneously the subject acting as the trainer providing critique and the object acting as the trainee being evaluated, all in the same virtual space.

So how on earth did the researchers design an experiment to isolate those kinds of psychological forces? They developed a custom VR application. It was called copycata. Copycata. And it was centered around a foundational karate sequence, something suitable for novices. The genius of the study was in its very clean experimental design. A two-by-two matrix that systematically varied two independent variables. In all conditions, participants were embodied as a generic first-person trainer avatar.

And their task was to give verbal feedback to a trainee avatar that they saw in front of them. Exactly. Continuous feedback. So let's isolate those two variables. They seem to define the core conflict here. What was variable number one? That was the avatar appearance. In the self-avatar conditions, the trainee they were criticizing was a personalized, photorealistic 3D scan of the participant. Of themselves.

Of themselves, yes. And in the other avatar conditions, the trainee was a generic avatar of the opposite gender, just to ensure a clear visual distinction. Okay, so that isolates the impact of just visual recognition. And variable number two, manipulating the source of the performance. That was the observed movement. Participants either evaluated their own movement. The ones they'd recorded earlier, flaws and all.

The very same. Pre-recorded actions from an initial imperfect training session. Or they evaluated other movements. These were expert movements pre-recorded by a martial arts master, representing, you know, peak proficiency. And that isolates the impact of motor congruence. Whether the movement feels like your own. Correct. So combining these creates four very distinct psychological scenarios. And these become critical reference points for the entire discussion.

Indeed. First, we have what they call dissonance. That's the other avatar performing the participant's own movements. You see your own motor skills in a foreign body. I can imagine. That's quite strange. Then you have what is maybe the most psychologically intense condition. Direct confrontation. That's the self-avatar performing their own movements. An unfiltered, ego-involved self-assessment. No escaping yourself there. None at all. Thur is the ideal self.

This is the self-avatar performing the expert's other movements. You see your own appearance achieving peak performance. A glimpse of your potential. In a way, yes. And finally, you have the baseline. The other avatar performing the other movements. This serves as the detached control condition. Just evaluating a stranger's expert performance. That gives us a clear framework. So before this VR confrontation even begins, we have to establish the starting line.

When a novice, someone lacking expertise, approaches a critique task like this. What is their inherent reflective baseline? The data paints a very clear picture here. Novices adopt a coping strategy. The coping strategy. Because they lack the deep internal schema of an expert, they can't diagnose subtle timing or weight distribution issues. They focus almost entirely on visually obvious concrete details. Ah. So they can't offer nuanced feedback on the quality of the movement, so they focus on the

presence of it. They gravitate toward tangible things they can just verify with their eyes. Exactly. This results in what's called descriptive reflection. It corresponds to the R0 or R1 levels in the Fleck and Fitzpatrick reflection framework. And for context, R0 and R1 is just simple description. What you see or naming an action. That's it. And the goal is to guide users toward R3, which is transformative reflection, understanding

why an error occurred, and then formulating a detailed strategy for corrective action. And their verbal feedback actually reflected this descriptive baseline. It did. The verbal data was dominated by high instances of body part focus and movement focus. They confidently used concrete nouns, descriptive words like kick and punch. But the moment they had to shift to true evaluation or diagnosis, they frequently used hedge words.

Things like, I think, or maybe. It expresses their fundamental uncertainty about the quality of what they were seeing. They knew what they saw, but not what it meant. Precisely. So this detached descriptive baseline is the setup for the psychological shift. When you introduce the self-avatar, when objective evaluation becomes an ego-involved self-assessment, what's the force that breaks through this descriptive detachment? The visual presence of the self-avatar acts as a significant psychological catalyst.

The past is immediately transformed. It moves away from just technical analysis into this emotionally charged process of self-assessment. It elicited strong feelings that participants described as spanning from embarrassing and ashamed or disappointed to, well, to simply finding it funny how poorly they performed. That's a profound emotional fallout. And just from seeing a virtual copy of yourself doing a novice action, why is a self-avatar so much more potent than just watching a 2D video of the same performance?

It's the immediacy of the confrontation. And this emotional shift is evident linguistically, too. In the other avatar conditions, speech was dominated by foreign reference. Meaning they spoke about the avatar. Yes. The trainee needs to lift its leg higher, that kind of thing. But this pattern completely inverted in the self-avatar conditions. The use of self-reference increased substantially. So they shifted from saying it seems to lack power to saying I should have put more power into my kick.

That is exactly it. They began talking to themselves or about themselves. They turned the external critique into an internal critical dialogue. This is the engine of that role conflict we mentioned. The battle between the observing self and the observed self. How did the researchers quantify this feeling? The sense of being two people at once. They looked at metrics that measured perceived role.

The visual presence of the self-avatar, maybe unsurprisingly, significantly increased the participants' sense of ownership over the trainer body, the avatar they inhabited in first person. That looks like me, so I own this role. But, and this is the pivotal finding, they were simultaneously feeling like the person being critiqued. Ah. Even as their sense of trainer ownership went up, participants reported feeling significantly more trainee, the individual being observed.

This confirms that the internal critical dialogue with the self, that ego-involved force, was psychologically superior. It superseded the avatar's external appearance or the first person perspective they were in. Their mind identified with the flawed performer. Their mind did. So, if the user's mind is rejecting the body they inhabit and identifying with the avatar they're criticizing, that forces us to redefine what we even mean by embodiment in a VR setting.

It forces us to challenge the simple binary view of embodiment. Yes. The intensive self-evaluation didn't lead to a simple feeling of identification or detachment. Instead, it led to a fluid, multifaceted sense of self that was constantly oscillating between the two roles. And the quantitative data shows that this multifaceted embodiment is driven by two distinct separable factors. This is where it gets highly specific for system design.

It is. The analysis of their avatar embodiment questionnaire showed two independent influences. First, avatar appearance was the significant factor driving the sense of ownership. Which is the cognitive connection. That looks like me, therefore I accept this as my body. The cognitive connection, correct. So, appearance drives the visual, ego-driven side of self-identification. What drives the physical sensory side? The observed movement. That was the significant factor driving the multi-sensory subscale of the embodiment questionnaire.

This means whether or not they were watching their own movements, the motor congruence, that was the key to physical and sensory integration, regardless of who the avatar looked like. So, it's a kinesthetic integration. The movement feels correct, even if the body looks wrong. Got it. Let's connect these two drivers back to those specific scenarios. In the direct confrontation condition, seeing your own face perform your own novice movements, is that where the mind fractures, and we see this split embodiment.

That is the pinnacle of the role conflict. Participants reported feeling a sense of bilocation, physically present as the trainer, but mentally and critically identified with the trainee. Their consciousness was effectively inhabiting both roles at the same time. It's an incredibly intense state of self-assessment. Okay, now consider the dissonance condition. The opposite. They saw their own movements, but performed by a foreign opposite-gender avatar.

This challenges that visual first assumption of VR. What psychological phenomenon emerged there? That's where the Puppet Master effect came out. Participants experienced a state of detached control over that foreign avatar. Detached control. They knew intellectually it was not their body, but the motor congruence, the fact that the movements matched their prior effort, it created a sensory integration strong enough to drive high multisensory scores.

So, motor congruence can override visual discrepancy. It demonstrates that the feeling of, this movement belongs to me, creates a form of sensory integration, even without the visual confirmation of a self-avatar. Did some participants just try to bypass this whole intense experience? They did. We saw reports of participants comparing the experience to watching a 2D video of themselves. This highlights a cognitive strategy, externalization.

They were attempting to transform the confrontational 3D VR experience into a more objective, detached view. Trying to return to that R0 or R1 descriptive baseline, just to cope with the emotional intensity. Exactly. This fragmented and fluid sense of self is perhaps the biggest takeaway then. Embodiment is a dynamic construct, shaped independently by visual identity and mode of performance. This knowledge isn't just descriptive, it has to be prescriptive.

How do we translate these deep psychological findings into concrete guidance for system designers, for people seeking to guide users toward that R3 transformative reflection? Well, first we have to overcome the foundational hurdle of novice uncertainty, their reliance on that R0, R1 descriptive feedback. Systems need to provide structured guidance and, importantly, display expert movements in parallel to the learners. This offers a concrete ground truth for comparison. It moves the user beyond simple description.

That gives the novice an immediate benchmark. It moves them past, I think the kick was weak, to the expert's leg rotation is clearly different from mine. What about leveraging the emotional power of the self-avatar without overwhelming the user? Since the self-avatar is such a potent psychological catalyst, but also emotionally challenging, the trainee representation must be a configurable parameter. A setting they can change?

Designers should enable users to consciously select an other avatar for objective, detached, technical comparison. Or, they can choose the self-avatar when they're prepared for that deeper, ego-involved self-assessment. They must be able to choose their psychological dosage. And finally, applying the embodiment findings. How do we use this duality of ownership and multisensory integration to refine the learning process itself? Designers now possess distinct controllers for embodiment, based on this data.

Visual appearance controls the sense of ownership, that emotional connection and acceptance of the avatar as oneself. But motor congruence controls multisensory integration, which enhances the sense of agency. This means, if you want a user to feel physically present in the system and feel the corrective actions are relevant, that's agency, you focus design efforts on ensuring motor accuracy and low latency. And if you want them to feel emotionally compelled to critique and change.

Then you focus on the photorealistic scan. You focus on ownership. So what we have established is that this confrontation with oneself in VR is far more than a high-fidelity viewing experience. Yeah. It's a profound, two-pronged psychological event. It reveals a core psychological duality. It confirms that the sense of self in a virtual environment is dynamic, constantly shaped separately by visual recognition and motor congruence.

This study focused on motor skill acquisition, using a karate sequence. But the methodology externalizing our internal self-talk through a virtual self avatar to facilitate deep evaluation, that's not limited to physical movement, is it? Not at all. If creating an embodied dialogue with our virtual self is effective for analyzing and refining complex motor skills, think about the new possibilities this opens up for emotional and psychological growth.

This methodology offers compelling avenues for therapeutic interventions or for self-counseling. If the visible psychological shift from subject to object can motivate better physical performance, what could it do for achieving emotional insight and personal transformation? It gives us a tangible way to talk to our own thoughts. Then we'll start about thinking about development and environmental 99's work to ensure that we don't understand what works to document as well.

Thank you. For those kind of African American movements with respect to gender expectations. For those kind of African language models to ensure that we need to increase our raw DNA both recent movements and anti-eeee and examples of relative väρο滑 products and ancient Britain's planning as well. Ones here we now look at the streets of 480 million devices that are highly discriminated bypyrod shoulderные combat degrees of wearing Downups to ensure our combined MassDe wave Smarios bestأ .

Ones here we would begin with which there are finally series of examples of non-cor emissions with yesterday's guidance on anything like which our stackings are fully qualified career goals in which are extremely busy climate changes that are designed on par models of natural environmental figured meet our expectations of negative