When a player straps on a virtual reality (VR) headset, their living room can turn into a virtual kitchen. They can reach for a mug on a table, and their hand will land exactly where it should. As they walk around the table, the mug stays put because the headset is tracking where they are in the real room.
Simultaneous localization and mapping (SLAM) makes that possible. The VR headset maps the room and finds its own place in it at the same time, picking out stable details the way people use landmarks to find their way around a new building. This is also called SLAM VR tracking. The same technology is now being used in robotics, where teams use consumer VR headsets to steer humanoids and record first-person demonstrations.
For instance, 1X’s home humanoid robot NEO has an Expert Mode that lets owners schedule a remote 1X Expert to guide it through chores it doesn’t know yet. Those operators wear Quest 3 headsets, turning each session into both a completed chore and a demonstration the robot can learn from.

VR headset tracking becomes especially critical once those sessions become training data. Each recording is a demonstration the robot can learn from, so it’s only as good as the tracking behind it.
SLAM VR tracking has to follow the operator’s head and hands throughout a session. Along the way, it records head pose, the position and angle of the wearer’s head. Head pose anchors egocentric and teleoperation data, because the rest of the recording is measured relative to the head.
If the tracking slips even slightly, the head pose drifts off, and every hand position built on it shifts along with it. Those errors are easy to miss in footage, but they still shape the data and the policies trained on it. Reliable SLAM tracking keeps that head pose trustworthy from the start.
In this article, we’ll explore how SLAM tracking works in VR headsets, where it can break down, and why those errors matter for robot training data. Let’s get started!
Before we see where VR headset tracking can go wrong, let’s cover how SLAM VR tracking works as you move.
Most standalone headsets use inside-out tracking, where cameras on the headset look out at the room. As you turn your head, those cameras pick up details that stay in place, such as the edge of a shelf, the corner of a table, or a pattern on a rug.
The VR headset keeps checking where those details appear from one frame to the next. That gives it two pieces of information at the same time: a rough 3D picture of the room and its own position inside that room. Handling both together is the core of SLAM tracking, and visual SLAM in robotics works on the same basic idea.
On their own, though, cameras can miss what happens between frames, especially during a quick head movement. An inertial measurement unit (IMU) helps fill that gap. It measures changes in acceleration and rotation much more frequently than the cameras capture images, giving the headset a faster read on how it’s moving.
The headset then combines what the cameras see with those IMU readings. This approach, called visual-inertial odometry, helps it keep track of movement even when your head turns quickly.
XREAL’s XR glasses are a good example. Two SLAM cameras on the sides of the frame detect feature points and track them over time, then combine that with IMU readings to estimate the glasses’ position and orientation. Together, these pieces give the headset 6DoF (six degrees of freedom) tracking, meaning it tracks the six ways your head can move. In simple terms, it knows where your head is and which way it is facing.

Three degrees cover movement up, down, forward, backward, and side to side. The other three cover rotation, such as nodding, turning left or right, and tilting your head. Combined, they let the headset track your full head movement as you move around the room.
VR headsets haven’t always tracked themselves. So, how does VR headset tracking compare with room-based tracking? With outside-in tracking, sensors around the room watch the headset as it moves.
On the other hand, inside-out tracking uses cameras on the headset to track its own position. Each method has its strengths, and the choice shapes where it can be used for data capture.
Here’s where the two approaches differ:

For data capture, the right setup depends on where the recording takes place. In a controlled lab, outside-in tracking still works well. For example, Berkeley Humanoid Lite uses SteamVR base stations to track handheld controllers during teleoperation.
Outside the lab, inside-out tracking makes it easier to capture robot demonstrations in real homes and work sites without a tracking rig. The tradeoff is that the headset relies on its own SLAM map, so any error in that map carries into the recorded head pose.
Now that we’ve seen how headsets track themselves and why capture moved to inside-out, let’s look at where that tracking starts to slip.
SLAM VR tracking depends on the cameras seeing enough stable detail to keep their position anchored. Blank walls, dim lighting, and reflective surfaces can leave the headset with fewer reference points to follow. Moving people create another problem, since the headset may treat them as fixed parts of the room.
On top of this, quick head turns can complicate things further. Fast movement can blur the camera frames, leaving the headset with fewer clear features to match from one frame to the next.
Recently, one research team put these weak spots to the test. They mounted several commercial headsets, including Meta Quest 3 and Apple Vision Pro, on a single helmet rig and compared their tracking against a motion capture system.
The headsets were tested in two versions of the same room. One had blank walls and plain furniture, giving the cameras little to work with. The other used projected brick patterns and everyday objects to add more visual detail.

The difference showed up clearly in the results. On one device, tracking errors during fast movement doubled from about 9 cm to more than 18 cm when the room offered fewer visual details. Faster movement also increased errors across the tested devices.
These problems grow during longer recordings. Small errors slowly add up into drift, pulling the recorded position away from where the headset really is. Losing and regaining tracking can also make the pose jump, and the hands add one more gap, since the headset can only track them while they stay in its cameras’ view.
Tracking errors become costly once a recording turns into training data. Take egocentric data collection, for example. The camera sits on the wearer’s head and sees the hands from their point of view, but the camera is moving too. As a result, a hand resting on a table can appear to move just because the wearer glances to the side.
Head pose helps separate these two movements. It puts the hands and objects in world coordinates, so their positions are measured against the room instead of the camera. This gives the data a fixed reference point, even when the wearer looks around. It also makes it easier to see how the hands actually moved.
An interesting example is the EgoVerse dataset, which shows how central this signal can be for robot learning. It contains 1,362 hours of demonstrations recorded on glasses, a head-strapped phone, and custom head-mounted rigs.

For every frame, the team pairs 3D hand keypoints with a 6DoF head pose from visual-inertial SLAM. Together, these two signals map how the person’s hands moved through the room, and a robot policy then learns to translate that motion into arm and gripper actions.
In egocentric robot datasets, a reliable head pose does more than keep hand positions in place. It determines whether recorded human hand motion can be turned into trajectories a robot arm can follow.
Here is a closer look at how head pose affects this kind of robot training data:
Those errors then carry into training. Imitation learning policies and vision-language-action models both learn tasks from recorded demonstrations, and the latter also follows written instructions. Both treat the recorded movement as ground truth, so drift gets learned as if it were real motion.
Teleoperation involves controlling a robot remotely, with a person guiding it through a task in real time. In these setups, the VR headset does more than show the operator what the robot sees. It tracks the operatorβs movements and sends them to the robot as commands. Those movements can also become part of a training dataset.
For example, Teleopit, a full-body humanoid teleoperation system, shows how much of the robot can depend on headset tracking. The operator wears a PICO headset that streams 6DoF head tracking, a 24-joint body skeleton, and 26 keypoints per hand to a humanoid robot. The robot’s body, hands, and camera then follow those streams in real time.

Head pose also controls where the robot looks by steering a two-axis camera. A dead zone and smoothing filter prevent small head movements from making the camera shake.
Beyond this, timing is critical. The robot responds about 0.05 to 0.15 seconds after the operator moves, so the system always acts on the newest head and hand pose instead of letting older data pile up. Every command the headset sends is also saved. Each session is logged at 30 Hz, and policies trained on 96 of those recordings completed a bottle-placement task 90 to 95% of the time.
Other open-source frameworks follow a similar approach. They use the operator’s head position as the reference point for arm commands, and they smooth the tracked joint data before sending it to the robot.
However, filtering only removes some of the noise before it reaches the robot. Any jitter or lag that slips through becomes part of the recorded signal, sitting alongside the operator’s intended movement.
A tracking problem is far cheaper to catch during capture than after you build a dataset. For instance, EgoKit, an open-source capture toolkit, logs a 6DoF head pose and 26-joint hand tracking for every frame on PICO 4 Ultra and Quest 3. It timestamps each pose to the millisecond and saves it alongside the video.
Even that level of careful logging leaves room for error. In one of the toolkit’s own PICO logs, the pose stream starts about 0.2 seconds after the video and ends about 0.3 seconds before it, so you have to align the two afterward.

Here are a few checks that you can build into every capture session:
Done during capture, these checks turn head pose from something a dataset simply assumes into something a team can verify. They work best when data experts build them into the capture process from the very first session.
At Objectways, we collect first-person egocentric video across everyday environments for training embodied AI. We also capture human-guided teleoperation demonstrations of robot manipulation at scale.
The real value of SLAM in VR headsets shows up in the data they capture. Every egocentric or teleoperation recording depends on knowing how the wearer’s head moved. The footage might look perfectly fine, but if SLAM VR tracking drifts or loses track, that mistake becomes part of what a robot learns.
So head tracking deserves the same scrutiny as the video itself. Logging tracking state, checking sync, and spot-checking sessions against a reference are small steps, but they can catch problems before they spread through a dataset. Good footage paired with verified head pose gives a robot a demonstration it can actually learn from.
Looking for egocentric data built around your robotics project? Reach out to Objectways to learn more about our data collection services.
No, they are different. SLAM is the software process that maps a space and works out a device’s location within it, while lidar is a physical sensor that measures distance using laser light.