Blog · Teleoperation

Quest controllers, headset off: how we built a teleop rig operators can use for hours

Robot learning runs on demonstrations, and demonstrations come from people. If you want a few hundred clean trials of a bimanual task, someone has to sit at the rig and do it, trial after trial, for most of a day. That makes the operator's comfort a data-quality problem, not just a nice-to-have.

This post walks through how we set up paddy, our teleoperation and recording stack for UFACTORY xArm arms, so that operators can keep going for hours. The short version: we use Meta Quest controllers, and we leave the headset on the desk.

The problem: headsets get tiring, and tired operators make noisy data

VR teleoperation is a popular way to collect manipulation data, and for good reason. A Quest controller gives you a tracked 6-DoF pose, an analog trigger and a handful of buttons for the price of a consumer headset. The catch is wearing it. After a long stretch with a headset on, people get warm, their eyes get tired, and the headset gets pushed up and adjusted between trials.

That fatigue shows up in the dataset. Late-session demonstrations get slower, hesitant or sloppier than early ones, and a policy trained on them inherits that inconsistency. In our own sessions, operators wearing the headset wanted a break after about 45 minutes; with the headset off they run sessions of 2-3 hours.

The idea: keep the controllers, put the headset on the desk

The controllers are what an operator needs. The headset's job in our setup is to be the tracking reference: it sits on the desk, facing along the arms' x axis, and the Quest2ROS app on it streams both controller poses and button states over Wi-Fi to the rig at about 50 Hz. The operator looks at the arms directly, and when the gripper is out of their line of sight, at the browser camera dashboard that shows each RealSense stream.

The headset is placed on a mount that keeps it stable and at a consistent height and orientation relative to the desk, ensuring reliable tracking throughout the session.

The hardware is deliberately ordinary:

  • Arms: one or two UFACTORY xArms (xArm5, xArm6, xArm7 or Lite6), with an xArm Gripper, Gripper G2 or BIO Gripper G2. Our reference rig is a dual xArm6.
  • Teleop: a Meta Quest 2 or 3 with both controllers.
  • Cameras: Intel RealSense D405 and D435, including wrist cameras.
  • E-stop: any USB foot switch that shows up as a keyboard, plus Shift in the operator GUI.

On a two-arm rig, mirror mode (on by default) has each arm driven by the opposite-hand controller, so the operator can work face to face with the robot.

Each controller maps the same way:

InputWhat it does
Grip trigger, heldDeadman: the arm follows only while it's held
Index triggerGripper, proportional: released is open, pressed is closed
Upper buttonToggle between the precise and fast scaling profiles

Releasing the grip trigger stops the arm at once. Gripping again resumes from wherever the arm is, so the operator can "clutch" to reposition their hand without moving the robot.

Scaling profiles: one button between precise and fast

Controller motion is scaled before it becomes an arm target, and the operator switches between two profiles with the upper button:

  • 0.7× (precise): hand motion is shrunk, so a small tremor or overshoot becomes an even smaller arm motion. This is the mode for fine manipulation: aligning a gripper on a well plate, inserting, placing.
  • 1.2× (fast): hand motion is amplified, so big moves across the workspace need less arm travel from the operator. This is the mode for repositioning between sub-tasks.

The toggle is designed so it never surprises anyone. Pressed with the grip trigger released, the new ratio applies at once and the GUI shows it. Pressed mid-motion, the change is queued and applies when the operator lets go, so the mapping never re-anchors while the arm is moving; pressing again before release cancels it. Taps closer than 0.2 s apart count as one. The scale in effect is published per arm, so the dataset always knows which mode a segment was recorded in.

Cartesian online planning: stream targets, let the controller plan

Teleoperation has no known goal. The operator decides where the arm goes next as they watch it, so there's nothing to plan ahead. That's why we don't pre-plan trajectories or put MoveIt in the loop.

Instead, each arm has a follower process that talks to the xArm controller directly through the xArm Python SDK. The controller runs in its online trajectory planning mode (mode 7) and solves inverse kinematics itself. The pipeline looks like this:

  1. The Quest pose arrives at about 50 Hz.
  2. The teleop controller applies anchoring, filtering and the deadman, then publishes a time-stamped Cartesian target.
  3. The follower validates every target: stale or out-of-order stamps, the workspace box, step size, velocity and tracking error.
  4. A 50 Hz command thread sends the newest valid target with a non-blocking set_position. It's a latest-wins mailbox, not a queue, because a control loop wants where the hand is now, not where it was.

For reactive tasks, this matters. If the operator changes their mind halfway through a reach, the next target simply replaces the old one and the controller's planner blends toward it. There's no trajectory to cancel and no backlog to drain. It also keeps the dependency surface small: no per-rig kinematics configuration, and everything site-specific lives in one validated rig profile.

Safety sits in the follower itself. Arms start with motion disabled, only move while the deadman is held, drop out of the armed state if the 20 Hz teleop heartbeat stops, and latch on any e-stop until someone resets them explicitly.

Everything is recorded

A demonstration is only as useful as what you can reconstruct from it later, so we record the operator's side and the robot's side together. Every trial is a folder of MCAP bags plus a metadata.json, and the streams include:

TopicWhat it holds
/paddy/quest/<hand>/{pose, inputs, twist}Raw controller poses, buttons and triggers
/paddy/<side>_arm/target_poseThe commanded TCP pose, while the deadman is held
/paddy/<side>_arm/actual_poseThe achieved TCP pose, at 50 Hz
/paddy/<side>_arm/joint_statesJoint positions per arm, at 100 Hz

Alongside these, we record gripper commands and status, the active scaling and control mode, safety state changes, TF and every RealSense color and depth stream. Having commanded and achieved poses side by side means you can measure tracking error per trial, spot segments where the arm lagged the operator, and choose whether to train on what the operator asked for or what the robot did.

Operators label each trial as they go (Success & Next is a single key), and paddy convert turns a session into a LeRobot v3 dataset.

What's next

  • Body-mounted tracking reference. Keep the headset on the operator's body, either with custom shoulder straps or mounted on a wearable laptop stand. The operator keeps the same comfort, but the tracking reference moves with them, so they can step around the cell for a better view of the gripper.
  • A headset-worn remote mode, for operators who aren't in the room with the robot and need the camera views in front of them.
  • More controller mappings, for tasks that need different button layouts or per-task scaling.
  • Gello support, for teams that prefer a joint-space leader arm to a VR controller.

See it for yourself

Want the full stack? See every component, topic and safety limit on our Technical page.

Collecting data? Reserve a Harvester or book lab time with our rig.