Vision and AprilTags

Module 9: Vision and AprilTags

Series: FRC Technical Foundations Unit: 4 — Software and Perception Session length: ~3 hours Prerequisites: Lesson 8

What you should be able to do after this lesson

  • Explain how exposure, resolution, and frame rate trade against each other.
  • Calibrate a camera and say what the calibration actually produces.
  • Detect AprilTags and use them to correct a pose estimate.
  • Quantify latency and decide whether it matters for a given task.
  • Design a fallback so that vision failure degrades performance instead of ending the match.

What vision is for

Lesson 8 ended on a problem: odometry drifts, and dead reckoning has no way to recover. Vision solves it by providing an absolute position reference — something fixed in the world whose location the robot knows.

That is the primary job. Vision in FRC is not general scene understanding. It is: find a known landmark, compute where I am relative to it, correct my belief about my position.

The landmarks are AprilTags — square fiducial markers placed at surveyed positions around the field. FIRST publishes the exact field coordinates of every tag each season. Because the tag’s position is known and the camera can measure its apparent size, orientation, and position in frame, you can compute the camera’s pose relative to the tag, and therefore the robot’s pose on the field.

Imaging fundamentals

Before any of the software works, the physical camera setup has to be right. Four parameters interact, and students consistently discover this the hard way.

Resolution. Higher resolution detects tags at greater distance but costs processing time per frame. Most teams run lower resolution than they expect — 640×480 is often plenty for tags at competition distances.

Frame rate. More frames per second means fresher data and more correction opportunities, but higher resolution reduces achievable frame rate. This is a direct trade, and you make it consciously.

Exposure. The single most important setting and the most commonly wrong one. Long exposure gathers more light but blurs while the robot is moving, and a blurred tag does not detect. Short exposure freezes motion but needs good lighting.

The counter-intuitive correct answer for AprilTag detection is almost always: set exposure low, and lock it. A tag detected in a dark, sharp image is worth far more than a tag lost in a bright, smeared one. Auto-exposure is actively harmful, because it changes settings as the robot turns toward a bright window, and your detection rate changes with it.

Mounting rigidity. If the camera mount vibrates, every pose estimate is noisy, and no amount of software filtering fully recovers it. This is Lesson 2’s content reappearing: the camera mount is a structural component with a tight requirement.

Rule 9.1 — Lock exposure low, lock gain, and mount the camera like it matters. Most “our vision is unreliable” problems are one of these three.

Camera calibration

A camera lens distorts. Straight lines bow, and the amount of bowing increases toward the edges of frame. Calibration measures this distortion so the software can undo it.

Calibration produces two things:

  • Intrinsic parameters — focal length and optical centre, describing how the lens projects the world onto the sensor.
  • Distortion coefficients — describing the lens’s specific bowing.

Without calibration, tag pose estimates are systematically wrong, and the error grows toward the image edges. With calibration, they are accurate to a few centimetres at typical ranges.

Calibration is done by showing the camera a known pattern — a chessboard or a Charuco board — from many angles. Both PhotonVision and Limelight include calibration workflows.

Two rules that matter:

  • Calibrate at the resolution you will run. Intrinsics are resolution-specific. A calibration done at 1280×720 is wrong at 640×480.
  • Recalibrate if you change the lens, the camera, or the mount. And record which calibration belongs to which camera, because you will have several.

You also need the camera-to-robot transform: where the camera sits relative to the robot’s centre, in three dimensions plus rotation. Measure this carefully. A 10 mm error in the measured camera position becomes a 10 mm error in every pose estimate forever.

The vision stack

Two mature options dominate FRC.

PhotonVision is free, open-source, and runs on a coprocessor you supply — a Raspberry Pi, an Orange Pi, or a mini PC. More setup work, more control, no hardware cost beyond the coprocessor and camera.

Limelight is a purpose-built camera and coprocessor in one unit with a polished configuration interface. Costs money, saves time.

Both provide WPILib libraries, both do AprilTag detection and pose estimation, and both integrate with SwerveDrivePoseEstimator. The choice is about your team’s resources, not capability.

Fusing vision with odometry

The correct architecture is not “use vision instead of odometry.” It is to fuse both.

WPILib’s pose estimators accept odometry updates continuously and vision measurements opportunistically, weighting each by a standard deviation representing how much you trust it. Odometry is smooth but drifts; vision is absolute but noisy and intermittent. Fused, you get a pose estimate that is both smooth and bounded.

Getting the trust weighting right is the real work:

  • Trust vision less at long range. Pose error grows with distance because the tag occupies fewer pixels. Scale your standard deviation with measured distance.
  • Trust multi-tag more. Seeing two or more tags simultaneously dramatically improves the estimate, because it constrains the solution far better than one tag can.
  • Reject implausible measurements. If vision claims the robot moved four metres since the last loop, it is wrong. Discard it. Also reject poses outside the field boundary and tag IDs that do not exist in this year’s field.

Rule 9.2 — Every vision measurement gets sanity-checked before it is trusted. One accepted bad pose can throw your estimate off and take seconds to recover, in the middle of an autonomous routine.

Latency

Every vision measurement describes a moment in the past. Capture, transfer, processing, and network transmission each add delay — typically 20–80 ms in total.

At 4 m/s, 60 ms of latency means the robot has travelled 24 cm since the image was taken. If you apply that measurement as if it were current, you are correcting to where you were, not where you are.

The solution is timestamped measurements. Both PhotonVision and Limelight report when a frame was captured, and WPILib’s pose estimator accepts a timestamp and inserts the measurement at the correct point in its history buffer. Use the timestamped API. This is not an optimisation; it is the difference between vision helping and vision oscillating.

Fallback: the part everyone skips

Vision will fail. A camera will be knocked out of alignment by a collision. A coprocessor will fail to boot. Another robot will park directly between yours and every visible tag. Lighting at one venue will be unlike anywhere you have tested.

The question is not whether this happens but what your robot does when it does.

Bad: the command waits for a tag that never arrives, and the robot sits motionless for the remaining ten seconds of autonomous.

Good: the command has a timeout, falls back to odometry-only navigation, completes a reduced version of the routine, and reports clearly on the dashboard that it ran in degraded mode.

Rule 9.3 — Every vision-dependent command needs a timeout and a non-vision fallback. Design the degraded mode deliberately, and test it by unplugging the camera.

This is also an engineering-ethics point worth making explicitly with students. A system that behaves unpredictably when its sensor fails is badly designed regardless of how well it performs when everything works. Documenting your system’s limitations — its false positive rate, the conditions where it is unreliable, what it does when it cannot see — is part of the engineering, not an afterthought.

[ANECDOTE SLOT] — A vision system failing at competition in a way you did not anticipate. Lighting, occlusion, a bumped camera, or a coprocessor that did not boot. The recovery is the interesting part.

🔧 Exercise 9.1 — Exposure and Motion Blur Study

Time: 60 minutes. Equipment: A camera/AprilTag station with controlled lighting, a drivable chassis with the camera mounted, printed AprilTags at measured distances.

  1. With the robot stationary, record detection rate at three exposure settings and three distances. Nine data points.
  2. Repeat with the robot driving past the tag at moderate speed.
  3. Repeat under reduced lighting.
  4. Find the exposure setting with the best detection rate across all conditions, not just the easiest one.

Evidence of learning: A detection-rate matrix across exposure, distance, and motion, plus a single recommended locked exposure value with justification.

🔧 Exercise 9.2 — Calibration and Pose Error Test

Time: 75 minutes. Equipment: Camera, calibration board, tape measure, marked floor positions.

  1. Run a full calibration at your chosen resolution.
  2. Mark five floor positions at measured distances and angles from a fixed tag.
  3. Place the robot at each and record the vision-reported pose against the measured truth.
  4. Compute error at each position.
  5. Then deliberately use a calibration from a different resolution and repeat. Compare.

Evidence of learning: A pose-error table for both calibrations, an error-versus-distance plot, and a calibration record stating camera, resolution, date, and computed intrinsics.

🔧 Exercise 9.3 — Occlusion and Fallback Stress Test

Time: 45 minutes. Equipment: A robot running a vision-assisted autonomous routine.

Run the routine five times under each condition:

  • A: normal.
  • B: a person standing between the robot and the tag for the first two seconds.
  • C: camera unplugged entirely.
  • D: camera deliberately rotated 15° from its calibrated mounting.

Record what the robot does in each case.

Condition D is the most instructive, because it is the one that produces confidently wrong answers rather than obviously missing ones. A system that fails silently and plausibly is more dangerous than one that fails loudly.

Evidence of learning: A results table for all four conditions, and a written fallback specification: what the robot should do, for how long, before giving up on vision.

Further reading

Next Lesson

Everything built so far can work. Lesson 10 — Reliability, Cycle Time and Competition is about making it work every time, for two full days, and knowing that it will.