Series: FRC Technical Foundations Unit: 4 — Software and Perception Session length: ~3 hours Prerequisites: Lesson 8
Lesson 8 ended on a problem: odometry drifts, and dead reckoning has no way to recover. Vision solves it by providing an absolute position reference — something fixed in the world whose location the robot knows.
That is the primary job. Vision in FRC is not general scene understanding. It is: find a known landmark, compute where I am relative to it, correct my belief about my position.
The landmarks are AprilTags — square fiducial markers placed at surveyed positions around the field. FIRST publishes the exact field coordinates of every tag each season. Because the tag’s position is known and the camera can measure its apparent size, orientation, and position in frame, you can compute the camera’s pose relative to the tag, and therefore the robot’s pose on the field.
Before any of the software works, the physical camera setup has to be right. Four parameters interact, and students consistently discover this the hard way.
Resolution. Higher resolution detects tags at greater distance but costs processing time per frame. Most teams run lower resolution than they expect — 640×480 is often plenty for tags at competition distances.
Frame rate. More frames per second means fresher data and more correction opportunities, but higher resolution reduces achievable frame rate. This is a direct trade, and you make it consciously.
Exposure. The single most important setting and the most commonly wrong one. Long exposure gathers more light but blurs while the robot is moving, and a blurred tag does not detect. Short exposure freezes motion but needs good lighting.
The counter-intuitive correct answer for AprilTag detection is almost always: set exposure low, and lock it. A tag detected in a dark, sharp image is worth far more than a tag lost in a bright, smeared one. Auto-exposure is actively harmful, because it changes settings as the robot turns toward a bright window, and your detection rate changes with it.
Mounting rigidity. If the camera mount vibrates, every pose estimate is noisy, and no amount of software filtering fully recovers it. This is Lesson 2’s content reappearing: the camera mount is a structural component with a tight requirement.
Rule 9.1 — Lock exposure low, lock gain, and mount the camera like it matters. Most “our vision is unreliable” problems are one of these three.
A camera lens distorts. Straight lines bow, and the amount of bowing increases toward the edges of frame. Calibration measures this distortion so the software can undo it.
Calibration produces two things:
Without calibration, tag pose estimates are systematically wrong, and the error grows toward the image edges. With calibration, they are accurate to a few centimetres at typical ranges.
Calibration is done by showing the camera a known pattern — a chessboard or a Charuco board — from many angles. Both PhotonVision and Limelight include calibration workflows.
Two rules that matter:
You also need the camera-to-robot transform: where the camera sits relative to the robot’s centre, in three dimensions plus rotation. Measure this carefully. A 10 mm error in the measured camera position becomes a 10 mm error in every pose estimate forever.
Two mature options dominate FRC.
PhotonVision is free, open-source, and runs on a coprocessor you supply — a Raspberry Pi, an Orange Pi, or a mini PC. More setup work, more control, no hardware cost beyond the coprocessor and camera.
Limelight is a purpose-built camera and coprocessor in one unit with a polished configuration interface. Costs money, saves time.
Both provide WPILib libraries, both do AprilTag detection and pose estimation, and both integrate with SwerveDrivePoseEstimator. The choice is about your team’s resources, not capability.
The correct architecture is not “use vision instead of odometry.” It is to fuse both.
WPILib’s pose estimators accept odometry updates continuously and vision measurements opportunistically, weighting each by a standard deviation representing how much you trust it. Odometry is smooth but drifts; vision is absolute but noisy and intermittent. Fused, you get a pose estimate that is both smooth and bounded.
Getting the trust weighting right is the real work:
Rule 9.2 — Every vision measurement gets sanity-checked before it is trusted. One accepted bad pose can throw your estimate off and take seconds to recover, in the middle of an autonomous routine.
Every vision measurement describes a moment in the past. Capture, transfer, processing, and network transmission each add delay — typically 20–80 ms in total.
At 4 m/s, 60 ms of latency means the robot has travelled 24 cm since the image was taken. If you apply that measurement as if it were current, you are correcting to where you were, not where you are.
The solution is timestamped measurements. Both PhotonVision and Limelight report when a frame was captured, and WPILib’s pose estimator accepts a timestamp and inserts the measurement at the correct point in its history buffer. Use the timestamped API. This is not an optimisation; it is the difference between vision helping and vision oscillating.
Vision will fail. A camera will be knocked out of alignment by a collision. A coprocessor will fail to boot. Another robot will park directly between yours and every visible tag. Lighting at one venue will be unlike anywhere you have tested.
The question is not whether this happens but what your robot does when it does.
Bad: the command waits for a tag that never arrives, and the robot sits motionless for the remaining ten seconds of autonomous.
Good: the command has a timeout, falls back to odometry-only navigation, completes a reduced version of the routine, and reports clearly on the dashboard that it ran in degraded mode.
Rule 9.3 — Every vision-dependent command needs a timeout and a non-vision fallback. Design the degraded mode deliberately, and test it by unplugging the camera.
This is also an engineering-ethics point worth making explicitly with students. A system that behaves unpredictably when its sensor fails is badly designed regardless of how well it performs when everything works. Documenting your system’s limitations — its false positive rate, the conditions where it is unreliable, what it does when it cannot see — is part of the engineering, not an afterthought.
[ANECDOTE SLOT] — A vision system failing at competition in a way you did not anticipate. Lighting, occlusion, a bumped camera, or a coprocessor that did not boot. The recovery is the interesting part.
Time: 60 minutes. Equipment: A camera/AprilTag station with controlled lighting, a drivable chassis with the camera mounted, printed AprilTags at measured distances.
Evidence of learning: A detection-rate matrix across exposure, distance, and motion, plus a single recommended locked exposure value with justification.
Time: 75 minutes. Equipment: Camera, calibration board, tape measure, marked floor positions.
Evidence of learning: A pose-error table for both calibrations, an error-versus-distance plot, and a calibration record stating camera, resolution, date, and computed intrinsics.
Time: 45 minutes. Equipment: A robot running a vision-assisted autonomous routine.
Run the routine five times under each condition:
Record what the robot does in each case.
Condition D is the most instructive, because it is the one that produces confidently wrong answers rather than obviously missing ones. A system that fails silently and plausibly is more dangerous than one that fails loudly.
Evidence of learning: A results table for all four conditions, and a written fallback specification: what the robot should do, for how long, before giving up on vision.
Everything built so far can work. Lesson 10 — Reliability, Cycle Time and Competition is about making it work every time, for two full days, and knowing that it will.



