To make your agent produce high-quality output, equip it with a way to check its work. Tests are part of that.
Here are some of the ones I think are worth investing in.
- Write end-to-end tests that simulate real user flows and act as the ground truth
- Use property-based testing to specify what should never happen and generate thousands of test cases that attempt to break those constraints
- If replacing a system, compare the old and new versions using randomly selected inputs
- Require tests to run quickly and deterministically, which reinforces the agent to retry when the loop is slow or unreliable (the agent may learn to retry, not fix the problem)
Example tests say what should happen; properties say what must not.
