Alluvia field notes · 30 Sep 2026 · Python + pytest
Does your pytest regression test catch the original bug?
A test that passes after a fix may have passed before it, too. Run the same test against both revisions to find out whether it catches the regression you care about.
This matters when you write the test after correcting the code, including when a coding agent writes it. A green result shows agreement with the current behavior. It does not tell you whether the test would have detected the original mistake.
A boundary bug you can reproduce
Suppose an upload must be smaller than 8 MiB. This implementation incorrectly accepts exactly 8 MiB:
def accepts_upload(size):
return size <= 8 * 1024 * 1024The fix changes one comparison:
def accepts_upload(size):
return size < 8 * 1024 * 1024A useful regression test checks the rejected boundary and a nearby accepted value:
from upload_policy import accepts_upload
def test_upload_limit_is_exclusive():
limit = 8 * 1024 * 1024
assert accepts_upload(limit - 1)
assert not accepts_upload(limit)In our reproducible example, that unchanged test fails by assertion against the old implementation and passes against the fixed implementation. This is a supplied example, not a customer result. The 48-second recording, full transcript, and JSON evidence let you inspect the run.
Keep the test constant
The comparison is meaningful only if both revisions receive the same test. Changing the assertion between runs answers a different question. You also need the expected project dependencies and an interpreter that imports the code from the revision under test.
You can arrange those runs yourself. Alluvia Protect handles the temporary Git snapshots, freezes the test bytes, records each outcome, and exports the test with its evidence. It leaves your checkout on its current revision.
Run the check with Alluvia
Install the CLI with uv tool install alluvia. It requires Python 3.12 or later. Your project needs Git, two committed revisions, and its existing pytest environment. Save the applicable correction in correction-source.json:
{
"kind": "user_instruction",
"text": "Uploads must be smaller than 8 MiB; exactly 8 MiB must be rejected."
}From the project root, select the single test. Replace the commit placeholders and interpreter path with your project's values, and choose a new output directory:
alluvia checks verify test_upload_policy.py::test_upload_limit_is_exclusive \
--before BUGGY_COMMIT --after FIXED_COMMIT --project . \
--requirement "Reject exactly 8 MiB; accept smaller uploads." \
--source-file correction-source.json --python /path/to/project/python \
--output ./verified-upload-limitThe complete example includes repository setup and exact commands. In Claude Code, the Protect skill can find the revisions, write the test, and run verification from your correction.
Read the failure before trusting the label
A verified result requires an assertion failure on the old revision and a pass on the fixed one. Check that the failure is caused by the behavior you intended to protect. A test can fail for an unrelated assertion; the before/after result alone cannot establish that your requirement is right.
Import errors, missing dependencies, setup failures, skipped tests, and timeouts are inconclusive. A passing test on both revisions has not caught this bug. One successful check does not prove complete application correctness.
The export contains an ordinary pytest file, a readable explanation, and a JSON report with the requirement, source, commit IDs, test hash, and results. Review the test and commit it to your repository. It runs without Alluvia afterward.
Alluvia is free and MIT licensed. Protect needs no Alluvia account or extra model key. Project tests execute locally with your permissions; temporary snapshots are not a security sandbox.