AI News technology

How to Run a Property Test After a Claude Code Edit

Give Claude Code a test that exposes the bug, then check its result: property tests led it to fix our planted bug in all three suite-pass runs.

Give Claude Code a test that exposes the bug, then check its result: property tests led it to fix our planted bug in all three suite-pass runs.

0 Comments

Claude Code 2.1.289Agent
0 of 3 fixesTold to pass the suite, examples only
3 of 3 fixesSame prompt, with property tests
541 of 1,000 differedOld vs. new

The short version

  • Three example tests passed even though the function had a planted bug.
  • When told to make the suite pass, Claude Code left the bug in place in three of three runs with those tests alone.
  • With property tests in the folder, it fixed the bug in three of three runs given the same instruction.

When you ask an AI coding agent to make tests pass, its checks shape which bugs it can see and fix. In our October 5, 2026 test, passing examples gave Claude Code no failure to act on; a property test exposed the planted bug.

How can I check whether Claude Code fixed a bug?

Give Claude Code a check that fails when the bug remains, then examine the result after its edit. A passing test shows that the code handled what that test checked. It cannot, by itself, establish that every input works.

Which four kinds of tests help Claude Code check its work?

Addy Osmani, who works on Claude Code at Anthropic, wrote on X on October 5, 2026 that four checks are worth giving an agent: user-flow tests, properties, comparisons with an older version, and tests that run quickly and consistently.

Four checks for an agent

Test What it checks What it gives the agent
End-to-end Whether a user can complete a real flow A result tied to visible behavior
Property-based Whether a rule holds across generated inputs Failures beyond the examples a person chose
Old against new Whether replacement code changes results for the same inputs Specific differences to investigate
Fast and repeatable Whether a check returns promptly and behaves consistently A result it can use while revising the code

He closed with: “Example tests say what should happen; properties say what must not.”

We ran it: Did property tests help Claude Code fix the planted bug?

Yes. With property tests in the folder, Claude Code fixed the bug in all three runs where it was told to make the suite pass.

On October 5, 2026, we ran the test on a Linux server with this setup:

  • Claude Code 2.1.289 with Claude Sonnet 5.5
  • Python 3.12.3
  • pytest 9.1.1
  • Hypothesis 6.168.4

The planted bug made the function shorten a time range when another range sat entirely inside it.

All three example tests passed on the buggy code. A property requiring every input range to remain covered failed and produced the smallest failing input it found, [(0, 2), (1, 1)]. The old and new functions returned different results on 541 of 1,000 random inputs. The full suite took about 2.2 seconds, and two runs reported the same failing input.

Bug fixes across twelve agent runs

Prompt Tests in the folder Runs that fixed the bug
Make the suite pass Three example tests 0 of 3
Make the suite pass Three example and two property tests 3 of 3
Check that the function is correct Three example tests 3 of 3
Check that the function is correct Three example and two property tests 3 of 3

The second instruction changed the outcome: Claude Code found the bug by reading the function in all three runs with example tests alone. This was one small function, one planted bug, one model, and three runs per setup.

How can property tests catch bugs that example tests miss?

They check a rule across generated inputs instead of checking only values someone picked in advance. In our test, the examples covered overlapping, touching, and separate ranges, but none put one range inside another.

End-to-end tests check a different kind of failure. The Playwright testing guide recommends checking what users see and do in the rendered application, such as whether an action produces the expected visible result.

For property-based testing, you describe a rule that every valid result must satisfy. The Hypothesis documentation says the library selects inputs from the range you define, including edge cases you may not have chosen yourself. Our rule was that merging ranges must never leave an original range uncovered.

Comparing old and new code checks both versions on identical inputs. In our comparison, 541 of 1,000 random inputs produced different answers, giving a concrete set of disagreements to inspect.

Fast, repeatable checks make another attempt useful. His post cautions that an unreliable check may lead an agent to rerun it instead of fixing the problem. Hypothesis offers derandomize=True for repeatable inputs; our full suite took about 2.2 seconds and exposed the same input twice.

How can I compare old and new code after a Claude Code change?

Keep the older function as a reference, feed both versions the same generated inputs, and inspect any different outputs. Our comparison found 541 disagreements among 1,000 inputs before the fix. After the runs in which Claude Code repaired the function, our comparison found none among those 1,000 inputs.

What to do now

Anthropic’s Claude Code guide has a section called “Give Claude a way to verify its work.” Its advice is direct: “Give Claude a check it can run: tests, a build, a screenshot to compare.”

Give your agent a command that returns pass or fail, and ask it to run the check after editing. In the project folder from our test, the property-test command was:

python -m pytest -q test_properties.py

Add a property test for the function you trust least. When replacing code, retain the older version long enough to compare both versions on the same random inputs.

FAQs

Which Claude Code version did NerdsChalk test?

We tested Claude Code 2.1.289 with Claude Sonnet 5.5 on a Linux server on October 5, 2026. The test used Python 3.12.3, pytest 9.1.1, and Hypothesis 6.168.4.

What input exposed the planted bug?

Hypothesis reported [(0, 2), (1, 1)]. The buggy function returned [(0, 1)], which no longer covered the original range (0, 2).

Did Claude Code find the bug when asked to check the function itself?

Yes. When asked to check whether the function was correct, Claude Code found and fixed the bug in three of three runs with example tests alone.

What command runs the property tests?

From the folder containing test_properties.py, run python -m pytest -q test_properties.py. In our test, it ran two properties and reported one failure on the buggy function.

How long did the full test suite take?

The three example tests and two property tests took about 2.2 seconds together in our October 5, 2026 run.

How limited was NerdsChalk’s Claude Code test?

It used one small function, one planted bug, one model, and three runs for each of four setups. The results do not measure how often an agent finds bugs in larger projects.

 

Leave a Reply

Your email address will not be published. Required fields are marked *