- The short version
- How can I check whether Claude Code fixed a bug?
- Which four kinds of tests help Claude Code check its work?
- We ran it: Did property tests help Claude Code fix the planted bug?
- How can property tests catch bugs that example tests miss?
- How can I compare old and new code after a Claude Code change?
- FAQs
The short version
- Three example tests passed even though the function had a planted bug.
- When told to make the suite pass, Claude Code left the bug in place in three of three runs with those tests alone.
- With property tests in the folder, it fixed the bug in three of three runs given the same instruction.
When you ask an AI coding agent to make tests pass, its checks shape which bugs it can see and fix. In our October 5, 2026 test, passing examples gave Claude Code no failure to act on; a property test exposed the planted bug.
How can I check whether Claude Code fixed a bug?
Give Claude Code a check that fails when the bug remains, then examine the result after its edit. A passing test shows that the code handled what that test checked. It cannot, by itself, establish that every input works.
Which four kinds of tests help Claude Code check its work?
Addy Osmani, who works on Claude Code at Anthropic, wrote on X on October 5, 2026 that four checks are worth giving an agent: user-flow tests, properties, comparisons with an older version, and tests that run quickly and consistently.
Four checks for an agent
| Test | What it checks | What it gives the agent |
|---|---|---|
| End-to-end | Whether a user can complete a real flow | A result tied to visible behavior |
| Property-based | Whether a rule holds across generated inputs | Failures beyond the examples a person chose |
| Old against new | Whether replacement code changes results for the same inputs | Specific differences to investigate |
| Fast and repeatable | Whether a check returns promptly and behaves consistently | A result it can use while revising the code |
He closed with: “Example tests say what should happen; properties say what must not.”
We ran it: Did property tests help Claude Code fix the planted bug?
Yes. With property tests in the folder, Claude Code fixed the bug in all three runs where it was told to make the suite pass.
On October 5, 2026, we ran the test on a Linux server with this setup:
- Claude Code 2.1.289 with Claude Sonnet 5.5
- Python 3.12.3
- pytest 9.1.1
- Hypothesis 6.168.4
The planted bug made the function shorten a time range when another range sat entirely inside it.
All three example tests passed on the buggy code. A property requiring every input range to remain covered failed and produced the smallest failing input it found, [(0, 2), (1, 1)]. The old and new functions returned different results on 541 of 1,000 random inputs. The full suite took about 2.2 seconds, and two runs reported the same failing input.
Bug fixes across twelve agent runs
| Prompt | Tests in the folder | Runs that fixed the bug |
|---|---|---|
| Make the suite pass | Three example tests | 0 of 3 |
| Make the suite pass | Three example and two property tests | 3 of 3 |
| Check that the function is correct | Three example tests | 3 of 3 |
| Check that the function is correct | Three example and two property tests | 3 of 3 |
The second instruction changed the outcome: Claude Code found the bug by reading the function in all three runs with example tests alone. This was one small function, one planted bug, one model, and three runs per setup.
How can property tests catch bugs that example tests miss?
They check a rule across generated inputs instead of checking only values someone picked in advance. In our test, the examples covered overlapping, touching, and separate ranges, but none put one range inside another.
End-to-end tests check a different kind of failure. The Playwright testing guide recommends checking what users see and do in the rendered application, such as whether an action produces the expected visible result.
For property-based testing, you describe a rule that every valid result must satisfy. The Hypothesis documentation says the library selects inputs from the range you define, including edge cases you may not have chosen yourself. Our rule was that merging ranges must never leave an original range uncovered.
Comparing old and new code checks both versions on identical inputs. In our comparison, 541 of 1,000 random inputs produced different answers, giving a concrete set of disagreements to inspect.
Fast, repeatable checks make another attempt useful. His post cautions that an unreliable check may lead an agent to rerun it instead of fixing the problem. Hypothesis offers derandomize=True for repeatable inputs; our full suite took about 2.2 seconds and exposed the same input twice.
How can I compare old and new code after a Claude Code change?
Keep the older function as a reference, feed both versions the same generated inputs, and inspect any different outputs. Our comparison found 541 disagreements among 1,000 inputs before the fix. After the runs in which Claude Code repaired the function, our comparison found none among those 1,000 inputs.
Anthropic’s Claude Code guide has a section called “Give Claude a way to verify its work.” Its advice is direct: “Give Claude a check it can run: tests, a build, a screenshot to compare.”
Give your agent a command that returns pass or fail, and ask it to run the check after editing. In the project folder from our test, the property-test command was:
python -m pytest -q test_properties.pyAdd a property test for the function you trust least. When replacing code, retain the older version long enough to compare both versions on the same random inputs.
FAQs
Which Claude Code version did NerdsChalk test?
We tested Claude Code 2.1.289 with Claude Sonnet 5.5 on a Linux server on October 5, 2026. The test used Python 3.12.3, pytest 9.1.1, and Hypothesis 6.168.4.
What input exposed the planted bug?
Hypothesis reported [(0, 2), (1, 1)]. The buggy function returned [(0, 1)], which no longer covered the original range (0, 2).
Did Claude Code find the bug when asked to check the function itself?
Yes. When asked to check whether the function was correct, Claude Code found and fixed the bug in three of three runs with example tests alone.
What command runs the property tests?
From the folder containing test_properties.py, run python -m pytest -q test_properties.py. In our test, it ran two properties and reported one failure on the buggy function.
How long did the full test suite take?
The three example tests and two property tests took about 2.2 seconds together in our October 5, 2026 run.
How limited was NerdsChalk’s Claude Code test?
It used one small function, one planted bug, one model, and three runs for each of four setups. The results do not measure how often an agent finds bugs in larger projects.







Leave a Reply