AI News technology

What Is Property-Based Testing for AI Coding Agents?

Property-based testing checks a rule against generated inputs, helping AI coding agents catch bugs that a few example tests can miss.

Property-based testing checks a rule against generated inputs, helping AI coding agents catch bugs that a few example tests can miss.

0 Comments

HypothesisPython library
fast-checkJavaScript/TypeScript
Within 18 casesBug found

The short version

  • Property-based testing checks a rule that should hold for every valid input against generated inputs.
  • Hypothesis is a Python library for this method; fast-check serves JavaScript and TypeScript.
  • Example tests check inputs you choose. Property tests can expose combinations you did not think to write.
  • In our small experiment, three example tests passed while a property test found a planted bug.

Property-based testing starts with a rule that must hold for every valid input, then a library generates inputs to try to break it. The library can try hundreds or thousands of cases, but it may stop sooner when it finds a failure. That gives AI coding agents a broader check to run before they treat passing tests as a sign that their work is done.

What is property-based testing, and how is it different from example tests?

A property-based test describes a behavior that should remain true across valid inputs. An example test supplies particular inputs and checks their expected results. Both can be useful: a familiar example makes an intended result clear, while a property can probe cases you did not choose by hand.

What each kind of test checks

Kind of test What you write What it can find
Example test Specific inputs and expected outputs Errors triggered by those chosen inputs
Property-based test A rule about the relationship between inputs and results Errors triggered by generated combinations and edge cases

For a function that merges time ranges, an example might check one pair of overlapping ranges. A property can check that every input range remains covered by the result, whatever valid ranges the library generates.

How do property-based tests generate inputs and shrink failures?

Property-based libraries generate inputs from a description of valid values, run the same rule against many of them, and report an input that breaks it. Hypothesis does this for Python; fast-check does it for JavaScript and TypeScript. Both can reduce a failing input to a smaller case that still fails, making the error easier to inspect.

Generation samples inputs rather than checking every possible value. You decide what values are valid and what rule the result must obey; the library tries cases within those bounds.

What did nerdschalk’s Hypothesis test catch that the example tests missed?

We ran it

We ran Python 3.12, pytest, and Hypothesis 6.168 on a Linux server on October 5, 2026. The project had a small function that merges time ranges and one planted bug: merging a range inside a larger one could shorten the larger range. All three example tests passed.

This is the property test from test_properties.py that checked whether every input range remained covered:

from hypothesis import given, settings, strategies as st

from intervals import merge_intervals

ranges = st.lists(
    st.tuples(st.integers(0, 50), st.integers(0, 50)).map(lambda p: (min(p), max(p))),
    max_size=8,
)


@settings(max_examples=1000, derandomize=True, deadline=None)
@given(ranges)
def test_no_input_range_is_ever_lost(intervals):
    merged = merge_intervals(intervals)
    for start, end in intervals:
        assert any(m_start <= start and end <= m_end for m_start, m_end in merged)

The property failed within the first 18 generated cases. Hypothesis shrank the failure to [(0, 2), (1, 1)]: the buggy function returned [(0, 1)], which no longer covered the first range.

These three lines from the pytest output show the failing input, the assertion, and the result of the two property tests:

intervals = [(0, 2), (1, 1)]
E           assert False
1 failed, 1 passed in 2.01s

Comparing the buggy function with an older version on 1,000 random inputs produced 541 different answers. The full suite of three example tests and two property tests took about 2.2 seconds, and a second run reported the same failing input.

Why do AI coding agents need property-based tests?

An AI coding agent can stop when the checks it was asked to run pass, so a suite with only a few examples may let a wrong answer through. Anthropic’s Claude Code guide puts the stopping behavior plainly: “Claude stops when the work looks done.” A property test gives the agent another pass-or-fail signal, including for inputs nobody wrote as examples.

In our log, when told to make the suite pass, Claude Code left the bug in place in 3 of 3 runs with example tests only and fixed it in 3 of 3 runs when the property tests were present.

What kinds of properties can I test first?

Start with a rule you can state independently of a particular answer. For data processing code, useful candidates are:

  • No input data is lost or invented in the result.
  • An output that should be sorted stays sorted.
  • Converting a value to another form and back returns the original input.
  • An old implementation and its replacement agree on the same inputs.

The range test checked whether any input range was lost. Our comparison of old and new functions showed how checking for agreement can reveal differences across generated inputs.

What are the limits?

A property test is only as useful as the rule you state and the inputs you allow it to generate. Our experiment involved one small function with one planted bug; its results do not measure how often agents find bugs in larger projects.

What to do now
  1. Keep your example tests, then identify a rule that should hold across valid inputs.
  2. Describe those inputs with Hypothesis for Python or fast-check for JavaScript and TypeScript, and run the property alongside your existing tests.
  3. When a failure appears, inspect the reduced input and keep it as a case your code must handle.

FAQs

Is Hypothesis for Python?

Yes. Hypothesis is a property-based testing library for Python. You describe valid inputs with strategies and write a rule for the generated cases to check.

Does fast-check work with JavaScript and TypeScript test frameworks?

Yes. fast-check supports property-based tests in JavaScript and TypeScript and can be used with major testing frameworks for those languages.

How many generated cases did Hypothesis need to find the bug?

Hypothesis found a failing input within its first 18 generated cases: 11 passed, 5 failed, and 2 were invalid.

What was the smallest failing input?

Hypothesis reduced the failure to [(0, 2), (1, 1)], the smallest failing input it found. The buggy merge returned [(0, 1)], losing coverage of (0, 2).

How often did the old and new functions disagree?

They returned different answers on 541 of the 1,000 random inputs in our comparison.

What are the limits of nerdschalk’s experiment?

It used one small function and one planted bug. The result shows what these checks caught in that setup, not how often property tests or AI coding agents will find bugs in other projects.

 

 

Leave a Reply

Your email address will not be published. Required fields are marked *