---
name: unit-testing-java-checks
description: 16 rules from the Noesa course "Unit testing with Java, day by day". For a Java developer who can read classes, methods and exceptions, and wants to design and review tests that deserve trust — especially when AI drafts the implementation or the test.
---

# Unit testing with Java, day by day — the rules

Use with: Claude Code or Claude (save as a skill), Cursor (save under .cursor/rules as .mdc), ChatGPT or any other assistant (paste the text below into custom instructions or a project's instructions).

16 rules, taken from the course at https://noesa.leafsoft.online/c/unit-testing-java

Each heading is one thing the course teaches. Most are checks to run on your own output before presenting it as done; a few are background you are expected to have. 8 also name a mistake models make by default, under "Watch for".

Apply these to the thing you are producing — the type, the schema, the query, the copy — not only to how you explain it. Where a rule names a field, a format or an identifier, that name belongs in the output.

## Build your test bench

Run one JUnit test from Maven and deliberately make it fail.

## Decide what one test proves

State the exact behavior a unit test proves and expose one behavior it leaves unprotected.

**Watch for:** If you asked an AI to test a Java delivery-fee rule, what would a plausible but weak suite likely miss?

It may mirror the common happy-path examples in the implementation or prompt and omit the exact boundary where `>` and `>=` differ. Ask which alternative implementation survives, then add the smallest example that distinguishes it.

## Name behavior before implementation

Structure a test around setup, action, and observation and give it a behavior-focused name.

**Watch for:** You ask an AI to strip the repetition out of a Java test. What can it hide, so the test later fails for the wrong reason?

It may move meaningful inputs into shared setup because shorter code looks cleaner. The tests still pass, but reviewers can no longer see which precondition distinguishes one behavior from another. Preserve visible differences even when a helper could hide them.

## Choose assertions that reveal defects

Select an assertion and observation that distinguish a correct result from a plausible wrong one.

**Watch for:** An AI-written test often checks only that something happened: not null, not empty, no error. How do you tell whether it would notice a wrong answer?

Propose a wrong result that still satisfies the assertion. An AI-generated test tends to mirror the shape of the code and assert activity rather than the rule. For the free-delivery rule, returning one cent is the useful challenge: it survives “less than 500” and “non-negative,” but it fails equality with zero. If such a case exists, strengthen the observation rather than adding unrelated checks.

## Test boundaries, not random examples

Derive boundary and equivalence-partition cases from a written rule.

## Test failures as behavior

Test exception type/message and prove the failing operation did not produce a forbidden result.

## Keep tests independent and deterministic

Remove shared mutable state, ordering, randomness, or environment dependence from a test.

**Watch for:** When asked to fix a flaky test, an AI adds a delay, retry, ordered execution, or broader tolerance. What must it explain before you accept the edit?

It must name the uncontrolled input and show how the change makes that input explicit. For these recorder tests, the input is a shared mutable list. Recreating it before each test removes suite history; ordering the tests preserves the defect. A quieter failure is not deterministic evidence.

## Parameterize rules without hiding cases

Convert repeated rule examples into a readable parameterized test while keeping boundary meaning visible.

## Choose the unit boundary

Decide whether behavior belongs in a unit test, broader test, both, or no direct test.

## Shape classes for unit testing

Refactor a mixed class into a decision core and side-effect shell with explicit dependencies and observable outcomes.

## Choose the lightest test double

Distinguish fake, stub, and spy and choose the least powerful double that answers the test question.

**Watch for:** You ask an AI to write a Java test with every dependency mocked. Why does the test end up checking which calls happened, instead of what the code actually produced?

Generated tests often follow the visible call structure because method names and collaborator calls are available in the code. Your requirement may care only about returned or stored outcomes. State the behavioral question first, then permit a spy only when an outbound interaction is the outcome you must prove.

## Test outcomes, not choreography

Reject interaction-heavy tests when a stable state or returned outcome proves the behavior.

**Watch for:** An AI writes tests by reading your refund method instead of the refund rule. What does it lock in that breaks the moment you tidy the code?

It can mirror the visible call sequence: repository lookup, save, then notifier invocation. Those assertions describe the current route and may agree with the same mistaken implementation. Give the AI explicit outcomes and review every interaction assertion against a stated requirement.

## Control time without sleeping

Inject `Clock` and test time-based rules without using the wall clock or `Thread.sleep`.

## Read coverage as a question

Use JaCoCo line and branch signals to ask for missing behavior without treating a percentage as quality.

## Review AI-generated tests

Audit generated tests for oracle copying, happy-path bias, implementation coupling, missing assertions, and tests that cannot fail.

**Watch for:** You hand an AI a fee method and ask for tests, without telling it the fee rules. What bug will the tests happily agree with?

It may treat the implementation as the specification, reproduce its conditions in the expected values, and favor ordinary inputs shown nearby. Supply the rule separately, inspect exact boundaries and invalid cases, and prove that a plausible defect makes each important test fail.

## Break the suite before trusting it

Complete a mutation-style capstone by inserting plausible defects, observing survivors, and strengthening the smallest useful tests.

**Watch for:** If you asked an AI to strengthen a Java suite after a surviving mutant, what durable mistake might it make?

It may generate several tests near the changed line without deciding which business behavior the mutation violates. Give it the rule and survivor, then require one defect-sensitive observation; you still judge whether the mutant matters and whether the unit boundary is appropriate.
