---
name: ai-for-builders-checks
description: 17 rules from the Noesa course "Using AI, day by day". For backend developers moving into a manager/integrator role — practical fluency to use and direct AI, not train models.
---

# Using AI, day by day — the rules

Use with: Claude Code or Claude (save as a skill), Cursor (save under .cursor/rules as .mdc), ChatGPT or any other assistant (paste the text below into custom instructions or a project's instructions).

17 rules, taken from the course at https://noesa.leafsoft.online/c/ai-for-builders

Each heading is one thing the course teaches. Most are checks to run on your own output before presenting it as done; a few are background you are expected to have. 10 also name a mistake models make by default, under "Watch for". "Wrong by default" lists 8 specific ones.

Apply these to the thing you are producing — the type, the schema, the query, the copy — not only to how you explain it. Where a rule names a field, a format or an identifier, that name belongs in the output.

## Set up your AI bench

Open Ari's scratch project and get one real model reply from Claude Code.

## See what an LLM predicts

Explain an LLM as a next-token predictor and connect that mechanism to both fluency and fabrication.

**Watch for:** If you asked an AI why it produced a particular answer, what could you mistake for evidence?

Its explanation can be another plausible continuation rather than a trace of how the answer was produced. Check the claim against supplied context, tool records, or external evidence; do not treat the model's self-description as provenance.

## Count the tokens

Estimate a request's token budget and explain how input and output tokens affect cost, speed, and context.

**Watch for:** If you asked an AI for the exact token cost of a request from the visible text alone, what could it miss?

It may estimate from familiar text patterns without the exact tokenizer, hidden instructions, tool schemas, history, or reserved output budget. Use the provider's tokenizer and measured usage metadata for the real count.

**Wrong by default:**
- The prompt is 1.4% of the input. The other 21,000 tokens — instructions, schemas, history, passages — are what you actually pay for and what fills the window. Budgeting the newest message is the error the day's failure table names first, and it is easy to make because the prompt is the only part you typed.
- 1,200 + 3,400 + 5,400 + 11,000 + 300 is exactly 21,300 — the reservation is on top, not inside. Providers report input and output separately, and treating them as one number is how you pack context until generation has nowhere left to go.
- 21,300 is an estimate of input. The invoice counts generated tokens too, and the provider's own usage report is the authoritative figure. An estimator is fine for capacity planning and wrong for money.

## Turn meaning into vectors

Compute cosine similarity over fixed embeddings and use the ranking to choose a relevant Ari passage.

**Watch for:** If you asked an AI to build semantic search, what control might a plausible implementation omit?

It may optimize for relevant-looking results and rank the whole corpus because your authorization model was not in its context. Scope candidates before or during retrieval, then test explicitly for cross-tenant leakage.

**Wrong by default:**
- Cosine measures the angle between two vectors. It is not a probability and not a confidence. Nothing in 0.994 says the passage contains the answer — only that its direction is close to the query's. Reading a similarity as a likelihood is how a retrieval score gets quoted as evidence.
- Retrieval and generation fail independently — today's lesson is explicit that a high similarity score cannot guarantee the model quotes the passage faithfully. Ranking chooses what is in front of the model; it does not constrain what comes out.

## Trace the assistant layers

Distinguish a base model, post-training, system instructions, and the product runtime around Claude.

## Steer output and verify limits

Write a bounded prompt contract and add verification for unsupported claims and nondeterministic output.

**Watch for:** If you asked the same AI to verify its own answer, why could the second response repeat the first mistake?

The second pass can follow the same context and common pattern, producing plausible agreement rather than independent evidence. Verify with allowed source IDs, deterministic rules, or a separate authoritative system.

## Run a model locally

Run one open-weight model with Ollama and decide whether local execution fits a workload.

## Ground answers with RAG

Assemble and evaluate a retrieve-then-answer flow over Ari's runbooks.

**Watch for:** If you asked an AI to add citations to a RAG answer, what could look correct while still being unsupported?

It may attach a real chunk ID because citations are expected, even when that passage does not entail the claim. Verify claim support and source authority, not only citation presence and format.

**Wrong by default:**
- That is chunk B, which is marked superseded. Retrieval did its job and generation quoted faithfully — the answer is perfectly grounded and currently wrong. This is stale authority, and it is the hardest RAG failure to see precisely because every other check passes.
- No chunk says anything about purging. Notice where it sits — last, after four sentences that were all traceable, in the same flat voice. A single ungrounded line inside an otherwise grounded answer is the shape to watch for, because the reader's trust was earned by the sentences above it.
- Also from the superseded chunk — the current policy parks the event instead. Two flaws from one stale source is normal: retrieving the wrong version does not go wrong once, it goes wrong everywhere that version is quoted.

## Let the model request tools

Implement the request-tool-result loop and keep execution authority outside the model.

**Watch for:** If you asked an AI to implement a tool from its JSON schema, what could schema-valid code still miss?

It may treat a well-formed identifier as authorized because the schema describes shape, not tenant ownership or user intent. Enforce identity, authorization, and approval in application code before execution.

## Connect capabilities with MCP

Explain MCP's host-client-server architecture and choose between local stdio and remote Streamable HTTP transport.

## Package capability as skills

Package Ari's incident-triage instructions and assets as a reusable, reviewable skill.

## Build an explicit agent loop

Describe and implement a bounded agent loop with tools, state, and a stop condition.

**Watch for:** If you asked an AI to write an agent loop, which stopping rule might the shortest working version omit?

It may cover the happy path and stop only when the model says it is done. Add hard iteration, time, token, repetition, cancellation, error, and success-verification stops in application code.

## Choose workflows and guard agents

Choose a fixed workflow or agent and specify guardrails against prompt injection and unsafe actions.

**Watch for:** If you asked an AI to defend an agent from prompt injection, why is a stronger instruction not enough?

Untrusted content still enters the same interpretation path and can influence the model. Contain the consequence with least-privilege tools, deterministic authorization, sandboxing, explicit approvals, and verification outside the prompt.

## Map Google's agent stack

Place Gemini, Vertex AI, ADK, and Agent Engine into model, development, and managed-runtime layers for a use case.

## Direct coding agents well

Give a coding agent a scoped task, review its evidence, and decide whether the change is ready.

**Watch for:** If a coding agent reports that every test passed, what important evidence could still be missing?

The existing suite may never exercise the requested behavior, while the agent's completion summary still sounds conclusive. Inspect the assertion, make it fail without the change, review the diff, and run the real acceptance path.

## Map the AI tool landscape

Classify an unfamiliar AI tool by layer, contract, state, authority, and evaluation burden.

## Scope Ari for production

Produce a build-ready AI feature plan that chooses prompt, RAG, workflow, or agent and defines cost, controls, and review.

**Watch for:** If you asked an AI to take an AI feature to production, what could a polished proposal hide behind one overall score?

It may default to a broad assistant and average unlike risks because your authority boundaries and release blockers are local business judgments. Keep the supported decision narrow and make security, isolation, and correct abstention non-negotiable gates.
