· 3 min read
Stop Vibe-Checking Prompts: Evals Are Your Unit Tests Now
Tweaking system prompts at 2 AM without a benchmark harness is just superstitious typing. Here is how to build deterministic eval suites that keep AI pipelines honest.
Here is the dirty secret of 80% of teams building with LLMs today:
An engineer sits in front of a web playground. They notice the model failed on a customer request about invoice reconciliation. They spend forty-five minutes carefully adjusting the system prompt:
“You are an expert accountant. You MUST NEVER format dates as MM/DD/YYYY unless explicitly asked. Pay special attention to VAT line items…”
They test it against two sample inputs. Both pass! They commit the prompt directly to main, deploy to production, and celebrate.
Three days later, customer support gets flooded with tickets because that prompt tweak silently degraded the model’s ability to parse European currency symbols by 22%.
In traditional software, we have a name for changing runtime logic without testing regressions: gross negligence. In AI engineering, people somehow call it “prompt crafting.”
It’s time to stop vibe-checking. Evals are your unit tests now.
┌────────────────┐ ┌─────────────────┐ ┌────────────────┐
│ Golden Dataset │ ────► │ LLM / Agent Run │ ────► │ Invariant Eval │
│ (100+ cases) │ │ Under Test │ │ (Assertion/LLM)│
└────────────────┘ └─────────────────┘ └───────┬────────┘
│
▼
[ Pass / Fail Score ]
[ Regression Delta ]
How to Build a Minimal Eval Suite in One Afternoon
You do not need a heavyweight enterprise observability suite with an annual contract to run rigorous evals. You can start with a simple script inside your existing test runner.
1. Build a Golden Dataset
Every time an LLM in your system produces an unacceptable response, an edge-case failure, or a brilliant answer, snapshot it. Store it in a version-controlled JSON or SQLite file:
[
{
"id": "tax-parse-014",
"input": "Invoice total €1,450.00 including 19% German MwSt.",
"expected": {
"total_cents": 145000,
"currency": "EUR",
"tax_cents": 23151,
"tax_rate_percent": 19
}
}
]
This dataset is your moat. As your application evolves, this dataset grows.
2. Tier Your Graders: Deterministic First, Semantic Second
Don’t use expensive reasoning models to grade things that a regex or a JSON schema can verify for free.
- Tier 1 (Deterministic): Did the output conform to the exact Zod schema? Is the JSON valid? Did it stay within latency and token budgets?
- Tier 2 (Assertion Rules): Did the response contain prohibited terms? Did calculated numbers match expected math?
- Tier 3 (LLM-as-Judge): For open-ended synthesis or tone, use a fast, calibrated model with an explicit rubric:
export async function gradeToneAndCompleteness(output: string, rubric: string): Promise<{ pass: boolean; reason: string }> {
const result = await judgeModel.generateObject({
schema: z.object({
pass: z.boolean(),
score: z.number().min(1).max(5),
reason: z.string(),
}),
prompt: `Evaluate the following output against the rubric.\n\nRubric: ${rubric}\n\nOutput: ${output}`,
});
return { pass: result.score >= 4, reason: result.reason };
}
3. Wire Evals Into Your CI Pipeline
Every time a pull request touches an agent prompt, tool definition, or system instruction, GitHub Actions runs the eval suite alongside your standard unit tests:
name: Model Evals
on: [pull_request]
jobs:
eval:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- run: npm ci
- run: npm run eval:fast
env:
OPENAI_API_KEY: ${{ secrets.EVAL_API_KEY }}
If accuracy drops below 95%, or if latency regresses by more than 200ms, the PR cannot be merged.
The Payoff
When you treat AI systems like software engineering instead of divination, the anxiety evaporates.
You can swap underlying foundation models without fear. You can optimize prompts with automated hill-climbing algorithms. You can sleep peacefully knowing that a midnight tweak didn’t quietly blow up your core business logic.
Stop guessing. Measure the diff.