AI Tech Magic
Back home

Spec + Guardrails Sprint

$7,50010 business days

The missing foundation your team will implement against — specs, evals, guardrails, and observability, delivered in two weeks.

Two weeks of focused delivery. We include everything in the Pilot Pre-Mortem, then build the artifacts your engineers actually implement against: system prompts, tool contracts, an eval harness with seed cases, guardrail configs, and an observability blueprint. Your team ships with something to point at.

What you get

Artifacts your team will use.

  1. Everything in the Pilot Pre-Mortem

    The full 15-20 page audit and 90-day roadmap are the starting point. We spend the first three days reproducing the Pre-Mortem, then use the remaining seven days to build what the audit recommends.

  2. Spec-engineering package

    System prompts written as documents (not code comments), tool contracts as JSON Schema, decision trees for ambiguous inputs, and an edge-case matrix. Delivered as a versioned repo your team owns from day one.

  3. Portable eval suite with 30+ seed test cases

    At least 30 seed test cases split across golden path, edge cases, and adversarial inputs, written as plain YAML in a repo you own. A thin runner adapter wires them to whatever you already use — Braintrust, LangSmith, promptfoo, or bare pytest in CI — so the tests outlive the tool. Configured to run on every prompt change.

  4. Guardrail configuration

    Input validation schemas, output constraints (JSON mode, structured output), cost and rate limits, safety filters. Written as configuration rather than one-off code so your team can extend them without our involvement.

  5. Logging and observability blueprint

    A written document naming exactly what to instrument, what dashboards to build, and what alerts to set. Compatible with whatever observability stack you already use (Datadog, Honeycomb, LangSmith, Braintrust).

  6. Two working sessions with your team

    Two 60-minute live sessions during the sprint. First one is mid-sprint (Day 5) to review specs and align on eval strategy. Second one is at handoff (Day 10) to walk through everything and hand off ownership.

  7. Two weeks of async follow-up

    Shared Slack channel for two weeks after handoff. Use it as your team starts implementing against the artifacts.

What we touch

Concrete, technical, in-scope.

  • system prompts
  • tool contracts
  • eval harness
  • seed test cases
  • guardrail configs
  • input validation
  • output constraints
  • cost limits
  • observability plan
  • handoff docs

Sample deliverable

Sample artifact bundle handed to your team

Every artifact is versioned in a repo you own, structured so your engineers can extend it without asking us. Names below are illustrative; actual filenames match your stack conventions.

  1. 01docs/spec-v1.md
  2. 02docs/edge-cases.md
  3. 03prompts/system.txt
  4. 04prompts/tool-contracts.schema.json
  5. 05evals/seed-cases.jsonl (30+ cases)
  6. 06evals/harness.config.ts
  7. 07guardrails/input-schema.json
  8. 08guardrails/output-constraints.json
  9. 09observability/blueprint.md
  10. 10README.md (ownership + how to extend)

Timeline

What happens when.

  1. Days 1-3
    Discovery and spec writing

    Same discovery flow as the Pre-Mortem, focused on producing the spec-engineering package. By end of Day 3, first draft of the system prompt, tool contracts, decision trees, and edge-case matrix.

  2. Days 4-7
    Eval and guardrail implementation

    Eval harness set up in your chosen tool, 30+ seed cases written, guardrail configuration authored, and the observability blueprint drafted. Mid-sprint working session on Day 5.

  3. Days 8-10
    Team handoff and iteration

    Live walkthrough of the artifact bundle, ownership transfer, and two days of iteration on feedback. Final commit hash tagged in the repo. Async follow-up window opens.

Not in scope

What this isn't.

Setting expectations up front so you can send us away quickly if it's the wrong fit.

  • We don't deploy to production. You own the deploy pipeline; we hand off artifacts your team ships.
  • We don't take on-call. The observability blueprint tells your team what to watch, but we're not the ones paged.
  • We don't build features. Spec, evals, and guardrails, not user-facing functionality.
  • We don't rewrite your agent's core logic. We audit it and give you the foundation to iterate on it. If a rewrite is warranted, that's a separate engagement.
  • We don't work in more than one repo per sprint. Multi-repo scope means multi-sprint sequencing.

Questions

Specific to this engagement.

Do you write code in this tier?

Yes, but scoped. Eval harness scaffolding, tool contract schemas, guardrail configs, and reference implementations where they clarify a spec. We don't ship your feature backlog or take PRs against product code.

What if we already have an eval harness?

Better. We audit what you have, add seed cases where coverage is thin, and integrate our new specs with your existing tooling. You get more depth for the same price because we skip the setup.

Can we split the sprint over more than two calendar weeks?

Yes, up to four weeks total with the same 10 billed days. Useful if your team has a launch mid-sprint or a holiday week to work around. Discuss on the Fit Call.

What if we find we need the Retainer after this?

About half of Sprint clients continue on Retainer. If it makes sense, we transition directly on Day 11. The Sprint artifacts become the starting point for ongoing spec and eval ownership.

Something else? info@aitechmagic.com

Give your team the spec they're missing.

20 minutes on a call tells you if the Sprint is the right fit. Same booking link as the other tiers.

Book your Pilot Fit Call

20 minutes · No slide deck · No pressure