Pilot Pre-Mortem
A written diagnosis of one stalled or pre-launch AI agent pilot, plus the 90-day roadmap to fix it.
You send us your repo, prompts, and eval data. We spend a focused week auditing where your agent will break in production and why. You get a 15-20 page written report with prioritized fixes, so your team can move without guessing.
What you get
Artifacts your team will use.
60-minute discovery call
We start with one call to understand the pilot: what it's supposed to do, what's stuck, and what your team has already tried. Async follow-up questions the same day.
Async access to code, prompts, and evals
Read-only access to the repo, any prompt files or datasets you can share, and whatever eval data exists. If nothing exists yet, we say so and adjust the audit to reflect that gap.
Failure-mode analysis across seven axes
We walk the pilot against a fixed rubric: spec clarity, eval coverage, tool boundaries, latency, cost per interaction, safety guardrails, and observability. Each axis gets a written finding with evidence and impact.
15-20 page written report
The core deliverable. Executive summary at the top, findings in the middle, prioritized 90-day roadmap at the end. Written to be forwardable to your CTO or board without additional context.
Prioritized 90-day roadmap
Every recommendation scored on effort (in engineer-days) and impact (measurable outcome, not vibes). Sequenced so your team knows exactly what to do first, second, third.
30-minute walkthrough call
We present the findings live, answer questions, and negotiate priorities. Recording provided so anyone on your team can catch up later.
Two weeks of async Slack follow-up
Shared channel or DMs for two weeks after handoff. Use it for clarifications, sanity checks on decisions, or rubber-ducking as your team starts implementing.
What we look at
Concrete, technical, in-scope.
- spec clarity
- eval coverage
- tool boundaries
- latency
- cost per interaction
- guardrails
- observability
- failure modes
Sample deliverable
Sample table of contents from a recent audit
Structure varies by pilot, but the anatomy is consistent. Every report opens with an executive summary short enough for a non-technical stakeholder, and closes with a roadmap that engineers can start on Monday.
- 01Executive summary (1 page, forwardable)
- 02Pilot overview and the question we were asked
- 03Failure modes ranked by production impact
- 04Spec audit findings
- 05Eval coverage gaps and recommended test cases
- 06Cost and latency analysis (current vs. projected at scale)
- 07Guardrail recommendations by category
- 0890-day roadmap with effort and impact scoring
Timeline
What happens when.
- Day 1 (Mon)Kickoff and access
60-min discovery call, access handoff for repo and eval data, and a shared Slack channel opens. We confirm which pilot we're auditing if you have more than one.
- Days 2-3 (Tue-Wed)Deep read
Async work through the code, prompts, and existing evals. We instrument what we can, replay traces where available, and log open questions.
- Day 4 (Thu)Findings drafted
First draft of all seven axes and the ranked failure-mode list. Internal review, then we send you the top three findings for a temperature check.
- Day 5 (Fri)Report delivered and walkthrough
Final PDF report delivered by end of day, followed by the 30-min walkthrough call. Recording and Slack follow-up window start immediately after.
Not in scope
What this isn't.
Setting expectations up front so you can send us away quickly if it's the wrong fit.
- We don't write production code. This is diagnosis, not implementation.
- We don't manage your engineering team or run standups.
- We don't do vendor selection or procurement. If the report recommends a new tool, the sourcing is on you.
- We don't cover more than one pilot in a single engagement. If you have three stuck pilots, we recommend three sequenced pre-mortems or the Sprint tier instead.
Questions
Specific to this engagement.
What if we don't have evals yet?
That's often the finding. The report will reflect the eval gap explicitly and the roadmap will include an eval-harness setup as an early priority. We still audit the seven axes; the eval axis just gets a longer section.
Do you need our production data?
No. Read access to code and prompts, plus whatever trace or eval data you can share without security review, is enough. If your pilot is fully pre-launch, we work off the design and prompts alone.
Can you audit more than one pilot in a week?
No. Every axis is examined in enough depth to produce evidence-backed findings, and that takes the full five days per pilot. Multiple pilots means sequenced Pre-Mortems or the Sprint tier.
What happens after the two-week follow-up window closes?
If you want ongoing support, the natural next step is the Sprint (to implement the roadmap) or the Retainer (for continuous senior judgment). If you're set, we're set. No pressure.
Get a straight answer on your pilot in five days.
20 minutes on a call tells you if the Pre-Mortem is the right fit. If it isn't, we'll say so and point you at what is.
Book your Pilot Fit Call →20 minutes · No slide deck · No pressure