Most engineering teams have a deployment checklist. It covers environment variables, database migrations, rollback procedures, and smoke tests against the API. That checklist is correct for software. It is incomplete for AI pipelines.
The six items below do not appear on standard checklists. Each one maps to a specific failure mode. Each failure mode has caused a real production incident for a team that assumed their existing process was sufficient.
The Six AI-Specific Checklist Items
1. Prompt Version Lock
What it prevents: Prompt drift between environments.
Prompts are code. If they are not pinned to a version hash or tagged release, a change in staging can silently propagate to production during a routine deploy. The failure mode: output format changes without a corresponding schema update, downstream parsers break, and the error surfaces three steps removed from the actual cause.
Checklist item: Confirm the prompt version hash in production matches the version tested in staging. Fail the deploy if they diverge.
2. Retrieval Smoke Test
What it prevents: Silent retrieval failure at the index level.
A retrieval system can return results without returning correct results. If the vector index was rebuilt with a schema change, embedding dimension mismatch, or stale document set, queries will resolve but answers will be wrong. The failure mode: the system appears healthy, but users get confidently wrong answers for 48 hours before anyone notices.
Checklist item: Run three known-answer queries against the production index before traffic is live. Require exact-match or threshold-match on expected document IDs. This is distinct from the retrieval regression probe covered in a prior post — this is a pre-traffic gate, not a regression suite.
3. Fallback Path Verification
What it prevents: Silent failure when the primary model or retrieval layer is unavailable.
Every AI pipeline should have a fallback: a simpler model, a cached response, or a graceful degradation message. The failure mode when this is skipped: the primary path goes down, the fallback has never been exercised in production, and it turns out the fallback route has a misconfigured API key or a timeout set to 2 seconds instead of 20.
Checklist item: Trigger the fallback path manually in production before enabling primary traffic. Confirm it returns a valid response within acceptable latency.
4. Output Schema Validation
What it prevents: Malformed model output breaking downstream consumers.
Models do not always return what you expect. A JSON field goes missing. A string field returns an integer. The failure mode: the downstream service that consumes model output throws an unhandled exception, and the error is logged as a generic 500 rather than an AI output error, making it hard to trace.
Checklist item: Run the model against five canonical inputs in production and validate output against the declared schema before routing live traffic. Use a strict validator, not a lenient one.
5. Confidence Threshold Confirmation
What it prevents: Low-confidence outputs reaching users without a guard.
Most pipelines set a confidence threshold during development and never verify it survived the deploy. Environment differences, model version changes, or index updates can shift score distributions. The failure mode: the threshold that filtered 15% of outputs in staging now filters 2% in production, and low-quality responses reach users at a higher rate than intended.
Checklist item: Sample 20 production requests immediately after deploy. Confirm the confidence score distribution matches the expected range from staging within an acceptable tolerance.
6. Alert Routing Check
What it prevents: AI-specific errors going undetected because they route to the wrong team or no team.
AI pipelines produce failure signals that standard application monitoring does not classify correctly. Retrieval misses, token limit breaches, and model timeout errors often land in a generic error bucket. The failure mode: an AI-specific degradation runs for hours because no alert fired, or an alert fired to a queue nobody watches.
Checklist item: Confirm that AI-specific error classes — retrieval failure, schema validation failure, confidence threshold breach, model timeout — each have a named owner and a tested alert path. Send a test alert before go-live.
Integrating This Into an Existing CI/CD Pipeline
None of these items require a new tool. They require a new step.
Add a post-deploy verification stage to your existing pipeline. This stage runs after infrastructure is live but before traffic is enabled. It executes the six checks above as scripts or test cases. If any check fails, the pipeline halts and rolls back.
The implementation pattern:
- Prompt version lock: A one-line hash comparison in a shell script.
- Retrieval smoke test: A Python script that queries the live index and asserts document IDs.
- Fallback path verification: An HTTP call to the fallback endpoint with a timeout assertion.
- Output schema validation: A JSON schema validator run against live model output.
- Confidence threshold confirmation: A sampling script that logs score distribution and fails outside tolerance.
- Alert routing check: A test event fired to your alerting system with a manual confirmation step.
Total added pipeline time: 3 to 8 minutes depending on model latency. That is a reasonable trade for catching the failure modes that cause the first production incident.
The teams that skip this step are not skipping it because they disagree with the logic. They skip it because nobody wrote the checklist before the first deploy. Write it before the first deploy.