Last reviewed: July 2026

The Kit Bag

What's actually in an AI consultant's tool bag, 2026

Every "best AI tools" listicle is affiliate bait. This one isn't — no affiliate links, and where a tool is ours we say so with a grin. It's what we and the practitioners around us actually use. The Stack Survey will correct us in public as the data comes in.

What our own logs prove

The receipts, not the vibes

We can't ask you to take this bag on faith, so here's what our own crawler logs actually show, live from the Bot Ledger. Caveats first, because they matter: our traffic is a floor not a census, our audience self-selects, a user-agent is a claim not an ID, and the sample is still small.

—
of crawler hits are agent-driven, not training scrapes (an assistant fetching for a human), 30 days
—
the assistant family reaching us most on behalf of users this month
—
the AI tool sending us the most humans (referral clicks)

Sample: loading…

✓ What this proves

The agent era is real and measurable, and it's assistant-mediated: the tools in categories 2 and 4 are the ones actually reaching the open web for your clients. That's why they're graded backed by our logs.

✗ What it can't

Which IDE or CLI belongs in your bag. A crawler can't see what a consultant runs locally, so the product picks stay honest opinion until the Stack Survey grades them. Add your stack →

1. The agentic workbench Practitioner pick

The consultant's primary surface in 2026, not a developer nicety. Deliverables — the scoping doc, the migration, the eval harness — increasingly get built in these, not typed into a blank page.

Claude CodeTerminal-native agentic coding. Where a lot of our own delivery work happens. Bites when you let it run unsupervised on a big change — review the diff, always.
CursorThe IDE-shaped entry point most teams meet first. Good for consultants embedding with an existing dev team on their editor.
Codex-class CLIsThe category, not one product. Pick the one your client's stack and security team will actually approve; that constraint matters more than the benchmark.

2. The connected workspace Backed by our logs

Agent workspaces that reach your tools, plus the MCP servers worth installing so your agent can query real data mid-engagement.

Cowork-class workspacesAgent workspaces that connect to your docs, mail and calendar. Useful for the non-code half of consulting; watch what you connect and what it can do.
The Project Wrecked MCP server OursYes, ours — disclosed with a grin. Read-only tools for model deprecations, rates, a red-flags checklist and our crawler stats, installable in any MCP client. /mcp-server.
Provider + vendor MCP serversInstall the ones that expose data you check often (issue trackers, docs, cloud consoles). Every install is an attack-surface and a dependency — curate, don't hoard.

3. Delivery & evidence From the wreck archive

The boring spine that separates a consultant from a demo. Unglamorous, and the first thing that saves you in a dispute.

Versioned docs & decision logsPlain, diffable, dated. The record of what was decided and why is worth more than any deck when a stakeholder's memory conveniently changes.
An eval harnessA few hundred representative inputs and approved outputs, built from day one. It's the difference between validating a model swap in days and arguing about vibes for a month.
A data audit, before you scopeNot a tool, a habit. The bodies are always in the data — go find them before you sign.

4. The commercial layer Backed by live data

Proposals, contracts, and pricing — the part that turns capability into an invoice. Where our own store products genuinely belong (disclosed as ours).

Proposal & SOW toolingWhatever you write contracts in, with a real model-deprecation clause in the template. We used to sell templates for this and stopped, because the clause itself is the valuable part and it is free to read.
A vendor-evaluation rubricWeight deprecation behaviour alongside capability and price. Build your own scorecard; the Platform Ledger supplies each provider's actual habits, live, and the Adoption Ledger shows which of their models the market is already leaving.
wreck-watch Ours LiveOur free CI check that fails the build when a pinned model ID is on the deprecation departure board. npx wreck-watch — the deprecation clause, enforced. /wreck-watch.

5. The watchlist Provisional

Explicitly provisional — what we're evaluating, not endorsing. Here so you can watch us change our minds in public.

Agent-eval & observability toolingThe category is moving fast and we don't have a confident pick yet. When we do, it'll show up here — and the survey will tell us if we're wrong.
🗳️

The green tags are backed by our logs; the rest is honest opinion, and we'd rather it were data. So tell us what's in your bag — the Stack Survey takes a minute, is anonymous, and once enough practitioners answer, its aggregates take over this page and correct us in public. Wondering whether a certification belongs in the bag? We graded those separately. No affiliate links here, ever — if that changes, we'll disclose it per link, in this same voice.