0/2 automated — the dogfooding scorecardPlaybooks

Every recurring process, its steps, and its automation status. Target state: everything runs on Alpha.

Cost-Shock Content Playbook (money-saved lane)

Surface: LinkedIn (top-of-funnel) → Arena. Job: make a technical buyer FEEL the surprise invoice, then route them to Arena for their own free number. Nothing else. GOVERNING RULE: "Cost gets them in the door. Control and compounding is what they pay for." On this surface lead with the cost shock + savings number ONLY. Do NOT pitch harness / operating layer / reliability / compounding (that's the product surface). If a draft explains what Alpha does, cut it. And beat the "models are getting cheap" objection every time: "Token prices fell 80%. Your agent bill didn't. The problem was never the model — it's the runs you can't control." THE 4-BEAT POST FORMULA (short posts = beats 1+3+4; teardowns = all 4): 1. HOOK — open with ONE visceral, specific number or the paradox. "$4,200 in one weekend" beats "agents are expensive." 2. MECHANISM — explain WHY the bill exploded (the loop, retries, context bloat, multi-model sprawl, expensive model on cheap task). This teaches → earns trust + shares. 3. TWIST — the counterintuitive turn: prices collapsed yet bills multiplied because agentic usage outran the price cuts. Makes the post un-dismissable by the "just use a cheaper model" crowd. 4. BRIDGE — one soft line to Arena's free, no-instrumentation waste number (<5 min). Frame as THEIR number, not the product. "Curious what your agents actually waste? Arena tells you free, no code — [link]." GUARDRAILS CHECKLIST (run before posting): - Opens with a specific number or the paradox, not a generalization - Beats the "cheap models" objection somewhere - ZERO product pitch — no harness/operating layer/reliability/compounding - Only CTA is a soft bridge to Arena's free number - Framed as THEIR bill, not "our tool" - Every claim traceable to a brain stat or attributed quote (no fabricated figures) - Any quote used is attributed to the named practitioner CADENCE: Don't burn Tier-1 (the paradox) twice a week — it's the spine, space it. Alternate original stat posts with react-and-reframe on fresh practitioner quotes (compound reach via author's network). ICP scanner surfaces new cost-blowout quotes nearly every run — pull the newest weekly. One number, one mechanism, one twist per post; don't stack stats.

STAT ARSENAL (tiered; rotate, don't burn the strongest in a week): TIER 1 — PARADOX STATS (signature beat): - Per-token prices fell ~80% in 12mo / ~1,000x since 2022, yet monthly bills keep multiplying (benchlm/cloudzero, Anthropic) → the core paradox post; lead the lane with it. - Same workload that cost ~$3,000/mo in 2024 now runs ~$150/mo, but nobody's bill went down → objection-killer. - Single agent task burns 50k–500k tokens across 10–20 calls vs 2k–4k for a chatbot (Research #1) → "why your agent costs 100x your chatbot." TIER 2 — SURPRISE-INVOICE STATS (emotional trigger): - Production agent workloads run 5–10x over pilot-budget projections (Gartner flag) → "your pilot lied." - $1k estimate → $3.8k invoice (canonical scale-wall number, positioning canon) → most on-message number we own. - 50% of AI product companies don't track LLM cost at all — one monthly Stripe charge (Mavvrik 2025) → "half of you are flying blind." - Teams waste ~40–60% of token spend on suboptimal implementations (Research #1) → direct savings hook. - Inference now ~85% of enterprise AI budgets (Anthropic) → "this isn't a line item, it's the budget." TIER 3 — AUTHORITY STATS (press-grade credibility): - Gartner: >40% of agentic AI projects cancelled by end 2027, driven by cost overruns (Jul 2026). - Gartner: $234B enterprise app spend "at risk" from agentic AI through 2030 (Jul 2026). - 86–88% of enterprise agent pilots never reach production (Research #4). - Enterprise LLM spend went $3.5B → $8.4B in six months (validation flag). VOC QUOTE BANK (real, attributed — react-and-reframe fuel): - "Uncontrolled LLM agent loops cost us $4,200 in one weekend." — Musa Usmani, Ihsan Systems - "Token costs don't surprise you in demo. They surprise you in production." — Mohit Sehgal - "Your 10-step agent costs $0.16 per run on paper. It costs $0.28 per successful run. That gap is the reliability tax — not on any pricing page." — Arvind R, Credit Saison - "We are 3x over our entire 2026 token budget and it's only April." — via FinOps Foundation / WorkOS - "I'm back to the drawing board, because the budget I thought I would need is blown away already." — Uber CTO (Claude Code 32%→84%, budget gone in 4 months) - "A single AI agent hitting an API can burn $100K+/year." — David Villalon, Maisa AI - "$47K eleven-day agent loop" postmortem — Erik Peterson, CloudZero - "You'll have to start explaining tokens to finance the way we first had to explain what an EC2 instance is." — Brian Gracely, Red Hat - "Which agentic workflow could blow up your AI bill today — and are you sure you'd see it before finance does?" — Yingzhao Ouyang - Glean production triage agent reportedly burning ~$1M/mo in tokens — Arvind Jain REACT-AND-REFRAME PATTERN: [Quote] → "This is the whole story of 2026 in one line." → paradox beat → "The bill isn't the model. It's the runs nobody's watching." → Arena bridge. FIVE READY ANGLES (pull stat → 4-beat formula → ship; comments mandatory): 1. Paradox Post (signature): prices fell 80%/your bill rose → usage outran cuts, 10–20 calls per task → "cheaper models" is a trap at the agent layer → Arena. 2. Pilot Lied Post: 5–10x over pilot budget → pilots test happy path, prod runs the loop → you mis-modeled, not mis-budgeted → Arena. 3. Flying-Blind Post: half don't track cost at all → no attribution to team/model/feature → can't cut what you can't see → Arena (<5 min, zero instrumentation). 4. Reliability-Tax Post (Arvind R react): the quote → cost-per-successful-run ≠ cost-per-run, failures compound → headline pricing is a marketing sheet → Arena. 5. 40% Waste Post: 40–60% of spend is waste → wrong model, oversized prompts, silent retries → the savings are already in your bill → Arena. WATCH-OUT: Experiment #2 shadow-savings meter projected (~$4.5k/mo) vs realized (~$1.3k/mo) gap is still unreconciled (Task #55). Until settled, anchor posted savings claims to RESEARCH stats (40–60% waste, 5–10x over pilot), NOT the Arena meter number — a disputable headline figure would undercut the shock. Full formatted playbook file generated 2026-07-29 (cost-shock-content-playbook.md).

Distribution: LinkedIn comment-warming → costly-signal teardown outreach

CHANNEL (finalized): LinkedIn comment-led warming into personalized outreach. FULL SEQUENCE: comment warms → teardown aches → HUD reveals → baseURL converts. DAILY CRANK (~30 min): 1. Morning (15 min): Agent surfaces fresh threads prospects engaged with. Pick 5-6 and write genuine comments BY HAND. 2. Midday (5 min): Check replies/reactions on yesterday's comments. Log warm signals into Alpha Brain. 3. End of day (10 min): For anyone at 2-3 touches WITH engagement back, queue a personalized teardown offer for tomorrow. TOUCH RULE: 2-3 meaningful comment touches before outreach. A reply or reaction back is a green light to reach out even at touch #2. Log which prospect was warmed and on which thread, so outreach references the exact thread. PROOF ASSET — the Cost Teardown (3 acts, told with real numbers): - Act 1, The Leak: Make the invisible visible. Hook = "Across all your tools, can you tell me right now what ONE agent run truly cost you, end to end?" They can't. Establish the fog. Name retries + unbounded context as EXAMPLES of costs nobody surfaces yet — honest framing. - Act 2, The Diagnosis: Provable spine TODAY is unoptimized spend — current LLM costs are higher than needed because there's no routing (premium models on trivial work) and no optimization layer. Reframe as a MISSING LAYER, not the buyer's fault. (Once retries ship, the retry blind spot becomes Act 1 + Act 2's hero.) - Act 3, The Path: Show the "after" with the delta the HUD actually quantifies. Stay diagnostician, not salesman: "Whether you build this yourself or use something like Alpha, here's what the layer needs to do." CONVERSION STEP — the cost HUD: The destination every warmed + teardown'd prospect is driven toward. Shows current LLM costs + routing/optimization savings. Point them at the Arena / HUD ungated after the teardown; the baseURL switch is the conversion event (aha). The teardown creates the ache, the HUD is the relief. TODAY: HUD does not yet show retries — don't send prospects looking for retry data. IN DEVELOPMENT: retry visibility is being built and will become the flagship reveal; re-anchor the teardown on it once shipped. HUD AS ASSET: A screenshot of the HUD surfacing real savings (and later, real retry waste) is itself proof content for comments, teardown visuals, and the public flagship. The HUD is also what eventually turns the live 30-min teardown into a self-serve aha, so human judgment shifts from "run every teardown" to "only close the ones worth a room." SEQUENCING: Run first 5 teardowns live (realistic scenario, e.g. a support agent at production volume — no client needed). Then anonymize the best one into a public flagship asset. ONE-BREATH SUMMARY: Warm the burned in LinkedIn comments → reach out after 2-3 touches with a personalized teardown offer → run leak/diagnosis/path → send them to the HUD → baseURL converts. Prove routing/optimization savings today; make retries the flagship reveal the moment they ship. Automate the finding, never the trust.

Founder-led distribution motion for thealpha.ai. Beachhead ICP: engineering leaders shipping agents to production who are blind on per-run cost, because their split observability + gateway tooling can't see the full picture of a run. Market splits into "burned" (in production, feeling pain — SELL) vs "pre-pain builders" (haven't hit the wall — TEACH via content). Prospect list is built from engagers on cost-related LinkedIn content, not authors (authors tend to be competitors). Core principle: automate the finding, never the trust — the agent surfaces threads, but the human writes every comment and delivers every teardown. Ceiling: distribution alone reaches early PMF + revenue; enterprise deals still close on trust earned in a room. PRODUCT REALITY (current, July 2026): The HUD shows current LLM costs and how much Alpha would save by routing + optimizing. It does NOT YET show agent retries — retry visibility is IN ACTIVE DEVELOPMENT and is intended to become the HUD's flagship differentiator. UNTIL RETRIES SHIP: teardowns prove routing/optimization savings; retries/context are named only as examples of costs nobody surfaces yet (honest framing, not a product promise), and prospects are never sent to the HUD looking for retry data. ONCE RETRIES SHIP: re-anchor the entire teardown on the retry blind spot — it's the sharpest wedge (the leak competitors' request-level tools structurally can't show) and now also the closeable payoff. That alignment (make the ache + be the only one who resolves it) is the strongest version of this motion.