Hypothesis → result → learning. Nothing evaporates.Experiments
concluded
Experiment #3: Bundled AI Credits Gateway ($99 → $30 credits, then BYOK)
Hypothesis: If the $99 Basic tier includes $30 of one-time bundled inference credits (hard-capped, routed through Alpha's upstream provider accounts), then ICP prospects (50–500 person SaaS shipping agents) will convert from Arena aha moment to paid signup at a materially higher rate than a BYOK-first ask, because bundled credits eliminate the security-review trust hurdle of routing their org's API keys through an unknown proxy on day one. Success criteria: paid conversions attributable to the credits path (each is a costly signal + demand ledger entry) and observed transition from credits to BYOK within the first billing cycle. Falsified if: ICP buyers ignore credits (they already have keys/negotiated rates) and conversion is unchanged, or credits attract only indie builders outside ICP. Design constraints: one-time credits (not recurring — avoid becoming a reseller, protect "cost gets them in the door, control and compounding is what they pay for"), hard cap with no overage, BYOK transition as a single setting flip, frame as "trial without the security review" not "cheap tokens." Margin: ~$69 per signup at zero markup, no arbitrage risk since credits < price.
Result: CONCLUDED-BY-DECISION 2026-08-26 (never run; zero signups, zero observations in 49 days). Retired by two independent decisions. Decision #404 names it a PLG-acquisition play to park. Decision #403 kills the mechanism on arithmetic: a one-time $30 credit bundle attached to a $30,000/year contract is 0.1% of ACV, and there is no version of it that moves a buyer at that price. Under founder-led sales the trust bridge the credits were designed to build is built for free in the conversation instead. ONE STRUCTURAL RISK CLOSES WITH IT, AT NO COST: Angle 5 established that OpenAI and Anthropic both prohibit reselling or leasing API access, and that routing bundled-credit traffic through Alpha's own provider accounts is precisely the architecture both tightened terms around. Retiring the mechanism permanently removes that exposure — the single most material risk identified in this experiment is now moot rather than pending legal review. PRESERVED AS GTM MATERIAL, WITH A DATED CAVEAT: Angle 2 found that at 50-500 employees there is usually no formal corporate security review, only individual engineer trust friction ("I don't want to route production keys through a proxy I haven't vetted"). That research was done for the $250/mo ICP. At a 20-agent, $30K buyer the formal review does appear — see flag #406 and the #50 re-score — so use Angle 2's framing for the engineer in the room and the compliance posture answer for the process behind them.
Learning: RESEARCH VALIDATION COMPLETE (Jul 9 2026) — 5-angle deep search across PLG precedents, mid-market security behavior, segment risk, LLM credit economics, and provider ToS. Summary verdict: the experiment structure is sound and the mechanism is validated by precedent, but three things need adjustment before treating conversion data as conclusive.
ANGLE 1 — CREDITS AS ONBOARDING (VALIDATED). Free credits at first signup is the canonical PLG activation pattern: Twilio ships $15.50 credits on free trial accounts (their #1 lead source), AWS gives $100 at signup, Cloudflare Startups runs $250K credit programs. The mechanism that makes it work is velocity to aha — developers who make a first API call within 10 minutes convert at 3–4x the rate of slower activators, and Twilio's credit-powered onboarding redesign produced 62% higher activation. The structural difference from Alpha's design: Twilio's credits remove friction to trial (free account), while Alpha's credits are bundled into the paid ($99) plan — this is actually a stronger signal because it requires purchase intent, but it's a different mechanism. The closer analog is "credits embedded in a paid tier to justify the buy and accelerate activation," not "free credits to reduce signup friction." This is sound; it just means the conversion lever is "does $30 de-risk the $99 commitment" rather than "does $30 get people to sign up who wouldn't have."
ANGLE 2 — TRUST HURDLE (VALIDATED BUT MISFRAMED). The BYOK individual trust barrier is real and documented. The key finding: at 50–500 employees, there usually is NOT a formal corporate security review blocking BYOK on day one — most mid-market companies "check if you have a privacy policy and maybe ask whether you're SOC 2 compliant." The formal pen-testing/procurement-delay security review is primarily an enterprise (>1000 employee) phenomenon. What exists at the ICP tier is individual engineer/team-lead friction: "I don't want to route our production API keys through a proxy I haven't vetted yet." This is a PERSONAL TRUST JUDGMENT, not a CORPORATE PROCUREMENT DELAY. Implication: the framing "trial without the security review" implies a formal organizational process that usually doesn't exist at 50–500. Better framing: "try it on our keys before committing yours" or "see it work without touching your API setup." The mechanism is correct — bundled credits eliminate the personal key-routing hesitation — but the language should match the actual blocker.
ANGLE 3 — SEGMENT RISK (MANAGEABLE). Free credits classically attract indie/hobbyist developers — documented across API products. However, the $99 price point is an effective ICP filter: indie builders don't pay $99 for a trial. The real segment risk isn't the wrong demographic signing up; it's credit-motivated purchase intent within the right demographic — mid-market teams buying the $99 plan primarily because it looks like "get $30 of tokens bundled," exhausting credits, and churning without ever flipping to BYOK. Twilio's no-CC free trial (their primary acquisition model) generates 27% more total paying customers than CC-required despite lower per-signup conversion — because the math works at volume. Alpha's CC-required $99 entry is the opposite bet: higher intent signal per signup, smaller top-of-funnel. This is probably right for the ICP but needs churn tracking by cohort (credit-exhausted-no-BYOK-flip vs. BYOK-converted). The BYOK flip being a "single settings flip" is the right design: it makes the behavioral test clean and low-friction.
ANGLE 4 — $30 CREDIT SIZING (CORRECT RATIONALE, NEEDS FRAMING DISCIPLINE). At July 2026 pricing: Claude Sonnet 4.6 is $3/M input / $15/M output. Typical complex agent task runs 50K–200K tokens across 10–20 model calls. $30 at Sonnet input pricing = ~10M tokens = roughly 50–200 production-quality agent task runs. For a team at $5k/month inference spend, $30 = ~0.6% of monthly spend (trivial money). For a $50k/month team, $30 is noise. The $30 is NOT meaningful as a cost saving for anyone in the ICP. It IS meaningful as "enough runway to run your actual agents against Alpha's stack and confirm it works before committing your own keys" — roughly 4-8 hours of a small deployment's traffic. This validates the sizing: it's enough to reach aha on real workflows, but not so much that it looks like token arbitrage or subsidized resale. The risk is in how it's communicated: never position as "save on inference" or "cheap tokens" — these framings (a) mismatch the economics for any team at ICP scale, and (b) pull in the wrong buyers who want credits, not platform. The design constraint "frame as trial without the security review, not cheap tokens" is correct and must be enforced in all copy.
ANGLE 5 — RESELLER TOS RISK (REAL, REQUIRES LEGAL REVIEW BEFORE SCALE). Both OpenAI and Anthropic explicitly prohibit reselling or leasing access to API accounts. OpenAI: "customers may not resell or lease access to their Account." Anthropic: "prohibits customers from accessing the Services to resell them, except as expressly approved by Anthropic." The "wrapper/proxy" architecture — where Alpha routes customer traffic through Alpha's own provider API keys — is precisely what both providers have tightened terms around since 2024-2025. Alpha's bundled-credits mechanism IS this architecture: Alpha pays upstream inference costs on behalf of customers during the trial period. The mitigating factors: (a) Alpha adds substantial value-add (routing, observability, operating layer — not pure API resale, which is where the legal distinction lives), (b) both providers have formal partner/reseller programs that larger operators access with negotiated agreements, (c) the one-time hard-cap design limits the surface area. This is the single most material structural risk in the experiment. At early-stage/low-volume, enforcement is unlikely. As the credits program scales, the architecture needs either (i) formal reseller/partner agreements with upstream providers, or (ii) routing bundled-credits traffic through AWS Bedrock or Google Vertex (which have clearer cloud-reseller provisions) rather than direct OpenAI/Anthropic API accounts. Action: get legal counsel review of the bundled-credits architecture before the program reaches meaningful volume. Do not scale this path without a clear ToS position.
OVERALL VERDICT: The experiment is structurally sound. The mechanism (credits bridge trust gap → faster BYOK activation → higher paid conversion) is supported by PLG precedent and the mid-market trust-behavior evidence. Three adjustments needed: (1) Reframe copy from "security review" to "no keys needed to see it work" — matches the actual blocker at ICP company size. (2) Track the credit-exhausted-no-BYOK-flip cohort as the key falsification signal — if >40% of credit users don't flip to BYOK within 30 days of exhaustion, the credits are attracting the wrong buyer intent, not eliminating trust friction. (3) Secure legal review of the upstream routing architecture before scaling — this is the one risk that can shut the mechanism down regardless of conversion results.
concluded
Passthrough proxy + team cost card + shadow-savings meter: does "deploy in 2 clicks, watch the waste accumulate, toggle on optimization" beat email capture as the post-aha conversion path?
Hypothesis: CORE HYPOTHESIS: If Alpha's free tier is a passthrough proxy that teams can deploy in ~2 clicks and that continuously shows a team-visible SHADOW-SAVINGS delta ("Spent: $X / With Alpha optimize+route: $Y"), then converting to paid becomes a toggle ("turn on optimization") rather than an integration ask — lifting activation and free→paid conversion vs. the email track-over-time cohort. Falsified if proxy-deploy CTA converts at parity or below email capture, or if teams in passthrough mode don't flip the toggle within 30 days despite an accumulating savings delta.
ARCHITECTURE (data plane / control plane split): Proxy = thin, stateless, boring binary (Go/Rust, single static image, ~10-20MB) doing passthrough + metering only. Fail-open by design: if Alpha control plane is unreachable, requests pass straight to the provider — this sentence kills the production-risk objection. All heavy components (Postgres, Qdrant, dashboards, compounding layer) stay in Alpha cloud, never in the request path.
DEPLOYMENT (one-click, user's cloud): /deploy page with marketplace buttons — Cloud Run (primary for ICP: scale-to-zero, ~60s to live), AWS CloudFormation quick-create (Fargate), Azure ARM template, Railway/Render/Fly for smaller teams. Button is clicked while signed in to Alpha, so template carries a pre-generated ALPHA_TOKEN env var. Hosted mode (switch baseURL to Alpha edge, zero deploy) remains the default alternative.
TEAM CARD MECHANICS: Proxy makes outbound-only metrics connection to Alpha cloud (model, tokens, cost, latency, run ID — payloads only on opt-in for compounding). No inbound rules, nothing for infra to approve. Alpha cloud is the aggregation point for hosted AND self-hosted, so Slack card / dashboard / alerts are identical in both modes. Slack app installed once at org level by one admin; every dev pointing at the internal proxy URL appears in the shared card automatically. TOKEN = IDENTITY, not the proxy: one org can run multiple proxies (staging/prod, per-team) and the card aggregates — later becomes the per-agent budget story. Phase 2 surface: IDE status-bar extension.
SHADOW-SAVINGS METER (the conversion engine — this is Arena running continuously on live traffic): control plane computes counterfactual cost alongside actual: (1) prompt/context optimization, (2) right-sized model routing per task, (3) fine-tune route once enough runs accumulate ("12,400 runs collected — enough to train a routed small model, est. additional 30% savings"; public copy = "memory/compounding layer", NEVER Trace-to-X names). Card always shows both numbers + cumulative "left on the table since connecting" delta. Fidelity ladder mirrors Arena: L1 static estimate (pricing math + routing heuristics, labeled "estimate" per honesty guardrail) → L2 sampled shadow replay (~1% of real runs replayed through optimized path, quality-parity verified, badge: "verified on your own traffic") — the badge is what makes the number credible to a CTO.
FULL FUNNEL: Arena aha (one-shot estimate) → one-click proxy deploy (passthrough, free) → live team card with shadow savings accumulating → "turn on optimization" toggle = paid conversion. Canon-consistent: cost gets them in the door; control and compounding is what the toggle activates. Slack card posts the monthly counterfactual to #eng, so the VP sees the waste number without anyone lobbying.
SUPERSEDED WITHIN THIS EXPERIMENT: Electron/desktop widget (per-dev friction), npx CLI (not team-visible), bundled local proxy in an app. Competitive context stands: Tokcat/token-monitor/CostGoat/TokenBar all meter the dev's own AI-tool usage via logs/billing APIs; nobody offers a team-visible live meter + counterfactual savings on the company's own application traffic via proxy.
Result: CONCLUDED-BY-DECISION 2026-08-26 (never deployed; zero observations in 50 days). Not falsified — retired. Decision #404 (Vishnu, 8/24) states plainly that Arena "is not a self-serve acquisition channel and is not part of the PLG funnel" and that "Experiments #2 (savings meter) and #3 (bundled credits) are both PLG-acquisition plays — deprioritize or park them." This experiment measures free-to-paid self-serve toggle conversion against an email-capture cohort; under founder-led sales neither cohort exists, so no result can ever be produced. Status changed from running to concluded so the portfolio stops reporting three live experiments when it has zero. NOTHING IS DELETED — the hypothesis, architecture and the Jul 11 design update are preserved in full above, and two pieces of it are now load-bearing under FLS rather than obsolete: (a) THE FAIL-OPEN SPLIT — "if the Alpha control plane is unreachable, requests pass straight to the provider" — is the one sentence that kills the production-risk objection on a founder call, and it now also answers a $30K buyer's security questionnaire; (b) THE ALPHA FLEET PANEL from the Jul 11 update, which that note itself calls the best single panel for cold outreach. Under #404 Arena's success metric is demo quality; that panel IS the demo and it already exists, which makes it the cheapest available version of Task #22. The unresolved projected-vs-realized reconciliation ($4.5K/mo projected vs ~$1.3K/mo realized) does NOT retire with the experiment — it moves to Task #55, because a wrong number said out loud on a sales call is worse than a wrong number on a page.
Learning: DESIGN UPDATE (Jul 11 2026) — shadow-savings meter UI iterated from 5 cost-first panels to a two-state, six-panel asset-first ladder. This is the surfaced form of the shadow-savings mechanic.
WHAT CHANGED. Original UI: every panel sold the same word — "save" (cut bill 24%, one key/cheaper setup, auto-swap cheaper, overflow cheaper, save $149/day). Five cost cards = strategically monotone; trains prospect to shop on price and files Alpha as a cost tool, the exact box a free gateway (Portkey OSS / LiteLLM / Cloudflare, per Entry #84) undercuts. Rebuilt as a ladder that climbs cost → capture → compounding → governance → ownership.
FINAL PANEL COPY (asset-first ladder):
1. Today/All Sources — "Every call is now a trace you own — cost down 24% too" (cost demoted to afterthought via "too").
2. By Provider — "One key captures every provider call as trace data" (reframes convenience feature as capture mechanism; strongest reframe line).
3. Top Models — "Your calls become eval data — routing keeps improving" (states the compounding loop plainly; "eval data" is IC-language, acceptable for ICP).
4. Local Agents — "Coding-agent traces captured — fuel to fine-tune & replay" (still under-points at the $3.9K API-equivalent value figure on-panel — open copy nit).
5. Alpha Fleet — "Every agent's traces captured — evals, budgets, alerts" (NEW panel; named agents + per-agent spend + red alert dot on one agent. Only panel showing the control plane vs the cost tool. Product in three words. Best single panel for cold outreach — proves what others claim.)
6. Savings — banner "Your traces working for you — savings are just the start" → button "Own your intelligence layer" / "Start compounding with thealpha.ai →".
KEY STRUCTURAL MOVE — two-state split (logged-in vs prospect): "your dashboard" panels = customer state, already bought the thesis, so copy leads with asset and treats savings as afterthought (reads as confidence). "Start compounding with thealpha.ai" panel = prospect state doing conversion work. Same surface, two jobs, cleanly separated. Resolves the long-standing tension: a logged-in user doesn't need selling, so copy can be pure ownership.
OPEN DESIGN QUESTION (flagged, unresolved) — PROJECTED vs REALIZED numbers don't reconcile. Prospect panel projects $149/day → $4.5K/mo → $187.7K 3-yr. Customer panel shows realized $1.3K/mo (~29% of projection). Both honestly labeled ("illustrative" vs "cost-per-success vs baseline") so not dishonest, but a buyer who converts on $4.5K and lives in $1.3K = churn + trust risk. Two fixes: (a) soften projection so realized beats expectation (under-promise), or (b) make realized panel show the compounding curve bending up over time so early $1.3K reads as "month one, climbing" not "projection was inflated." Asset story supports (b): savings compound as corpus grows, so an early realized number SHOULD be low — panel just needs to say so. RESOLVE BEFORE POINTING OUTREACH AT THIS FUNNEL.
COSTLY-SIGNAL REQUIREMENT (unchanged, load-bearing): "Start compounding" is the one prospect-facing button. If it collects name + company, a click writes a real buyer into the demand ledger and the whole ladder pays off as measurable intent. If it drops into an anonymous free install, the ladder still works but generates zero ledger signal — best line wired to a cheap signal. Success metric for this experiment stays: costly signals captured on the prospect-state button, tagged with role + company.
NEXT: (1) resolve projected/realized reconciliation; (2) confirm prospect button captures identity; (3) decide cold first-touch panel — Alpha Fleet is strongest, but "trace" vocabulary is assumed across all panels, so first-touch must teach the word once before later panels spend it.
concluded
People want to reduce their LLM costs
Hypothesis: Everyone building agents want to reduce their LLM costs, move to open source where possible but do not know how to
Result: Validated with segmentation caveat: cost pain is severe and real at production scale (Gartner 5-30x agentic token multiplier; Stanford: 62% of agent bills = re-sent context; enterprise spend 3x over token allocations per FinOps Foundation 2026; Uber burned annual AI budget by April 2026; Gartner projects 40% of agent projects cancelled by 2027 due to cost overruns). Total enterprise AI spend is rising despite 94.5% per-token price drop — Jevons Paradox confirmed. However, LangChain State of AI Agents survey (1,340 respondents, Dec 2025) found cost is NOT the top stated barrier — quality (32%) and latency (20%) rank higher. Reconciliation: cost pain arrives at scale, not at pilot stage. "Move to open source" is the WRONG mechanism — open weights hold only 11% enterprise share (down from 19%); break-even for self-hosting is 10-30M tokens/day requiring 0.5-1.0 FTE MLOps. The winning cost-reduction pattern is multi-model routing (70-80% to managed cheap models like DeepSeek V4 at $0.14/M, Gemini Flash) not self-hosting. "Doesn't know how to" is partially wrong — LiteLLM/Portkey/OpenRouter are well-known; the real gap is cost architecture: teams can't attribute spend at task level, don't implement prompt caching correctly (Stanford: 62% of agent bills = re-sent context; PwC: caching reduces costs 41-80%), and can't diagnose which agentic loop patterns are wasteful.
Learning: Refined hypothesis: Rephrase from "reduce LLM costs / move to open source / doesn't know how" to "teams running agents in production have severe, growing cost-and-control problems driven by agentic architecture failures, not per-token rates." The entry pain (Arena's job) is the surprise bill — but the product must be framed as cost ARCHITECTURE and CONTROL, not cost savings. The mechanism teams want is routing + budget enforcement + per-task observability, not open-source migration paths. The "open source" angle is a liability: it implies operational burden that the ICP doesn't want. The stronger wedge is: "you're wasting 40-60% of your token budget on context inflation and retry storms — Alpha shows you exactly where and enforces budgets per agent." Quality and latency are now equal or higher stated concerns than cost (LangChain survey), which reinforces the "control layer" framing: Alpha's value is agents that run reliably within budget, not just cheaper agents. Sharpest product wedge: "run more agents for the same budget" rather than "cut your API bill." Sources: Cockroach Labs, Oplexa, The Deployment Layer, LangChain State of AI Agents, Digital Applied, BenchLM.ai, Epoch AI, FinOps Foundation 2026, Marka Development, Stanford Digital Economy Lab.