NextUp SoCal · Aug 21, 2026

Reducing AI Operating Costs

The full deck, laid out to read on a phone. Nick McCarty, Upskilled Consulting.

Tap any slide to open it full screen, then pinch to zoom. Turning your phone sideways helps on the charts.
01 · NextUp SoCal · Virtual Event

Reducing AI Operating Costs

What the tokens actually cost, why the invoice grows faster than the usage, and the five levers that bring it back down.

Slide 1: Reducing AI Operating Costs Tap to enlarge
02 · Where we left off · June 5

In June we mapped the levels. Today: what each one costs to run.

Slide 2: In June we mapped the levels. Today: what each one costs to run. Tap to enlarge

Adoption figure from the June 5, 2026 NextUp SoCal session (Deloitte).

03 · The hour

From “it’s cheap” to a number you can budget.

Slide 3: From “it’s cheap” to a number you can budget. Tap to enlarge
04 · 01 · The unit

You’re not billed per prompt. You’re billed per token.

Slide 4: You’re not billed per prompt. You’re billed per token. Tap to enlarge

Token counts approximate; every model tokenizes slightly differently. Prices: GPT-5 standard tier, $1.25 / 1M input, $10 / 1M output.

05 · 01 · The unit

Your work, converted into the billable unit.

Every artifact your team touches has a token weight.

Slide 5: Your work, converted into the billable unit. Tap to enlarge

Speech figure assumes 130–160 wpm. Image tokens ≈ (width × height) ÷ 750 for common vision models.

06 · 01 · The unit

Here is what one of everything costs.

Read it in, hand a summary back. Common artifacts, priced.

Slide 6: Here is what one of everything costs. Tap to enlarge

GPT-5 standard tier ($1.25 / $10 per 1M). Summary assumed at 10% of input. Small-model tier drawn at 1/10 the token price.

07 · 01 · The unit

For perspective: what everyday work costs.

Slide 7: For perspective: what everyday work costs. Tap to enlarge

Derived from the per-artifact costs on the previous slide: input tokens plus a summary at 10% of input length.

08 · 02 · The multiplier

Nobody's budget breaks on price. It breaks on volume.

Slide 8: Nobody's budget breaks on price. It breaks on volume. Tap to enlarge
09 · Scenario 1 of 3 · Level 1

The assistant. Chat, summarize, draft.

Slide 9: The assistant. Chat, summarize, draft. Tap to enlarge

600 employees × 250 working days. GPT-5 standard tier: $1.25 / 1M input, $10 / 1M output.

10 · Scenario 2 of 3 · Levels 2–3

The agent. Agents read far more than they write.

Slide 10: The agent. Agents read far more than they write. Tap to enlarge

600 employees × 250 working days. GPT-5 standard tier: $1.25 / 1M input, $10 / 1M output.

11 · Scenario 3 of 3 · Past Level 3

Always-on automation. Agents that never stop watching.

Slide 11: Always-on automation. Agents that never stop watching. Tap to enlarge

600 employees × 250 working days. GPT-5 standard tier: $1.25 / 1M input, $10 / 1M output.

12 · 02 · The multiplier

Same 600 people. Same 250 days. Same price list. 47× apart.

Slide 12: Same 600 people. Same 250 days. Same price list. 47× apart. Tap to enlarge

Annual inference cost at GPT-5 standard-tier pricing. Linear scale — scenario 1 is drawn to the same scale as the others.

13 · 03 · The levers

Five levers. Same capability, a fraction of the bill.

Slide 13: Five levers. Same capability, a fraction of the bill. Tap to enlarge

Illustrative model. Each cut applies to the running remainder, not the original total. Your mix will differ.

14 · Lever 1 · Right-size the model

We’re paying frontier prices for a 1.7% edge.

Slide 14: We’re paying frontier prices for a 1.7% edge. Tap to enlarge

Benchmark gap: Stanford HAI, 2025 AI Index Report. Routing math assumes the small tier is ~10× cheaper.

15 · Lever 2 · Script the deterministic steps

Most of what you're automating doesn't need to think.

Slide 15: Most of what you're automating doesn't need to think. Tap to enlarge

Illustrative: $0.50 per agent run, a one-time build of ~$400 in analyst time, $1,200/yr maintenance.

16 · Lever 3 · Engineer the context

Your agent is re-reading the same page all day long.

Slide 16: Your agent is re-reading the same page all day long. Tap to enlarge

Illustrative: a 20-turn loop adding ~2,000 tokens per turn. Engineered = rolling window plus a running summary.

17 · Lever 4 · Stop redoing work

The chat window is a shredder with a very good memory of nothing.

Slide 17: The chat window is a shredder with a very good memory of nothing. Tap to enlarge

Rework share shown at 30% — an assumption. Measure yours before budgeting against it.

18 · Lever 5 · Run commodity work locally

Smaller models, your hardware, nothing leaves the building.

Slide 18: Smaller models, your hardware, nothing leaves the building. Tap to enlarge

Memory = parameters × bytes per parameter (FP16 2, INT8 1, INT4 0.5). Break-even assumes premium-tier volume; it does not hold at commodity-tier pricing.

19 · 04 · What you don't spend

What you don't put in the prompt is a line on the budget too.

Slide 19: What you don't put in the prompt is a line on the budget too. Tap to enlarge

Categories shown are the four that show up most often in enterprise prompt audits.

20 · 06 · Who gets the lever

The most expensive line item is the one that isn't on the invoice.

Slide 20: The most expensive line item is the one that isn't on the invoice. Tap to enlarge

Pew Research Center, 5,119 U.S. adults, surveyed Feb 17–23 2026. Corroborated by Lean In’s 2026 workplace survey (via Forbes).

21 · 06 · Who gets the lever

The appetite is already there. The training is the missing line item.

Slide 21: The appetite is already there. The training is the missing line item. Tap to enlarge

63%: Skillsoft Women in Tech Report. 46%: BCG, surveying C-suite expectations for reskilling over three years.

22 · Where to start

Four moves you can make without a budget request.

Slide 22: Four moves you can make without a budget request. Tap to enlarge
23 · Live demo

This is what “it reads constantly” looks like.

Slide 23: This is what “it reads constantly” looks like. Tap to enlarge

ctarp, an agentic security scanner, localizing CVE-2024-2952 against the commit before its fix — no advisory, no hint, just tokens. Read/write ratio shown at left uses Scenario 2's figures.

24 · Wrapping up

The takeaways worth keeping.

Slide 24: The takeaways worth keeping. Tap to enlarge
Pinch to zoom · tap anywhere to close