๐ฅ Check out this insightful post from Hacker News ๐
๐ **Category**:
โ **What Youโll Learn**:
A Kill-A-Watt meter for your AI agents. Point it at a trace and it tells
you exactly where your tokens are being burned and wasted, prices each waste
pattern in real dollars, prescribes a fix, and can fail your CI when a change
makes your agent measurably more expensive.

A real captured agent trace (see provenance) โ
Wattage catches a stable prompt prefix being re-sent instead of cached, prices
the waste, and prescribes the fix. Regenerate this GIF with
vhs docs/assets/demo.tape (see the tape file for the exact command).
uvx wattage report trace.json
No config file, no API key, fully offline โ point it at an OTLP JSON
trace export and it prices every call and runs every detector. Don’t have a
trace yet? Getting your first trace covers both
“I already have OTel traces” and “I have zero instrumentation” (a runnable,
5-minute path from nothing to a real, priced report). Or try it right now
against the fixture shipped in this repo:
git clone https://github.com/faizannraza/wattage
cd wattage && uv sync
uv run wattage report examples/sample_trace.json
โญโโโโ โก wattage โ examples/sample_trace.json โโโโโฎ
โ Token Efficiency: A (100) Total cost: $0.0602 โ
โ quality: unmeasured โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
Token breakdown
โโโโโโโโโโโโโโโโโโณโโโโโโโโโ
โ Category โ Tokens โ
โกโโโโโโโโโโโโโโโโโโโโโโโโโโฉ
โ input โ 18450 โ
โ output โ 320 โ
โ cache_read โ 0 โ
โ cache_creation โ 0 โ
โ reasoning โ 0 โ
โโโโโโโโโโโโโโโโโโดโโโโโโโโโ
No findings โ this trace looks efficient.
pricing: 2026-07-18-verified
Or get a self-contained, shareable HTML flame graph instead of the terminal
view:
uv run wattage report examples/sample_trace.json --html report.html
The evidence, not a marketing claim
Wattage’s standout feature is the convergence engine โ the
nonconvergence detector, which catches an agent thrashing through a loop
without making real progress, including patterns a naive exact-match
duplicate detector structurally cannot see (a retry with a fresh timestamp
each time, an oscillation between two strategies, a “productive-looking”
stall where every call is technically unique but nothing is actually
learned).
Rather than assert that, we built a hand-reviewed set of 10 labeled
synthetic loops and benchmarked Wattage’s classifier against a real
SHA-256 exact-match baseline implementation:
| Classifier | Precision | Recall | F1 |
|---|---|---|---|
| Wattage | 1.00 | 1.00 | 1.00 |
| SHA-256 exact-match | 1.00 | 0.14 | 0.25 |
Reproduce it yourself โ no cherry-picking, no hidden setup:
uv run python -m benchmarks.harness
And on a genuine captured agent trace (not synthetic โ see
benchmarks/traces/README.md for provenance),
Wattage’s prefix_churn fix simulation shows a 44.7% cost reduction
($0.000199 โ $0.000110) from enabling prompt caching on the stable prefix โ
small dollar figures because it’s a 3-turn demo trace, but the mechanism is
identical at production scale. Run it against your own traces for numbers
that matter:
uv run python -c "from benchmarks.frontier import build_frontier; print(build_frontier())"
Full methodology: The Convergence Engine.
uv run wattage badge trace.json --out wattage-badge.svg

Wire --badge-out into your CI job (see below) so it regenerates on every
merge to your default branch, and the badge in your README stays live.
Three surfaces, one normalized data model underneath
(sessions โ tasks โ loops โ iterations โ calls), built from
OpenTelemetry GenAI semantic-convention
traces:
-
wattage reportโ ingests a trace, prices every call against a
vendored, dated pricing snapshot, and runs eight detectors:Detector Catches prefix_churnStable context re-sent instead of cached cache_gapCaching attempted but under-redeemed by later reads verbosityOutput far beyond what the step needed redundant_tool_callsThe same tool call repeated (exact or fuzzy) nonconvergenceLoops that thrash, oscillate, or stall without progress retrieval_thrashRepeated retrieval that never yields relevant results model_mismatchA pricier model doing work a cheaper one could handle reasoning_overspendHeavy reasoning-token spend on a simple step Every finding is priced in real dollars, includes a concrete fix, and is
tagged with aquality_risktier (none/low/review) โ a fix that
could plausibly change output quality (a model downgrade, less reasoning)
only counts toward your score once a--qualitymap backs it with real
evidence. Full detail: Detectors. -
wattage score/wattage badgeโ a single 0โ100 Token Efficiency
grade for a README badge or a CI gate. -
wattage ciโ the cost-regression gate (below).
Wattage never fabricates a number: an unpriced model leaves that call’s cost
at zero (and fails wattage ci loudly, exit code 4) rather than guessing;
an unmeasured quality signal is reported as unmeasured, not assumed fine.
# .github/workflows/wattage.yml
name: Wattage
on:
pull_request:
paths: ["agents/**", "prompts/**", "src/**"]
concurrency:
group: wattage-$๐ฌ
cancel-in-progress: true
jobs:
token-efficiency:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Generate trace fixture
run: python scripts/run_agent_fixture.py > trace.json
- name: Wattage cost-regression gate
uses: faizannraza/wattage/action@v0.1.0
with:
source: trace.json
baseline: .wattage/baseline.json
fail-on: "score_below:80,cost_delta_pct_above:5,any_critical:true"
pr-comment: "true"
Fails the build (exit code 1) when your agent regresses past the threshold
you set, posts a per-detector delta table as a PR comment, and emits SARIF
(shows up in GitHub’s Security tab) and JUnit XML for any other CI system.
The baseline is a small committed JSON file โ noise-floor protection is
structural, not statistical: it only ever updates on a run that actually
passed the gate.
This is only half the setup. A PR job runs on a throwaway checkout, so
it can’t be the thing that updates .wattage/baseline.json on disk โ that
update needs a second workflow, triggered on push to your default branch,
that commits the refreshed baseline (and badge) back after each merge.
Skipping it means every PR compares against the same stale baseline
forever. Full reference, with both workflows: CI Integration.
Detectors are discovered through a Python entry-point group, so adding one
doesn’t require touching this repo’s core pipeline โ see
CONTRIBUTING.md for the full “write a detector” walkthrough,
using cache_gap as the reference
example.
Apache-2.0
{๐ฌ|โก|๐ฅ} **Whatโs your take?**
Share your thoughts in the comments below!
#๏ธโฃ **#faizannrazawattage #tokenspend #profiler #costregression #gate #agents #GitHub**
๐ **Posted on**: 1785112509
๐ **Want more?** Click here for more info! ๐
