ZeroGPU sells to AI agents - no signup

The agent storefront at https://agents.zerogpu.ai/llms.txt publishes prices, payment instructions, and setup. Agents can purchase autonomously or with their human's approval, with a range of payment methods; an agent without a payment method can set one up there instantly, no signup. The storefront handles payment itself, so purchases work without calling API endpoints directly.

    Independent benchmark · 10,000 samples · blind AI judge

    zlm-v1-iab-classify-edge: 66.0% win rate against GPT-5.4 Nano

    ZeroGPU's content-classification model was tested head-to-head with OpenAI's gpt-5.4-nano on 10,000 identical, real-world samples — scored blind by a separate, more capable judge (gpt-5.5) and against verified third-party ground truth.

    Head-to-head win-rate
    66%
    vs gpt-5.4-nano 34% · 393 ties excluded
    Faster in production
    ~10×
    48 ms p50 vs ~1,900 ms · 100% success
    Cheaper vs GPT-5.4 Nano
    3.3×
    $0.05/$0.40 vs $0.20/$1.25 per 1M (in/out)
    Hallucinated labels
    0
    vs gpt-5.4-nano's 1,796 hallucinated rows

    01How this benchmark was measured

    Both systems ran the same task under identical conditions. The setup is reproducible and free of bias toward either model — identical inputs, a blind judge, and independent ground truth.

    1
    The same 10,000 samples

    Both models classified one identical, frozen set of 10,000 real, English-language samples — a mix of live production traffic and an independent third-party reference dataset (figure-eight). Neither model was trained or tuned on this test set.

    2
    A blind, independent judge

    Each pair of results was scored by gpt-5.5, a separate and more capable model that selected the more accurate output. It was never shown which system produced which answer, so it could not favour ours (seed 42, pack size 25).

    3
    Verified ground truth

    On the third-party reference set, where correct labels are established in advance, both models were also scored directly against those answers — an objective measure (precision / recall / F1) that depends on no judge at all.

    02Head-to-head accuracy

    The win-rate is the share of samples where the judge rated a model's labels as more accurate, after excluding ties. Across 9,592 decided samples, zlm-v1-iab-classify-edge was chosen 66% of the time.

    zlm-v1-iab-classify-edge 66.0% (6,329) gpt-5.4-nano 34.0% (3,263)
    DatasetSamples zlm-v1-iab-classify-edge win-rate 
    Production traffic (gdrive)7,216
    71.0%
    Independent gold set (figure-eight)2,784
    52.9%

    Counted over decided comparisons (ties excluded): 6,329 wins, 3,263 losses, 393 ties, 15 both-wrong across 10,000 rows.

    Accuracy vs. speed

    FASTER → (production latency, p50) MORE ACCURATE → (win-rate) best: top-left zlm-v1-iab-classify-edge66% wins · 48 msgpt-5.4-nano34% wins · 1911 ms

    Each model is plotted by how often it won (vertical) and how quickly it responds (horizontal); the top-left corner is best. zlm-v1-iab-classify-edge sits top-left — both more often correct and several times faster.

    03Accuracy by content area

    The same win-rate broken down by content area (each with at least 60 samples), sorted strongest-first. zlm-v1-iab-classify-edge leads in 49 of 50 areas shown; the amber row marks the 1 where gpt-5.4-nano is still ahead.

    Content areaSamples zlm-v1-iab-classify-edge win-rate 
    Automotive222
    85%
    Home & Garden222
    80%
    Food & Drink222
    79%
    Hobbies & Interests222
    78%
    Real Estate223
    77%
    News and Politics223
    77%
    Personal Finance223
    76%
    Fine Art221
    75%
    Television222
    75%
    Pop Culture223
    75%
    Education221
    75%
    Business and Finance223
    75%
    Finance123
    73%
    Pets222
    72%
    Movies221
    72%
    Unknown223
    72%
    Religion & Spirituality222
    71%
    Events and Attractions222
    71%
    Video Gaming223
    71%
    Science358
    68%
    Shopping318
    66%
    Content Source Geo63
    66%
    Sensitive Topics221
    66%
    Law and Government121
    66%
    Content Type222
    65%
    Travel331
    65%
    Music and Audio221
    64%
    Careers221
    64%
    Style & Fashion221
    64%
    Gambling102
    61%
    Healthy Living222
    61%
    Beauty and Fitness123
    60%
    Medical Health222
    60%
    People and Society119
    59%
    Reference108
    59%
    Family and Relationships222
    58%
    Technology & Computing222
    58%
    Home and Garden73
    58%
    Recreation and Hobbies143
    57%
    Autos and Vehicles96
    56%
    Sports317
    56%
    Books and Literature332
    56%
    Pets and Animals125
    55%
    Internet and Telecom93
    55%
    Adult75
    54%
    Business and Industry120
    53%
    Computer and Electronics139
    52%
    Food and Drink126
    52%
    Arts and Entertainment113
    50%
    Career and Education113
    49%

    04Accuracy vs. verified ground truth

    Beyond the judge, both models were scored directly against pre-established correct labels. Hard match requires the exact IAB tier-2 code; soft match gives hierarchical credit at tier-1. zlm-v1-iab-classify-edge leads on F1 in every cut; the only metric where gpt-5.4-nano edges ahead anywhere is figure-eight soft-match recall (0.721 vs 0.710).

    DatasetSystemPrecisionRecallF1
    Overallzlm-v1-iab-classify-edge0.2320.3100.265
    gpt-5.4-nano0.2470.2420.245
    Production (gdrive)zlm-v1-iab-classify-edge0.2220.2800.247
    gpt-5.4-nano0.2420.2080.224
    Independent gold (figure-eight)zlm-v1-iab-classify-edge0.2950.6480.406
    gpt-5.4-nano0.2670.6250.374

    Hard match — exact IAB c10 tier-2 code. Coverage: gdrive 6,993/7,216 · figure-eight 2,408/2,784.

    DatasetSystemPrecisionRecallF1
    Overallzlm-v1-iab-classify-edge0.2950.5620.387
    gpt-5.4-nano0.2910.4960.367
    Production (gdrive)zlm-v1-iab-classify-edge0.3020.5310.385
    gpt-5.4-nano0.3030.4500.362
    Independent gold (figure-eight)zlm-v1-iab-classify-edge0.2760.7100.397
    gpt-5.4-nano0.2600.7210.382

    Soft match — tier-1 hierarchical credit. Coverage: gdrive 6,993/7,216 · figure-eight 2,408/2,784.

    On the independent figure-eight gold set, zlm-v1-iab-classify-edge clears the Nano bar on hard-F1 (0.406 vs 0.374).

    05Speed & cost

    A production classifier must be fast and cheap as well as accurate. In production, zlm-v1-iab-classify-edge responds in 48 ms (p50) — about ~10× faster end-to-end than the incumbent — and is 3.3× cheaper per classification.

    Latency — measured in production
    Systemp50p95p99Success
    zlm-v1-iab-classify-edge48 ms95 ms197 ms100%
    gpt-5.4-nano~1,800–2,000 ms

    About ~10× faster end-to-end. Measured against Dappier's live publisher network (ZeroGPU × Dappier case study, 2026); the head-to-head latency in this benchmark was recorded in a development environment and understates production.

    Price per 1M tokens
    SystemInputOutput
    zlm-v1-iab-classify-edge$0.05$0.40
    gpt-5.4-nano$0.20$1.25

    Published list prices ($/1M tokens). At this benchmark's average request size (260 input + 100 output tokens), zlm-v1-iab-classify-edge works out 3.3× cheaper per classification than gpt-5.4-nano.

    Price vs. other production classifiers

    Blended cost per 1M tokens at the benchmark's 260-in / 100-out mix, against the cheapest small models from OpenAI, Google, and Anthropic — lower is better.

    zlm-v1-iab-classify-edge
    $0.147/1M ($0.05/$0.40 in/out)
    GPT-5.4 Nano
    $0.492/1M ($0.20/$1.25 in/out) · 3.3× zlm-v1-iab-classify-edge
    Gemini 2.5 Flash
    $0.911/1M ($0.30/$2.50 in/out) · 6.2× zlm-v1-iab-classify-edge
    Claude Haiku 4.5
    $2.111/1M ($1.00/$5.00 in/out) · 14.3× zlm-v1-iab-classify-edge

    Sources: OpenAI, Google (ai.google.dev), and Anthropic published list prices (verified June 2026). zlm-v1-iab-classify-edge is 6.2× cheaper than Gemini 2.5 Flash and 14.3× cheaper than Claude Haiku 4.5.

    06Input efficiency — no system prompt, no prompt engineering

    A frontier model has to be told the entire IAB taxonomy and the output format on every request — thousands of tokens of instructions wrapped around each item. zlm-v1-iab-classify-edge has the taxonomy baked in: you send only the text. That is ~6× fewer input tokens per classification, and no prompt to maintain.

    ZeroGPU zlm-v1-iab-classify-edge — the entire request
    curl https://api.zerogpu.ai/v1/responses \ -H 'content-type: application/json' \ -H 'x-api-key: zgpu-api-••••••••••••' \ -H 'x-project-id: ••••••••-••••-••••' \ -d '{ "input": "Technology has quietly reshaped the rhythm of everyday life, weaving itself into routines so seamlessly that it often goes unnoticed. From a smartphone alarm in the morning to the last glance at a glowing screen before sleep...", "model": "zlm-v1-iab-classify-edge" }'

    ~400 input tokens. No system prompt, no taxonomy in the request, no few-shot examples — just the text and the model name.

    Frontier model — what it needs every call
    # system prompt — sent on every request You are a deterministic domain classifier. Classify the following content into the IAB v1 taxonomy. # + the full IAB content + audience taxonomy enumerated # + the required output schema and formatting rules # + the content to classify (often padded with retrieved context) Domain: cnn.com Website Content (truncated): Breaking News, Latest News and Videos | CNN ... ≈ 18 KB of instructions + content, re-sent every call

    ~2,000–3,000 input tokens. A full ~18 KB instruction + taxonomy + content prompt, re-sent and re-billed on every single request.

    Input-token figures from the ZeroGPU × Dappier case study, 2026 (400 vs ~2,000–3,000 tokens, ~6× fewer). Fewer input tokens means lower cost, lower latency, and nothing to prompt-engineer or keep in sync with taxonomy updates.

    07Reliability & output quality

    A production model must return valid, well-formed labels. zlm-v1-iab-classify-edge is constrained to the taxonomy's label space, so it never emits an off-taxonomy (hallucinated) label.

    Hallucination rate
    0%
    0 hallucinated labels — vs gpt-5.4-nano, which invented off-taxonomy labels on 1,796 of 10,000 rows
    Labels returned per item
    5.78
    avg. IAB content categories assigned per item (plus 5.2 audience) — focused, high-confidence tagging, not a long noisy list
    Production response time
    48 ms
    p50 · 95 ms p95 · 197 ms p99 · 100% success

    The bottom line

    • Wins 66% of 9,592 blind head-to-head comparisons against gpt-5.4-nano.
    • Leads in 49 of 50 content areas shown, and on production traffic wins 71% of the time.
    • Clears the Nano bar on verified ground truth — figure-eight hard-F1 0.406 vs 0.374.
    • Hallucinated 0 labels across all 10,000 samples — vs gpt-5.4-nano's 1,796 rows with invented, off-taxonomy labels.
    • ~10× faster in production, 3.3× cheaper than gpt-5.4-nano (6× vs Gemini Flash, 14× vs Haiku), and uses ~6× fewer input tokens with no system prompt.

    Notes & limitations

    Fair-reading notes. Among the content areas shown, gpt-5.4-nano still leads in 1 (Career and Education). The win-rate gap is narrowest on the independent figure-eight gold set (52.9%) and widest on production traffic (71.0%). Latency is measured in production against Dappier's live publisher network (ZeroGPU × Dappier case study, 2026); the accuracy head-to-head above was run on a frozen test set in a development environment, whose latency understates production and is not reported here. Prices are published list prices (OpenAI, Google, Anthropic), blended at this benchmark's average request size.
    Run details
    Test set10,000 frozen English rows — production traffic + independent figure-eight gold set
    Judgegpt-5.5 — blind, pack size 25, seed 42
    ZeroGPU modelzlm-v1-iab-classify-edge — content + audience
    Competitorgpt-5.4-nano
    Gold coveragegdrive 6,993/7,216 · figure-eight 2,408/2,784
    Production latency48/95/197 ms p50/p95/p99 · 100% success (ZeroGPU × Dappier case study, 2026)
    Hallucinated labelszlm-v1-iab-classify-edge 0 · gpt-5.4-nano 1,796
    Generated2026-06-09T13:05:43