Subscribe to get high-signal insights on how modern fintech is built.

engineering

The Token Arbitrage Economy

Stolen API keys, mass-registered accounts, and USDT resale markets. AI compute fraud looks nothing like payment fraud, and the defenses are different too.

By Alex Kugell ·

Someone steals your credit card and buys a television. The bank reverses the charge. The retailer files an insurance claim. The TV might even get recovered. Fifty years of chargeback infrastructure exists for exactly this scenario.

Someone steals your API key and runs 10,000 inference requests through a frontier model. There is no reversal. The GPU cycles are spent. The electricity is consumed. And the bill lands on your account three days later, $82,000 heavier than it should be.

That $82,000 figure comes from a real incident. A developer's Gemini API key leaked from a misconfigured server. The attacker ran inference for 48 hours before anyone noticed.

A worse case: METR, an AI safety testing nonprofit, lost $600,000 in API credits over three weeks. An attacker exploited a fail-open authentication bug on a public EC2 instance, added an SSH key for persistence, and ran inference against their account. METR's normal evaluation workloads are high-volume, so the usage spike blended in. No alert fired.

Both incidents are symptoms of a supply chain that has industrialized AI compute theft.

The three-layer supply chain

The underground market for stolen AI compute operates in three tiers, each feeding the next.

Raw credential merchants sit at the bottom. They harvest API keys from public GitHub repositories, exposed environment files, misconfigured cloud instances, and phished developer accounts. Automated scanners scrape every new public commit looking for patterns that match OpenAI, Anthropic, Google, and AWS credential formats. Sysdig's Threat Research Team, which coined the term "LLMjacking" in May 2024, estimated exposure of $46,000 per day for a single compromised Claude-class credential.

Aggregation operators sit in the middle. They pool stolen credentials using open-source API proxy software, primarily projects called one-api and new-api. These are legitimate self-hosted tools designed to present a unified OpenAI-compatible API across multiple providers.

Repurposed by operators, they load-balance stolen keys, rotate credentials when one gets revoked, and present a single clean endpoint to downstream buyers. From the buyer's perspective, it looks like a normal API.

Retail frontends sit on top. Consumer-facing storefronts wrapped around pooled access, with pricing at steep discounts to official rates. Settlement in USDT, because stablecoins offer no chargebacks, pseudonymous transactions, and instant finality. The buyer has no legal recourse if the access disappears. One documented case involved a user who lost roughly $25,000 through an unofficial API gateway with no path to recovery.

The economics work because GPU inference is expensive at retail and nearly free when stolen. A buyer who pays 10% of the official token price is still profitable for the operator, who paid nothing for the compute.

Why traditional fraud tools miss this

Stripe Radar, Sift, and similar fraud prevention systems were built for a specific problem: is this payment transaction legitimate? They score a single event. A card number, a billing address, a device fingerprint, a purchase amount. The decision is binary and happens once at checkout.

AI compute fraud doesn't have a checkout. It has a continuous stream of API calls, each one individually legitimate. A single request for 2,000 input tokens and 500 output tokens looks exactly like any other API call. The fraud signal is in the sequence, not the event.

What distinguishes a stolen key from a legitimate developer? Traffic from a single API key spanning dozens of geographic regions and IP ranges. Model selection patterns inconsistent with a single application, because relay operators route across many models while a real app uses one or two. Request timing that shows multiple independent usage cadences on the same key, the fingerprint of many downstream buyers sharing pooled access.

Account farming shows a different signature. Legitimate users browse documentation before making their first API call. Farming operations hit endpoints immediately after registration. The time between account creation and first API call is one of the strongest signals, and most platforms don't measure it.

What you can build

The defenses split into two categories: things you check before inference runs, and things you detect after.

Pre-inference checks sit in the API request path. These have to be fast, because every millisecond of billing latency is a millisecond added to the inference response.

Per-key spend caps set at creation, not retrofitted after a breach. Short-lived API credentials with least-privilege scoping, so a leaked key expires before an attacker can extract significant value. Budget reservation, where the system estimates the request cost and holds it against the account balance before dispatching to the GPU. If the balance is insufficient, the request gets a 429 and the GPU never fires.

Post-inference detection runs asynchronously against the usage event stream. The signals are behavioral, not transactional.

How fast is an account burning through its balance compared to its historical baseline? A developer who normally uses $50/day suddenly consuming $5,000/hour is a signal even if each individual request looks normal.

Do five accounts share a device fingerprint, a payment instrument, or an IP address? Individually they look like separate developers. Together they look like a farming operation harvesting promotional credits.

Is a single key making requests from São Paulo, Frankfurt, and Singapore within the same minute? Has a dormant key suddenly started generating requests at 3am in the account holder's timezone?

And how many models is a single key calling? A legitimate application calls one or two consistently. A relay operator routes to whatever model the downstream buyer requests. High variance in model selection from a single key is a strong signal.

The structural asymmetry

Payment fraud and compute fraud share a word but not a structure. When someone steals a credit card number, the financial system has mechanisms to unwind the transaction. Chargebacks, provisional credits, merchant liability. The money can be recovered, or at least redistributed.

Compute theft has no unwind mechanism. The GPU ran, the watts were consumed, and the cost is sunk the moment inference begins.

Detection latency maps directly to financial loss in a way that payment fraud doesn't. Payment fraud has a recovery path. Compute fraud doesn't.

This asymmetry is why pre-inference gating matters more than post-inference detection. A payment fraud system that catches 95% of fraudulent transactions within 24 hours is effective, because the 5% that slip through can still be charged back. A compute fraud system that catches 95% of stolen-key abuse within 24 hours has already paid the full cost of the 5% that got through. Plus the full cost of the 95% it caught, because detection happened after the GPU ran.

The only way to avoid paying for stolen compute is to stop the request before it reaches the GPU. The billing gate described in "When the Billing System Decides Who Gets the GPU" is the enforcement point. Pre-inference budget reservation, entitlement verification, and anomaly scoring, all synchronous, all before the expensive part happens.

What the platform can't do alone

Platform-level defenses catch the patterns that show up in usage data. They can't catch the patterns that show up in how credentials were obtained.

Secret scanning is a developer tooling problem, not a billing problem. GitHub, GitGuardian, and TruffleHog scan repositories for exposed credentials. But they only cover public repositories, and they only catch credential formats they recognize. AI API keys don't always follow the patterns these scanners were trained on.

Credential rotation and short-lived tokens are an authentication architecture decision. If your API keys are long-lived bearer tokens with no scope restrictions, a single leak exposes your entire account. If they're scoped to specific models, rate-limited by default, and expire after 24 hours, the blast radius of a compromise shrinks dramatically.

No amount of pre-inference gating helps if the key was committed to a public repo three hours ago.

Sources

Frequently Asked Questions

What is LLMjacking?
LLMjacking is the theft and resale of AI API credentials. Attackers harvest API keys from public repos, exposed environment files, and phished accounts, then pool them through proxy software and resell access at steep discounts, typically settled in USDT.
How much can a stolen AI API key cost the victim?
A single compromised Claude-class credential can expose $46,000 per day in compute costs. In one documented case, a developer's Gemini key was exploited for $82,000 in 48 hours. METR, an AI safety nonprofit, lost $600,000 over three weeks.
Why don't traditional fraud tools catch AI compute theft?
Tools like Stripe Radar score single events at checkout. AI compute fraud is a continuous stream of individually legitimate API calls. The fraud signal is in the sequence, not the event: geographic spread, model selection inconsistency, and multiple usage cadences on one key.

Built by Trio, a fintech-native engineering partner helping teams build the next generation of financial technology and infrastructure.

Subscribe to Ledger Drift for high-signal insights into how modern fintech is built, from systems to code to teams.

Keep reading

engineeringYour AI Agent Has a Credit Card and No Spending LimitI built a policy engine for agent payments. Here's what the problem actually looks like when you try to solve it.
fintechWhen the Billing System Decides Who Gets the GPUAt AI-company scale, the billing system sits in the inference hot path. Every API request passes through it before the G...
analysisThe 20-Point Gap Between Stablecoin Adoption and Stablecoin ProtectionVisa asked Americans if they'd use stablecoins with bank-level protections. Adoption jumped 20 points. The GENIUS Act ju...
View more ›