The Model Gateway Is a Compliance Boundary
Every company routing LLM calls across providers builds a gateway. In financial services, that gateway is a regulatory surface whether you designed it as one or not.
If you're running LLM calls in production, you've probably built something that looks like a gateway. A layer between your application and the model providers that handles routing, fallback, maybe some cost optimization. Every team building on AI eventually writes one, or adopts one.
The standard version is well-understood. Route to the cheapest model that meets the latency target. Fall back when a provider goes down. Cache identical prompts. Log requests for debugging.
But if regulated data touches those prompts, the gateway you built for cost optimization just became something else entirely. And the distance between "model router" and "compliance boundary" is measured in the regulatory obligations you didn't know you were triggering.
The inventory problem
In April 2026, the Fed and OCC issued SR 26-2, extending the existing SR 11-7 model risk management framework explicitly to AI systems. SR 11-7 has been around since 2011. It defines "model" broadly: anything that takes inputs and produces quantitative output for decision-making. LLMs that influence consumer-facing outcomes fit.
The requirement is specific. Every model must be inventoried. Each must undergo independent validation before deployment and ongoing monitoring afterward. The institution needs to know which models are in use, which version of each model produced which decision, and whether performance has degraded since validation.
Without a centralized gateway, teams call models directly. An engineer in the fraud team uses Claude. The lending team calls GPT-4o. The customer support team experiments with Gemini. Each integration is a separate API client, a separate logging setup, a separate shadow on the model inventory.
Think of it like a construction site where every subcontractor sources their own materials. The general contractor has no bill of materials. The building inspector shows up and asks what's holding up the third floor, and nobody can answer from a single document.
A centralized gateway doesn't just route calls. It becomes the model registry. Every request passes through one layer that records which model, which version, which provider, when. The inventory assembles itself from traffic rather than from a spreadsheet someone updates quarterly.
Audit trail depth
Generic gateways log what you'd expect: timestamp, model, tokens used, latency, status code. Enough for debugging and cost attribution.
Financial regulators want more. For broker-dealers whose AI outputs reach clients, SEC Rule 17a-4 and FINRA Rule 4511 require immutable, auditable records. The obligation activates when AI-generated content is transmitted externally: a recommendation, a risk assessment, a compliance decision communicated to a customer.
The record must capture the output, the input data that produced it, and the model configuration at the time of generation. Retention runs three to six years depending on the record type. The storage itself must be tamper-evident: write-once or with a verified audit trail.
A generic gateway logging request and response as JSON to CloudWatch doesn't meet this bar. The logging layer needs to capture the full prompt (including system prompts and any retrieved context), the complete response, the model identifier down to the version or checkpoint, the user or system that initiated the request, and enough metadata to reconstruct why that model was selected over alternatives.
That's roughly twelve fields per decision, written to immutable storage, indexed for retrieval under regulatory examination. Not a logging enhancement. A different storage architecture bolted onto the routing layer.
Where the prompt can't go
A customer service agent summarizes a dispute. The prompt includes the cardholder's name, the last four digits of their card, and the transaction amount. Routine.
PCI DSS 4.0.1 doesn't mention AI. It doesn't need to. It applies to any system that stores, processes, or transmits cardholder data. The moment that dispute summary enters a prompt, the model endpoint, the network path to it, and every log that captured the request are in PCI scope. Eight of the twelve PCI requirement families apply directly.
In a generic gateway, the prompt passes through to whatever model the routing logic selects. In a financial-grade gateway, a data loss prevention layer sits between the application and the router. It inspects outbound prompts for cardholder data, account numbers, Social Security numbers. Anything matching a regulated pattern gets redacted or tokenized before the prompt reaches the model.
Same DLP pattern enterprises use for email and file sharing, applied to a new surface. The difference is latency budget. Email DLP can take seconds. A model gateway adding DLP to the hot path needs to do it in single-digit milliseconds or the downstream latency SLA breaks.
Data residency adds another routing constraint. A prompt containing EU customer data can't route to inference infrastructure outside approved jurisdictions under GDPR.
The EU AI Act, fully enforceable since August 2026, classifies AI systems used in credit scoring and insurance pricing as high-risk, with penalties reaching 7% of global turnover. DORA, the EU's digital operational resilience regulation, requires explicit contractual provisions covering where third-party inference runs.
According to Redress Compliance, 67% of organizations have no contractually enforceable residency restriction covering inference. They assume their cloud provider's region settings handle it. But region settings govern storage, not necessarily inference. A prompt sent to a US-region API endpoint may still be processed on hardware in a different jurisdiction depending on the provider's load balancing.
So the gateway doesn't just pick the cheapest model. It picks the cheapest model that's approved for this data classification, running in an approved jurisdiction, with a provider whose contractual terms cover the regulatory regime the data falls under. The routing decision tree goes from two variables (cost, latency) to five or six.
Who's already building this way
JP Morgan built LLM Suite, a centralized internal platform available to 250,000 employees. Roughly half use it daily. It routes across multiple model providers with an eight-week update cycle for new models. Teresa Heitsenrether, JP Morgan's Chief Data and Analytics Officer, described the intent as routing fluidly across providers while maintaining centralized control over data filtering, audit logging, and security.
A bank building an internal gateway where the compliance requirements came first and the cost optimization came second.
Ramp took the opposite path. They built model routing internally to manage their own AI spend, ran it for three years, then launched it publicly as Router. It processes 2.75 trillion tokens per month and advertises 40% cost reduction on inference. The product pitch is cost optimization, but the architecture had to satisfy Ramp's own compliance requirements before it could serve external customers.
Stripe's acquisition of OpenRouter for roughly $7 billion is the market-level signal. I covered the business strategy in "What Stripe Gets from OpenRouter": meter usage, route the model call, collect the payment. But the engineering underneath has to handle every constraint in this article the moment a Stripe customer in financial services routes a single prompt containing regulated data through the platform.
Three companies, three different starting points, all converging on the same architecture: a routing layer that enforces compliance at the point where prompts leave the building.
The five decisions
A generic model gateway makes two decisions per request: which model and which provider.
A financial-grade model gateway makes at least five. Which models are approved for this use case under the institution's SR 11-7 inventory. Which providers operate inference infrastructure in jurisdictions approved for this data classification. Whether the prompt contains regulated data that needs redaction before it leaves the gateway.
What fields to log, at what granularity, to what storage tier, for how long. And then, after all of that, which of the remaining eligible models is cheapest at acceptable latency.
Cost optimization is the last decision, not the first. That ordering is what makes this a different system, not a configured version of the same one.
Sources
- OCC/Fed SR 11-7: Supervisory Guidance on Model Risk Management - Original 2011 guidance defining model inventory, validation, and monitoring requirements
- SR 26-2: Extension of Model Risk Management to AI Systems - April 2026 guidance explicitly extending SR 11-7 to AI systems including LLMs
- JP Morgan LLM Suite - Architecture and adoption details of JP Morgan's centralized AI routing platform
- Ramp Router - Ramp's model routing infrastructure processing 2.75T tokens/month, built from three years of internal use
- VGS: AI and PCI Compliance 2026 - Analysis of how PCI DSS 4.0.1 applies to AI systems processing cardholder data
- Redress Compliance: AI Data Residency - Survey finding 67% of organizations lack contractual residency restrictions for inference
- EU AI Act: High-Risk AI in Financial Services - Classification of AI in credit scoring and insurance as high-risk, penalties up to 7% global turnover
Frequently Asked Questions
Built by Trio, a fintech-native engineering partner helping teams build the next generation of financial technology and infrastructure.
Subscribe to Ledger Drift for high-signal insights into how modern fintech is built, from systems to code to teams.