Skip to content

Success story · AI model gateway

AIRouterIn five weeks, one person launched an OpenRouter-class AI gateway that has since processed 168 billion+ tokens

AIRouter sends requests to many AI providers through one API and does everything OpenRouter does. On top of that it runs inference inside the customer's country, bills in local currency and issues local closing documents, so companies get the best models, stay compliant with personal-data law and pay 20% less than they would buying from the providers directly. I designed the architecture from the data center up, and the service launched in five weeks with a team of one, where a platform of this scope usually takes a full team and several quarters. It has since processed more than 168 billion tokens in more than 29.5 million API requests for 50+ client companies.

Client
AIRouter
Product
AI model gateway, OpenRouter class
Customers
50+ client companies
Traffic
168 billion+ tokens in 29.5 million+ API requests
Infrastructure
Instances in several countries, designed from the data center up
My role
Fractional CTO
Team
1 person, where such platforms usually need a team
AIRouter: the product

In numbers

tokens processed in 29.5M+ API requests, counted live on its site
168B+
to launch with a team of one, where such platforms take quarters
5 weeks
what a client pays compared with buying from providers directly
−20%
client companies reach 547 models under one contract
50+
of router overhead, with automatic fallback across 14 providers
<50 ms
inference in the country, so the data stays home and the law is met
Local

The product

The public site: every leading model under one company agreement, a catalogue of 547 models, dedicated GPU capacity in Kazakhstan and live traffic statistics.

  • The model catalogue with 547 models, prices per million tokens and capabilities
  • Dedicated GPU capacity inside Kazakhstan on NVIDIA Blackwell, Hopper and Ada Lovelace
  • Live traffic statistics: tokens processed, API requests, active models and providers
01

The starting point

Companies want the best AI models from many providers through one entry point. In many countries they also need the data to stay inside the country, to pay in their own currency, and to get closing documents their accounting will accept. Global gateways cover the first need and leave the rest to the customer.

AIRouter set out to give customers all of that in one service that is easy to use, and brought me in to design it and get it to market fast.

02

What we did

  1. One API in place of 14 separate integrations, under 50 ms added

    Requests go to different AI providers through one API, and the router picks the provider automatically, the way OpenRouter does, adding under 50 ms to a request. The catalogue lists 547 models, so a customer's developers integrate once and switch models without rewriting their code.

  2. Everything OpenRouter offers, from launch day

    When a model or provider fails, the request falls back to another one, so one vendor's outage does not become the customer's outage. Providers are chosen by price and speed, which keeps spend down without anyone tuning it by hand. The API takes images, PDFs, audio and video, supports tool calling and structured outputs, and caches responses. Teams get workspaces with budgets and guardrails for sensitive data.

  3. Data stays in the country, and the compliance risk goes away

    Inference runs locally in several countries, so a customer's data stays in its own country and the service complies with personal-data law. Companies that cannot send data to a foreign gateway can now use leading models without legal exposure.

  4. Billing the finance department accepts on day one

    Customers in each country pay in their local currency and receive local closing documents for their accounting. In Kazakhstan that means invoices in tenge, paid by bank transfer, with no currency conversion and no arguments with accounting over foreign invoices.

  5. Designed from the data center up, so a new market is a new instance

    I designed the architecture as many instances placed by country, and I was responsible for the design from the data center up. Entering another market means placing another instance of a design that already works.

  6. One person took the service to launch in five weeks

    The team was one person, and the service launched five weeks after the start. A gateway of this scope usually takes a team and several quarters to reach the market.

  7. 168 billion+ tokens for 50+ companies, at 20% below direct prices

    The service's own site counts its traffic live: more than 168 billion tokens processed in more than 29.5 million API requests. More than 50 client companies use it, and each pays 20% less than it would buying the same models from the providers directly.

03

The result

One person launched AIRouter in five weeks, and it has since processed more than 168 billion tokens for 50+ client companies, each paying 20% less than buying from the providers directly. Local markets got an OpenRouter-class service they use under their own law, pay for in their own currency and put in their own books. Infrastructure launches rarely reach the market at that speed and cost.

oleg.isSent on request

Full case · PDF

AIRouter

In five weeks, one person launched an OpenRouter-class AI gateway that has since processed 168 billion+ tokens

  1. 01The architecture: instances by country
  2. 02Data-center design for local inference
  3. 03Routing and fallback between providers, by price and speed
  4. 04Compliance with personal-data law
  5. 05Local-currency billing and closing documents
  6. 06The five-week launch plan, delivered by one person
Prepared by Oleg SotnikovPDF

The full AIRouter case

The public story stops here. The PDF has the rest: the architecture before and after, the migration plan, how the team works with AI agents, what it all costs to run and where the savings came from.

I send every case myself and use your email for nothing else.

More stories

All stories
01AulaHousing and utilities · KazakhstanFractional CTOThree releases a day instead of one in four months, 3 people doing the work of 16, and a cloud bill down 70%Aula is one app for apartment residents and the 109 companies that manage their buildings, with 247,000 accounts. As fractional CTO I rebuilt how the product is made and run: development and testing went from 16 people to 3, the servers from 18 to 3, the cloud bill fell by 70% and errors in the mobile app by 97%. A task now reaches production in a day instead of 2 to 3 weeks, and the product ships three times a day at 99.99% uptime, a pace most in-house teams never reach.Read the story02MetizPromServiceMetalworking and parts manufacturing · KazakhstanFractional CTOA factory saves $300,000+ a year and halved its IT costs after two engineers I hired replaced its contractorsMetizPromService machines parts to customers' drawings in Kazakhstan. When I joined, contractors ran its IT and its accounting, production and lathe-control systems, and the company paid for outside software and licences. I built its own infrastructure, hired and trained two engineers, and we replaced every paid production system with an ERP and CRM of our own, with the ERP live in 3 months. The company now saves more than $300,000 a year on software, spends over 50% less on IT than the contractors cost, salaries included, and its CNC machines stand idle 23% less, with half as many unplanned repairs.Read the story03CV RocketJob search and AI CVs · USA and worldwideFractional CTOOne engineer and a fractional CTO took CV Rocket from zero to sales in five weeks and to 1,000+ customersCV Rocket writes a CV for one job posting in 15 to 50 minutes and brings the US and world job market into one place: more than 1.5 million postings from 89,000+ job sites. I designed the architecture, the cloud services and the payment integration, and one engineer and I built the product in four weeks, a launch that usually takes a full team several quarters. It sold from week five and now has more than 1,000 customers.Read the story

Want the next story to be about your company?

In a 30-minute call we pick the first task to hand to AI and estimate what it will save you.

Book a callThe call is free, and you talk to me directly.