Skip to content

Success story · AI chat on abliterated models · Worldwide

Submissive AIFour people took an AI product from idea to launch in 9 weeks, with inference 60%+ cheaper than API providers

Submissive AI is a usage-based chat on abliterated models, meaning open-weight models with their refusals removed, plus models we fine-tuned ourselves, for red-team and professional use. I was hired as fractional CTO to get the product off the ground, and it went from idea to launch in 9 weeks. I designed the whole system and its inference engine, and I run the distributed data centers, the billing pipelines and the cost of inference: our inference costs 60%+ less than buying it from API providers, and GPU utilisation stays above 80%. Four people cover work that usually needs separate ML, infrastructure and billing teams.

Client
Submissive AI
Product
Usage-based AI chat on abliterated models
Use
Red teaming and professional work
My role
Fractional CTO, hired to launch the product
Team
4 people
Idea to launch
9 weeks
Inference
60%+ cheaper than API providers, GPUs 80%+ utilised
Submissive AI: the product

In numbers

from idea to launch, for one of the hardest kinds of AI product
9 weeks
cheaper inference than API providers, on our own engine
60%+
GPU utilisation, so the hardware we pay for is busy most of the time
80%+
people run what usually takes separate ML and infra teams
4
of models: open-weight abliterated, and our own fine-tuned
2 kinds
countries host our inference, with load overflowing between them
Several

The product

A usage-based chat with a choice of abliterated models, paid from a balance that tops itself up. Pictures from submissive-ai.com.

  • The chat: conversations on the left, a choice of abliterated models above the answer
  • Billing: the balance, card payments and automatic top-ups
01

The starting point

Red teams, and professionals who test and work with AI, need models that answer the question as it was asked, without refusals and without rewriting the prompt. That means running abliterated models: open-weight models with their refusal behaviour removed, and models fine-tuned for the job.

Technically this is one of the hardest kinds of AI product to build. It takes custom models, fine-tuning, test harnesses and inference you host yourself across several countries, where every GPU-hour costs money whether anyone uses it or not. Get the architecture wrong and the GPU bill outgrows the revenue.

02

What we did

  1. Hired to launch it, live in 9 weeks

    The company brought me in as fractional CTO specifically to get this product off the ground, and we went from idea to launch in 9 weeks. I lead a team of four and set how it works.

  2. The system and the inference engine

    I designed the whole system and the inference engine that serves the abliterated models. With its own engine the company pays for compute, and no other provider takes a margin on every token, so our inference costs 60%+ less than buying it from API providers.

  3. Open-weight and our own models

    We run ready open-weight models alongside the ones we fine-tuned ourselves, with harnesses that test and evaluate each of them, so a model reaches customers only after it has proven itself.

  4. Data centers in several countries

    Inference runs in our data centers in several countries around the world, and I manage the contractors who provide them. The company gets capacity around the world without building an infrastructure department.

  5. Inference cost under control, GPUs 80%+ busy

    I optimise the cost of inference: which model runs where, on what hardware, and how much each answer costs us. GPU utilisation stays above 80%, so the GPU bill follows the revenue instead of running ahead of it.

  6. Overflow and billing pipelines

    I build the overflow and billing pipelines that keep the product correct across distributed inference data centers. Load that one site cannot take moves to another, and every request is billed to the customer's balance exactly once, so no request is lost and no revenue leaks through missed or double charges.

03

The result

A working AI product on abliterated and custom-trained models, launched 9 weeks after the idea and served by its own inference engine from data centers in several countries. Its inference costs 60%+ less than buying it from API providers and its GPUs run above 80% utilisation, so a team of four keeps cost and billing under control, which is exactly where AI products of this kind usually lose their money.

oleg.isSent on request

Full case · PDF

Submissive AI

Four people took an AI product from idea to launch in 9 weeks, with inference 60%+ cheaper than API providers

  1. 01The architecture of the system and the inference engine
  2. 02Abliterated and fine-tuned models: the pipeline
  3. 03Harnesses and evaluation
  4. 04Inference across data centers in several countries
  5. 05Optimising the cost of every answer
  6. 06Overflow and billing pipelines: no lost requests, no double charges
Prepared by Oleg SotnikovPDF

The full Submissive AI case

The public story stops here. The PDF has the rest: the architecture before and after, the migration plan, how the team works with AI agents, what it all costs to run and where the savings came from.

I send every case myself and use your email for nothing else.

More stories

All stories
01AulaHousing and utilities · KazakhstanFractional CTOThree releases a day instead of one in four months, 3 people doing the work of 16, and a cloud bill down 70%Aula is one app for apartment residents and the 109 companies that manage their buildings, with 247,000 accounts. As fractional CTO I rebuilt how the product is made and run: development and testing went from 16 people to 3, the servers from 18 to 3, the cloud bill fell by 70% and errors in the mobile app by 97%. A task now reaches production in a day instead of 2 to 3 weeks, and the product ships three times a day at 99.99% uptime, a pace most in-house teams never reach.Read the story02MetizPromServiceMetalworking and parts manufacturing · KazakhstanFractional CTOA factory saves $300,000+ a year and halved its IT costs after two engineers I hired replaced its contractorsMetizPromService machines parts to customers' drawings in Kazakhstan. When I joined, contractors ran its IT and its accounting, production and lathe-control systems, and the company paid for outside software and licences. I built its own infrastructure, hired and trained two engineers, and we replaced every paid production system with an ERP and CRM of our own, with the ERP live in 3 months. The company now saves more than $300,000 a year on software, spends over 50% less on IT than the contractors cost, salaries included, and its CNC machines stand idle 23% less, with half as many unplanned repairs.Read the story03CV RocketJob search and AI CVs · USA and worldwideFractional CTOOne engineer and a fractional CTO took CV Rocket from zero to sales in five weeks and to 1,000+ customersCV Rocket writes a CV for one job posting in 15 to 50 minutes and brings the US and world job market into one place: more than 1.5 million postings from 89,000+ job sites. I designed the architecture, the cloud services and the payment integration, and one engineer and I built the product in four weeks, a launch that usually takes a full team several quarters. It sold from week five and now has more than 1,000 customers.Read the story

Want the next story to be about your company?

In a 30-minute call we pick the first task to hand to AI and estimate what it will save you.

Book a callThe call is free, and you talk to me directly.