Success story · AI chat on abliterated models · Worldwide
Submissive AIFour people took an AI product from idea to launch in 9 weeks, with inference 60%+ cheaper than API providers
Submissive AI is a usage-based chat on abliterated models, meaning open-weight models with their refusals removed, plus models we fine-tuned ourselves, for red-team and professional use. I was hired as fractional CTO to get the product off the ground, and it went from idea to launch in 9 weeks. I designed the whole system and its inference engine, and I run the distributed data centers, the billing pipelines and the cost of inference: our inference costs 60%+ less than buying it from API providers, and GPU utilisation stays above 80%. Four people cover work that usually needs separate ML, infrastructure and billing teams.
- Client
- Submissive AI
- Product
- Usage-based AI chat on abliterated models
- Use
- Red teaming and professional work
- My role
- Fractional CTO, hired to launch the product
- Team
- 4 people
- Idea to launch
- 9 weeks
- Inference
- 60%+ cheaper than API providers, GPUs 80%+ utilised
- Website
- submissive-ai.com

In numbers
- from idea to launch, for one of the hardest kinds of AI product
- 9 weeks
- cheaper inference than API providers, on our own engine
- 60%+
- GPU utilisation, so the hardware we pay for is busy most of the time
- 80%+
- people run what usually takes separate ML and infra teams
- 4
- of models: open-weight abliterated, and our own fine-tuned
- 2 kinds
- countries host our inference, with load overflowing between them
- Several
The product
A usage-based chat with a choice of abliterated models, paid from a balance that tops itself up. Pictures from submissive-ai.com.
The starting point
Red teams, and professionals who test and work with AI, need models that answer the question as it was asked, without refusals and without rewriting the prompt. That means running abliterated models: open-weight models with their refusal behaviour removed, and models fine-tuned for the job.
Technically this is one of the hardest kinds of AI product to build. It takes custom models, fine-tuning, test harnesses and inference you host yourself across several countries, where every GPU-hour costs money whether anyone uses it or not. Get the architecture wrong and the GPU bill outgrows the revenue.
What we did
Hired to launch it, live in 9 weeks
The company brought me in as fractional CTO specifically to get this product off the ground, and we went from idea to launch in 9 weeks. I lead a team of four and set how it works.
The system and the inference engine
I designed the whole system and the inference engine that serves the abliterated models. With its own engine the company pays for compute, and no other provider takes a margin on every token, so our inference costs 60%+ less than buying it from API providers.
Open-weight and our own models
We run ready open-weight models alongside the ones we fine-tuned ourselves, with harnesses that test and evaluate each of them, so a model reaches customers only after it has proven itself.
Data centers in several countries
Inference runs in our data centers in several countries around the world, and I manage the contractors who provide them. The company gets capacity around the world without building an infrastructure department.
Inference cost under control, GPUs 80%+ busy
I optimise the cost of inference: which model runs where, on what hardware, and how much each answer costs us. GPU utilisation stays above 80%, so the GPU bill follows the revenue instead of running ahead of it.
Overflow and billing pipelines
I build the overflow and billing pipelines that keep the product correct across distributed inference data centers. Load that one site cannot take moves to another, and every request is billed to the customer's balance exactly once, so no request is lost and no revenue leaks through missed or double charges.
The result
A working AI product on abliterated and custom-trained models, launched 9 weeks after the idea and served by its own inference engine from data centers in several countries. Its inference costs 60%+ less than buying it from API providers and its GPUs run above 80% utilisation, so a team of four keeps cost and billing under control, which is exactly where AI products of this kind usually lose their money.
Full case · PDF
Submissive AI
Four people took an AI product from idea to launch in 9 weeks, with inference 60%+ cheaper than API providers
- 01The architecture of the system and the inference engine
- 02Abliterated and fine-tuned models: the pipeline
- 03Harnesses and evaluation
- 04Inference across data centers in several countries
- 05Optimising the cost of every answer
- 06Overflow and billing pipelines: no lost requests, no double charges
The full Submissive AI case
The public story stops here. The PDF has the rest: the architecture before and after, the migration plan, how the team works with AI agents, what it all costs to run and where the savings came from.
More stories
All storiesWant the next story to be about your company?
In a 30-minute call we pick the first task to hand to AI and estimate what it will save you.


