How should you compare fractional CTO proposals?
Learn to compare fractional CTO proposals on scope, baselines, authority, savings, and the full cost of the first 90 days.

Table of Contents
Two fractional CTO proposals can quote the same monthly fee and describe completely different jobs. One may put an experienced operator in charge of engineering decisions, delivery measurement, and team changes. The other may buy a few advisory calls, a tool shortlist, and a strategy deck. Comparing the retainers without normalizing the work is how founders choose the cheaper proposal and discover, six weeks later, that nobody had authority to change anything.
I compare fractional CTO proposals on one sheet with five scored categories: scope clarity, delivery baselines, decision authority, savings assumptions, and total cost for the first 90 days. The sheet forces every bidder to answer the same questions in the same units. It also exposes a proposal that depends on unpriced implementation work, heroic savings estimates, or access the founder has not agreed to grant.
The scoring method below is deliberately plain. It will not tell you which person you trust. It will tell you whether the commercial promises, operating plan, and decision rights fit together. That is the minimum evidence needed before trust should influence a purchase of this size.
Normalize the outcome before comparing the prose
A fair comparison starts with one written outcome that every proposal must price. If bidders solve different problems, the scoring sheet will produce tidy numbers and a false conclusion.
Write the outcome in operational terms. "Adopt AI across engineering" is too vague to price or verify. A useful outcome might be: reduce the time from an approved product change to a production release while maintaining the current service targets, then identify which engineering roles, workflows, and tools should change. The outcome names the flow, protects an existing constraint, and requires an organizational decision.
Give each bidder the same short fact pack before asking for a final proposal. Include the current team roster and loaded monthly cost, the last eight to twelve weeks of delivery records, the release process, the incident history, the tool contracts, the product commitments already sold, and the founder's nonnegotiable constraints. Redact personal data if necessary, but do not hide the inconvenient backlog or the contractor who owns half the deployment process. A proposal built from a polished org chart will fail against the real organization.
Ask every bidder to state what will be observably different on day 30, day 60, and day 90. "AI strategy completed" does not count. A day 30 result could be a measured baseline for two delivery flows and a ranked set of experiments. A day 60 result could be one production workflow running with agreed review controls. A day 90 result could be a staffing and operating plan supported by measured changes in cycle time, escaped defects, and support load.
The comparison unit is therefore not "fractional CTO for three months." It is a defined change, under defined constraints, with evidence at three decision points. Once that unit is fixed, the proposals become much easier to challenge.
Scope clarity needs exclusions and acceptance tests
A clear scope says who will do each piece of work, what artifact or operating change will result, how you will accept it, and what the fractional CTO will not do. A long task list without those four details only creates the appearance of precision.
Separate the work into discovery, implementation, and operating change. Discovery covers access, interviews, baseline measurement, architecture review, and workflow mapping. Implementation covers tool setup, repository changes, automation, tests, security controls, and production rollout. Operating change covers decision forums, role changes, hiring or reduction recommendations, training, and the reporting cadence. Many proposals price discovery and advice while describing implementation benefits in the same paragraph.
For every deliverable, add an acceptance test. A "current-state assessment" is accepted when it documents the delivery flow, names the source systems used for each baseline, records data gaps, and has been reviewed by the engineering lead and founder. An "AI development pilot" is accepted when a named workflow has run on production-bound work, its review rules are documented, and its results appear beside the baseline. Acceptance tests should verify evidence, not force a promised business result that depends on your team.
Then mark ownership with four simple labels: CTO does, client does, vendor does, or shared. "Implement coding agents" becomes a different commitment when the CTO configures the tools and changes repositories versus telling the client's engineers how to do it. Shared work needs an estimate of client hours and the role required. Ten hours from a founder and ten hours from a platform engineer do not have the same cost or availability.
Exclusions deserve their own row. Ask whether the fee excludes security review, data cleanup, cloud changes, tool licenses, recruiting, termination support, after-hours incident work, travel, or hands-on programming. An exclusion is not automatically bad. Hidden dependence is bad. If the engagement needs a cloud migration but the proposal excludes cloud implementation, add the missing vendor or employee cost to the first 90 days.
Score scope clarity from zero to five. Give zero when the proposal lists themes. Give one when it lists activities but no owners. Give two when it names deliverables. Give three when most deliverables have owners and dates. Give four when acceptance tests and exclusions are explicit. Give five only when dependencies, client effort, and change control are also priced or bounded.
A delivery baseline must survive contact with the data
A proposal cannot credibly promise faster delivery or a smaller team until it defines the present delivery system with data you can reproduce. The baseline is not a workshop opinion and it is not story points.
Ask which flow the bidder will measure. Feature delivery, production incidents, customer support fixes, and infrastructure changes often move through different queues. Combining them can make the average look stable while one critical flow deteriorates. The proposal should name one or two flows tied to the business outcome and define their start and end events.
For feature delivery, a useful start event might be "work accepted for implementation," and the end event might be "change verified in production." The bidder should name the source for each timestamp, such as the issue tracker, source control system, deployment log, or incident tool. If those systems disagree, the proposal should explain which record wins and why. Manual sampling can be acceptable for a first pass, but the sample rule must be written before anyone sees the results.
A practical baseline usually pairs speed with quality and demand. Cycle time without work volume can improve because the team ships fewer changes. Deployment frequency without escaped defects can reward smaller, riskier releases. Defect counts without severity or customer effect can punish teams that improve detection. Ask the bidder to define a compact set that can answer four questions: how long work waits, how much completes, how often production suffers, and how much unplanned work interrupts the team.
The DORA metrics are useful here, but they are not a universal scorecard. DORA defines deployment frequency, lead time for changes, change failure rate, and time to restore service. Those measures work well for software delivery performance. They do not reveal whether engineers are building the right product, whether support demand is consuming the team, or whether AI review time shifted work from programmers to senior reviewers. A serious proposal adapts the measurement set to the decision instead of pasting four familiar names into a slide.
Score the baseline from zero to five. Zero means no current measurement. One means the bidder plans interviews only. Two means metrics are named but events and sources are not. Three means flows, events, and systems are defined. Four adds quality, demand, and a treatment for missing data. Five adds a reproducible query or export, an owner, and a schedule for checking whether the metric still reflects real work.
Decision authority determines whether the plan can move
Decision authority must be explicit because an advisor who can recommend a change is not the same as an interim executive who can make it happen. Founders routinely buy the second outcome under the first governance model.
List the decisions the transformation is likely to require: engineering workflow, approved AI tools, data access, repository policy, architecture standards, vendor spend, role definitions, hiring, performance management, and production release controls. For each decision, name who proposes, who decides, who must be consulted, and who executes. You do not need a complex governance framework. You need one person in the deciding column and a deadline for a disputed decision.
Pay special attention to personnel and security. A fractional CTO may be asked to redesign roles but lack access to compensation, performance history, or employment terms. They may be accountable for AI adoption but unable to approve a tool that sends source code to an external service. Those limits can be reasonable. The proposal must adjust the outcome and schedule to match them.
Authority also has a time budget. If the founder retains every decision but can only meet twice a month, the engagement will wait. Require a cadence for operating reviews, a response time for approvals, and an escalation path when the engineering lead and fractional CTO disagree. Put the client's obligations in the proposal, including access dates and executive availability.
Score authority from zero to five. Give zero when the title implies authority but the contract says nothing. Give one for advisory access to the founder. Give two when decision areas are named. Give three when decision owners and meeting cadence are explicit. Give four when access, response times, and escalation are included. Give five when the authority model also matches the promised deliverables, especially staffing, security, and production changes.
Savings assumptions need a ledger, not a percentage
Treat every savings claim as a small financial model with named inputs. A percentage alone hides when savings begin, which costs disappear, and which new costs replace them.
Start with the loaded cost of the affected team, not base salary. Include payroll taxes, benefits, contractors, recruiting amortization if you already track it, software tied to headcount, and management time. Keep one-time severance or contract exit fees separate because they affect cash in the first 90 days but not the recurring run rate. Do not invent a universal burden rate. Finance should provide the company's actual convention.
Next, classify each claimed benefit. Cash savings remove or reduce an invoice or payroll obligation. Avoided cost prevents a planned hire or contract. Capacity gain lets the same team complete more useful work. Revenue effect depends on selling or retaining more business. These categories are not interchangeable. Only cash savings and avoided costs should reduce the cost line in a conservative approval case. Capacity and revenue belong in an upside case with a separate owner.
Ask for the mechanism behind every number. If the proposal assumes fewer engineers, it should state which work disappears, which work becomes automated, which work moves to the remaining team, and who handles review and incidents. If it assumes faster releases, it should state which waiting stage changes. If it assumes fewer defects, it should name the control that prevents or catches them. "AI productivity" is not a mechanism.
Timing matters as much as amount. A role cannot produce payroll savings while the employee remains on payroll. A contractor contract may require notice. An annual tool may not be refundable. Training and migration can reduce capacity before they increase it. Put a start month and confidence level beside each line, then show expected, conservative, and adverse cases.
I reject proposals that calculate savings by multiplying an assumed productivity percentage by the whole engineering payroll. The method is popular because it produces a large number in five minutes. It is wrong because time saved does not become cash unless management removes a cost, avoids a committed cost, or converts the capacity into work the business actually needs.
Score savings assumptions from zero to five. Zero means an unsupported percentage. One means a payroll total and target percentage. Two names affected roles or contracts. Three separates cash, avoided cost, capacity, and revenue. Four includes timing, replacement costs, and scenario ranges. Five connects every material line to a mechanism, an owner, and evidence that can be checked during the engagement.
The first 90 days cost more than the retainer
The comparable price is the full cash and internal labor cost for the first 90 days, including plausible change orders. Monthly retainers alone are almost never comparable.
Build the cost view by week or month. Include the fractional CTO fee, assessment fee, onboarding charge, travel, software and model usage, implementation vendors, recruiting, legal or security review, severance, and contract termination fees. Add internal labor for the founder, engineering lead, engineers, security, finance, and people operations using a consistent loaded hourly rate. Mark refundable deposits and annual prepayments so cash timing remains visible.
Separate committed, likely, and optional cost. Committed cost appears in the signed scope. Likely cost follows from a dependency even if the proposal leaves it out. Optional cost buys an extension or additional outcome. When a bidder says implementation depends on a specialist, ask for a range and put its midpoint in the expected case and its upper bound in the adverse case.
Change control can dominate the comparison. Record the rate for work outside scope, who can approve it, the minimum billing unit, and any cap. A lower retainer with an open-ended implementation rate may cost more than an inclusive proposal. Ask what happens when the baseline data is unusable, a tool fails security review, or the pilot exposes architectural work. Those are foreseeable branches, not surprises.
Do not subtract annualized savings from the first 90 days of cost. Show three numbers beside each other: cash required during the period, recurring monthly cost at day 90, and verified monthly savings active at day 90. A proposal may have a sound annual return and still create a cash squeeze the company cannot accept.
Score cost transparency from zero to five. Zero means a retainer only. One adds a term and payment schedule. Two names major extras. Three includes internal labor and licenses. Four includes dependency ranges and change control. Five shows committed, expected, and adverse totals with cash timing and the recurring position at day 90.
Put every proposal on this one comparison sheet
The sheet should preserve facts and judgment separately. Enter proposal facts first, resolve questions with each bidder, then score. If you score while reading sales prose, a confident writing style will leak into every category.
Use these columns in a spreadsheet:
- Category: Scope, baseline, authority, savings, or first 90 days cost
- Proposal fact: Exact commitment, assumption, exclusion, fee, or date
- Evidence: Contract section, appendix, data sample, or bidder answer
- Gap: Missing owner, test, source, range, or approval right
- Clarification: Question sent and dated answer received
These columns preserve the evidence and the gaps before scoring. Add four more columns for the judgment:
- Score: Integer from 0 to 5 using the rubrics in this article
- Weight: Your agreed importance before opening proposals
- Weighted score: Score multiplied by weight
- Red flag: Yes or no, with the condition that triggered it
Give the five categories weights totaling 100. A sensible default for an execution-heavy AI transformation is scope 25, baseline 20, authority 20, savings 20, and cost 15. Change them before seeing final prices. A company facing an immediate cash constraint may raise the cost weight. A company with a resistant leadership team may raise authority. Do not change weights after one bidder scores badly.
Calculate the result with this compact formula:
weighted_total = SUM(category_score / 5 * category_weight)
The result runs from zero to 100. Keep price out of the category weights beyond the cost-transparency score. Price is a separate decision variable, not evidence of proposal quality. Place expected cash cost for the first 90 days, adverse cash cost, and active monthly savings beside the weighted total.
Add pass or fail gates below the score. I usually require a named executive decision owner, permission to inspect source systems, written data-handling boundaries, a cap or approval rule for extra work, and no staffing savings counted before an executable staffing decision. A proposal that fails a gate does not win because its other categories average well.
Finally, run a sensitivity check. Increase and decrease each category weight by ten points, balancing another category so the total remains 100. If the winner changes under small, reasonable weight adjustments, the proposals are effectively tied on the evidence. Choose based on references, working chemistry, or a smaller paid diagnostic, not a decimal place in the spreadsheet.
Due diligence should test the operating claims
References should confirm how the fractional CTO worked when the plan met resistance, weak data, and production risk. Generic praise about intelligence or communication tells you very little.
Ask a reference what the CTO personally changed, what the client's team had to supply, which promised result did not arrive, and how the engagement handled disagreement. Ask whether costs outside the retainer appeared and whether the reference would use the same authority model again. A credible reference can describe tradeoffs and unfinished work. A polished success story with no friction is hard to use as evidence.
Review sample artifacts with sensitive details removed. Look for a baseline definition, a decision log, a weekly operating report, a risk register, and a before-and-after workflow. Judge whether the artifacts caused decisions. A beautiful deck that records no owner, date, or threshold will not run your engineering team.
For AI work, ask how the bidder handles source code, credentials, customer data, generated changes, and human review. NIST's AI Risk Management Framework organizes work around Govern, Map, Measure, and Manage. That is a useful reminder that model selection sits inside a broader management system. A proposal does not need NIST terminology, but it should assign ownership, identify the context of use, measure outcomes and failure, and specify what happens when a control does not hold.
Run the same clarification meeting with every finalist. Send questions in advance, keep the same time limit, and record written amendments afterward. Score only commitments that enter the proposal or contract. An impressive answer on a call has no delivery value if the signed scope still contradicts it.
A worked comparison exposes the hidden bargain
Consider a 14-person software company comparing two three-month proposals. Proposal A charges $8,000 per month. It promises an assessment, an AI tool rollout, team coaching, and a staffing recommendation. Proposal B charges $11,000 per month and includes baseline extraction, two production workflow pilots, repository changes, manager training, and a day 75 operating-plan decision.
The founder first sees $24,000 versus $33,000 and prefers A. The sheet changes the comparison. A assigns tool configuration to the client's engineering lead, estimates 120 internal engineering hours, excludes security review, and prices implementation after discovery. B estimates 55 internal hours, includes configuration for the two pilots, and caps extra work at $5,000 without a new approval.
Using the company's internal planning rate of $90 per engineering hour, A consumes $10,800 of internal engineering time and B consumes $4,950. A likely needs a $4,000 security review and carries an unpriced implementation range of $8,000 to $20,000. B includes security work for the named pilots but requires $2,500 of model and tool budget. The expected 90-day costs become $46,800 for A if implementation lands at the $8,000 lower estimate, and $40,450 for B. The cheaper retainer was not the cheaper engagement.
The quality scores explain the commercial difference. Suppose A scores scope 2, baseline 1, authority 2, savings 1, and cost transparency 2. With the default weights, its weighted total is 32. B scores 4, 4, 4, 3, and 4, for a total of 76. B loses one point on savings because its capacity claim has not yet become a cash decision. That restraint makes the proposal more credible, not less attractive.
Now inspect authority. A calls the CTO an advisor and requires founder approval for tools, workflow, and staffing, but provides no response deadline. B grants authority over engineering workflow and pilot tooling within an approved budget, while the founder retains personnel decisions at scheduled day 45 and day 75 reviews. B's plan can operate between meetings. A's schedule assumes decisions its governance does not permit.
This example does not prove that a higher fee wins. It shows why the fee cannot settle the question. If A revised its implementation range, reduced client labor, defined a baseline, and added decision deadlines, it might become the better buy. The sheet gives the founder exact amendments to request instead of a vague demand for "more detail."
Contract the evidence and the exit
The final contract should carry the commitments that produced the winning score. Attach the scope, acceptance tests, authority table, baseline definitions, cost assumptions, and clarification answers. If procurement replaces them with a short services description, the comparison work has been discarded.
Tie payment dates to time and accepted evidence without making the fractional CTO guarantee outcomes outside their control. For example, the first payment can cover access and baseline work, the second can follow delivery of reproducible baseline data and a pilot design, and the third can cover implementation and the operating decision. A client who delays access should not pretend the original dates still apply, so define schedule relief as well.
Include an exit that protects continuity. Require return or deletion of confidential data, transfer of accounts and configurations, delivery of decision and risk logs, and a short handover to the named internal owner. State which tool subscriptions, prompts, automation, and repository changes belong to the client. Define notice, final invoice treatment, and support for a production issue discovered during handover.
A paid Team & AI Audit can be a sensible first contract when the proposals remain too assumption-heavy to compare. The fixed diagnostic should produce baselines, savings evidence, and an implementation choice, not obligate the company to a longer transformation.
Choose the proposal whose scope, measurements, authority, economics, and price describe the same engagement. If one of those pieces tells a different story, ask for a revision before signing. A fractional CTO can fix a weak operating system, but no operator can execute authority the founder never granted or bank savings the company never decided to realize.
Frequently Asked Questions
What should a fractional CTO proposal include?
It should define the outcome, deliverables, owners, acceptance tests, exclusions, decision rights, schedule, fees, and change-control rules. For an AI transformation, it should also name the baseline data, production pilots, data boundaries, review controls, and the evidence expected at 30, 60, and 90 days.
How many fractional CTO proposals should I compare?
Three serious proposals usually give enough contrast without turning the process into unpaid consulting. If only two candidates fit, use the same fact pack and scoring sheet for both, then test the result with references and a structured clarification call.
How much does a fractional CTO cost for an AI transformation?
The website context for this article places fractional CTO leadership for AI team transformation at $5,000 to $10,000 per month. Your comparable cost must also include internal labor, tools, security review, implementation help, and any personnel or contract changes during the first 90 days.
Should I choose a fixed project or a monthly retainer?
Use a fixed project when the diagnostic output and acceptance tests are clear. Use a retainer when the CTO must make recurring decisions and adapt implementation, but cap or approve work outside scope and keep 30, 60, and 90 day evidence in the agreement.
How do I verify promised engineering savings?
Require a ledger that separates cash savings, avoided cost, capacity, and possible revenue. Each material line needs a mechanism, owner, start date, replacement cost, confidence range, and evidence that you can check during the engagement.
What authority should a fractional CTO have?
Grant authority that matches the promised outcome, with explicit limits for budget, personnel, security, architecture, and production. Name one decision owner per area, response deadlines, and an escalation path so the work does not wait for an unavailable founder.
Which engineering metrics belong in the baseline?
Choose metrics for the actual delivery flow and pair speed with quality, volume, and unplanned demand. DORA measures can help with software delivery, but add product or support measures when the business decision depends on work those four measures do not capture.
Can I compare proposals using only their monthly fees?
No. Add all cash costs, internal labor, tools, specialists, change orders, and likely implementation dependencies for the first 90 days. Keep recurring monthly cost and verified active savings beside that total rather than mixing annual projections into near-term cash.
What is a red flag in an AI transformation proposal?
Watch for savings stated only as a percentage, implementation benefits paired with advisory-only scope, unnamed baseline sources, and authority implied by a title rather than granted in writing. Uncapped extra work and staffing savings counted before a staffing decision also deserve a stop.
When should I start with an audit instead of a transformation?
Start with an audit when baseline data, savings potential, technical constraints, or leadership authority remain too uncertain to price implementation. The audit should end with reproducible evidence and a decision, while leaving you free to choose who performs the next phase.


