How an AI adoption roadmap works for a 100-person company
Build an AI adoption roadmap for a 100-person company with quarterly goals for literacy, pilots, platform choices, governance, and measurable returns.

Table of Contents
An AI adoption roadmap for a 100-person company should change how work gets done within a year, not produce a stack of licenses and a presentation about innovation. The sequence matters: teach people enough to judge the tools, prove a few workflows, choose shared infrastructure from evidence, then make governance part of normal management.
At this size, the company is large enough for uncontrolled use to create real exposure and small enough for a program weighed down by committees to stall everything. I have watched both failures. The sensible plan gives one executive the outcome, puts business owners on every pilot, and uses quarterly gates that can stop weak ideas before they become permanent costs.
The plan below assumes a company with several functions, existing software procurement, and no dedicated AI department. If you already have machine learning specialists, they can help, but they should not own every use case. The people who own sales, support, operations, finance, product, and engineering must own the changes inside their work.
Start with operating facts, not an AI wish list
The first job is to establish where time, money, errors, and sensitive information sit today. Without that baseline, every later claim about productivity becomes a story told by the person who bought the tool.
Name one executive sponsor who can resolve conflicts across functions. Give a program lead about half of a role for the first two quarters, plus named representatives from security or IT, legal or compliance, finance, and the business teams running pilots. A company of 100 people does not need an AI center of excellence with its own hierarchy. It needs clear decision rights and enough time assigned to do the work.
Build a workflow inventory, not an application inventory. Ask each function to identify repeated work that consumes material staff time, has a visible output, and can be checked by a qualified person. Record current volume, time per item, wait time, error or rework rate, systems touched, data classes used, and the person accountable for the final result. That separates a useful candidate such as drafting a support reply from a vague ambition such as "use AI in customer service."
Set three company outcomes for the year. Natural candidates are hours returned to teams, cycle time reduced, and avoidable software or contractor cost removed. Keep quality and risk as gates, not as numbers to trade away. A faster process that makes more refund errors or exposes customer records has failed even if its time metric looks excellent.
Freeze the baseline for every accepted pilot before anyone changes the workflow. Use four weeks of normal operating data where volume varies, or a representative batch where work is event driven. Finance should approve the cost model. The business owner should approve the quality measure. Security should approve the data classification. These approvals prevent a pilot team from changing its definition of success after seeing mediocre results.
The yearly budget should also have an exit rule. Reserve money for training, controlled experiments, integration work, security review, and usage charges, but release it by quarter. Do not prepay a broad suite for the whole company because a vendor offers a discount. A discount on unused software is still waste.
Tell staff what the program can and cannot decide about jobs. Silence makes every workflow interview sound like a disguised headcount exercise, and employees then protect the information you need. State who decides staffing, when affected teams will hear about changes, and whether the year's goal is capacity, cost reduction, growth without hiring, or some mix. Do not promise that roles will never change if management has not made that commitment. Do promise that the company will not judge people by raw prompt counts or quietly use pilot participation as an individual performance score.
Managers also need a way to handle returned capacity. Before approving a pilot, ask where saved time would go if the trial works. The answer might be an existing backlog, faster customer response, more sales coverage, or removal of contractor work. If the manager has no answer, the pilot can still test quality, but its financial case remains unproven. This conversation turns "productivity" into an operating choice rather than a hopeful percentage.
Quarter 1 builds judgment before scale
The first quarter should produce a literate workforce, an enforceable use policy, a workflow inventory, and a short pilot queue. It should not aim for automation across the whole company.
Run literacy sessions tailored to every employee's role, including managers and the people who approve risk. A general session should explain what generative models do, why fluent output can be wrong, what happens when context is missing, and why confidential data cannot be pasted into an unapproved service. Then use examples from each function. Finance needs to test spreadsheet explanations and document extraction. Sales needs to inspect invented account facts. Engineers need to review generated changes and dependency claims. Managers need to understand that output volume is not evidence of impact.
Make employees practice verification. Give them a plausible but flawed output and ask them to find unsupported claims, missing constraints, confidential inputs, and steps that cannot be audited. People remember a confident error they caught better than another slide about hallucinations.
Publish acceptable use rules that fit on one page, with four data classes: public, internal, confidential, and restricted. For each class, state which approved tools may receive it, whether the provider may retain prompts, who can authorize an exception, and how staff report a mistake. Cover customer data, employee records, source code, financial information, and regulated material explicitly. "Use good judgment" is not a control.
Create a simple intake record that every experiment must use. The format matters less than forcing the owner to answer the same questions:
pilot_id: SUPPORT-01
business_owner: Head of Support
workflow: Draft first response from approved knowledge articles
baseline:
items_per_week: 420
minutes_per_item: 11
success_gate:
median_minutes_per_item: 8
factual_acceptance_rate: ">= 0.95"
data_class: internal
human_approval: required_before_send
approved_systems:
- support_sandbox
stop_conditions:
- customer_data_exposed
- acceptance_rate_below_gate
review_date: YYYY-MM-DD
The expected output of intake is not permission to deploy. It is a ranked queue with an owner, a baseline, a testable gate, and a stop condition for each candidate. Select three to five pilots for quarter 2. Reject ideas with no measurable output, no accountable owner, or no safe test environment.
End the quarter with a short practical assessment, not attendance certificates. Staff who will use an approved assistant should be able to classify an input, explain how they will check an answer, and report an incident. Managers should be able to distinguish saved effort from shifted effort. If they cannot, repeat the training before opening more access.
Quarter 2 turns pilots into evidence
The second quarter should test a small portfolio against real baselines, with human approval and limited data access. A pilot exists to answer a decision, not to make the sponsor look right.
Choose varied workflows so the company learns about different failure modes. One pilot might retrieve approved knowledge and draft support responses. Another could summarize internal sales notes into a fixed template. A third might help engineers write tests for a bounded service. Avoid the most sensitive process and the most politically visible process first. You need useful evidence, not a heroic rescue.
Each pilot needs a business owner, a workflow operator, a technical contact, and a risk reviewer. The owner decides whether the result helps the function. The operator records where the new process saves or adds work. The technical contact controls access and logs. The risk reviewer checks whether the agreed boundaries hold. A vendor success manager cannot fill any of these roles.
Run the old and new process on comparable work for long enough to include ordinary variation. Measure total elapsed time and human handling time separately. AI often reduces drafting time while adding review, correction, and waiting time elsewhere. Count rejected outputs, material errors, escalations, and rework. Ask users about friction, but do not substitute a satisfaction poll for operating data.
Score every pilot on four gates, each from zero to three: business impact, output quality, adoption in the intended group, and control effectiveness. A pilot must score at least two on every gate before expansion. This prevents a large time saving from hiding unacceptable quality, or enthusiastic use from hiding weak economics.
For example, suppose a support pilot cuts median handling time from 11 minutes to 8, passes 96 percent of sampled factual checks, is used on 70 percent of eligible cases, and produces complete review logs. It can advance to a larger controlled group. If the same pilot achieves 6 minutes but factual acceptance falls to 89 percent, it stops. The team then narrows retrieval sources, changes instructions, or abandons the use case. Averaging those results into one attractive score would conceal the failure.
Track the full cost: licenses, model usage, integration, security work, training, review time, and the program lead's time. Convert saved hours into money only when the team can remove external spend, avoid planned hiring, increase handled volume, or move people to work the company would otherwise pay for. "We saved 500 hours" means little if nobody can say what changed because of them.
At the quarter gate, expand, revise, or stop each pilot. Preserve the record for stopped pilots. A clean rejection teaches platform and governance requirements, and it can save another team from repeating the same experiment six months later.
Quarter 3 chooses platforms from proven needs
The third quarter should standardize the capabilities that successful pilots share, then expand only the workflows that passed every gate. Platform selection comes after pilots because real use exposes requirements that demos hide.
Translate pilot evidence into a capability matrix. Typical rows include identity integration, group access, retention controls, regional processing, model choice, retrieval from approved sources, audit events, usage limits, evaluation support, and an export path. Weight each row using the data and workflow risks you actually observed. Do not add features because they appeared in a vendor presentation.
Run a structured comparison with your own tasks and redacted or synthetic data. Give vendors the same scenarios, scoring rules, and contract questions. Record answer quality, latency that users experience, administrative effort, traceability, and unit cost at expected volume. A polished chat interface deserves almost no weight if the planned workflows use APIs and background jobs.
Procure in layers. Most companies need an employee assistant for general work, an application layer for approved workflow integrations, model access beneath it, and controls around identity, data, evaluation, logging, and spend. One vendor can supply several layers, but document where each layer begins and how data crosses it. This makes replacement possible and clarifies who owns an incident.
Negotiate the terms your policy requires. Check whether submitted data trains provider models, how long prompts and outputs remain available, which subprocessors handle them, how deletion works, where processing occurs, how the service reports incidents, and whether administrators can export audit records. Legal language and product settings must agree. A contract promise does not help if an administrator leaves the wrong retention option enabled.
Expand passed pilots in stages. Move from the original group to perhaps one function, verify the gates again under higher volume, then add adjacent groups. Assign a trained owner for prompts, knowledge sources, evaluation examples, and change review. Models and connected sources change, so a workflow that passed in May can regress in September.
This is also the quarter to remove duplicate experiments. Teams often buy several assistants because procurement sees each request as small. Consolidation is sensible when requirements match, but do not force a specialized engineering or regulated workflow into a general assistant merely to reduce the vendor count. Standardize controls first and interfaces second.
Quarter 4 makes governance routine
The fourth quarter should move AI decisions into existing budgeting, security, legal, operations, and performance rhythms. A separate AI committee that reviews every minor change will either become a queue or get bypassed.
Create risk tiers based on consequence and autonomy. A low tier might cover drafting internal text from inputs that are not sensitive, with human review. A medium tier could include recommendations based on internal records or content that reaches customers after approval. A high tier includes decisions affecting employment, credit, safety, regulated obligations, or actions taken without a person confirming them. Define examples that match your company instead of relying on abstract labels.
Route decisions by tier. An experiment in the low tier can use a preapproved pattern and owner signoff. Medium use should add security or privacy review, documented evaluations, and monitoring. Use in the high tier needs executive and legal approval, a deeper impact assessment, clear appeal or fallback paths, and evidence that the benefit justifies the exposure. Some ideas in the high tier should remain prohibited.
Put recurring reviews on calendars. Business owners review quality and impact monthly. IT or security reviews access, logs, incidents, and unapproved services. Finance compares actual cost and realized capacity against the approved case each quarter. Procurement checks renewals against active use. The executive sponsor reviews the portfolio and resolves ownership gaps.
Define incident handling before the first incident. Staff need one reporting route for confidential data entered into the wrong service, harmful or materially false output, unintended external action, access outside the approved group, or missing audit evidence. The response should contain the workflow, preserve relevant logs, notify the right owners, assess affected people and obligations, correct the process, and decide whether the workflow can resume.
Update job expectations carefully. If an approved workflow changes a role, managers should state which tasks changed, what review remains mandatory, how quality is judged, and where returned time goes. Do not set an arbitrary target that every employee must use AI. That target rewards visible activity, encourages unsuitable uses, and tells staff that management cares more about adoption charts than their work.
Close the year by publishing the approved service catalog, risk routes, active workflow owners, and retirement list. New proposals then enter a known process rather than restarting the same debate in every department.
Platform choice is an architecture decision
A company should choose replaceable layers and explicit data boundaries, not search for one permanent AI vendor. Model quality moves, prices change, and different work needs different controls.
Start with identity. Every approved service should use company accounts, access based on groups, rapid removal when someone leaves, and administrator visibility. Shared consumer accounts destroy attribution and make offboarding unreliable. If a service cannot fit your identity and access process, keep sensitive work out of it.
Separate the user experience from the model provider where practical. An employee may use a managed assistant, while a production workflow calls a model through an application layer that enforces templates, retrieval, validation, and logs. These are different procurement decisions. A good employee chat tool does not automatically provide the controls needed for a process that reaches customers.
Keep authoritative data in systems that already own it. Retrieval should fetch the smallest relevant context from approved sources, enforce the user's permissions, and record source versions. Copying an entire knowledge base into an unmanaged index creates a second permission system and a problem with stale data. Test whether deleted or restricted material disappears from results when expected.
Plan an exit before signing. Export prompts, evaluation cases, configuration, usage records, and workflow definitions in usable formats. Keep business logic outside builders controlled by one vendor when the process matters enough that replacement would hurt. You do not need theoretical portability across every model. You need a credible way to move the few workflows that matter.
Treat model hosting as a risk and economics choice, not an ideological one. A managed model may reduce the work of deployment, scaling, updates, and abuse controls. A model hosted by your own team may offer tighter infrastructure control but makes that team responsible for patching, capacity, monitoring, and model lifecycle. Open weights do not make a system private by themselves. Privacy depends on where requests travel, who can access logs and storage, which tools the model can call, and how the whole service is operated.
Run a sensitivity check before choosing the more complex option. Estimate normal and peak volume, context size, response targets, evaluation load, engineering support, and the cost of an outage. Then vary the two assumptions most likely to be wrong. A model hosted by your own team that looks cheaper only at continuous high utilization may cost more for a small, uneven workload. A managed service that begins cheaply may become expensive when a workflow sends large documents repeatedly. Use your measured pilot traffic instead of a vendor's sample calculator.
Build only the parts that express a rule particular to the company or provide control you cannot buy. Authentication, general chat interfaces, basic model routing, and standard document connectors rarely distinguish the business. A pricing approval rule, a specialized quality check, or access logic tied to your customer contracts might. Every custom component creates maintenance, evaluation, and incident work, so make its owner and replacement cost part of the platform decision.
Avoid building a universal internal AI platform in year one. A 100-person company rarely has enough stable use cases to justify that engineering load. Buy the ordinary capabilities, write thin integration around proven workflows, and retain control over identity, data access, evaluation, and business rules. Revisit a larger platform only when repeated needs make its economics obvious.
Governance should follow consequence and ownership
Good governance tells an employee who may decide, what evidence that person needs, and when a decision must move upward. A long policy that nobody can apply during real work offers little protection.
Use the NIST AI Risk Management Framework as a vocabulary, not a certificate. Its Govern, Map, Measure, and Manage functions make a useful test: have you assigned accountability, described the context and affected people, measured performance and risk, and acted on the results? The framework deliberately allows different organizations to adapt it. That flexibility is helpful, but it means management still has to set thresholds.
ISO/IEC 42001 addresses an organizational management system for AI. It fits companies that need formal policies, assigned responsibilities, internal review, and continual improvement across a wider program. Do not blur that with evaluating one model output. A management system can make responsibilities repeatable, while a workflow evaluation determines whether a specific use performs acceptably. You may need both, and passing one says little about the other.
Keep one register of active AI uses. For each, record the owner, purpose, users, affected parties, data classes, connected systems, provider and model where relevant, human decision point, evaluation set, review date, incidents, and current status. Shadow use becomes easier to address when the approved route is faster than hiding.
Give employees a safe disclosure path. People will experiment before policy catches up. Punishing every first disclosure pushes use underground. Reserve discipline for deliberate evasion, repeated disregard, or harmful conduct. Treat an early report of an honest mistake as information that improves training or controls.
Governance also needs retirement authority. An owner should close a workflow when quality falls, source data changes, cost exceeds the case, the provider changes terms, the process no longer exists, or nobody remains accountable. Dormant integrations with broad access are liabilities, not harmless leftovers.
Adoption metrics must prove an operating change
The useful metrics connect AI use to business output, quality, cost, and risk. Login counts and prompt counts can diagnose access problems, but they cannot prove a return.
Build a small measurement chain for each workflow. Start with eligible work volume, then measure the share processed through the approved workflow, human handling time, total cycle time, acceptance without material correction, rework or escalation, incidents, and full operating cost. Compare cohorts or periods with similar work. Record major process changes so you do not credit AI for a staffing or policy change.
Measure realized capacity with the business owner and finance. Returned time can increase volume, reduce backlog, improve service levels, avoid a hire, end contractor spend, or support a deliberate staffing change. Name which outcome occurred. If the answer is "people had more time," keep the benefit as capacity until the company can show what that capacity produced.
Assign one owner to each claimed benefit and set the date when finance will verify it. The workflow owner can confirm saved handling time, but only the budget owner can confirm that a planned hire was avoided or a contract ended. Revenue claims need a comparison that accounts for other changes in pricing, staffing, and demand. Without that separation, the same returned hour often appears once as cost savings, again as extra capacity, and a third time inside a revenue estimate.
Watch distribution, not only averages. A tool can help experienced staff while slowing new hires, or work well in English while failing another market. Sample difficult and ordinary cases. Review failure categories, not just a single quality percentage. One severe privacy event matters more than many clean drafts.
Survey employees about trust, training, and friction at the workflow level. Ask whether they know when to use it, how to check it, and how to stop or report it. Anonymous comments often reveal unofficial workarounds before logs do. Still, treat sentiment as diagnostic evidence, not the business case.
Report a portfolio view on one page each quarter: workflows proposed, testing, scaled, paused, and retired; realized annualized value; current run cost; material incidents; quality against gates; and decisions required. Do not combine every benefit into a speculative grand total. Separate verified savings, avoided cost with an approved budget, revenue effects with credible attribution, and capacity that has not yet turned into money.
The year ends with a portfolio decision
At the end of quarter 4, the company should have a small set of scaled workflows, a larger set of documented rejections, shared platform decisions, trained owners, and a governance process that staff can actually use. If all you have is high license adoption, the roadmap did not change the company.
Review every workflow and platform against the same gates used during pilots. Expand items with verified impact, acceptable quality, sustained use in the eligible group, and effective controls. Repair items with a clear, bounded problem. Retire items whose economics depend on optimistic time conversion, whose owners left, or whose risk cannot be reduced enough.
Then decide what capability the second year requires. You may need stronger internal integration, a dedicated evaluation function, more formal management controls, or fewer vendors. Hire specialists only for repeated work that justifies a role. Do not create an AI department to compensate for business owners who refuse to own their processes.
For founders who want an external baseline before committing a year of budget, the Team & AI Audit on oleg.is takes five business days and costs $5,000. It identifies at least $50,000 per year in savings or the audit is free, which makes it a practical way to test the cost case before selecting platforms or reorganizing teams.
The strongest sign of maturity is ordinary behavior: staff know which tools are approved, owners can show whether a workflow works, finance can trace the return, and someone can shut it down. Approve next year's money only for the parts that meet that standard.
Frequently Asked Questions
How long should an AI adoption roadmap take?
Use a yearlong roadmap with decisions at the end of every quarter. That is long enough to gather operating evidence and short enough to stop weak tools before they become embedded costs.
Who should own AI adoption in a 100-person company?
One executive should own the company outcome, while business leaders own workflows inside their functions. A program lead coordinates the work, and IT, security, legal, finance, and procurement supply the controls they already own.
How much should a company budget for AI adoption?
Build the budget from training, pilots, integration, review time, model usage, security work, and licenses. Release funds by quarter and require evidence before expansion; a fixed percentage of revenue or payroll ignores the actual use cases.
How many AI pilots should run at once?
Three to five varied pilots are enough for a 100-person company in the second quarter. More pilots usually spread technical and review capacity too thin, while one pilot teaches too little about different data and workflow risks.
Which AI use cases should a company try first?
Choose repeated work with measurable volume, a checkable output, a willing business owner, and a safe test environment. Avoid the most sensitive and politically visible processes until the company can evaluate and contain failures.
Should employees use free public AI tools at work?
Do not allow confidential or restricted company information in unapproved public services. Give employees approved options and a simple data policy, because a ban with no usable alternative tends to push experimentation out of sight.
How do you measure return on an AI pilot?
Compare total cost with realized changes in cycle time, quality, handled volume, contractor spend, backlog, or approved hiring. Saved minutes count as capacity until the business owner and finance can show what the company did with them.
Does a small company need an AI governance committee?
It needs assigned decisions, not necessarily a standing committee. Route low, medium, and high consequence uses through existing owners, and convene a group only when a decision crosses their authority.
When should an AI pilot be stopped?
Stop when it misses any mandatory quality or control gate, exposes data, lacks an accountable owner, or costs more than the credible benefit. A fast workflow with unacceptable errors has not earned more time.
Should a company standardize on one AI vendor?
Standardize controls and common capabilities, but keep important layers replaceable. One vendor may cover several needs, yet identity, data boundaries, evaluations, logs, and exit options should remain explicit.


