Enterprise buyers have a procurement function that asks awkward questions on their behalf. Teams of five to fifty do not. They have someone in operations, a corporate card, and a demo that went well.
That gap is where most bad AI purchases happen — not through negligence, but because nobody knew which questions carried real consequences and which were procurement theatre. This checklist is the short version: nine questions that have actually cost small teams money when they went unasked, in the order you should ask them.
Work through it before the trial ends, not after.
1. What happens to our data, specifically?
“We take privacy seriously” is not an answer. Three separate questions hide inside it, and vendors routinely answer one while leaving the others ambiguous.
- Is our content used to train models? Get the answer in writing, and check whether it differs by plan tier. Several vendors train on free-tier traffic and not on paid traffic, which means the answer changes the moment someone signs up with a personal account.
- How long is data retained after deletion? “Deleted immediately” and “deleted from active systems within 30 days, backups within 90” are very different commitments. Both may be fine. Only one is a real statement.
- Which sub-processors see the content? A platform routing to third-party model providers passes your text to those providers. That is not inherently a problem, but you need the list, because your own customer commitments may depend on it.
If the vendor cannot answer all three in a single email, treat that as the answer.
2. Which models are included, and who decides when that changes?
Model rosters are marketing surfaces, and they change without notice. The questions that matter:
- Is the roster contractual, or a webpage the vendor can edit on a Tuesday?
- What is the typical lag between a lab releasing a new model and it appearing here?
- What happens when a model you depend on is retired — do you get notice, and how much?
That last one bites hardest. Teams build prompts, evaluation sets, and internal documentation around specific model behaviour. When a model is deprecated with two weeks’ notice, that work needs redoing. Ask what the deprecation policy is, and be suspicious if there is not one.
3. What does it cost when we are actually using it?
Headline pricing describes a customer who barely uses the product. You need the number for a customer who does.
Model the two ends honestly. For a light user, does the entry plan’s allowance cover a normal month, or does it run out in week three? For your heaviest user — there is always one, usually an engineer or a researcher — what happens when they exceed the included allowance? Overage rates, hard cutoffs, and forced upgrades are all common, and they are all disclosed somewhere unglamorous.
Get the tier structure explicitly. Most AI platforms publish plan tiers with their included allowances rather than making you ask, which is a reasonable early signal about how the rest of the relationship will go. If the pricing page requires a sales call to decode, the invoices will too.
Then check for the four charges that are usually excluded from the headline number:
| Charge | Why it is missed |
|---|---|
| Overage on credits or tokens | Advertised as “flexible”, priced as a penalty |
| Premium model surcharges | Best model is not on the base tier |
| Seat minimums on team plans | Five-seat floor for a four-person team |
| API access as a paid add-on | Assumed to be included; often is not |
4. How does this compare on the specific things we do?
Vendor comparison tables are written by the vendor. They select the axes on which that vendor wins. This is not dishonest, exactly, but it is not evidence either.
Build your own axis list first, before looking at anyone’s table. For most small teams the list is short: context window, cost per unit of real work, the specific capability your workflow depends on, and refusal behaviour on your actual content. Then check independently. Neutral side-by-side model comparison tools that put pricing, context windows, and benchmark scores next to each other are more useful than any vendor’s own chart, precisely because you choose which models to line up.
Better still: run the same five real prompts from your own work through each candidate. Twenty minutes of that beats a week of reading comparisons.
5. What is the exit cost?
Ask this before you sign, because it is the only moment you have leverage.
- Can we export conversation history and any custom configuration? In what format?
- Are prompts, agents, or custom instructions portable, or vendor-locked?
- What happens to data on cancellation — deleted, retained, or held pending payment?
The exit-cost answer determines how much you should invest in vendor-specific tooling. A platform you can leave in a day earns the benefit of the doubt. One that holds two years of institutional knowledge in a proprietary format needs to be much better to justify the risk.
6. What is the actual uptime commitment?
There is a real difference between a published SLA with service credits and a status page with good intentions.
For most small teams this is genuinely low-stakes: if the assistant is down for two hours, people do something else. It becomes high-stakes the moment AI sits in a customer-facing path — support triage, onboarding, anything a customer waits on. Decide honestly which side of that line you are on, and only demand an SLA if you are on the second one. Demanding enterprise terms for a non-critical tool costs you negotiating capital you will want later.
7. Who is accountable internally?
Not a vendor question, but the one most likely to cause the failure.
Name an owner before purchase. That person holds the credentials, watches the usage, handles the renewal, and answers “can I get access?” Unowned subscriptions do not get audited, do not get downgraded when usage falls, and do not get cancelled when the person who championed them leaves.
8. What does the free tier or trial actually prove?
Trials are designed to demonstrate the happy path. Deliberately test the unhappy one:
- Your longest real document, not a sample
- Your most sensitive category of content, to see the refusal behaviour
- A prompt you know is ambiguous, to see how it handles uncertainty
- Concurrent use by three people at once, to see rate limits
A tool that performs well on curated demos and badly on your actual worst case is a tool you will be re-evaluating in six months.
9. What is the renewal trap?
Read three specific clauses:
- Auto-renewal window. Thirty days’ notice is standard; ninety exists and is easy to miss.
- Price protection. Is your rate locked for the term, or can it move mid-contract?
- Annual discount conditions. A 20% annual discount that requires a twelve-month commitment in a category this volatile is often a bad trade. Model pricing has moved repeatedly in the last two years, and mostly downward.
Put the notice deadline in a shared calendar the day you sign. Not the renewal date — the notice deadline, which is earlier and is the one that actually matters.
The one-page version
If you only get ten minutes with a vendor, ask these five:
- Is our data used for training, and does that answer change by plan?
- What is your model deprecation notice period?
- What does a heavy user cost, all-in?
- Can we export everything, and in what format?
- What is the cancellation notice period?
An honest vendor answers all five without hedging. Hesitation on any of them is information.
Frequently asked questions
Do small teams need a formal AI procurement process? Not a formal one, but they need a consistent one. A nine-question checklist applied every time beats an elaborate process applied once and then skipped.
Should we buy annual or monthly for AI tools? Monthly, in most cases. AI pricing and capability are both moving quickly enough that a twelve-month lock rarely pays for itself. Revisit once the category stabilises.
How do we evaluate AI tools without technical staff? Use your own real work as the test. Take five tasks you actually do, run them through each candidate, and compare outputs. That requires domain knowledge, not technical knowledge, and it is more predictive than any benchmark.
What is the most commonly missed cost in AI procurement? Overage. The headline plan is priced for average use, and the person driving the purchase is usually a heavy user whose real consumption sits well above it.
Before you sign
The purchase itself is rarely the mistake. The mistake is the assumption underneath it — that data is not used for training, that the model roster is stable, that the price is the price. Nine questions surface all three in under an hour, and every one of them is cheaper to ask now than to discover later.

