The problem
You run a freight brokerage with 25 people. Three AI sellers have pitched you a quoting helper this quarter. Each one sent a deck. Each deck says "secure", "accurate" and "enterprise-grade". None of them says what happens when a quote is wrong.
Your insurer gave you a vendor questionnaire years ago. You pull it out. Half the questions are about server rooms and password rules. None of them ask what the helper is allowed to do in your load board, or who checks its work, or how you get your money back.
So you do what most owners do. You pick the seller whose salesperson answered emails fastest.
Nobody asked the first question either: does this job need an AI helper at all?
"Many use cases positioned as agentic today don't require agentic implementations."
That is the problem. You have no checklist for judging these sellers, and the old software checklist does not fit. AI helpers act, spend and touch your systems. Software mostly sat there.
Why this keeps happening
The software buying process was built for a product you install and staff use. The questions were about the vendor's company: their security, their uptime, their backups. Those still matter. They just miss the new risks.
An AI helper is different in four ways. It does work on its own, so the result can be wrong without anyone noticing. It can reach into your systems and change things. It runs up a bill one job at a time. And it usually depends on other sellers underneath it, each with their own terms.
Some of what you are shown is not new at all. As of June 2025, Gartner warns of "agent washing": old chatbots and automation tools relabelled as AI agents without much new inside. A written list of what goes in, what comes out and what it touches is the quickest test.
Big platforms have started to filter sellers for you. As of September 2026, AWS Marketplace lists AI helpers from registered sellers, and new listings start with limited visibility. Salesforce's AgentExchange accepts partners only, after a security review. Open directories list whatever is submitted. Filtering helps you shortlist. It does not tell you what one specific service will do in your business.
Here is what skipping the review costs a company your size.
| What you did not check | What it costs you |
|---|---|
| What the service may touch | A quoting tool that can also edit customer records |
| Whether you could try it first | A year's contract for something that did not fit |
| What happens to a wrong result | You eat the loss on a bad quote and pay for it too |
| Who owns the output and your data | Your rate history is training someone else's product |
| How you leave | Months of notice, no export, a login left open |
How to fix it
If you searched for an ai agent vendor evaluation checklist 2026 or an ai vendor due diligence questionnaire template, use this. Seven items. Ask them of every AI seller, and write down where each answer came from. An answer on the product page is worth more than an answer in an email.
- Is the offer written down in full? The result, the price and its unit, what goes in, what comes out, what it needs from you. If any of that is "contact us", the review stops here.
- What exactly may it touch in my systems? Named actions, one per line. "Read loads" is an answer. "Integration with your load board" is not. Ask which actions need a person to approve first. The least-privilege guide is the standard to hold them to.
- Can I try it on sample data before paying? A good trial states how many free runs you get and needs none of your systems connected. A trial that needs your live systems first is asking for access before it has earned the sale.
- What happens to a wrong result? Who decides it was wrong, how fast, and what do you get back. You want a named person who reviews, and a refund when they reject.
- Who owns the output and my data? Three questions. Do I own what it produces? What does the seller keep afterwards, and for how long? Can my data be used to improve their product? Get the answers in the terms, not in an email.
- Where does it run, and who holds the keys? On the seller's machines or on the platform's. Whose uptime you depend on follows from that. So does who holds any login to a data source underneath.
- How do I leave? Can I cancel and cut off access myself? Do I keep the records of what ran? Is anything left behind?
Item five deserves extra care. Lawyers who write these deals say the usual "you own the work product" clause was written for human consultants.
"Standard work-product assignment clauses drafted for consulting and systems integration engagements may not map cleanly onto outputs from AI agents because those clauses often assume human authorship."
Mayer Brown, Key Contract Issues in Agentic AI Implementation and Integration Deals (June 2026)
So do not assume the boilerplate covers you. Ask for a sentence that names AI output.
On the ai agent sla question, be specific. Uptime is the least useful promise for an AI helper. Ask instead for four things. Will the service be reachable when my helper calls it? How fast does a person review a flagged result? Do I get a record for every job? Will I be told before the version I depend on changes? The reliability guide shows how to test the first one yourself. As of June 2026, Mayer Brown suggests measuring an AI helper by its completion rate, how often it hands off to a person, and how often its work needs redoing. Those are plain numbers you can ask any seller for.
Fold these into your ai vendor contract requirements checklist and keep the old software questions too. They still cover the seller's company. These cover the service.
What BlueBear's marketplace does about it
BlueBear's marketplace is built so most of the seven items are answered before you talk to anyone. It is a pilot today. Publishing is by invitation, and BlueBear lists each service on the seller's behalf after reading what it does. That means someone has read the offer before you do. It does not replace your review. It changes how much of it you can do from your desk.
Items one, two, three and six are on the public page. Every service has a page that states the result and the price in credits, where one credit is one US dollar. The same page lists what it needs from you, the exact actions it may take, any trial and its limits, and where it runs. A gap on that page is an answer too.
Item four is handled by a named person and a refund rule. Results that need judgment go to a reviewer queue. A named reviewer checks the result, fixes it or releases it. If they reject it, you are refunded. Ask the seller how fast their reviewers turn things around, because the seller chooses who sits in that queue.
Item seven is yours to control. Paying for a service and letting it reach your systems are two separate switches, and you hold both. No seller holds your login. And there is a receipt for every job, kept with your other records. It is signed, so it cannot be edited, and it stays valid after you cancel. If you already keep an audit trail for your AI helpers, that receipt is one more entry in it. The guide on making AI actions auditable sets the bar the receipt is built to meet.
Item five is still answered in writing by the seller. Ownership of outputs, retention and training use live in the seller's terms. Ask directly before buying anything that touches confidential data.
None of this replaces a contract. Ownership and uptime still belong in writing with the seller. The platform answers the items a contract handles badly: what ran, what it cost, and whether a person checked it.
What the pilot does not offer: BlueBear does not publish a standard service-level promise across sellers. The platform guarantees the receipt, the refund on reject and the two switches. Uptime and review speed are the seller's promises, so write them into your own agreement. There is no league table of sellers, and sellers are paid by hand.
What to do next
Take the seven items to the three decks on your desk. Score each seller on where their answers came from. Public page, terms, email, or nowhere. The seller with the most answers in public is usually the seller with the least to hide.
Keep the finished checklist next to the version of the offer you approved. When the service updates, re-check only the items that changed. To see what a listing that answers most of the checklist looks like, open the public marketplace and read a service page before you read a sales deck. If a service passes, the guide to letting your helpers buy within a budget is the next step.
Questions people actually search for
- what should i ask an ai seller before signing
Seven things. Is the offer written down in full, with price and unit? What exactly may it touch in my systems, action by action? Can I try it on sample data first? What happens to a wrong result, and do I get a refund? Who owns the output and my data? Where does it run and who holds the keys? How do I leave? Note where each answer came from; a public page beats an email.
- what should an ai agent sla actually cover
Not just uptime. Ask for four promises. The service is reachable when your helper calls it. A person reviews a flagged result within a stated time. You get a record for every job. You are told before the version you depend on changes. On BlueBear's marketplace the platform guarantees the receipt and the refund on reject; uptime and review speed are the seller's promises, so put them in your contract.
- should i try an ai service before paying for it
Yes, and look at what the trial asks of you. A fair trial runs on sample data, states how many free runs you get, and needs none of your systems connected. A trial that needs your live systems first is asking for access before it has earned the sale. Treat the trial as a test of the offer's claims about inputs, outputs and price, not as a demo.
- who owns the work an ai service produces
Usually you do, but only if the seller's terms say so. Ask three questions before you buy. Do I own what it produces? What does the seller keep after the job, and for how long? Can my data be used to improve their product? On BlueBear's marketplace these answers still live in the seller's written terms, so ask directly before buying anything that touches confidential data.